Codex CLI 0.160.0 Arrives: Workspace-Default Sessions, Saved Permissions, and Guardian Handoff Context
OpenAI’s Codex changelog records Codex command-line interface (CLI)A text-based interface for running commands and tools. Open glossary entry 0.160.0 on 1 October 2026. The release entry identifies three related changes: sessions started outside a project can use workspace defaults when policy permits; saved permissions can be restored when a task or session is resumed; and Guardian, an optional automated reviewer for approval requests, can receive earlier user instructions and selected agent-handoff context when that review is enabled. This is a dated CLI release-note event, not evidence that every Codex surface, account, workspace or operating system gained a generally available capability on that day. For users and administrators, the immediate job is to determine whether those three behaviours affect existing control boundaries before adopting the version.
Evidence checkpoints
Documented point: The OpenAI changelog records Codex CLI 0.160.0 and its published npm install command; npm is the standard package manager for Node.js. [OpenAI documentation: Changelog Node.js Learn: An introduction to the npm package manager]
Documented point: OpenAI’s implementation checks local execution, configuration and managed policy before applying projectless workspace defaults. [OpenAI GitHub pull request #49160]
Documented point: OpenAI’s Guardian implementation identifies guardian_root_handoff_context as disabled by default. [OpenAI GitHub pull request #49057]
Documented point: OpenAI’s review implementation explains why earlier instructions, restrictions and revoked permissions may matter before side-effecting actions are approved. This describes the cited Guardian implementation path; verify the active client and workspace controls before relying on it. [OpenAI GitHub pull request #49036]
Documented point: OpenAI documents /permissions, /model and /review as Codex CLI controls; the 0.160.0 release does not announce a model launch. Check which commands the installed Codex command-line interface and workspace policy actually support before relying on them. [OpenAI documentation: Codex CLI]
Documented point: Current defaults recommend the Auto mode with workspace-write permissions and on-request approvals for version-controlled folders, but managed policy can constrain behaviour. [OpenAI documentation: Agent approvals and security]
What Codex CLI 0.160.0 changes
The three changes belong to different stages of a Codex workflow and should not be treated as one broad permissions feature. Projectless defaults affect how a new session begins outside a project. Permission restoration affects how an existing saved task or session resumes. Guardian context affects what an optional automated reviewer may receive when assessing an eligible approval request. Keeping those stages separate prevents a successful check in one area from being mistaken for evidence about another. Model Context Protocol (MCP)A protocol for connecting artificial intelligence applications with tools and data sources through defined interfaces. Open glossary entry connection security is discussed in a separate earlier release guide linked below, not listed as one of these three 0.160.0 changes.
- Starting outside a project: the changelog says workspace defaults can apply to sessions started outside a project. OpenAI’s merged pull request (PR)A proposed set of repository changes submitted for review before integration. Open glossary entry #49160 narrows that behaviour: workspace-write permissions and granular approval defaults apply only where local execution, configuration and managed policy allow them. On Windows, the implementation can also require a sandbox-setup prompt.
- Resuming with saved permissions: the release restores saved task permissions on resume unless the operator explicitly overrides them. Restoration concerns a previously saved permission selection; it is not documented as an expansion to unrestricted access, a way around current policy or a guarantee covering every legacy session.
- Supplying context to Guardian: the release adds opt-in retrieval of earlier user instructions and selected context associated with agent handoffs. The upstream implementations are disabled by default and impose conditions and bounds. They do not provide Guardian with an assured, complete transcript.
A practical first step is to map each change to an owner and a decision. The CLI user can validate session start and resume behaviour in a disposable repository or non-sensitive directory. A workspace administrator should verify the applicable managed policy and whether optional Guardian facilities are authorised. A human code owner or deployment approver must still inspect proposed changes and remain accountable for external side effects and production release decisions.
Example mapping, not a product guarantee: a team might assign “projectless start” to its developer-tools maintainer, “resumed permissions” to the person responsible for local Codex policy, and “Guardian context” to the workspace administrator and security reviewer jointly. The output could be a three-row review record containing the expected policy, the observed prompt or permission state, and the human decision. It should not contain credentials, private source material or copied conversation history.
Decision rule: upgrade directly only if the organisation understands all three affected paths and can test them without relaxing its existing sandbox, approval or managed-policy settings. If the team depends on projectless starts, long-lived saved sessions or optional Guardian review but cannot yet inspect the relevant controls, pinning the existing version while preparing a bounded validation is the more conservative trade-off. That delays access to the release changes but avoids adopting altered session behaviour without an accountable review.
The exact release scope—and what sits outside it
The assigned 0.160.0 release sources establish the projectless-session default, resumed-permission restoration and bounded Guardian context changes. Evaluate those three behaviours against separate acceptance criteria rather than blending them into a general claim that 0.160.0 is “safer”, “more autonomous” or universally preferable. No assigned source supplies a release-specific performance result or says these changes fix an earlier settings-persistence defect.
The changelog publishes the npm installation command npm install -g @openai/[email protected]. That command identifies the requested package version; it does not establish that a particular account, workspace or optional feature is entitled, configured or available. Before running it, record the currently approved version and determine how the installation is managed. In a centrally administered environment, follow the organisation’s package and change-control process rather than replacing a managed installation locally.
The 1 October entry does not announce a named model, a model-selector change, a new default model, a price change, a usage allowance, a regional list, a staged-rollout timetable or a comprehensive operating-system matrix. Current Codex CLI documentation exposes commands such as /model, /permissions and /review, but their presence does not turn this release into a model launch. OpenAI’s current plan documentation, accessed on 4 October 2026, also distinguishes local Codex access from Codex Cloud eligibility and says usage and workspace conditions vary.
Procedure: prepare a release-scope note before installation. Under “included”, list only the three behaviours stated in the changelog. Under “not established”, list model availability, pricing, rollout, platform coverage and plan entitlement. Then compare that note with the actual reason for upgrading. If the request is really seeking a new model or additional usage, 0.160.0’s release notes do not support the request and a separate official account or product check is required.
Worked fictional example: suppose an administrator receives the request, “Install 0.160.0 so everyone gets Guardian and a newer model.” The source-faithful rewrite would be: “Assess 0.160.0 for conditional projectless workspace defaults, restoration of saved permissions on resume, and opt-in Guardian context. Confirm separately whether our workspace can use the optional reviewer and which models our current account exposes.” The rewrite narrows the change request without predicting availability or an outcome.
Decision rule: approve the upgrade rationale only when it names a documented 0.160.0 change and identifies the policy owner who will validate it. Reject or return a rationale based on an undocumented model, price, plan, rollout or performance assumption. The trade-off is administrative effort versus traceability: a short source-scoped record takes longer than an informal update, but it makes later permission observations easier to interpret.
For a controlled upgrade on an earlier, different release—with background server startup, conversation forking, remote import, draft recovery and network-policy changes rather than the 0.160.0 workspace defaults—consult Upgrade to Codex CLI 0.157 and Test the New Workflow, then compare that release’s testing procedure without treating its feature list as evidence for this one.
Projectless sessions are conditional, not an unrestricted workspace mode
“Outside a project” is essential to the first feature’s meaning. It describes a session that does not begin in a recognised project context; it does not mean Codex may treat every arbitrary directory as writable. PR #49160 says implicit workspace-write permissions and granular approval defaults are conditional on local execution, configuration and managed policy. A trust or sandbox setup requirement can therefore remain relevant, and an administrator’s restrictions can still constrain the result.
This distinction matters because a project directory and a workspace permission are different concepts. A project can provide a bounded version-controlled context in which changes are readily reviewed. A projectless start lacks that starting context, while workspace-write describes an execution permission. The release can choose applicable workspace defaults for the former without abolishing the controls governing the latter.
OpenAI’s current approvals guidance, accessed on 4 October 2026, recommends the Auto mode with workspace-write permissions and on-request approvals for version-controlled folders, while noting that managed policy can constrain behaviour. It also states that approval requests route to the user by default. This guidance does not justify copying the same permission setting into an unversioned, sensitive or broadly mounted directory. The directory’s contents and the organisation’s policy remain part of the decision.
Bounded validation procedure:
- Create or select a disposable directory containing only synthetic files. Do not use a home directory, production checkout, mounted secrets directory or folder containing customer data.
- Record the applicable local configuration and managed policy without pasting either confidential policy text or authentication material into a prompt.
- Start Codex CLI 0.160.0 outside a project and inspect the permission state through the documented
/permissionssurface. - If the client presents a trust, approval or Windows sandbox-setup prompt, handle it according to existing policy. Do not select a broader option merely to make the test proceed.
- Ask for a harmless, local proposal such as creating a file named
validation-note.txtcontaining non-sensitive sample text. Review any requested write and the resulting diff or file manually. - End the session and retain only a sanitised observation: starting directory type, expected policy class, permission state shown, prompts encountered and reviewer decision.
Fictional input example: “In this disposable directory, propose creating validation-note.txt with the text ‘0.160 projectless validation example’. Do not access parent directories, the network, environment variables or credentials. Ask before writing if the active policy requires approval.” This is a suggested method, not a promise that a particular interface sequence or approval will appear.
Example observation: “Project context: none; directory: synthetic; expected: managed policy remains controlling; observed permission state: [human records the displayed value]; additional prompt: [recorded if present]; file and diff reviewed: yes/no.” The bracketed fields must be completed from the actual installation. Fabricating them would turn a validation record into an unsupported test result.
Decision rule: accept the projectless-start behaviour only if the displayed permissions and any requested action remain within the pre-approved local and managed-policy boundary. Stop if the session appears broader than expected, if the effective policy cannot be determined, or if the directory includes consequential data. Convenience is the benefit of workspace defaults; reduced visibility into why a default was selected is the corresponding operational risk.
Saved permissions apply to resumption, with an explicit override path
The second feature concerns continuity. PR #49160 says saved task permissions are restored when the task is resumed unless they are explicitly overridden. This differs from projectless defaults because it starts with a stored session state rather than selecting defaults for a new session. It also differs from an escalation: a restored selection is still subject to present configuration, trust requirements, local execution conditions and managed policy.
An operator should not infer that every old session will resume identically. The cited upstream documentation does not provide a complete compatibility matrix for sessions created by all earlier versions, nor does it specify every outcome after policy or configuration changes. A meaningful validation therefore needs both a stable-policy case and a changed-policy case, with the latter performed only by an authorised administrator.
Bounded resume procedure:
- In a synthetic version-controlled folder, start a disposable session and use
/permissionsto select an option already permitted by policy. - Record the selected profile in the test note, then save or leave the session by the normal supported route. Do not place secrets or untrusted external text in the conversation.
- Resume that same saved session and inspect
/permissionsbefore requesting any file modification or command. - Explicitly override the resumed permission selection with another policy-compliant option, then inspect it again. This checks the documented override path without seeking elevated access.
- If authorised to test a policy-change case, have the administrator tighten a disposable test policy through the approved management process. Resume a separate saved session and verify that current constraints still govern. Do not weaken production policy to create the test.
- Have a human review the session identity, active directory, permission display and any resulting file change before marking the case complete.
Fictional example: a saved session was created with a policy-permitted workspace-write profile. On resumption, the tester first records whatever /permissions displays, then explicitly chooses a more restrictive read-only option for the next request. The relevant evidence is the observed state before and after the explicit override. The tester must not pre-fill the record with “restored successfully”, because no hands-on test has been reported here.
Decision rule: treat a resumed session as unverified until a human has checked its identity, directory and active permissions. Continue only if the displayed state is expected and currently allowed. If a saved selection conflicts with a newer managed requirement, the newer control must win; pause and escalate if the effective result is ambiguous. Restoring a familiar profile saves setup time, but it can also make an old assumption less visible, so inspection before consequential work is the appropriate trade-off.
For the earlier Codex CLI 0.158 release’s distinct security changes—MCP connection secrets, WebSocket bearer tokens, elevated approvals and sandbox tests—see Upgrade to Codex CLI 0.158. That earlier guide covers its own release; verify the later 0.160.0 workspace-default and resumed-permission changes separately using this article’s checks.
Guardian receives selected context only through opt-in, conditional paths
The Guardian additions address missing review context rather than granting Guardian complete memory. PR #49036 explains that a transcript supplied to Guardian may omit earlier instructions, restrictions or revoked permissions that matter before a side-effecting action is considered. Its conversation-history retrieval is disabled by default and depends on Apps being enabled, along with the parent session’s live Apps connection and conversation identity. Current policy is checked for every call, and the PR describes an estimated default output budget of 4,000 tokens.
PR #49057 separately adds the disabled-by-default guardian_root_handoff_context feature. It selects worker-specific evidence from recorded handoff calls: three root messages preceding each relevant handoff and the three latest root messages, subject to existing shared caps and documented fallbacks. That bounded selection is neither a complete conversation nor proof that every instruction was captured, attributed correctly or still applies.
The distinction between the two Guardian paths is operationally important. Conversation-history retrieval seeks earlier instructions through a conditional Apps-connected route. Root handoff context selects limited root messages associated with handoffs. Enabling or observing one does not establish that the other is enabled, available or complete. The cited upstream documentation does not provide a stable, reader-ready administrator procedure for exposing either feature through an end-user control, so operators should not invent configuration names or assume that an upstream feature flag is supported in their workspace.
Suggested validation design: first ask the workspace administrator whether optional Guardian review is authorised and whether an official control for either context path is available in the deployed environment. If it is not verifiably available, record “not assessed” rather than attempting unsupported configuration. If it is available, use a synthetic conversation containing an early restriction, a later harmless task and a clearly labelled fictional agent handoff. Submit only an eligible, reversible request and inspect the review material that the supported interface actually exposes.
Fictional prompt sequence: an early message says, “Do not use the network or modify files outside ./sandbox.” A later message asks a worker to propose editing ./sandbox/example.txt. A handoff note identifies that bounded task. A human then checks whether the resulting review preserves the restriction sufficiently to support a decision. This sequence is an example of test construction, not evidence that Guardian will retrieve a particular message or produce a particular judgement.
Keep untrusted repository text, copied web content, personal data, access tokens, private keys and production credentials out of prompts and test conversations. Earlier context becoming available to a reviewer does not establish a new retention or training policy. OpenAI’s plan documentation, accessed on 4 October 2026, distinguishes consumer data controls from business terms; organisations needing a privacy assurance specific to this selected Guardian context should obtain an applicable official statement rather than infer one from the release note.
Decision rule: use Guardian context only if an authorised administrator can verify the supported opt-in path, the test data is non-sensitive and a human remains responsible for the final decision. Reject any workflow that assumes Guardian has the full transcript or can replace code review, approval governance or deployment controls. More context may help an optional reviewer evaluate an eligible request, but selective retrieval introduces incompleteness and additional data exposure that must be weighed deliberately.
A source-bounded upgrade gate
A defensible 0.160.0 decision needs three independent pass conditions. First, a projectless synthetic session must remain within current local configuration and managed policy. Second, a resumed test session must present an expected, inspectable permission state and permit an explicit policy-compliant override. Third, any Guardian assessment must use a supported opt-in path and be treated as bounded context for an automated reviewer, not as accountable human approval.
Record each condition as pass, fail or not assessed, with a human reviewer and date. “Not assessed” is appropriate when the optional Guardian control is unavailable or the organisation does not use it; it must not be silently converted into a pass. For consequential code changes, external side effects and production deployment, require separate human review regardless of the validation status.
Example decision: “Approve 0.160.0 for a limited developer-tools cohort after projectless and resume checks; leave Guardian context unassessed and disabled pending an administrator-supported control.” An equally valid outcome could be to retain the earlier approved version because permission ownership is unclear. Neither result can be inferred from the changelog alone.
Final rule for this release event: upgrade for the documented CLI behaviours only when the organisation can observe them without weakening approvals or policy. Defer where the effective permission state, optional Guardian configuration or human ownership is uncertain. This preserves the practical value of session continuity and richer review context while keeping current controls—not the convenience of a new version—as the deciding authority.
Run a bounded, no-write upgrade validation
The operator’s job is not to prove that Codex CLI 0.160.0 has broader autonomy. It is to determine whether three control-sensitive behaviours differ from the team’s approved baseline: starting outside a project, restoring permissions when a saved session is resumed, and supplying selected context to Guardian review. OpenAI’s changelog dates 0.160.0 to 1 October 2026, but it does not provide a platform matrix, entitlement table or migration guarantee for older sessions. The procedure below is therefore a suggested validation method, not a product guarantee or a report of hands-on testing.
Use a disposable directory containing no repository, deployment configuration, credentials, customer data or production connection. Keep untrusted data and secrets out of prompts. Do not let the validation modify files, install dependencies, call external services, open pull requests, change infrastructure or deploy software. A human reviewer should record the observed status and permission displays, compare them with the organisation’s approved policy, and decide whether the upgrade can proceed.
1. Establish the administrative boundary before installing
Use the projectless-session conditions established above as the acceptance boundary. In this check, “outside a project” means Codex does not treat the test directory as a project; it does not suspend organisational controls. Record the displayed profile and any restriction before proceeding.
Before changing the installed version, ask the workspace administrator to identify the applicable managed policy, the expected local-execution state and the approved permission profile for both project and projectless use. Also record whether the operating environment requires an additional trust or sandbox step. The upstream implementation notes that Windows can require a sandbox-setup prompt, so the absence or presence of that prompt must not be generalised to every operating system.
A practical pre-install record can contain the following fields:
- the currently installed CLI version and the proposed target, 0.160.0;
- the account and workspace category used for the check, without recording credentials or tokens;
- whether local Codex use is enabled by the workspace administrator;
- the expected projectless permission mode and approval route;
- the local configuration files that administrators consider authoritative, recorded by path only if that path reveals no secret;
- whether managed policy is expected to constrain or override a local preference;
- the human owner who will approve or reject the upgrade.
Decision rule: do not continue if the team cannot state what permissions and approval routing should apply. Observing a setting without an approved baseline cannot establish whether restoration or default selection is correct. The trade-off is a slower upgrade, but it avoids treating whatever appears in the user interface (UI)The controls and visual surfaces through which a person interacts with software. Open glossary entry as implicitly authorised.
Fictional example: Northstar Engineering records that its disposable projectless directory should not receive network access, that approval requests should continue to route to the operator, and that no deployment command may run. “Workspace-write” may be an expected workspace default, but the test still prohibits writing. This separates the product’s displayed permission profile from the narrower actions allowed by the validation plan.
2. Pin the version and preserve a comparison point
The 1 October changelog publishes npm install -g @openai/[email protected]. An administrator may use that pinned command if npm is the team’s approved package-management route. Do not infer from the changelog that every distribution channel or platform build became available simultaneously. If the organisation uses a managed package, image or internal mirror, follow that route instead and verify that its resolved package is 0.160.0 before assessing release behaviour.
First capture the current version through the organisation’s normal package or CLI inspection method. Then retain the existing configuration and managed-policy state unchanged. Install the pinned version only in the designated test environment. Do not loosen approvals merely to make the new behaviour visible: a blocked default can be evidence that policy is working as intended.
After installation, verify the resolved version without asking Codex to inspect or edit a repository. If the reported version is not 0.160.0, stop and investigate package resolution rather than continuing with an ambiguous installation. Do not convert this step into a model test: OpenAI documents /model as a CLI control, but the 0.160.0 entry does not announce a model-selector or default-model change.
Decision rule: proceed only when the package version is known, existing policy remains in force and rollback is available through the team’s approved software-management process. The meaningful trade-off is reproducibility versus convenience: pinning makes the observation attributable to 0.160.0, whereas an unpinned install may resolve a later release.
Fictional example: Northstar’s administrator approves the published pinned npm command for an isolated workstation. The operator records “target package: 0.160.0” and checks the installed version. They do not record an access token, paste configuration contents into a prompt or ask Codex to diagnose the company’s policy service.
3. Observe a projectless start without granting it work
Create or select an empty disposable directory that is deliberately outside a project. Confirm that it contains no version-control metadata, source tree, environment file, cloud configuration, deployment manifest or mounted production data. Starting in a normal repository would test project behaviour and could conceal the projectless default that this release changes.
Start a fresh CLI session from that directory. Do not ask the agent to create a file as a quick proof of write access. Instead, inspect the effective permission and approval state through documented controls. OpenAI’s current CLI documentation, accessed 4 October 2026, documents /permissions; if the installed client shows a non-sensitive session identifier or workspace indication through a supported display, record it separately without asserting an undocumented command exists. Have a human transcribe only the effective mode, approval posture and relevant session context into the validation log.
Compare the display with the step-one baseline. A restrictive result may reflect managed policy rather than a defect; do not bypass a trust or sandbox prompt to force the expected profile. Log the actual result and seek the policy owner’s interpretation.
Use a harmless prompt such as: This is a validation example. Describe the current directory at a high level, but do not create, edit or delete files; do not run commands; do not use the network.
Treat any response as illustrative rather than proof of enforcement. The authoritative evidence for this stage is the displayed effective status, permissions and approval route, checked by a human against policy—not the model’s prose about what it believes it may do.
Decision rule: accept the projectless observation only if the displayed profile is allowed by managed policy and the session has performed no write, command, network or deployment action. Escalate an unexpectedly permissive display to the administrator before trying to exercise it. The trade-off is that a no-action check cannot demonstrate every sandbox boundary, but it can identify a configuration mismatch without creating side effects.
Fictional example: Northstar starts in an empty folder named codex-0160-observation. The operator records the visible status and permissions, then exits without asking Codex to create a sample file. If workspace-write is displayed despite the test’s no-write rule, the operator records that fact but does not use the permission. Human review determines whether the display matches Northstar’s approved workspace default.
4. Build a safe saved-session fixture
Resumption uses previously saved permissions, subject to the current controls described above; this is not a test of new projectless-session defaults. Choose a known test session, record its previous selection and compare it with the resumed display.
In the disposable location, select only a profile that the administrator has approved for this validation. Use /permissions to inspect the effective choice, and do not select full access or weaken an approval requirement for test convenience. Record the permission display and any verified non-sensitive session identity immediately before saving or leaving the session. Give the session a benign instruction that requires no action, such as: Validation example: remember that this session must not write files, execute commands, access a network or initiate a deployment.
Exit through the normal saved-chat workflow documented by OpenAI rather than terminating the process in a way intended to simulate failure. The aim is to compare a supported return to a saved chat, not to test crash recovery. Record a non-sensitive session identifier only if the CLI exposes one and organisational policy permits it; do not place conversation identifiers, account metadata or internal paths into an external prompt.
Decision rule: the fixture is usable only if its pre-exit status and permission profile were recorded and no side-effecting operation occurred. If the operator changed permissions but did not capture the final state, create a new fixture rather than guessing. The trade-off is additional setup time in exchange for an attributable resume comparison.
Fictional example: Northstar creates “Fixture A” with the organisation’s ordinary approval routing and a strict test instruction. Its log contains a human-transcribed /permissions snapshot labelled “fresh session before exit” and, if shown through a supported display, a separately checked non-sensitive session label. The log does not contain prompt secrets or a complete conversation transcript.
Restoring permissions on resume is a different question from managing accumulated patches and terminal output in a long-running session. For that separate context-management problem, Codex Automatic Recaps vs Experimental Context Management examines long-session continuity; its guidance does not establish which permissions 0.160.0 restores.
5. Compare status and permissions immediately after resumption
Resume the saved fixture without supplying a permission override. Before entering any task prompt, inspect /permissions again and independently confirm that the intended saved session was resumed through the supported interface available in that installation. Compare each relevant observation with the pre-exit record. This ordering matters: if the operator first changes permissions or begins work, the team cannot distinguish restoration behaviour from a subsequent explicit choice.
The comparison should classify—not merely note—each difference:
- Unchanged and policy-compliant: the resumed profile matches the saved record and remains allowed by current controls.
- Changed because current policy is more restrictive: retain the restriction and ask an administrator to verify the policy path; do not attempt to restore the older, broader state manually.
- Unexpectedly broader: stop without exercising the permission, preserve non-sensitive evidence and obtain administrative review.
- Unavailable or reset: record the result as an unresolved compatibility observation rather than claiming that restoration is universally absent.
- Explicitly overridden by the operator: treat this as a separate test and do not mix it with the default-resume result.
Why pair the permission display with a session-identity check? /permissions is the documented operator surface for the permission choice, while a supported session display or saved-chat selection can help a human confirm that the intended fixture was resumed. Neither observation should be treated in isolation as authorisation to perform a consequential action. The administrator must reconcile the observations with managed policy and the approved validation boundary.
Next, exit without doing work and resume the same fixture using an explicit, approved permission override if the CLI’s available workflow permits it. Choose a profile no broader than the baseline. Capture the permission display and verified session identity once more. This distinguishes “restored because no override was supplied” from “operator deliberately selected a different profile”. The official implementation says restoration can be explicitly overridden; it does not say an override defeats current policy, trust checks or approvals.
Decision rule: recommend the upgrade only if a no-override resume produces a policy-compliant state, an explicit override remains bounded by current controls, and a human can explain every difference. Reject or pause the upgrade if a broader profile appears unexpectedly, even if no command has yet been run. The trade-off is that this procedure validates state selection and visibility, not every possible legacy-session migration.
Fictional example: Northstar’s Fixture A is resumed first without an override. The operator captures the permission display and checks the session label before typing any task. They then close it and perform a separately labelled “Fixture A—explicit restrictive override” observation. Sample log entries might read “matches saved display”, “restricted by current policy” or “requires administrator investigation”; these are example classifications, not predicted product outputs.
6. Check approval routing without triggering an approval
Permissions and approvals are related but distinct. A sandbox or workspace permission describes what execution scope may be available; approval routing determines who is asked about eligible actions. OpenAI’s approvals guidance, accessed 4 October 2026, says approval requests route to the user by default, while auto-review is an optional reviewer for eligible interactive requests. It does not review every action and is not a human sign-off.
Inspect the current permission and approval configuration through the documented CLI controls, but do not deliberately request a risky command merely to make an approval box appear. Record whether the resumed session’s displayed routing matches the fresh-session baseline and the managed workspace policy. If a difference cannot be explained, stop before any action that would invoke it.
A low-risk example prompt is: Validation example only: state that no command, file change, network request or deployment is authorised in this session. Do not attempt any of them.
Again, this is not an enforcement test. It helps keep the conversational task bounded while humans inspect the actual settings.
Decision rule: no release candidate should pass solely because Guardian or another automated reviewer is configured. A named person must remain accountable for accepting code changes, external side effects and production deployment. The trade-off is reduced automation during validation, balanced against preserving the team’s established approval chain.
Fictional example: Northstar sees that approvals continue to route to the operator. It records that configuration but does not provoke a shell command. If an optional auto-review setting is present, the team documents it separately and still requires the release engineer to approve any future production change.
7. Treat Guardian context as a separate, opt-in assessment
Saved permissions do not establish whether either separate Guardian context path is enabled. Run the optional Guardian check only with approval, using the prerequisites and disabled-by-default conditions stated in the source-boundary section above. Do not guess an administrator setting.
Accordingly, the bounded procedure is administrative rather than activation-led. Ask the workspace owner whether either feature is approved and exposed through a supported control in that environment. If the answer is unknown, leave it disabled and record “not assessed”. Do not edit undocumented feature flags or lower policy gates to force a result.
If the organisation already has an approved, supported configuration, prepare a synthetic conversation containing no secrets or untrusted text. Include an early restriction, a later worker handoff and a final harmless review question. A suitable fictional sequence is: “Do not write files”; “Worker A may analyse naming only”; “Handoff to Worker B for a review of the proposed explanation”; then request review without authorising execution. A human should inspect what context is reported as available and verify that no side-effecting action follows.
For the bounded handoff windows and history-output limit, see the Guardian source-boundary section above. Record whether a synthetic restriction was actually visible; missing context remains possible, so retain human review and do not infer comprehensive memory.
Decision rule: use Guardian context only as supplementary review evidence. If a consequential decision depends on an earlier restriction, revoked permission or provenance detail, a human must inspect the authoritative conversation and current policy directly. The trade-off is context relevance versus completeness: selective evidence may reduce irrelevant material, but it cannot guarantee perfect memory or provenance.
Fictional example: Northstar’s synthetic handoff places “no writes” near an earlier handoff. Even if that instruction appears in Guardian’s selected context, the release engineer still checks the root conversation and current permission display. If it does not appear, the team records the documented bounded-selection limitation rather than concluding that Guardian ignored policy.
The new opt-in handoff context is supplementary review evidence, not permission for automated approval or an unreviewed pull request. For a general Codex Guardian workflow rather than 0.160.0 release instructions, Codex Guardian Auto-Review Playbook discusses auto-review around pull requests; a human code owner must still assess the proposed changes and any deployment decision.
8. Apply the no-write, no-deploy acceptance checklist
Conclude with a human-reviewed checklist, not an autonomy claim. Mark each item “pass”, “fail” or “not assessed”, and attach only non-sensitive observations:
- the installed CLI resolves to 0.160.0 through the approved distribution route;
- the test starts outside a project in an empty, disposable directory;
- local execution, configuration, managed policy and any trust or sandbox prompt remain unchanged;
/permissionsis captured for the fresh projectless session, alongside any independently verified non-sensitive session identity that the installed client exposes;- the saved fixture’s status and permissions are captured immediately before exit;
- the permission display and saved-session identity are checked immediately after resume and before any task;
- any explicit override is tested separately and is no broader than the approved baseline;
- approval routing remains attributable to the expected user or approved reviewer configuration;
- no file is created, edited or deleted;
- no shell command, dependency installation, network request or external side effect is initiated;
- no pull request, infrastructure change or deployment is attempted;
- Guardian features remain disabled unless an administrator confirms a supported, approved opt-in path;
- any Guardian assessment uses synthetic, non-sensitive context and is treated as incomplete evidence;
- a human reviewer signs the upgrade decision and records unresolved differences.
Final decision rule: upgrade only where the observed projectless default and resumed profile both conform to current policy, explicit overrides remain constrained, and the approval chain still has an accountable human. Pause where the state is unexpectedly broader, the applicable policy is unclear, or the team would need to weaken a control to complete the check. A “not assessed” Guardian result need not block an upgrade if those opt-in features are not part of the intended workflow, but it must not be presented as proof that they are unavailable or safe.
Fictional example: Northstar’s review board can approve 0.160.0 for ordinary CLI use while leaving Guardian context “not assessed” because no supported workspace control has been confirmed. It could instead pause deployment if the resumed permission display is broader than the recorded baseline. In either case, the conclusion is about compatibility with Northstar’s controls—not a general claim that 0.160.0 is more autonomous, universally available or safe without human oversight.
Set the boundary for Guardian-assisted approval
Guardian’s context changes in Codex CLI 0.160.0 should be assessed as two separate, opt-in review inputs rather than one general memory feature. OpenAI’s changelog, dated 1 October 2026, describes retrieval of earlier user instructions and context from agent handoffs. The underlying PRs show that conversation-history retrieval and handoff-aware root context are disabled by default. Enabling one does not establish that the other is enabled, available or suitable for a particular workspace.
The practical decision is therefore not simply whether Guardian should be “on”. An administrator must decide which source of context an eligible automated review may inspect, under which policy, and whether the resulting evidence is sufficient for the proposed action. A bounded validation should preserve existing sandboxing, approval routing and managed policy. It should use synthetic, non-sensitive instructions and stop before any external side effect. Human review remains necessary for accepting code, changing external systems or deploying to production.
Separate history retrieval from handoff-root selection
Conversation-history retrieval follows the opt-in, Apps-connected prerequisites and live policy checks established above. Here the relevant question is whether a specifically approved synthetic restriction survives that route; a failed prerequisite must be logged rather than treated as a model finding.
Handoff-aware root context is a separate opt-in route governed by the bounded message-window conditions above. Test it with a synthetic recorded handoff and log which instruction, if any, appears in the review evidence; do not infer that a complete transcript was supplied.
To validate the distinction, create two separate test cases in an approved non-production workspace. In the first, place an earlier restriction in the parent conversation but do not create an agent handoff. In the second, create a permitted handoff around a synthetic task and place distinct, harmless constraints immediately before it. Ask the reviewer to identify which constraints are visible, but do not ask it to execute a command. Record the enabled feature, connection state, current policy and selected evidence for each case. Do not infer that a result from one path proves the other path works.
Example, not a product guarantee: a synthetic parent instruction might say, “Review proposed edits only; do not run commands or access a network.” A later synthetic handoff might say, “Worker B may inspect fixtures/example.json but must not modify it.” The expected human task is to compare Guardian’s cited basis with the known fixture, not merely to accept a safe-looking verdict. If the first restriction is absent, investigate whether history retrieval was disabled, unavailable, disconnected or denied by current policy. If the handoff restriction is absent, investigate whether the relevant handoff was recorded and fell within the bounded selection. Do not respond by widening permissions.
Decision rule: enable only the narrower context path needed for an identified approval problem. If earlier parent instructions are essential, assess conditional history retrieval. If worker-specific handoff evidence is the issue, assess handoff-root selection. If neither path can reliably expose the controlling restriction in a synthetic validation, retain direct human approval and redesign the task so that the decisive constraint is present in the immediate approval request.
Check live connection and current policy independently
Apps availability and policy authorisation answer different questions. A live parent Apps connection and usable conversation identity can make retrieval technically possible; the per-call current-policy check determines whether a particular lookup remains permitted. A previously successful lookup does not establish continuing authority. Conversely, a policy that allows retrieval does not repair a missing or expired live connection. Treat connection state as transport and identity state, and policy as the present authorisation boundary.
Before a controlled assessment, ask the workspace administrator to confirm whether Apps are enabled for the relevant account or workspace and whether Guardian history retrieval is permitted. Then confirm that the parent conversation remains the live connected conversation used by the test. Immediately before each request, recheck the applicable managed policy rather than relying on the policy recorded when the session began. The source material does not provide a stable, reader-ready end-user switch or administrative interface for these upstream flags, so use only documented controls available in the actual workspace and do not guess configuration names.
Worked fictional example: a parent session contains a synthetic restriction at 10:00. At 10:15, an administrator changes the workspace policy so history retrieval is no longer allowed. A Guardian request at 10:20 must be assessed against the current policy check described by OpenAI, not against the earlier session state. The operator should record the expected denial boundary and verify that no approval proceeds on the assumption that the old policy still applies. This example illustrates a validation method; it does not predict a particular interface message or response.
Decision rule: fail closed when either dependency is uncertain. If the parent connection, conversation identity or current policy cannot be verified, do not treat missing history as evidence that no earlier restriction exists. Route the decision to a person with access to the authoritative task record. The trade-off is additional delay, but accepting an automated conclusion based on unknown context would make the approval basis unverifiable.
Interpret bounded evidence as partial evidence
The bounded windows can show nearby instructions without proving completeness. An earlier or differently branched instruction may remain outside the selected evidence. Record this limitation separately from a model’s conclusion and keep the human approval checkpoint.
Build a context map before relying on automated review. For each proposed side effect, list the immediate request, relevant parent restriction, applicable handoff, current permission state and authoritative policy. Mark which items Guardian actually received, which were requested but unavailable, and which remain known only to the human reviewer. This turns “Guardian reviewed it” into a testable statement about specific evidence.
| Evidence item | Operator check | Consequence if absent |
|---|---|---|
| Immediate approval request | Confirm the requested action, target and side effect are explicit. | Reject or rewrite the request; context retrieval cannot repair an ambiguous action. |
| Earlier parent restriction | Verify whether history retrieval was allowed and whether the restriction was returned. | Require human comparison with the authoritative conversation record. |
| Worker-specific handoff | Confirm the recorded handoff and inspect the bounded root evidence selected around it. | Do not infer worker authority from the worker’s current message alone. |
| Current policy | Check the policy applicable at review time. | Deny or escalate; historical permission is not current authorisation. |
| Permission revocation | Locate the latest explicit revocation and verify that it governs the proposed action. | Keep the permission revoked until a responsible person grants a new, scoped permission. |
Example, not a guaranteed output: suppose a root conversation grants a worker permission to edit a fixture, then revokes that permission six messages later. A subsequent handoff asks the worker to “finish the change”. If the selected handoff window contains the grant but not the revocation, its evidence is incomplete even though every displayed message is authentic. The human reviewer must inspect the authoritative sequence and apply the latest controlling instruction. Guardian’s selection can assist discovery; it cannot establish complete provenance merely because a message appeared in context.
Decision rule: automated review may inform an approval only when the immediate request is unambiguous and all controlling evidence is either present or independently verified. If the reviewer reports uncertainty, a fallback is used, a cap may have omitted material, or the operator cannot locate the latest restriction, require human adjudication. The meaningful trade-off is between concise, relevant context and transcript completeness; the documented design favours bounded selection, so the workflow must compensate with explicit verification.
This 0.160.0 CLI release does not announce a new model or grant account entitlement. For a separate administrator-facing checklist on model access, group provisioning and Codex audit evidence, see Model Access, Group Provisioning, Codex Audits which provides prompts for human-administered reviews. It does not imply that version 0.160.0 changes those settings; verify availability in the actual workspace.
Treat revoked permission as controlling context
A revoked permission is not merely historical commentary. PR #49036 specifically explains that omission of revoked permissions can matter before side-effecting actions are approved. This concern is distinct from Codex 0.160.0 restoring saved permissions when a task or session resumes: restoration is conditional, can be explicitly overridden, and does not displace present policy or a later restriction. Guardian context should help expose conflicts, not convert an old grant into renewed authority.
For a bounded check, create a synthetic sequence with an initial narrow grant, an explicit revocation and a later request that would need the revoked capability. Keep all resources local and disposable, and do not perform the action. Ask Guardian to state the latest relevant permission and quote or identify the evidence it used. A person should then compare that answer with the known sequence and the current /permissions state. The CLI documentation identifies /permissions as an operator-facing control, but the displayed session state must still be interpreted alongside managed policy and the latest instruction.
Fictional example: “You may edit scratch/report.txt” is followed by “Permission to edit files is withdrawn; review only.” After resumption, a request proposes appending a line to that file. Even if a saved permission profile is restored, the operator should preserve the no-write instruction and refuse the edit unless a responsible person issues a new, explicit and appropriately scoped grant when policy permits. The verification target is the conflict handling, not whether the system can be induced to write.
Decision rule: where an earlier grant conflicts with a later revocation, apply the later revocation and current policy. Do not accept an automated approval based solely on restored session permissions or a selected older message. Escalate only to an authorised human who can verify the task’s purpose and issue a fresh permission; do not ask Guardian to infer consent.
Keep sensitive material out of the review path
Opt-in access to earlier instructions and handoff context increases the material that may be made available to an automated reviewer. It does not establish a change to retention, training or audit policy. OpenAI’s plan guidance, accessed on 4 October 2026, distinguishes consumer data controls from business terms and notes that consumer Plus and Pro conversations may be used for improvement unless training is disabled. Workspace and account conditions differ, so an administrator must verify the applicable controls rather than transferring assumptions between plans.
Before enabling either context source, classify the proposed fixture. Remove credentials, access tokens, personal records, unreleased commercial information and production data. Replace them with obvious placeholders such as EXAMPLE_TOKEN_NOT_REAL and synthetic identifiers. Do not paste secrets into prompts, parent conversations, approval requests or handoff messages. Keep untrusted repository content separate from instructions: quote only the minimum excerpt needed for review and label it as data that must not direct the reviewer.
Example: instead of providing a real deployment configuration, create a synthetic file containing endpoint = "https://example.invalid" and token = "EXAMPLE_TOKEN_NOT_REAL". Ask Guardian whether a proposed parser edit stays within a read-only review boundary. A human must inspect the prompt and selected context before submission, then inspect any suggested output for accidental inclusion of unrelated material. This procedure tests contextual reasoning without making a privacy assurance or exposing an operational secret.
Decision rule: if useful review requires secrets or regulated, personal or otherwise sensitive context, do not use the validation fixture until the organisation’s authorised data owner has confirmed the applicable account, workspace and data controls. Prefer a minimised synthetic reproduction. The trade-off is reduced realism, but that is preferable to broadening contextual access merely to test an optional reviewer.
Preserve accountable human approval
Guardian review and human approval are different control layers. OpenAI’s current approvals guidance, accessed on 4 October 2026, says approval requests route to the user by default. Auto-review is an optional reviewer for eligible interactive approval requests; it does not review every action and is not a replacement for sandbox controls, managed policy or human governance. A favourable automated assessment therefore remains advice within an approval process, not accountable sign-off.
Assign three roles before enabling a production-adjacent workflow: the task owner defines permitted outcomes; the workspace administrator maintains policy and reviewer configuration; and the human approver accepts or rejects consequential actions. One person may hold more than one role where organisational rules allow, but accountability must be explicit. Record the proposed action, Guardian’s evidence basis, any omitted or fallback context, the current permission state and the human decision.
Fictional approval example: Guardian reports that a proposed command appears consistent with the visible handoff. The command would nevertheless publish an artefact to an external registry. The human approver must verify the destination, package contents, credentials boundary, version and rollback plan, then approve or refuse through the organisation’s established process. Guardian’s statement cannot authorise publication because external side effects require accountable review beyond contextual consistency.
Decision rule: require direct human approval for accepting code changes, creating external side effects and deploying to production, regardless of an automated review result. Where Guardian’s evidence is complete and consistent, it may reduce the material the person must locate manually. Where evidence is partial or conflicting, it should increase scrutiny rather than lower the approval threshold.
Use a stop-or-proceed record for the upgrade decision
Conclude the validation with an evidence record, not a general judgement that Guardian is “safe”. Note the installed CLI version, workspace type, relevant managed policy, whether Apps and the live parent connection were verified, which of the two disabled-by-default context paths was intentionally enabled, and whether the synthetic restrictions were visible. Also record caps, fallbacks, missing context and the retained human approval route. Do not record secrets or copy unnecessary conversation content into the log.
- Proceed with the upgrade while leaving Guardian context disabled if the team needs the other 0.160.0 changes but has no reviewed use case for these optional context paths.
- Proceed with a narrowly enabled context path only if its dependencies and current policy are verified, the synthetic fixture exposes the controlling restriction, and human approval remains in place.
- Pause Guardian enablement if Apps or live-connection state is unclear, current policy cannot be confirmed, or bounded context omits decisive instructions.
- Stop the consequential action if a permission was revoked, evidence conflicts, sensitive material would need to enter the prompt, or no accountable human can review the result.
This gate deliberately separates upgrading the CLI from trusting optional automated review. As of 4 October 2026, the assigned OpenAI sources do not provide a release-specific entitlement table, regional guarantee, rollout schedule or stable reader-facing setup procedure for the upstream Guardian flags. Absence of the feature in a particular account or workspace should therefore be treated as an availability question, not as a reason to weaken policy or substitute an undocumented configuration.
Availability is not a single yes-or-no entitlement
Codex CLI access, Codex Cloud access and access to a particular 0.160.0 behaviour are separate questions. OpenAI’s plan documentation, accessed on 4 October 2026, says Codex is included across ChatGPT plans, with usage limits that vary. It separately describes Codex Cloud as available to eligible accounts subject to rollout and workspace settings. Managed workspaces can also control Codex Local and Codex Cloud independently. Consequently, successful use of the local CLI does not prove that Codex Cloud is enabled, and Cloud eligibility does not prove that every local feature or optional Guardian path is active.
Check availability in layers rather than inferring it from one successful sign-in. First, identify whether the intended workflow runs locally through the CLI, delegates work to Codex Cloud, or combines both. Second, have a workspace administrator verify whether the relevant surface is enabled for that account and workspace. Third, confirm the installed CLI version. Finally, inspect the effective model, permissions and approval route in the client before allowing any work. The current CLI documentation identifies model, permission and review controls, but the 0.160.0 changelog does not turn those controls into an entitlement table.
Example availability record: a fictional team might record “local CLI required; Cloud not required; managed workspace applies; the intended account has authenticated; the installed version is still to be checked; optional Guardian context remains unapproved”. This is a sample administrative record, not evidence that the feature is available to that team. A human workspace administrator must verify each entry against the actual account, policy and client.
The decision rule is to proceed only when the required product surface, account, workspace setting and client version are all confirmed independently. Stop if the test depends on Codex Cloud but only local CLI access has been established, or if a personal account appears to behave differently from the managed workspace in which deployment will occur. Testing with a personal account may be quicker, but it can hide managed-policy constraints and therefore provides weak evidence for an organisational upgrade.
Plan, rollout and usage conditions remain independent
The 1 October 2026 changelog entry records a CLI release; it is not evidence that every account received a new general capability on that date. As of 4 October 2026, no release-specific regional list, staged-rollout timetable, operating-system matrix, plan gate or pricing change was found in the assigned OpenAI sources. The plan article states that usage limits vary, while Cloud availability remains subject to eligibility, rollout and workspace settings. Those conditions can produce different observations for two otherwise similar users without establishing a defect in 0.160.0.
Before upgrading a shared environment, create a short entitlement worksheet with separate fields for plan, account type, managed-workspace status, local access, Cloud requirement, administrator setting, usage capacity and installed version. Obtain current answers from the account owner or workspace administrator rather than copying assumptions from another user. If a capability is missing, check these fields before attributing the absence to the release.
Worked fictional example: Operator A can open the local CLI, while Operator B can also submit an eligible Cloud task. The correct conclusion is only that the observed surfaces differ. It would be unsound to conclude that 0.160.0 “failed to roll out” to Operator A without checking workspace controls, account eligibility, usage conditions and whether the workflow actually invokes Cloud. Human review should classify the difference before any support escalation or policy change.
Use a conservative decision rule: absence is a release blocker only when the required capability is documented, the correct surface and version are in use, account and workspace prerequisites are verified, sufficient usage capacity exists, and the behaviour is still unavailable. The trade-off is additional administrative work before diagnosis, but it avoids weakening workspace controls merely to make a test resemble another account.
Do not infer a model change from a client update
Version 0.160.0 and the model selected for a session are different variables. The CLI documentation says users can choose a model, reasoning effort, permissions and commands, and exposes a /model control. The release entry does not announce a named model, a model-selector redesign, a changed default model or a release-specific model allowance. No release-specific model announcement was found in the assigned official sources as of 4 October 2026.
Record the effective model separately for the comparison client and the candidate client. If organisational policy allows model selection, keep that choice constant while comparing session and permission behaviour. If policy or account conditions select a different model, record the difference and do not treat output variation as proof of a CLI regression. Never paste credentials, private source code, customer records or other untrusted data into a prompt merely to reproduce a model-dependent result.
Example comparison design: “Candidate client: 0.160.0; model: operator-verified value; permissions: observed value; test prompt: synthetic repository description.” The corresponding baseline should use the same harmless prompt and, where available and authorised, the same model. This sample describes a method only; it does not predict equivalent output because model availability and selection can depend on current account, workspace and service conditions.
The decision rule is to isolate the upgrade variable. If the model cannot be held constant, restrict acceptance to deterministic control observations such as displayed version, effective permission state, changed files and approval routing. Postpone subjective output comparisons. Holding the model constant improves attribution, while accepting an administrator-selected model may better represent production; the release owner should state which objective takes priority.
Before deploying the new workspace-default sessions and saved permissions, teams may want to inspect an unfamiliar codebase without changing anything, and one option is to Build a Read-Only Codex Repository Atlas, which keeps the mapping work inside read-only boundaries while findings are verified before any human-approved change.
Build an upgrade packet that can support rollback
A safe upgrade decision needs more than confirmation that the new executable starts. It needs a retrievable baseline, a bounded candidate environment and a rollback path that does not depend on remembering previous settings. OpenAI’s changelog publishes npm install -g @openai/[email protected] for the pinned release. Use that command only through the organisation’s approved npm package-management process and when policy permits. Do not replace package-governance, integrity or change-management procedures with an article command.
Before installation, record the currently approved exact version, installation channel, relevant non-secret configuration location, workspace policy owner and the session identifiers selected for non-destructive validation. Back up only configuration material that policy permits the operator to retain. Do not copy authentication tokens, application secrets or sensitive conversation content into tickets, prompts or upgrade notes. Where configuration is centrally managed, record the policy revision or administrator confirmation instead of attempting to duplicate protected settings.
Prepare rollback before testing the candidate. Preserve access to the organisation’s known-good pinned package through its normal artefact or package process, document who may reinstall it, and define what happens to sessions created or resumed under 0.160.0. Do not assume that downgrading guarantees compatibility with every state written by a newer client; if session-state compatibility is not established, quarantine the test sessions and revert the executable and approved configuration only.
Example rollback packet: a fictional release owner could list “baseline exact version; candidate 0.160.0; approved installation source; configuration owner; test-session labels; rollback operator; stop conditions; evidence directory”. It should not contain secrets or copied production transcripts. A second human should check that the known-good package remains obtainable before the candidate replaces a shared installation.
Rollback readiness is a gate, not an aspiration. Do not upgrade a shared workstation, build agent or administrator image if the team cannot identify the baseline version, authorised rollback operator and treatment of candidate-created session state. The trade-off is that a disposable test host takes time to prepare, but it prevents package rollback from becoming an improvised recovery exercise.
Use read-only evidence before permitting edits or commands
The candidate should first be assessed with observations that cannot alter a repository or external system. Confirm the installed version; start in a disposable location outside a project; inspect the effective permissions; open only synthetic or non-sensitive content; and compare the resulting state with the approved expectation. On Windows, account for the possibility of a sandbox-setup prompt described by the projectless-session pull request. Treat that prompt as a separate environment prerequisite, not as evidence that workspace-write has already been granted.
For a resumed-session check, use a purpose-built session containing no secrets and no production instructions. Record its approved permission profile before closing it, then resume it under 0.160.0 and inspect the effective state before issuing any command. Saved permissions are documented as restorable unless explicitly overridden, but current configuration and managed policy still constrain operation. An unexpected value should trigger investigation, not an attempt to broaden permissions until the expected label appears.
Do not include the optional Guardian paths in the basic CLI-availability gate. If an owner separately authorises their tests, apply the opt-in prerequisites, bounded-evidence limits and human review steps documented above; an unavailable optional path does not prove the core CLI upgrade failed.
Example no-write sequence: a fictional operator records version and workspace, opens an empty disposable directory, observes the initial permission state, exits without asking for a file change, resumes a synthetic saved session, observes its permission state, and closes it. The expected output is a completed observation table, not a claim that commands, edits or Guardian review are safe. A human reviewer compares that table with managed policy.
The stop rule is immediate if the candidate requests unexplained broader access, the observed approval route conflicts with policy, the workspace cannot be identified, a resumed state cannot be reconciled, or testing would require sensitive data. Read-only validation offers less functional coverage than an end-to-end run, but it establishes whether the control boundary is intelligible before any side effect is allowed.
Promote only after reviewing the actual diff and tests
Once the control observations pass, the next stage should use a disposable, version-controlled fixture with no production credentials or network side effects. The purpose is to inspect what the candidate proposes and changes, not to demonstrate that an artificial task predicts every production workload. Keep existing sandbox, trust, approval and managed-policy controls intact. Default approval requests route to the user according to OpenAI’s current approvals documentation; optional auto-review for eligible requests is not a substitute for that accountable human.
Give the candidate a small, reversible task with an explicit file boundary and a locally available test command. Before allowing execution, state that it must not access external services, install dependencies, alter policy files or touch paths outside the fixture. After the task, inspect the complete version-control diff, list every changed file and run only the pre-approved tests. Review generated files and configuration changes as carefully as source changes. Reject unrelated formatting churn if it obscures the substantive patch.
Worked fictional example: the fixture contains a small function and an existing unit test. The sample request is to change one synthetic return value and update only the corresponding test, without network access. The operator then reviews the proposed diff and runs the repository’s already approved local test command. This is an example validation pattern, not a product guarantee or a reported hands-on result. If the candidate modifies an unrelated configuration file, the correct response is to reject or revert the patch and investigate scope control.
Use three acceptance questions. Did the effective permissions and approval route remain within policy? Does every changed line correspond to the authorised task? Do the approved tests and human inspection support the change without unexplained side effects? A “no” or “unknown” to any question blocks promotion. Passing tests alone is insufficient because they may not cover unrelated edits; a clean-looking diff alone is insufficient because behaviour may still be wrong.
The trade-off is between breadth and attribution. A large integration task may exercise more code, but it makes it harder to distinguish a session-state change from model variation, repository complexity or external-service behaviour. Start with the smallest useful fixture, then widen coverage only after the release owner accepts the preceding evidence.
Make rollback criteria explicit before shared deployment
Define rollback triggers in advance: unexplained permission broadening; approval routing inconsistent with managed policy; a projectless session receiving defaults that the administrator did not authorise; resumed state that cannot be explained by saved permissions, explicit override or current policy; unexpected files in the diff; failure of approved tests; or reliance on optional Guardian context that cannot be verified. Also roll back or pause if support for a required account, workspace, operating environment or product surface remains uncertain.
If a trigger occurs, stop the candidate session, preserve non-sensitive diagnostic evidence, revert the fixture, reinstall the approved known-good version through the authorised package process, and verify its version before further work. Keep candidate-created test sessions isolated until compatibility is understood. Escalate the discrepancy with version, platform, workspace class, effective policy and minimal reproduction details, but exclude secrets and sensitive transcripts.
Example incident note: “Candidate 0.160.0; synthetic fixture; no production data; expected approval route: user; observed state: pending human classification; writes stopped; fixture reverted; baseline restoration assigned.” The word “pending” matters: the operator should not label the event a security bypass or software defect until a qualified human has compared it with current configuration and managed policy.
The decision rule is to favour rollback when the control state is ambiguous and further diagnosis would require broader access. Continuing may collect more evidence, but it can also turn a bounded validation discrepancy into a consequential change. The release owner may approve a temporary pause instead of rollback when the candidate is isolated and cannot affect shared work.
Require a named human release owner to sign off
The final decision belongs to an accountable human, not to Guardian, auto-review, a passing test suite or the operator who performed the upgrade. Guardian may receive selected earlier instructions or handoff evidence under optional, conditional paths, but that evidence can be incomplete and is not human judgement. The release owner must decide whether the collected evidence is sufficient for the organisation’s intended use.
Provide the owner with a compact sign-off pack: source release date of 1 October 2026; candidate and baseline versions; required local or Cloud surface; plan and workspace verification; model observation; effective projectless and resumed permission states; approval route; Guardian feature status, if relevant; complete fixture diff; test output; unresolved deviations; rollback readiness; and the proposed deployment scope. Mark inferred or unavailable information explicitly rather than filling gaps with assumptions.
Example decision: “Approve 0.160.0 for two named development workstations only; retain existing managed policy; Guardian context remains disabled; production deployment prohibited; review after bounded use.” An alternative sample is “Defer because Cloud eligibility and resumed-session behaviour have not been verified in the managed workspace.” These are examples of decision wording, not recommendations for a particular organisation.
Approval should be withheld if release-specific regional, model, rollout or entitlement assumptions are necessary to justify deployment, because the cited 0.160.0 material does not supply those guarantees. It should also be withheld if the team cannot review diffs and tests, protect prompts from secrets and untrusted data, or restore the known-good client. The cost of a staged release is slower adoption; its benefit is that projectless defaults, resumed permissions and optional reviewer context are evaluated under the same controls that will govern real work.
The final rule is straightforward: promote only the tested client version, to the verified account and workspace scope, under unchanged approvals and managed policy, with rollback available and a named human release owner signing the record. Any later expansion to Cloud use, additional workspaces, broader permissions, optional Guardian context or production deployment requires a fresh human decision.
Access 40,000+ AI Prompts for ChatGPT, Claude & Codex — Free!
Subscribe to get instant access to our complete Notion Prompt Library — the largest curated collection of prompts for ChatGPT, Claude, OpenAI Codex, and other leading AI models. Optimized for real-world workflows across coding, research, content creation, and business.
Useful Links
- ChatGPT & Codex changelog — Codex CLI 0.160.0
- Support projectless TUI sessions with workspace defaults
- Add handoff-aware root context for Guardian reviews
- Add opt-in conversation history retrieval to Guardian reviews
- Codex CLI — Inspect, edit, and run code from your terminal
- Agent approvals & security
- Using Codex with your ChatGPT plan
- Node.js: An introduction to the npm package manager
