Codex Refactor Delegation Playbook: Plan Locally, Implement in Cloud, Review Before PR

Original conceptual editorial illustration showing two isolated block workspaces linked through approval for a human-reviewed workflow

A refactor is ready for Codex Cloud only when local investigation has reduced it to one behaviour-preserving milestone with explicit boundaries, evidence and a recovery route. Your job at this stage is not to ask artificial intelligence (AI)Computer systems designed to perform tasks that normally require human intelligence, such as understanding language, recognising patterns, or making predictions. Open glossary entry to “clean up” a subsystem. It is to prepare a delegation packet that another engineer could execute without guessing: intended structure, unchanged behaviour, target files, forbidden changes, validation commands, review evidence and rollback criteria. This is the foundation for the later choice between requesting revisions, pulling the work locally or creating a pull request (PR)A proposed set of repository changes submitted for review before integration. Open glossary entry. It does not transfer technical accountability to Codex: consequential decisions require human review.

Original conceptual editorial illustration showing two isolated block workspaces linked through approval for a human-reviewed workflow
Plan a bounded Codex refactor locally before delegating one milestone to Cloud.

Evidence checkpoints

Documented point: OpenAI’s Codex prompting guidance documents the “Delegate refactor to the cloud” workflow. [OpenAI Codex prompting guidance]

Documented point: OpenAI requires a published environment before a new Codex Cloud task begins. [OpenAI documentation: Cloud]

Documented point: OpenAI says cloud tasks run on managed computers after the project environment is prepared. [OpenAI documentation: Using codex cloud]

Documented point: Codex Cloud is available to eligible Plus, all Pro tiers, Business, Enterprise, Healthcare and Education accounts, subject to rollout and workspace settings; Free and Go are excluded. Check plan, workspace and rollout conditions before relying on this. [Using Codex with your ChatGPT plan]

Documented point: New tasks start from the published environment’s prepared filesystem; existing tasks retain their own state, including uncommitted changes and installed tools. [OpenAI documentation: Cloud environments]

Documented point: The legacy documentation explains the container execution lifecycle: repository checkout, setup, internet settings, a command-and-test loop, diff presentation and an optional pull request. [OpenAI documentation: Cloud environment]

Use a readiness gate, not a refactoring wish list

OpenAI’s Codex Prompting guide, accessed on 4 October 2026, documents a deliberate sequence: commit or stash current work, investigate and revise a plan locally, choose one milestone for a cloud task, review the resulting diff and then either create a PR or pull the changes locally to test and finish. The important distinction is between local discovery and cloud implementation. Discovery resolves ambiguity about behaviour and boundaries; implementation carries out an already bounded change.

Apply this procedure before dispatch:

  1. State the observable behaviour that must remain unchanged.
  2. Name one structural change that can be reviewed independently.
  3. List the files expected to change and the files or interfaces that must not change.
  4. Record repository-specific validation commands and any manual checks.
  5. Define what evidence would make you accept, revise or abandon the result.
  6. Create a clean, recoverable local comparison baseline.
  7. Confirm that Cloud access, a published environment and suitable repository permissions actually exist.

Dispatch only if every item can be answered concretely. If the plan cannot name a bounded milestone or verification method, keep investigating locally rather than dispatching Cloud. The trade-off is intentional: more local investigation delays delegation, but dispatching an ambiguous architectural goal makes it harder to distinguish an implementation defect from a misunderstood requirement.

Fictional example: “Improve authentication architecture” is not ready. “Move token parsing from src/auth/handler.ts into src/auth/token-parser.ts, retain the exported authenticate(request) signature and existing error mapping, update only directly affected imports and tests, and run the project’s documented authentication test and type-check commands” is potentially ready. This example describes a proposed scope, not a tested result or product guarantee.

Write a single behaviour-preserving milestone

A milestone is smaller than the overall refactor objective. “Separate the authentication subsystem” may involve token parsing, session lookup, policy evaluation, logging and public interfaces; putting all of that into one task creates several independent reasons for a diff to be rejected. A bounded milestone changes one structural seam while preserving externally observable behaviour.

Draft the milestone with five clauses:

  • Action: the one structural operation, such as extracting a parser or moving a dependency behind an existing interface.
  • Behaviour invariant: what callers must continue to observe.
  • Scope: files or modules that may change.
  • Exclusions: related work that is deliberately postponed.
  • Evidence: commands and inspection points used to evaluate the change.

A usable fictional statement is:

Example milestone: Extract bearer-token parsing from src/auth/handler.ts into an internal src/auth/token-parser.ts module. Preserve the existing authenticate(request) export, accepted header formats, missing-token handling and error translation. Modify the handler, the new internal module and directly associated unit tests only. Do not change session lookup, authorisation policy, logging, dependency versions or public documentation. Validate with the repository’s recorded unit-test, type-check and lint commands, then inspect the diff for public-interface changes.

This formulation distinguishes a structural outcome from an aspirational one. It does not say “make authentication cleaner”, because cleanliness has no objective completion point. It also avoids prescribing every line of implementation, leaving Codex room to perform the mechanical extraction.

Choose a milestone that can be accepted or rejected without waiting for later stages. If extracting the parser necessarily requires changing the public authentication contract at the same time, either negotiate that contract locally first or split out preparatory work. The trade-off is that a narrower milestone may require more tasks, but it produces a smaller diff and a clearer causal link between the requested change and any behavioural regression.

Turn “no behaviour change” into observable invariants

“Behaviour-preserving” is a constraint, not a validation method. It becomes useful only when translated into observable invariants. Public output, error type, ordering, side effects and compatibility may all matter even when function names stay the same.

Build an invariant table from existing code, tests and documented contracts. For each relevant path, record the input condition, current observable result and evidence source. Do not paste secrets, production payloads or untrusted external data into a prompt. Use sanitised structures or repository-local fixtures that your organisation permits.

Fictional path Invariant to preserve Proposed verification
Valid bearer header The existing handler passes the parsed token into the unchanged session lookup path. Run the existing focused unit test and inspect the call boundary in the diff.
Missing header The same existing application error is returned through the same public handler. Run the current missing-header test and inspect error mapping.
Malformed scheme The current rejection path remains in effect; no new accepted format is introduced. Run the malformed-header cases and compare affected test assertions.
Public import Callers continue importing authenticate from the established module. Search tracked source for imports and inspect exported symbols.

The table is an example preparation method, not evidence that these tests exist in any particular repository. Insert the project’s real behaviours and checks. Where current behaviour is unclear, reproduce it locally or consult the responsible owner before delegating.

An invariant is adequate only if a reviewer can point to a check, inspection or authoritative contract that can reveal a violation. If “works as before” depends solely on intuition, the task remains in discovery. The meaningful trade-off is coverage versus task size: record invariants relevant to the selected seam, rather than attempting to specify the entire subsystem, but do not omit a known externally observable effect merely to keep the packet short.

Draw a public-interface boundary and a non-goal boundary

Public-interface constraints and non-goals serve different purposes. A public-interface constraint protects something callers rely upon, such as an exported function, command-line interface (CLI)A text-based interface for running commands and tools. Open glossary entry option, application programming interface (API)A documented way for software systems to exchange requests and results. Open glossary entry response or persisted data shape. A non-goal prevents adjacent but unnecessary work, such as renaming modules, upgrading dependencies or redesigning error messages. Treating them as one vague “out of scope” list can hide the difference between compatibility and convenience.

Inspect the selected seam locally:

  1. Find exports, external entry points, routes, configuration names, serialised formats and imports outside the target directory.
  2. Mark which of these must remain byte-for-byte or semantically compatible.
  3. List tempting adjacent changes revealed during investigation.
  4. Classify each adjacent change as a future milestone, an explicit non-goal or a prerequisite that blocks delegation.
  5. Ask the API or component owner to confirm uncertain compatibility assumptions.

Fictional example: the authentication handler’s exported name and returned error categories are public boundaries. Moving a private parsing helper is the milestone. Rewording errors, adding token formats, changing logs and converting the whole directory to a different module convention are non-goals. If external consumers import the supposedly private helper, that discovery is not a minor implementation detail; it blocks the extraction until ownership and compatibility are resolved.

If the change affects a public boundary, do not label it behaviour-preserving merely because the intended business outcome is similar. Reclassify it as a contract change, obtain the appropriate approval and design migration evidence. The trade-off is between architectural neatness and compatibility: retaining a forwarding export may look less tidy, but it can keep a structural milestone independent from a later, explicitly reviewed API migration.

Name target files without pretending the list is infallible

A target-file list gives the cloud task a review perimeter. It is not a guarantee that no other file can legitimately change. Generated files, snapshots, import maps or nearby tests may need updates, while an unexpected dependency manifest or deployment file may indicate scope drift.

Prepare three file groups:

  • Expected: files that should change, including the proposed new module.
  • Conditionally allowed: files that may change only for a stated reason, such as a test index if the project requires explicit registration.
  • Protected: public contracts, dependency manifests, database migrations, deployment configuration or unrelated modules that must remain untouched for this milestone.

Fictional example:

Expected:
- src/auth/handler.ts
- src/auth/token-parser.ts
- tests/auth/handler.test.ts

Conditionally allowed:
- tests/auth/index.ts — only if required to register the new test file

Protected:
- package.json and lockfile
- src/auth/session-store.ts
- src/policy/**
- database migrations
- public API documentation

This is a sample packet format, not a prediction of Codex’s output. After implementation, every changed file still needs inspection. An isolated cloud workspace separates one task’s working files from another task’s files, but isolation does not prove that the task stayed within scope or prevent effects on repositories or external services when permissions, credentials, network access and write actions are available.

An unexpected file is a review event. Accept it only when its necessity can be tied directly to the milestone and independently validated; otherwise request a revision or pull the work locally for controlled adjustment. Overly rigid file prohibition can obstruct legitimate repository conventions, while an unrestricted list makes scope creep difficult to detect.

Record validation commands and what each one can prove

Do not tell Codex merely to “run the tests”. Record exact repository commands from project documentation or established scripts, the intended working directory, required setup and the expected category of evidence. OpenAI says Codex attempts validation and presents results, but no assigned source establishes a universal command set or guarantees that agent-reported checks prove correctness.

Use a validation ledger. The commands below are illustrative only for the fictional TypeScript authentication example; they were not executed and must be replaced with commands verified in the reader’s actual repository:

Fictional example command Purpose Limitation requiring review
npm test -- tests/auth/handler.test.ts Exercise behaviour nearest the extracted seam. May not cover unknown callers or integration behaviour.
npx tsc --noEmit Detect incompatible types and unresolved imports. Does not establish runtime equivalence.
npx eslint src/auth/handler.ts src/auth/token-parser.ts Check configured static conventions. Does not establish architectural correctness.
npm test Look for regressions outside the focused module. May be unavailable or incomplete in the prepared environment.

Use the local repository’s verified scripts rather than copying the fictional commands above. Run the most relevant checks locally before delegation where practical, and record pre-existing failures separately. Otherwise, a cloud failure cannot reliably be attributed to the refactor.

Also specify manual inspection. For the fictional extraction, a reviewer should compare exported symbols, trace malformed input to the same error mapping, check that session lookup remains unchanged and confirm no new package was added. These checks complement automated commands rather than duplicating them.

Do not dispatch when the only proposed validation is that the project builds, unless compilation is genuinely the complete accepted contract for that component and an accountable owner confirms it. If an important invariant cannot be automated, name a reproducible manual check and its reviewer. The trade-off is speed versus evidence depth: a focused command gives fast feedback, whereas broader checks detect wider coupling; a consequential refactor normally needs both proportionate automated evidence and human inspection.

Define success, revision and rollback before implementation

Success criteria describe an acceptable result. Rollback criteria describe when not to keep it. They should be written before the cloud task so that an attractive implementation does not silently lower the standard.

For the fictional token-parser milestone, sample success criteria could be:

  • The parser exists behind an internal module boundary.
  • The existing public handler export and its relevant behaviour remain unchanged.
  • Only expected or justified conditional files change.
  • The recorded validation commands complete with reviewable output, or any environment limitation is clearly reported rather than concealed.
  • The diff introduces no dependency, configuration, network or data-format change.

Sample revision triggers could include an unnecessarily broad rename, duplicated parsing logic or a changed error path. Sample rollback triggers could include an unresolved public-interface change, inability to execute or reproduce the minimum verification, unexplained protected-file changes, or evidence that the extraction cannot be separated from a larger redesign.

Rollback at this stage means discarding or reverting the candidate change and returning to the clean baseline; it does not mean claiming that every external effect can automatically be undone. Repository writes, PR actions or external-service calls require their own controls.

Request a revision when the milestone remains valid and the defect is bounded; abandon or re-plan when the evidence invalidates the milestone’s assumptions. Pull locally when the change is useful but needs investigation that depends on local tools or context. Create no PR until the later diff-and-evidence gate is satisfied by a responsible person.

Create a clean local baseline

Before local planning and handoff, commit the intended baseline or stash unrelated work, as directed by OpenAI’s Prompting guide. This gives you a known comparison point and reduces accidental inclusion of unrelated edits. It does not guarantee freedom from merge conflicts, omitted files or mistakes.

Use this repository-neutral procedure:

  1. Inspect repository status and identify tracked, untracked and ignored material.
  2. Remove secrets and prohibited data from the prospective task context; do not use a stash as a substitute for proper secret handling.
  3. Commit reviewed work that belongs in the baseline, or stash unrelated local edits according to team practice.
  4. Confirm that the remaining diff contains only deliberate planning artefacts or no changes at all.
  5. Record the baseline branch and commit identifier in your private work record where team policy permits.
  6. Run the selected pre-refactor validation commands and note any existing failures without inventing or suppressing results.

Example baseline note: “Start from the reviewed commit recorded in the task tracker; local unrelated work is stashed. The focused authentication checks are the minimum comparison set. One known failure, if genuinely observed, must be documented with its actual command and output rather than paraphrased as passing.” This wording is illustrative and supplies no claimed test result.

The current Cloud Environments guide, checked on 4 October 2026, says new tasks begin from a published environment’s prepared filesystem, while existing tasks retain their own state, including uncommitted changes and installed tools. It also warns that saved state does not replace source control. OpenAI’s current Cloud Help article likewise says a new task does not recover another task’s uncommitted changes. Therefore, do not rely on an earlier cloud task or local working tree as the only copy of important work.

If you cannot identify the intended baseline or separate unrelated edits, stop. Cleaning the baseline costs time, but it gives reviewers a meaningful diff; preserving a mixed working tree may feel faster while making attribution and rollback substantially harder.

Run an access, environment and data preflight

A delegation-ready code plan is not the same as an executable cloud task. As of 4 October 2026, OpenAI’s plan Help article said Codex Cloud was available to eligible Plus, all Pro tiers, Business, Enterprise, Healthcare and Education accounts, subject to rollout and workspace settings; Free and Go did not include Cloud. Guest, K–12 and Enterprise view-only seats could not create cloud environments, and Enterprise cloud access was separately controlled and off by default where it had not been enabled. No assigned source provides a fixed country or region list, so verify the options visible in the actual account and workspace rather than inferring access.

Before handoff, confirm:

  • the user and workspace currently expose Codex Cloud;
  • an administrator has enabled any required workspace access;
  • a suitable environment has been published, because OpenAI’s Cloud documentation requires one for a new task;
  • the environment points to the intended repository and prepared filesystem;
  • repository permissions are no broader than the intended workflow requires;
  • prepared files, tools, credentials, outbound domains and sharing controls have been reviewed by their owners;
  • the task prompt contains no secrets, production credentials or untrusted data.

OpenAI’s current Cloud Environments guide distinguishes environment variables from network secrets and makes network policy review relevant. Allowing a domain does not itself supply credentials or service permission. Conversely, a separate task workspace must not be treated as a complete safety boundary once access and write capabilities are present.

The Cloud Help article checked on 4 October 2026 states that Codex Cloud is not covered by OpenAI’s business associate agreement (BAA)A contract that establishes required safeguards and responsibilities when a covered entity or business associate handles protected health information. Open glossary entry and must not process protected health information (PHI)Individually identifiable health information protected under applicable United States health-information rules. Open glossary entry. Exclude PHI. Route all regulated or sensitive-data decisions through the organisation’s approved policy and controls rather than improvising a prompt-level workaround.

If a model choice is offered, use only options visible to that account and workspace. OpenAI’s model-selection documentation says availability, tools, reasoning settings and usage limits differ by product and model version. A model named in documentation is therefore an example, not evidence that it is selectable everywhere.

No published environment, no approved access or no safe data path means no cloud dispatch. Continue locally or ask the responsible administrator or repository owner to resolve the prerequisite. The trade-off is between convenience and controlled capability: broader credentials or outbound access might make setup easier, but they also enlarge what a mistaken task could affect.

Assemble the delegation packet and apply the final stop check

Combine the approved material into a short packet, keeping evidence references close to the instruction they support. The filled example below belongs to the fictional authentication repository: adjust it only after checking the real repository and obtaining the responsible owner’s approval.

Milestone:
Extract bearer-token parsing from src/auth/handler.ts into src/auth/token-parser.ts; leave session lookup unchanged.

Behaviour to preserve:
Preserve authenticate(request), accepted header formats, missing-token handling and existing error mapping.

Expected files:
src/auth/handler.ts; src/auth/token-parser.ts; tests/auth/handler.test.ts.

Conditionally allowed files:
tests/auth/index.ts only if the fictional test registration requires it.

Protected interfaces and files:
Public handler export; src/auth/session-store.ts; package.json; lockfile; deployment settings.

Non-goals:
No new token formats, dependency upgrades, logging changes or policy redesign.

Validation:
Replace the illustrative test/type/lint commands above with this repository’s verified scripts; inspect the public export and error mapping manually.

Success:
Focused tests and broader checks are reported with actual results; the diff is scoped; a human reviews the invariants.

Request revision if:
The changed-file list includes unexplained adjacent cleanup, or a focused test fails for a bounded reason.

Abandon or re-plan if:
The public handler contract must change, the baseline cannot be identified, or minimum verification is impossible.

Baseline:
The source-control reference recorded and verified by the responsible project owner before delegation.

Keep sensitive records outside the prompt. Do not paste API keys, credentials, PHI, confidential production samples or untrusted issue content into the packet. If external material is necessary, sanitise it and have an accountable person confirm that its use complies with organisational controls.

Final decision rule: dispatch one cloud task only when the packet names one independently reviewable milestone, a clean source-control baseline, behaviour and API boundaries, expected files, exact verification methods and an authorised environment. Otherwise, continue local investigation. Passing this gate means only that the work is sufficiently bounded to attempt; it does not establish that the eventual implementation will be correct, secure, compliant or ready to merge.

Negotiate the refactor locally before delegating any implementation

Original conceptual editorial illustration showing blank modular blocks inside a scope boundary for a human-reviewed workflow
Name the limited refactor milestone and its verification path before delegation.

Use Codex locally—in an integrated development environment (IDE)A software application combining tools for writing, building, testing and debugging code. Open glossary entry or command-line interface (CLI)—to investigate the repository and produce a reviewable plan, not to begin the refactor. This distinction matters: a planning request asks Codex to inspect, trace and propose; an implementation request authorises file changes. OpenAI’s Prompting guide, accessed on 4 October 2026, documents this local-plan-to-cloud sequence and calls for reviewing and negotiating the plan before sending one milestone to a cloud task.

Your job at this stage is to decide whether the proposed module boundary is technically coherent, whether the first migration step is independently verifiable, and whether its rollback is credible. Codex can help map evidence, but it does not own architectural intent or approval. Keep untrusted issue text, logs, copied web content and secrets out of the prompt. Do not include application programming interface (API) keys, production credentials, protected health information or other regulated data. The earlier access-and-data preflight gives the source-backed prohibition on protected health information in Codex Cloud; other sensitive-data decisions must follow the organisation’s approved policy.

Separate repository discovery from permission to modify

A useful first prompt explicitly limits Codex to read-only investigation. This is different from saying “refactor the authentication package”, which combines diagnosis, design and implementation in one ambiguous instruction. Ask for repository evidence and a proposed sequence while prohibiting edits, generated files, formatting changes and dependency installation.

Example planning request, not a product guarantee:

Planning only. Do not edit files, install dependencies, create commits or run
commands that write to the repository.

Investigate how authentication responsibilities are divided across:
- src/auth/session.py
- src/auth/tokens.py
- src/http/middleware.py
- tests/auth/
- tests/http/

Trace imports, public call sites, test seams and configuration dependencies.
Identify any circular-import risk if token parsing moves into
src/auth/token_service.py.

Return:
1. current responsibilities, with file and symbol references;
2. inbound and outbound dependencies;
3. direct and indirect call sites;
4. a migration order split into independently reviewable milestones;
5. tests or checks relevant to each milestone;
6. risks, assumptions and rollback points;
7. unresolved questions requiring a maintainer’s decision.

Before accepting the answer, inspect the working tree and compare it with the clean baseline prepared earlier. The official prompting workflow advises committing or at least stashing current work before local planning. That makes unexpected edits easier to identify and preserves a comparison point; it does not prevent every conflict, accidental write or incorrect conclusion.

Continue only if the output is a proposal grounded in named repository locations and the working tree contains no unexplained changes. If Codex edited code, discard or isolate those changes, restate the read-only boundary and repeat the investigation. The trade-off is an extra planning pass, but it prevents accidental implementation from becoming unreviewed evidence for its own design.

Build a responsibility map from symbols, not filenames alone

A file list shows location; a responsibility map shows why moving code may alter behaviour. For each candidate module, record the symbols it owns, the state it reads or writes, the errors it emits, and the callers that depend on those details. Ask Codex to distinguish declared responsibilities from incidental co-location. Two functions living in the same file do not necessarily belong to the same abstraction.

Use a compact table and verify every important row by opening the referenced definitions and representative callers. For the fictional authentication split, a sample planning artefact could look like this:

Current symbol Observed responsibility Dependencies to verify Proposed owner
tokens.parse_bearer() Parses an Hypertext Transfer Protocol (HTTP)A standard request-and-response protocol for exchanging representations and related metadata between web clients and servers. Open glossary entry header into token text Header format, error type, middleware callers auth/token_service.py, subject to interface review
tokens.verify_claims() Checks token claims against configuration Clock source, issuer settings, exception mapping auth/token_service.py
middleware.authenticate() Adapts an HTTP request to authentication inputs Request object, response handling, logging Remain in http/middleware.py

This is an example of a proposed map, not a claim about any real repository. A maintainer must confirm whether parsing an HTTP header belongs in the authentication service or should remain in the HTTP adapter. That decision affects dependency direction: moving transport-specific parsing into a core module may simplify one caller while coupling the core to HTTP conventions.

Assign a symbol to a target module only when its responsibility and dependency direction agree. If the proposed owner would need to import a higher-level transport or application module, either retain the symbol, introduce a narrow input type, or defer it to another milestone. Prefer a slightly smaller first move over a “cleaner” diagram that conceals a new layer violation.

Trace call sites beyond direct imports

Direct import searches are necessary but incomplete. A function may be re-exported, registered in a framework, injected through a constructor, referenced by a string, wrapped by a decorator or replaced in tests. Ask Codex to separate confirmed call sites from possible dynamic references, then inspect the uncertain cases manually.

A practical procedure is to request four inventories: definitions and re-exports; static callers; registrations or factories; and tests, fixtures or mocks that patch the old path. In the fictional example, moving verify_claims could require updating a test that patches src.auth.tokens.verify_claims, even if production callers import it through src.auth. The plan should name both the public import path and patch points rather than merely state “update imports”.

Ask for evidence in a form such as:

For each call site, report:
- file and symbol;
- import path actually used;
- whether the reference is direct, re-exported, injected or dynamic;
- behaviour relied upon, including exceptions and return shape;
- confidence and the repository evidence supporting it.

Do not infer dynamic callers as facts. Put uncertain references in a separate list.

Manually check entry points, dependency-injection wiring, test configuration and package export files. Search results can establish that text exists; they do not establish that all runtime paths have been found. Conversely, an apparently unused export may be consumed outside the repository.

Approve a move only when known in-repository callers and compatibility obligations are named. If external consumption is plausible but cannot be established locally, preserve the old import path with a reviewed compatibility shim or stop for an owner decision. A shim reduces immediate migration risk but prolongs dual interfaces and needs an explicit removal milestone.

Expose circular-import risk before choosing migration order

A circular import is not merely a graph blemish. Moving a symbol can change module initialisation order, expose partially initialised modules or force a lower layer to depend on a higher one. Ask Codex to draw the relevant import edges before and after the proposed move and to label edges introduced by type annotations, package initialisers, test helpers and configuration modules.

For the fictional split, suppose http/middleware.py imports auth/tokens.py, while auth/tokens.py imports an error adapter from http/errors.py. Moving more token logic into authentication without removing the error-adapter dependency would retain an inward and outward edge between the packages. The plan should not hide this by suggesting local imports unless a maintainer has deliberately accepted that compromise.

A better planning question is: “Which dependency must be inverted before any symbol move?” One possible example is to define an authentication-domain exception in auth/errors.py, leave HTTP response conversion in http/errors.py, and make middleware translate between them. That sequence changes ownership before moving implementation.

If Milestone 1 introduces or preserves a cycle across the intended boundary, revise the order. First extract the lowest-level contract or error type that breaks the edge. Deferred imports may avoid an immediate import-time failure, but they obscure the architecture and shift failure to runtime; use them only when the repository owner records that trade-off explicitly.

Turn the map into dependency-ordered milestones

A refactor plan is not a chronological list of all desired edits. It is a dependency-ordered set of checkpoints in which each milestone has one architectural purpose, a bounded file set, observable invariants and a rollback point. Codex Cloud should receive only the first approved milestone, not the entire roadmap.

For the fictional authentication work, a proposed sequence might be:

  1. Milestone 1: introduce src/auth/errors.py; change token verification to raise the domain exception; adapt existing middleware to preserve its current HTTP response behaviour; update only directly affected tests.
  2. Milestone 2: create src/auth/token_service.py and move claim-verification logic while preserving the existing package export.
  3. Milestone 3: migrate internal callers to the new import path and retain a compatibility export for any unresolved consumers.
  4. Milestone 4: remove the compatibility export only after repository owners confirm the supported interface and downstream migration.

This example deliberately places dependency inversion before the larger file move. It also separates compatibility removal from implementation movement, because those actions have different failure modes and review owners.

For every milestone, require the plan to state its prerequisites, exact intended files, prohibited files, expected import changes, relevant validation commands, manual checks and reversal procedure. Treat commands as project-specific inputs: OpenAI documents that Codex can work through commands and tests, but the assigned sources provide no universal lint, type-check, build or test command and no guarantee that a reported pass establishes correctness.

Decision rule: Milestone 1 is delegable only if it can be reviewed and reverted without completing Milestones 2–4. If an intermediate state cannot run, cannot preserve the agreed interface or requires simultaneous edits across an unbounded area, split it differently or keep the work local. Smaller milestones increase coordination overhead; larger ones weaken attribution when behaviour changes.

Identify test seams without turning tests into proof

A test seam is a controllable boundary at which the refactor’s behaviour can be observed: a public function, adapter, injected clock, fixture, command entry point or error conversion layer. It differs from test coverage. Coverage may reveal exercised lines, while a seam explains where old and new structures can be compared meaningfully.

Ask Codex to map each invariant to existing tests and to flag gaps rather than inventing confidence. In the example, domain-exception tests can check claim validation independently, while middleware tests can check that the same authentication failure is converted into the same HTTP response. A package-import test might detect a broken re-export, but it would not demonstrate equivalent claim validation.

A sample planning output could state:

  • Domain seam: invoke claim verification with the project’s existing fixtures and controlled clock; compare return shape and documented exception categories.
  • Adapter seam: invoke middleware with existing request fixtures; check the project-defined response behaviour for missing, malformed and rejected credentials.
  • Import seam: import the supported package path used by current callers; check that compatibility remains intentionally available.
  • Manual review seam: inspect logging and error translation to ensure sensitive token material is not newly exposed.

These are example methods, not guaranteed commands or sufficient evidence. The maintainer must replace them with the repository’s actual test names, fixtures and approved manual checks. Passing tests can fail to cover dynamic registrations, unsupported consumers or changed operational behaviour.

Do not approve a milestone whose central invariant has no identifiable observation point. Either add a narrowly scoped characterisation test before delegation, define a safe manual check, or reduce the move. Adding a test costs time and may encode accidental behaviour; omitting it leaves the reviewer unable to distinguish preservation from coincidence.

Make risks and assumptions falsifiable

Generic risks such as “imports may break” are too vague to guide review. Rewrite each risk as a condition that a developer can inspect, a consequence, and a safeguard. Likewise, separate verified repository facts from assumptions requiring an owner’s answer.

Weak statement Reviewable replacement Required action
Imports could break src/auth/__init__.py re-exports verify_claims; callers may rely on that path Preserve the export in Milestone 1 and search named callers
Tests might need updates Two fixtures patch the implementation module rather than the public package path Decide whether patch targets are contractual or test coupling
Rollback is easy The milestone changes exception ownership and middleware translation together Revert the milestone commit and rerun the named domain and adapter checks

The entries above are illustrative. In a real plan, Codex should cite actual files and symbols, and a human should open them. Ask it to place uncertain claims under “Questions” rather than filling gaps with plausible architecture.

Block delegation when an unanswered assumption could change the milestone’s public interface, data handling, authorisation behaviour or rollback path. Minor naming questions can remain bounded instructions for the implementer. Consequential decisions require human review by the relevant maintainer, code owner or policy owner.

Rewrite the proposal as an exact Milestone 1 contract

After investigation, negotiate the plan by challenging it. Ask why each file must change, what remains untouched, which dependency edge improves, and how the old behaviour remains observable. Then edit the proposal yourself. The final contract should name exact moves and safeguards without pretending that the file list is infallible.

Example approved-plan wording:

Milestone 1 only — proposed implementation contract

Purpose:
Break the auth-to-HTTP exception dependency before moving token verification.

Expected files:
- add src/auth/errors.py
- modify src/auth/tokens.py
- modify src/http/middleware.py
- modify the specifically identified auth and middleware tests

Required changes:
- define the domain authentication exception in src/auth/errors.py;
- make the existing verification path raise that exception;
- translate it at the middleware boundary into existing project behaviour;
- preserve supported package exports and function signatures.

Safeguards:
- do not move token parsing or claim-verification functions yet;
- do not change dependency versions, unrelated formatting, logging policy,
  persistence, deployment files or public API names;
- do not add network access or credentials;
- stop and report if another production file is required.

Evidence to return:
- changed-file list and rationale;
- exact commands attempted and their unedited outcomes;
- tests not run and why;
- remaining assumptions and follow-up suggestions.

Rollback:
Revert this milestone as one isolated change and run the named pre-existing checks.

This remains a plan until it is explicitly sent as an implementation task. Labels such as “required changes” inside a reviewed planning document do not themselves grant permission to edit locally. Preserve that boundary by ending the local interaction with a statement such as: “Do not implement; return the final negotiated plan only.”

Approve only Milestone 1 when every required edit supports its single purpose and every safeguard is enforceable in review. If the plan says “update related files as needed”, replace that with a stop condition: the implementer may identify additional files but must request approval before expanding consequential scope. This may trigger another round trip, but it protects the bounded handoff.

Apply the local approval gate

Before delegation, conduct a final human review of the plan against the repository rather than reviewing prose alone. Open the named definitions, inspect representative callers, confirm package exports, examine relevant tests, and verify the proposed dependency direction. Check that the working tree still matches the intended baseline and that the prompt contains neither secrets nor untrusted material.

Record one of three outcomes:

  • Revise locally: repository evidence contradicts the map, the file scope is uncertain, a cycle remains, or an important assumption lacks an owner.
  • Keep local: the milestone depends on interactive product judgement, sensitive data, unavailable cloud controls or an intermediate state that cannot be safely bounded.
  • Approve Milestone 1 for delegation: purpose, exact expected files, safeguards, evidence requirements, stop conditions and rollback are all reviewable.

Do not approve later milestones “in principle”. Repository state and discoveries from Milestone 1 may invalidate their order or scope. Approval applies to the packet as reviewed, not to every change that appears useful during implementation.

As of 4 October 2026, OpenAI’s documented workflow supports carrying the existing plan and local context into a new cloud chat from the IDE. That context transfer is useful, but the explicit Milestone 1 contract should still stand on its own: copied context can contain exploration, rejected options and assumptions that are not permissions. The next stage should send only the approved milestone as the operative task and treat all wider roadmap material as background.

Final decision rule: delegate only when a reviewer can answer four questions without inference: what may change, what must not change, what evidence must return, and what condition stops the task. Otherwise, continue negotiating locally. Planning effort is cheaper than reviewing an unbounded diff, but excessive decomposition can make a coherent dependency change impractical; choose the smallest milestone that leaves the repository in an intentionally reviewable state.

Delegate one approved milestone under explicit cloud controls

Original conceptual editorial illustration showing two separate blank change sets before final review for a human-reviewed workflow
Inspect the cloud diff and test evidence before deciding on a pull request.

The cloud hand-off is a controlled transition from an approved local plan to one implementation task, not permission to execute the entire refactor. As of 4 October 2026, OpenAI’s Prompting guide documents this sequence: select a cloud environment, delegate one milestone with the existing context, inspect the resulting diff and test evidence, and only then decide whether to request changes, pull the work locally or create a PR. The developer or tech lead remains responsible for confirming access, constraining the task, protecting sensitive data and conducting human review.

Delegate only when the milestone is independently reviewable and the required cloud controls are known. If the implementation would need unresolved design decisions, broad repository exploration or several coupled milestones, return to local planning.

Confirm that the account can create the intended task

Codex availability and repository authority are separate questions. Use the earlier “Run an access, environment and data preflight” section to confirm this account’s plan, seat and administrator settings rather than repeating the eligibility list. A user may have local Codex access without permission to create a cloud environment or write to the intended repository. Check the actual identity and permissions for this milestone.

Run the preflight with the identity that will start the task:

  1. Confirm that Codex Cloud is visible and enabled for the relevant account and workspace. Do not infer eligibility from access to local Codex, another ChatGPT feature or another member’s account.
  2. Check whether the seat can use an existing published environment or is expected to create or publish one. If an administrator must enable Cloud or an owner must publish the environment, stop and obtain that action rather than substituting a personal workspace.
  3. Verify access to the exact repository and the baseline branch named in the milestone. Read access may support inspection but not later commit or PR actions.
  4. Determine whether the task needs only a patch for local retrieval, permission to commit, or permission to open a PR. Granting write authority merely because it might be convenient expands the consequence of an error.
  5. Identify the person who can approve any later PR. Task creation and merge approval should not be treated as equivalent authority.

Fictional example: Priya can use Codex locally and can see the team workspace, but her seat cannot create cloud environments. A published environment maintained by the repository owners is available and she has repository read access. Her approved milestone only requires producing a reviewable diff that she will pull locally. The appropriate action is to use that published environment without requesting write access. If the workflow instead requires Codex to open a PR, Priya must first establish that the task identity has the necessary repository permission and that the team permits that action.

Use the least authority that completes the approved hand-off. If Cloud, the published environment or required repository access is unavailable, do not improvise with credentials in the prompt. Keep the milestone local or ask an authorised owner to correct the access path. Availability is also subject to rollout and workspace settings; the assigned official sources do not establish a universal country or region list.

Select a published environment that matches the milestone baseline

A cloud task cannot be started from an unpublished draft configuration. OpenAI’s Codex Cloud documentation says a published environment is required for a new task, and the current Cloud Environments guide says new tasks start from that environment’s prepared filesystem. Existing tasks retain their own state, including installed tools and uncommitted changes, but that retained state is not a substitute for source control.

Before selection, compare the candidate environment with the approved delegation packet:

  • Repository: confirm the repository identity rather than relying on a similar display name or fork.
  • Starting point: confirm the intended baseline branch or committed revision can be obtained. The local clean baseline remains the comparison reference; it does not guarantee the cloud workspace has identical state.
  • Prepared files: inspect what the published filesystem adds or generates, particularly configuration files that could alter builds, dependency resolution or tests.
  • Toolchain: check that the environment supplies the runtime, package tooling and project commands required by the milestone. Do not invent substitute validation commands during hand-off.
  • Access configuration: establish which repositories and external services the environment can reach and whether any of them permit writes.

Fictional example: the approved milestone extracts token parsing from src/auth/session.ts into src/auth/token-parser.ts, while preserving the exported createSession() interface. “Auth maintenance” is not enough to choose an environment. The reviewer checks that the published environment prepares the correct repository, can reproduce the project’s documented dependency installation and contains the branch against which the local plan was approved. An environment prepared for a neighbouring identity-service repository fails this gate even if it has a compatible language toolchain.

Prefer a newly started task for the newly approved milestone rather than reusing an older task merely because it contains useful uncommitted work. Each task has its own workspace; a new task will not recover another task’s uncommitted changes. The clean-start trade-off is that setup may be repeated, but the task’s provenance and diff boundary are easier to inspect. If previous work is genuinely required, put it into an approved source-control baseline or revise the milestone explicitly before delegation.

Inspect environment state without mistaking isolation for harmlessness

Workspace isolation separates one cloud task’s working files and changes from another task’s workspace. It does not make all actions read-only, prevent repository writes or neutralise external side effects. If credentials, network access and service permissions are available, a task may still be capable of affecting a repository or external service when instructed to perform a write action.

Apply a two-part inspection. First, review the environment as a starting image: prepared repository files, installed tools, environment variables, network-secret configuration, repository permissions and saved network policy. Secondly, review what the proposed task may do: which commands it may run, which services those commands contact and whether any requested operation writes outside the task workspace.

Example inspection note:

Environment purpose: auth-library maintenance
Repository required: example/auth-library
External writes required: none
Repository write required during implementation: no
Prepared tools required: project runtime and existing test runner
Allowed outbound use required: package source only if setup cannot use prepared state
Prohibited effects: publishing packages, updating tickets, rotating credentials,
writing to production services, opening a PR before review

This is an example control record, not a product-generated guarantee. Verify the actual environment rather than assuming the note is accurate. If a command would invoke an integration test against a live service, omit it from the cloud task until the service owner has approved a safe target and the organisation’s controls permit the action.

Isolation is sufficient for file separation, not for authorising external consequences. If the milestone can be implemented with repository-local edits and non-destructive validation, exclude external writes. If an external write is essential, treat it as a separate approval decision rather than silently broadening the refactor task.

Verify domains, credentials and the internet boundary

Internet access during the agent phase is conditional, not assumed. OpenAI’s Prompting guide says it is off unless enabled for the environment. The current Cloud Environments guide describes domain-configurable internet access and distinguishes direct environment variables from network secrets. Allowing a domain establishes a possible network route; it does not itself supply credentials or grant permission at the destination.

Review network requirements by purpose rather than opening broad access:

  1. List each external host the documented setup or validation procedure genuinely requires.
  2. For every host, name the operation: for example, retrieve an approved dependency, read documentation or call a non-production test endpoint.
  3. Check whether prepared environment state makes the request unnecessary.
  4. Confirm that the saved network policy permits only the required destination set.
  5. Determine whether authentication is needed and whether an approved environment-owned mechanism exists.
  6. Remove any destination or credential unrelated to this milestone before starting it, where the relevant administrator or environment owner controls that configuration.

Fictional example: the token-parser extraction uses dependencies already present in the published environment and its validation commands operate entirely on repository fixtures. The correct domain list for the agent phase may therefore be empty. Enabling unrestricted internet “in case the build needs it” would increase exposure without advancing the approved work. Conversely, if the prepared setup legitimately needs an approved package host, permit that specific requirement under the workspace’s policy; do not interpret the domain permission as authorisation to publish a package.

Do not paste an application programming interface (API) key, production token or other secret into the hand-off prompt, chat context or repository. A prompt is task instruction, not a secret-distribution channel. Do not ask Codex to print environment variables, inspect credential stores or report secret values as evidence that access works. A suitable verification asks whether an authorised command completed, while logs and diffs must still be checked for accidental disclosure.

Network access can enable dependency retrieval or an approved remote check, but it also adds changing external inputs and possible side effects. Keep it disabled when prepared state supports the task. Where access is necessary, restrict it to the approved purpose and require a reviewer to examine commands, output and changed files for unexpected network-dependent behaviour.

Exclude protected and unnecessary data before context transfer

The cloud chat can receive existing context, including the negotiated plan and local source changes, but context carry-over is not a reason to send everything from the local investigation. Treat each plan excerpt, source fragment, log and fixture as data leaving the local task boundary for cloud processing.

Apply the earlier access-and-data preflight’s prohibition on protected health information in Codex Cloud to every excerpt and fixture sent in the hand-off. Route other regulated or sensitive-data decisions through the organisation’s approved policy; neither an isolated workspace nor a published environment overrides it.

Use a content-screening procedure before hand-off:

  • Remove secrets, authentication headers, private keys, session values and production connection strings.
  • Replace real customer, employee or patient records with approved synthetic fixtures where those fixtures are sufficient for the milestone.
  • Exclude incident logs or database exports that contain unrelated personal or confidential information.
  • Check local uncommitted changes carried in context for temporary debugging output, copied credentials and generated data.
  • Include only source and documentation necessary to understand the approved milestone.
  • Record any redaction that changes an assumption the implementation must respect.

Fictional example: local investigation found a token-parsing failure in a production support log. The cloud hand-off must not include the real token or user record. The delegation packet instead states the structural invariant—“reject a token whose issuer field is absent”—and points to an approved synthetic fixture in the repository. If the behaviour cannot be represented without prohibited data, the task is not eligible for this cloud hand-off.

Include data only when it is required for the bounded implementation and permitted by the organisation. Redaction can reduce diagnostic detail, but that limitation should lead to a narrower task or local completion, not to bypassing the data boundary.

Send a cloud hand-off prompt that freezes the approved boundary

The new cloud chat receives existing context from the local conversation, including the plan and local source changes described by OpenAI’s Prompting guide, but it starts a separate isolated task workspace. Context continuity preserves useful reasoning; workspace separation means the developer must still state the baseline, milestone and acceptance boundary explicitly. Do not assume that carried context is current, complete or authoritative.

Construct the prompt from the approved packet in this order:

  1. Name exactly one milestone and its intended behavioural outcome.
  2. State the committed baseline or other approved comparison point.
  3. List the interfaces and observable behaviours that must remain unchanged.
  4. Identify likely target files while permitting discovery of directly necessary files.
  5. Specify non-goals and prohibited side effects.
  6. Provide the project’s actual validation commands and what each can establish.
  7. Require reporting of changed files, commands run, results, omissions and unresolved assumptions.
  8. Tell the task to stop rather than broaden scope when it encounters an unapproved architectural decision, missing access, sensitive data or a required external write.
  9. Prohibit commits, PR creation, publication and deployment unless separately authorised after review.

Example hand-off prompt:

Implement only approved Milestone 1: extract token parsing from
src/auth/session.ts into a focused internal module.

Baseline: use the repository and revision prepared in the selected published
environment. Report the resolved revision before editing.

Preserve:
- the public createSession() export and its call signature;
- existing acceptance and rejection behaviour represented by repository tests;
- current error categories exposed to callers.

Likely files:
- src/auth/session.ts
- src/auth/token-parser.ts
- directly relevant unit tests

Non-goals:
- no authentication policy changes;
- no dependency upgrades;
- no session-storage redesign;
- no production-service calls;
- no package publication or PR creation.

Validation:
Run only the repository-approved lint, type-check and focused test commands
listed in the approved plan. Report each exact command and result. If a command
cannot run, explain why; do not replace it with an unrelated check.

Stop and report if the change requires another milestone, a public-interface
change, unavailable credentials, sensitive data, a new outbound domain or an
external write.

Return a summary, changed-file list, diff, test evidence, unresolved assumptions
and any follow-up recommendation. Do not claim production readiness.

This sample is an example, not a guarantee of implementation quality or task behaviour. Replace paths, invariants and commands with repository-verified details. Do not copy placeholders into a live task. If the selected product surface offers a model choice, use only options visible and permitted in the account and workspace. OpenAI’s model-selection guidance says availability, tools, reasoning settings and usage limits differ by product and model version; a model named in documentation is not necessarily selectable in every workspace.

The final delegation rule is strict: send the task only if every sentence can be traced to the approved milestone or an established control. If the prompt contains “also”, “while you are there” or permission to refactor adjacent systems, split or renegotiate it. The benefit of richer context must be balanced against scope drift: include enough evidence to implement the milestone, but not unresolved alternatives that invite an architectural choice the task was never approved to make.

Record the hand-off so the next gate can reconstruct it

Before starting the cloud run, preserve a compact delegation record outside the task’s mutable workspace. Record the local baseline, selected published environment, milestone text, relevant permissions, required network policy, sensitive-data exclusions, exact prompt and named reviewer. This record distinguishes what was authorised from what the task later reports doing.

Example delegation record:

Milestone: AUTH-M1 token-parser extraction
Local comparison baseline: approved committed revision
Cloud environment: team-published auth maintenance environment
Repository authority: read and task workspace changes; no approved PR action
Agent internet requirement: none
Sensitive-data check: synthetic fixtures only; no PHI or production tokens
External writes: prohibited
Reviewer: repository maintainer
Permitted outcome: reviewable diff and validation evidence
Stop conditions: interface change, additional domain, credential need,
cross-milestone edit or unavailable project command

This is a suggested audit aid, not evidence that the environment enforced every statement. After task completion, compare the actual changed files, command transcript and reported assumptions with this record. If they diverge, the next action is revision or local inspection—not automatic acceptance.

Start the cloud task only when access, environment, data, network and scope checks all pass. Any unresolved item remains a stop condition. This may delay implementation, but it prevents an access workaround or ambiguous instruction from becoming part of the refactor’s unreviewed technical design.

Build the review packet before deciding what happens next

The cloud task’s answer is not the review packet. Treat it as a claim about work performed, then reconstruct the evidence needed for a PR decision. As of 4 October 2026, OpenAI’s Prompting guide documents the sequence as reviewing the cloud diff, iterating when needed, and only then creating a PR or pulling the changes locally to test and finish. OpenAI’s Codex Cloud documentation likewise instructs the user to inspect changed files and results before committing or opening a PR.

Require the task to return six concise items: a changed-file map, the constraints it was meant to preserve, commands and reported results, checks it did not run, remaining test gaps, and rollback notes. These items answer different questions. The file map describes scope; preserved constraints describe intent; command evidence describes attempted validation; unrun checks and test gaps expose uncertainty; rollback notes describe how to remove the milestone without improvising during a failure.

  1. Copy the approved milestone contract beside the cloud response.
  2. Record the task’s starting commit or other agreed source-control baseline.
  3. Extract the six evidence items without rewriting ambiguous claims into stronger ones.
  4. Mark each item as reported by task, verified in diff, or verified locally.
  5. Stop if the baseline, changed files or validation status cannot be reconstructed.

Example review-packet heading, not a product-generated guarantee: “Milestone 1: move token parsing behind the existing authentication module boundary; baseline <commit>; five files reported changed; unit test command reported successful; repository-wide integration suite not run; no database or public interface changes intended.” A reviewer must replace placeholders and verify every statement against the repository and task evidence.

Proceed to detailed diff inspection only when the packet distinguishes facts from assumptions and explicitly names missing evidence. A shorter packet with an honest “not run” is preferable to a polished narrative that conceals uncertainty.

Reconcile the changed-file map with the approved boundary

A changed-file map is not merely a list from git diff --name-only. It should state why each file changed and whether the change was authorised, incidental or generated. This distinguishes legitimate implementation spread from scope drift. A refactor may need an additional import, fixture or build manifest change, but an unexplained file remains unexplained even if the overall patch looks plausible.

File or path Claimed purpose Reviewer check Disposition
src/auth/token_parser.* New internal extraction Confirm logic came from the approved module and exposes no new public contract Expected, revise or reject
src/auth/index.* Delegation to extracted code Compare inputs, outputs and error propagation with the baseline Expected, revise or reject
tests/auth/* Coverage of preserved behaviour Confirm assertions test behaviour rather than implementation shape alone Expected, revise or reject
package-lock.* or equivalent Dependency metadata Determine whether the milestone authorised any dependency change Justify, revert or reject

Worked fictional example: the approved authentication extraction permits changes under src/auth/ and corresponding tests, but the diff also changes a deployment manifest. Do not infer that the manifest update is harmless. Ask why it changed, inspect the exact lines, and require its removal unless it is necessary and separately approved. If it is genuinely necessary, revise the milestone record before proceeding rather than quietly expanding the original mandate.

Also inspect for missing files. A changed implementation without an expected test update may reveal a gap; a moved module without corresponding export or ownership metadata may leave the repository inconsistent. Conversely, unchanged tests are not automatically a defect if existing tests exercise the preserved behaviour. The reviewer’s job is to determine whether the evidence covers the refactor’s risk, not to reward file churn.

Request a follow-up for a small, explainable boundary correction. Pull locally when the scope is valid but repository-specific checks cannot run credibly in the cloud environment. Reject or restart the milestone when changes cross into unrelated subsystems, dependencies, deployment configuration or public interfaces without prior approval.

Check preserved constraints line by line

“No behaviour change” is an objective, not evidence. Revisit each invariant from the local plan and connect it to both code and validation. Separate structural success—such as moving a parser—from behavioural preservation—such as retaining accepted inputs, error categories and call ordering. A structurally tidy patch can still violate the behavioural contract.

  1. List each approved invariant in its original wording.
  2. Locate the changed code paths that could affect it.
  3. Identify the test, static check or manual examination that bears on it.
  4. Record whether the evidence is direct, partial or absent.
  5. Escalate any newly discovered invariant to the responsible human rather than silently defining expected behaviour during review.

Fictional example matrix:

Preserved constraint Possible evidence Human verification
Callers continue using the existing authentication entry point Diff contains no caller migration outside the approved module Search the repository for direct imports of the new internal parser
Malformed tokens retain the existing error category Existing or added behavioural assertions Compare error construction and propagation with the baseline
No new package is introduced Dependency manifests and lockfiles remain unchanged Inspect the full changed-file set, not only source files
Logging does not gain token material Diff examination of log arguments Reject any exposure; do not place real tokens or secrets in a prompt or test fixture

Inspect the changed code and its relevant seams; reviewing only the newly extracted file misses altered call sites and error paths. Prioritise semantic seams: public entry points, data transformations, side effects, exception handling, concurrency boundaries and configuration reads.

A constraint passes this gate only when a reviewer can name the relevant code path and the supporting evidence. “Codex said it preserved behaviour” and “the tests passed” are not substitutes for that mapping. Consequential acceptance requires human review.

Audit commands, reported results and everything not run

Command evidence has three parts: the exact command, where it ran, and the reported outcome. Keep those distinct from independent confirmation. OpenAI’s current Help article, checked on 4 October 2026, says users should review changes and test results before using cloud work; it does not say an agent-reported result proves correctness or production readiness.

For each command, capture the working directory, relevant non-secret configuration assumptions, whether it targeted the whole repository or a subset, and whether the result came from the cloud task or a later local run. Do not paste credentials, application programming interface (API) keys, production data or other secrets into prompts or review notes. Keep untrusted issue text, logs and repository data out of prompts unless they have been reviewed and reduced to the minimum necessary context.

Example evidence ledger, with fictional commands:

Reported by cloud task:
- Command: <project unit-test command for auth package>
- Scope: authentication package only
- Outcome: <copy the task's exact reported outcome>

Not run:
- repository integration suite
- platform-specific checks
- build using release configuration
- manual compatibility check by an owning team

To verify locally:
- <project lint command>
- <project type-check command>
- <project full-test command>

Never convert a vague statement such as “tests look good” into a successful result. Ask for the command and output summary. If a command failed, preserve the failure and distinguish a pre-existing failure from one introduced by the patch; establishing that distinction may require checking the clean baseline locally. If neither baseline nor patch can complete a required check, status is unknown—not passed.

Test gaps deserve their own list because “not run” and “not tested” differ. An integration test may exist but remain unrun because the cloud environment lacks a service. A compatibility case may have no test at all. The former suggests local execution; the latter may require a new test, targeted manual verification, or explicit risk acceptance by the appropriate owner.

Request a cloud follow-up when the missing check is available in the published environment and remains within scope. Pull locally when the required service, platform, toolchain or trusted fixture exists only in the approved local environment. Do not create a PR when a required check is unknown and no accountable reviewer has accepted that uncertainty.

Inspect the diff independently, not through the completion narrative

Read the actual patch before reading the task’s rationale in depth. This reduces the risk of adopting its framing and overlooking contradictory changes. Start with the summary and file list, then inspect the complete diff, and finally compare high-risk functions against the baseline. A separate cloud workspace helps isolate task working files, but isolation does not establish that repository writes, external actions or credential-backed operations are harmless.

  1. Confirm the comparison base matches the delegated milestone.
  2. Review deletions before additions so that removed guards and side effects are visible.
  3. Trace changed imports, exports, call sites and error paths.
  4. Search for duplicate old and new implementations that could diverge.
  5. Inspect generated files separately and reproduce them with the project’s approved procedure where required.
  6. Check for debug output, temporary bypasses, broadened network access, test-only shortcuts and accidental data exposure.
  7. Review comments and names for claims that the code does not enforce.

Worked fictional example: a new parseToken() helper reproduces the success path but changes a catch-all exception into a narrower exception. The file move itself matches the milestone, yet the error boundary may not. Compare the old and new control flow, identify callers that depend on the former error handling, and request a correction or explicit plan revision. Do not approve merely because existing happy-path tests remain green.

Use follow-up tasks for bounded corrections such as reverting an unauthorised manifest edit, restoring a lost guard or adding a clearly specified behavioural test. Avoid an open-ended instruction such as “clean up anything else you notice”; that creates a new redesign task with no stable review boundary.

If the diff can be corrected without changing the approved architecture or acceptance criteria, request a precise follow-up. If review reveals that the milestone itself was based on a false dependency assumption, stop cloud implementation and return to local planning.

Choose between revision, local completion and a PR

These are three distinct outcomes, not stages that every task must traverse. A cloud revision is suitable when the defect is bounded and the environment can verify it. Local completion is suitable when trusted services, platform-specific tooling or a fuller repository baseline are required. PR creation is suitable only when the evidence is coherent and an accountable reviewer is ready for the normal repository review process.

Request a follow-up
State the observed discrepancy, the exact required change, prohibited collateral changes and the command or examination needed afterward. Example: “Restore the existing malformed-token error category; change only the parser extraction and its behavioural test; do not alter callers or dependencies; report the targeted command and any unrun checks.”
Pull changes locally
Use the repository’s approved source-control procedure, confirm the baseline and inspect the resulting diff before running commands. OpenAI’s Prompting guide explicitly presents pulling changes locally to test and finish as an alternative to creating a PR directly from cloud. Do not assume uncommitted state from another cloud task will appear: OpenAI’s current environment guidance says tasks retain separate state and warns that saved state does not replace source control.
Create a PR
Use the reviewed milestone wording, changed-file map, validation ledger, known gaps and rollback note in the PR description. PR creation starts another review surface; it is not approval to merge or deploy. Apply repository ownership, continuous integration and release controls as separate gates.

Fictional decision example: the authentication extraction is within scope, but its integration suite requires an internal service unavailable to the cloud task. If policy permits testing that branch locally, pull it, verify the source-control base, run the approved integration procedure and record the actual result. If local testing exposes a defect, either correct it under the same bounded milestone with review or request a precise follow-up. Create the PR only after the evidence ledger reflects what actually happened.

Choose the least expansive action that resolves the uncertainty. Do not use a new cloud iteration to avoid a required local check, and do not create a PR merely to discover whether the patch is reviewable.

Write rollback notes that match the refactor’s real failure modes

Rollback is not simply “revert the commit”. A refactor may be safely removable when it changes only internal code, but rollback becomes more complex if generated artefacts, migrations, external writes or public contracts are involved. The note should distinguish source rollback from operational recovery and should not claim reversibility that has not been examined.

For a bounded, behaviour-preserving extraction, record the commits or patch boundary to revert, any generated files to regenerate, the validation to repeat after reversal, and whether the old implementation remains available. Do not include sensitive values or production credentials. If the task performed an external write, altered repository settings or opened a PR, record that separately; workspace isolation does not undo those actions.

Example rollback note: “Example only: revert the milestone’s source and test changes as one unit; confirm the original authentication entry point is restored; rerun the project’s approved authentication checks; no database migration or intended external state change is part of this milestone. A human must verify those assumptions before relying on this note.”

Approve a simple source rollback only when the diff confirms that the task introduced no migration, irreversible transformation, external side effect or public compatibility commitment. Otherwise, require a service owner’s recovery plan before PR creation.

Apply the final go/no-go gate

Use this checklist after diff inspection and any local testing. It is a decision record, not a scoring system: one unresolved mandatory item is enough for “no-go”.

  • Baseline: the reviewed diff is against the intended commit or clean, documented source-control state.
  • Scope: every changed file has an understood purpose, and unauthorised changes have been removed or separately approved.
  • Constraints: each preserved behaviour and non-goal has code-level and evidential mapping.
  • Commands: exact commands, execution locations and reported outcomes are recorded without upgrading claims.
  • Unrun checks: required checks not run in cloud have been run in an approved environment or explicitly blocked.
  • Test gaps: absent coverage and manual verification needs have named owners and dispositions.
  • Diff: a qualified person has independently inspected deletions, additions, call sites, side effects and error paths.
  • Data: no protected, unnecessary or secret data was placed in prompts, fixtures or review artefacts.
  • Rollback: the reversal boundary and post-reversal validation are credible for the actual changes.
  • Permissions: repository write and PR actions are intentional and within the reviewer’s authority.
  • Approval: the designated human reviewer accepts the remaining uncertainty and the normal repository controls still apply.

Go: create the PR when all mandatory evidence is present, discrepancies are resolved, and the responsible reviewer is satisfied. No-go: request a bounded revision, pull locally for missing verification, or return to planning. A passing task-reported test set, an isolated workspace or a polished explanation cannot independently establish correctness, security, compliance or production readiness.

Do not delegate these cases

  • Protected health information: do not use Codex Cloud to process it. Follow the earlier “Run an access, environment and data preflight” prohibition; route other regulated-data questions through the organisation’s approved policy and controls.
  • Unbounded redesigns: do not delegate a vague request such as “modernise the subsystem” without specifying stable interfaces, explicit non-goals and a reviewable milestone. Investigate and negotiate the architecture locally first.
  • Unknown test baselines: do not delegate implementation when the team cannot distinguish existing failures from regressions. Establish or explicitly adjudicate the clean baseline before asking cloud work to preserve it.
  • High-impact credentials: do not expose production keys, privileged repository credentials or credentials capable of consequential external writes. Redesign the task around least-privileged, approved access or keep it in a controlled human-run process.

Access 40,000+ AI Prompts for ChatGPT, Claude & Codex — Free!

Subscribe to get instant access to our complete Notion Prompt Library — the largest curated collection of prompts for ChatGPT, Claude, OpenAI Codex, and other leading AI models. Optimized for real-world workflows across coding, research, content creation, and business.

Access Free Prompt Library

Useful Links

Get Free Access to 40,000+ AI Prompts for ChatGPT, Claude & Codex

Subscribe for instant access to the largest curated Notion Prompt Library for AI workflows.

More on this