Codex Git Worktrees Guide: Run Two Bounded Bug Fixes in Parallel Without Polluting Your Main Checkout
Source date: 5 October 2026. The practical answer is to use two separate Codex Worktree chats only after establishing, as far as inspection allows, that the two bug fixes (Fix A and Fix B) are bounded, independent changes and that the chosen base branch and commit are known. Give each chat one defect, one focused reproduction and one explicit “done when” check. If the likely edits overlap in a module, migration, generated output, lockfile or stateful external service, do not run the fixes in parallel: sequence them from the same verified base instead.
Evidence checkpoints
Documented point: Worktrees let Codex run multiple independent chats in the same project without interfering with each other, while the repository, worktree and commands remain on the computer or remote development environment containing the project. Source accessed 5 October 2026. [OpenAI documentation: Git worktrees]
Documented point: Local environments configure worktree setup steps and common project actions. Source accessed 5 October 2026. [OpenAI documentation: Local environment]
Documented point: The desktop app distinguishes Local (current project directory), Worktree (isolated Git-worktree changes) and Cloud (a remote task workspace from a published reusable environment). Source accessed 5 October 2026. [OpenAI documentation: Modes]
Documented point: OpenAI advises keeping one Codex chat per coherent unit of work. In this guide, the two fictional bug fixes are treated as separate, bounded tasks whose results require independent human review. Source accessed 5 October 2026. [OpenAI documentation: Best practices]
Documented point: In the app, /review can assess a base branch or uncommitted changes and report prioritised findings without changing the working tree. Source accessed 5 October 2026. [OpenAI documentation: Code review]
Documented point: By default, the agent has network access turned off; local execution uses an operating-system-enforced sandbox plus an approval policy. Source accessed 5 October 2026. [OpenAI documentation: Agent approvals security]
Documented point: Each desktop-app chat has a terminal scoped to its current project or worktree. Source accessed 5 October 2026. [OpenAI documentation: Integrated terminal]
Documented point: For Git repositories, a scheduled task can run in Local or a new dedicated background worktree; worktrees separate scheduled-task changes from unfinished local work. Source accessed 5 October 2026. [OpenAI documentation: Automations]
1. The exact fit: two unrelated, reviewable bug fixes
A Git worktree is a separate checkout with its own files while sharing the repository’s Git metadata. In the ChatGPT desktop app, Codex Worktree mode is intended to isolate changes from the current checkout; OpenAI’s environment documentation describes it as “Worktree: isolate changes in a Git worktree”. This is different from Local mode, which operates in the current project directory, and from Cloud mode, which is outside this guide’s local two-lane workflow. These distinctions are documented by OpenAI’s Codex environments guidance, accessed on 5 October 2026.
The fit is narrow but useful: two defects can be investigated concurrently when each has a small, intelligible change surface, a separate acceptance check and no expected dependency on the other fix. Worktrees isolate the Git checkouts; they do not isolate every database, port, cache, queue, credential, browser session or third-party service used by the application.
For longer-term task management, including archiving or deleting worktrees, see the Codex 0.155 Managed Task Lifecycle Playbook. Keep those lifecycle decisions separate from the two local fixes in this guide.
1.1 Apply a five-part independence test
Before opening either chat, write a one-line description of each defect and compare them across five boundaries. Do not rely on ticket titles alone: two differently worded tickets can still converge on the same implementation file or shared runtime state.
- Code boundary: identify the modules, packages or components most likely to contain the root cause. Parallel work is reasonable when the likely edit sets are separate. Stop if both investigations may alter the same module or a shared helper beneath both call paths.
- Schema boundary: ask whether either fix can require a database migration, schema declaration, fixture format or serialisation contract. If both may change one of these shared structures, sequence the work.
- Dependency boundary: determine whether either fix is likely to modify a dependency manifest or shared lockfile. A worktree gives each investigation its own file copy, but two independently regenerated lockfiles can still conflict or encode incompatible dependency choices.
- Generated-output boundary: identify generated clients, snapshots, compiled assets or checked-in artefacts. If both fixes regenerate the same output, do not treat separate source edits as sufficient independence.
- External-state boundary: list shared services used during reproduction, such as one development database, message broker, test account or fixed network port. Separate checkouts do not prevent one lane from changing data or consuming messages needed by the other.
Decision rule: proceed with two Worktree chats only if all five boundaries are separate or can be made separate without broadening permissions or changing the intended fixes. One uncertain shared boundary is enough to pause. The trade-off is straightforward: parallel investigation may reduce waiting time, but sequencing gives a reliable order of changes and avoids reconciling two answers to the same underlying problem.
1.2 Worked classification example
Consider an illustrative repository with two reports:
- Fix A: a text parser mishandles an empty token in one reproducible input case.
- Fix B: a separate cache-expiry boundary retains an entry beyond its expected lifetime.
A preliminary path search suggests that Fix A belongs in the parser implementation and its focused regression test, while Fix B belongs in the cache-expiry module and its own unit test. Neither appears to require a schema change, dependency update, generated client or shared external service. That is a plausible two-lane candidate, subject to repository inspection. In the rest of this guide, each fix and its own Worktree chat form a lane: Lane A handles Fix A and Lane B handles Fix B.
Now change the example slightly: suppose both defects turn out to depend on a shared serialisation helper. The tickets still describe different symptoms, but the implementation boundary is shared. Sequence them. Similarly, if both fixes require updating the same snapshot bundle or running a dependency installer that rewrites one lockfile, keep them in one ordered lane unless a human maintainer has deliberately separated the dependency decision.
This assessment is not proof of independence. It is a stop/go screen. Ask a developer familiar with the repository to review the predicted paths, shared dependencies and runtime services before two chats begin. If early investigation in either lane discovers a shared-risk area, stop both implementation paths, preserve the evidence and decide which fix should go first.
1.3 Keep one coherent fix in each chat
OpenAI’s Codex best-practices guidance, accessed on 5 October 2026, says: “Keep one chat per coherent unit of work.” For this guide, one coherent unit means one reproducible defect and the smallest supporting test or check needed to establish the repair. Do not ask one lane to clean nearby code, modernise dependencies, rename unrelated interfaces or absorb the other bug because it appears convenient.
Use a bounded bug card before creating a chat. The following is an example structure, not a guarantee of the product’s output:
Goal: Fix the parser when one reproducible input contains an empty token.
Context:
- Reproduction: parse a permitted input with an empty token between delimiters.
- Suspected paths: parser/tokenise.ts and its focused tests.
Constraints:
- Keep the diff minimal.
- Do not refactor unrelated parsers or the cache module.
- Do not edit database migrations, dependency files or generated artefacts.
- Stop and report if the root cause requires a shared serialisation change.
Done when:
- The smallest focused regression test for this reproduction passes.
- Run the relevant type check if available.
- Report changed files, root cause, exact commands, failures or skipped checks,
remaining uncertainty, and whether the lane is ready for human review.
Fix B needs its own card, with its own paths and check. For example, its stop condition might prohibit changes to the shared schema, localisation catalogue and lockfile. Do not copy Fix A’s terminal output, assumptions or acceptance result into Fix B’s chat. A successful test in one worktree says nothing about the other checkout.
Failure handling: if a chat reports that its minimal repair must touch an excluded shared area, treat that report as a useful investigation result rather than permission to continue. Ask it to summarise the discovered dependency and stop editing. A human should inspect the reported paths and decide whether to sequence the fixes, revise their scope or reject the proposed direction.
1.4 Understand what isolation does—and does not—provide
According to OpenAI’s Worktrees documentation, accessed on 5 October 2026, Codex can use separate worktrees for independent chats in the same project. Each checkout has separate repository files, while Git metadata is shared. This protects the main checkout from ordinary file edits made in the two worktrees and prevents the chats from writing into the same checked-out file tree.
It does not promise conflict-free integration. Both lanes can independently edit equivalent lines, choose incompatible assumptions or generate conflicting artefacts. Nor does checkout isolation allocate separate ports, clone a database or create distinct third-party accounts. If both test suites start a server on port 3000, for example, one may fail because the other already owns the port. The remedy is a documented project-specific port override or sequential execution—not an assumption that Worktree mode virtualises the runtime.
Use least privilege while assessing or running either lane. Do not place access tokens, production records, private customer material or other secrets in prompts. Keep untrusted logs and issue text out of instructions unless they have been reviewed and reduced to the minimum safe reproduction details. If a setup or test command genuinely needs network access, a human should evaluate and approve that need for the particular investigation rather than broadening permissions merely to make both lanes run.
Human verification: before parallel work starts, one developer should sign off the two bug cards and the independence table. The review should answer three questions: are the predicted edit surfaces separate; can the tests run without shared destructive state; and does each lane have a focused, observable acceptance condition? If any answer is “unknown”, investigate that uncertainty first.
2. Preflight the main checkout before selecting a base
The base decision determines what both investigations see. Do not create worktrees first and inspect the repository afterwards. OpenAI’s Worktrees documentation says a new Worktree chat starts from the selected branch’s HEAD, and selecting a branch with unstaged local changes applies those changes to the new worktree. A dirty checkout can therefore inject unrelated edits into what appears to be an isolated investigation.
2.1 Confirm that the directory is a Git repository
From the intended main checkout, run read-only commands before changing anything:
git rev-parse --show-toplevel
git status --short --branch
git branch --show-current
git rev-parse HEAD
git rev-parse --show-toplevel should identify the repository root. git status --short --branch gives a compact view of the current branch and changed or untracked paths. git branch --show-current records the named branch when one is checked out, while git rev-parse HEAD records the exact commit identifier.
These commands are examples of a preflight, not evidence that the repository is suitable. If the first command fails, stop: Worktree mode requires a Git repository. If the reported root is a parent monorepository rather than the component you expected, assess status and test commands at that actual boundary. Running from a nested directory does not make Git forget changes elsewhere in the repository.
Decision rule: continue only when the reported root is the intended project and HEAD resolves. If the repository is mid-merge, mid-rebase or otherwise in an unfinished Git operation, complete or abort that operation deliberately before creating either lane. Do not ask Codex to infer ownership of an interrupted operation.
2.2 Classify every status entry
Read the status output path by path. Group entries into four categories:
- Intended base content: committed files already present at the selected
HEAD. - Developer work in progress: tracked modifications, staged changes or untracked source files belonging to another task.
- Local tooling state: ignored configuration, caches or environment files used only to run the project.
- Unexpected residue: files whose origin or ownership is not understood.
Only the first category is automatically a sound base. Developer work in progress must not be silently carried into both fixes. Unexpected residue requires investigation. Local tooling state needs a separate setup decision; being present in the main checkout does not mean every ignored or untracked file will appear in a managed worktree.
For example, suppose status reports:
## feature/payment-copy
M web/checkout/banner.ts
?? notes/payment-investigation.txt
If neither file belongs to Fix A or Fix B, do not select this dirty branch merely because its committed history is convenient. Selecting a branch with unstaged changes can apply those changes to the new worktree under the documented flow, contaminating the investigation and later diff. The safe choices are human-owned: finish and commit that work on an appropriate branch, stash it with a clear recovery plan if repository policy allows, or choose a separate clean and known base. This guide does not prescribe stashing automatically because doing so can conceal work or mishandle untracked files.
If status is clean but the branch contains an unreviewed commit from another feature, “clean” is not synonymous with “correct base”. Compare the branch’s purpose and commit history with the target fixes. A release-maintenance defect might need the maintained release branch; a defect intended for current development might start from the repository’s agreed integration branch. Use the project’s contribution policy rather than guessing from branch names.
2.3 Record a reproducible base manifest
Create a small human-readable preflight record outside the prompts if it contains sensitive repository information. At minimum, record:
- repository root;
- selected base branch;
- exact base commit from
git rev-parse HEAD; - whether status was clean at inspection time;
- the time of inspection;
- Fix A’s reproduction and acceptance command;
- Fix B’s reproduction and acceptance command;
- known excluded paths and shared-state stop conditions.
An illustrative record could read:
Base branch: develop
Base commit: <record the actual commit identifier>
Status: clean at preflight
Fix A reproduction: focused parser test for an empty token
Fix A done when: that regression test passes; a relevant type check passes if available, or its absence is recorded for human review
Fix B reproduction: focused cache-expiry test at the lifetime boundary
Fix B done when: that regression test passes; a relevant lint check passes if available, or its absence is recorded for human review
Stop either lane if: shared schema, migration, lockfile, generated client,
localisation catalogue or common external test account must change
Do not fabricate a commit identifier or claim that a command passed before running it. If no focused test exists, define the smallest deterministic reproduction that a human can inspect, then ask the lane to add an appropriately scoped regression test if repository conventions permit. A screenshot, anecdotal manual click or “the output looks right” is weaker than a repeatable check and should be labelled accordingly.
2.4 Define “done when” as an observable test
A useful acceptance statement identifies an input, expected behaviour and command or procedure. “Fix the cache” is not enough. “Given an entry at the documented lifetime boundary, the focused cache-expiry test confirms its removal at the expected point” is bounded and reviewable. Add a broader lint, type-check or build command only when it is relevant and practical; do not let a large unrelated failure obscure the focused reproduction.
Record baseline behaviour where feasible. Run the smallest reproduction on the chosen base before opening the worktree chats. Its purpose is to confirm that the defect exists at that exact commit and that the check can distinguish failure from success. Report the command, exit status and material output without secrets. If the defect cannot be reproduced, stop and refine the conditions rather than instructing two chats to repair an unverified symptom.
Baseline failures also need classification. If Fix A’s focused test fails for the expected assertion, it is a useful baseline. If it fails because dependencies are absent, setup must be resolved first. If it calls a shared unstable service, it is not yet a dependable lane-level acceptance test. The trade-off is between speed and interpretability: a narrow test may miss wider regressions, while a full suite may be slow or noisy. Use the narrow test for the defect signal, then specify the smallest relevant broader check as supporting evidence.
2.5 Decide how dependencies will be provisioned
As of 5 October 2026, OpenAI’s Local environments documentation says: “Setup scripts run automatically when Codex creates a new worktree at the start of a new chat.” The documentation scopes local environments to Codex in the ChatGPT desktop app. A setup script can install dependencies or perform an initial build, but successful setup is not acceptance of either fix.
Inspect the intended setup script before assigning it to both lanes. Determine whether it is repeatable, whether it needs network access, whether it modifies tracked files and whether two executions can contend for a global cache, fixed port or external service. Prefer deterministic, repository-defined setup over an improvised prompt that asks the agent to discover credentials or fetch arbitrary tools.
For example, a repository might define an illustrative setup action that installs locked dependencies and performs a non-destructive initial build. Before using it twice, verify that it respects the lockfile rather than rewriting it, does not start a persistent server on a shared fixed port and does not require production credentials. If it would alter the lockfile in both lanes, stop and correct the setup procedure or sequence the work.
A .worktreeinclude file requires deliberate review. In the documented local managed-worktree case, it copies only matching ignored paths, skips source symlinks and does not overwrite files already present. This behaviour is not documented for remote or command-line-created worktrees. Do not assume that every ignored or untracked local file is copied. More importantly, do not add secret-bearing paths merely for convenience. Prefer a safe, documented way to create lane-specific development configuration, and have a human review every include pattern and the resulting destination files.
2.6 Reserve branch names; do not occupy them prematurely
A Codex-managed worktree begins in detached HEAD by default. That is useful during investigation because it avoids prematurely occupying a named branch. Create or assign a named branch when a change is ready to persist or publish under the repository’s workflow, not merely to label an experiment.
Git’s occupancy rule still applies. OpenAI’s Worktrees guide states: “Git only allows a branch to be checked out in one place at a time.” In operational terms, a branch cannot be checked out in both Local and a worktree simultaneously. Do not create fix/parser-empty-token in a worktree and then try to force Local onto that same branch. Reserve distinct descriptive names for the two prospective fixes, but leave them unoccupied until needed.
When one worktree already occupies a branch, finish that fix and write down the handoff clearly before moving on; for a larger change, the guidance in Codex Refactor Delegation Playbook describes planning one bounded refactor locally, handing an approved milestone to Cloud, and inspecting its diff and tests.
An example naming plan might reserve fix/parser-empty-token for Fix A and fix/cache-expiry-boundary for Fix B, subject to the repository’s conventions. Distinct names reduce ambiguity; they do not prove that the changes are independent. If one completed lane later needs the usual integrated development environment (IDE)A software application combining tools for writing, building, testing and debugging code. Open glossary entry, an existing development server or direct collaboration in the foreground checkout, use the documented Handoff flow rather than attempting a second checkout of its occupied branch. Handoff is not a merge, rebase or conflict-resolution mechanism.
2.7 Final human gate before opening two chats
Have a human reviewer compare the preflight record with live Git output immediately before starting. Confirm that the branch and commit have not changed, status remains understood, both baseline reproductions are valid, and setup does not need unjustified secrets, workspace expansion or network access.
Proceed when the base is explicit, status is clean or every carried change is deliberately owned, both defects have separate predicted edit surfaces, each has an observable “done when” check, and dependencies can be provisioned without a shared collision. Stop and sequence when either investigation may touch a common module, migration, generated artefact, lockfile or externally stateful service, or when the baseline cannot reliably distinguish the defect.
The preflight record informs the decision but does not make it. A person remains responsible for deciding whether the two fixes may run concurrently and, later, whether either resulting diff is correct enough to commit or propose in a pull request.
3. Prepare repeatable worktree setup
Configure the environment before opening either bug-fix chat. The objective is not to make every worktree reproduce an entire production system; it is to give both lanes the same minimal, reviewable path from a clean checkout to the focused validation command. This preparation belongs in Codex in the ChatGPT desktop app: OpenAI’s local-environment documentation, accessed 5 October 2026, scopes local environments to that desktop workflow and states that setup scripts run automatically when a new worktree is created at the start of a chat. Availability can still depend on the reader’s plan, workspace controls, role, client version and rollout.
3.1 Define one desktop local environment for both lanes
A local environment differs from a bug prompt. The environment supplies repeatable setup steps and common commands; the prompt defines the defect, constraints and acceptance evidence. Keeping those responsibilities separate prevents Lane A from acquiring undocumented bootstrap instructions that Lane B lacks.
- Open the project in the Codex view of the ChatGPT desktop app.
- Locate the project’s local-environment configuration and create or select the environment intended for these two worktrees.
- Configure only the setup required by a fresh checkout: dependency installation, generated local prerequisites that are safe to recreate, and an initial build only if later tests require it.
- Add reusable actions for the smallest useful test and build commands.
- Save the environment, but do not launch the two chats until the script has been inspected for side effects and secret handling.
For example, a TypeScript repository might need dependency installation before either defect can be reproduced. An illustrative setup script could be:
set -eu
npm ci
This is an example, not a guaranteed command for every repository. Use the package manager and lockfile policy already established by the project. Do not silently replace a locked installation with an updating install, and do not add operating-system package installation merely because it is convenient. If the repository requires a private registry, first determine whether dependency installation genuinely needs network access and how credentials are supplied without copying them into a prompt.
The decision rule is to automate a step only when both lanes need it, it is deterministic enough to repeat, and its side effects remain within the intended workspace or an explicitly reviewed dependency cache. Leave bug-specific fixture creation, destructive database resets and external-service mutations out of the shared script. Those steps belong in the relevant lane, with explicit approval and a clear rollback procedure.
If the setup fails, preserve the error output and classify the failure before editing application code. A missing runtime is an environment prerequisite; a lockfile mismatch may be a repository-state problem; a registry authentication failure is an access problem. Do not ask Codex to work around any of these by weakening verification or embedding a credential. A human should review the exact failing command, confirm that the script matches repository policy, and decide whether the lane may proceed.
Worktree setup scripts that install dependencies, copy environment files or start services form part of the local security surface; review how to Harden Codex Local Projects before relying on them.
3.2 Keep the setup script minimal and non-destructive
A minimal script establishes prerequisites; it does not alter the defect’s evidence. This distinction matters because an eager bootstrap can regenerate tracked files, rewrite a shared lockfile or initialise an externally stateful service before either investigation begins. Two isolated Git checkouts do not isolate a shared database, registry account, network service, port or cache.
Review the script line by line using this procedure:
- Mark each command as read-only, workspace-writing, network-using or externally state-changing.
- Remove commands unrelated to reproducing or testing both fixes.
- Replace broad commands with repository-supported, locked equivalents where available.
- Check whether two simultaneous executions contend for the same port, database schema, generated destination or shared temporary path.
- Make failure explicit so that a failed prerequisite does not fall through to later commands.
Suppose a project ordinarily starts a development server after installation. Do not put that long-running process in the automatic setup merely because both bugs are web-facing. Define it as a reusable action instead. That keeps worktree creation finite and lets the developer choose whether Lane A or Lane B owns the usual port. If both fixes require a live server on one fixed port and the project offers no supported isolation, run those checks sequentially.
Apply the same caution to generated files. If an initial build writes a tracked bundle or shared lockfile, either exclude the build from setup or stop the parallel plan. The trade-off is a slightly slower first validation versus avoiding unexplained changes in both diffs. Prefer the slower but attributable workflow.
OpenAI’s agent approvals and security guidance, accessed 5 October 2026, says local agent network access is off by default and that monitoring does not replace sandboxing, permissions or review of the result. Do not broaden network or workspace permissions simply to make setup frictionless. Approve access only when the named dependency or reproduction step requires it, and keep secrets and untrusted issue content out of prompts.
For human verification, run or inspect the setup in a disposable clean context permitted by the project, then compare git status --short before and after. The expected result should be stated in advance—for example, “dependencies may appear only in ignored paths; tracked files must remain unchanged”. If tracked files change, investigate before launching the worktrees rather than accepting the setup as successful.
3.3 Create reusable Test and Build actions with different purposes
A Test action answers whether specified behaviour meets an assertion. A Build action answers whether the relevant project target compiles or packages under the repository’s rules. They provide different evidence and should not be collapsed into a vague “check everything” action.
Configure actions that run the project’s existing commands in the integrated terminal. OpenAI’s integrated-terminal documentation, accessed 5 October 2026, says each desktop-app chat has a terminal scoped to its current project or worktree. That scope helps direct commands to the intended checkout, but it does not guarantee that a command is correct, complete or free from external side effects.
An illustrative action set for a TypeScript project might be:
Test focused:
npm test -- --runInBand path/to/relevant.test.ts
Build:
npm run build
These are sample commands, not product guarantees or universal syntax. Replace them with repository-documented commands. If the test runner does not accept the shown arguments, do not improvise until a command exits successfully; consult project configuration and record the actual invocation.
Use this action-design procedure:
- Create a fast, focused test action or document the parameter that each lane must supply.
- Create a build or type-check action only when it provides distinct evidence relevant to the changed code.
- Keep linting separate if it can rewrite files; prefer a check-only mode where the project supports one.
- Name any prerequisites, such as a service or fixture, rather than hiding them in a compound shell command.
- Require the lane summary to report the command, exit state, failed or skipped checks, and relevant output—not merely “tests passed”.
For example, Lane A might run only the parser regression test followed by a type-check, while Lane B runs the cache-expiry unit test and the same build. The shared action is the invocation mechanism; the prompt still identifies the specific test target. The decision rule is to run the smallest test that reproduces the defect first, then add the narrowest broader check justified by the changed surface. A full suite may be appropriate later, but it should not conceal whether the original reproduction was tested.
If an action fails in both untouched worktrees, treat that as possible baseline evidence rather than evidence that both fixes are wrong. Record the command and failure, compare it with the selected starting branch, and pause implementation if the baseline cannot support a meaningful conclusion. A human must verify that the reported command ran in the correct chat-scoped terminal and that skipped tests, warnings and non-zero exits have not been summarised away.
3.4 Decide deliberately whether to use .worktreeinclude
.worktreeinclude is a selective copying mechanism, not a general clone of local ignored or untracked state. According to OpenAI’s worktree documentation, accessed 5 October 2026, matching ignored files can be copied into local Codex-managed worktrees. The documented behaviour does not overwrite existing files and skips symlinks found at source paths. The documentation does not describe this behaviour for remote worktrees or worktrees created manually on the command line.
Begin with no entries. Add a path only after answering all of these questions:
- Is the path ignored by Git and genuinely required to bootstrap each local managed worktree?
- Can its contents safely be duplicated into two worktree directories?
- Does it contain credentials, tokens, personal data, environment-specific endpoints or other secret-bearing values?
- Could a copied value cause both lanes to mutate the same external resource?
- Would generating the file independently be safer and more reproducible?
For example, a reviewed, non-secret local tool configuration might be a candidate if the project cannot generate it and both lanes require exactly the same settings. By contrast, an .env file containing credentials should not be added merely to make setup convenient. Even a credential-free environment file may point both lanes at the same database or service, creating a runtime collision despite separate checkouts.
The decision rule is conservative: prefer independent generation or an explicit per-lane configuration; use .worktreeinclude only for narrowly matched ignored paths whose contents and side effects have been reviewed. Never use a broad pattern to sweep an ignored directory into every worktree. Do not place untrusted files or secrets into the Codex prompt as a substitute for proper provisioning.
After editing the file, have a human inspect each pattern and the source path it matches. Create one worktree first, verify which files appeared, confirm that no symlink or secret-bearing path was expected to transfer, and check that setup still behaves correctly. If a required file is absent, do not infer that every ignored file should have been copied; either correct the precise pattern or choose an explicit generation step.
3.5 Freeze the environment contract before opening two chats
The final preparation gate separates environmental readiness from implementation readiness. Record the selected local environment, setup script, expected setup writes, Test action, Build action and any approved .worktreeinclude entries. This short contract prevents one lane from quietly changing the shared bootstrap while the other is already investigating.
For example, the contract might state: “Both lanes install from the committed lockfile; setup must not alter tracked files; no server starts automatically; focused tests run per lane; the build action is manual; no ignored secret files are copied.” This is an example of a review checklist, not proof that the repository is safe to parallelise.
If either fix later requires a shared lockfile update, database migration, generated tracked artefact or common external service mutation, stop and sequence the work. The trade-off is lost concurrency in exchange for attributable state and a reviewable diff. A human should approve the frozen contract and re-run git status --short in Local before proceeding, because choosing a base branch with unstaged changes can apply those changes to a newly created worktree.
4. Launch the two bounded lanes
4.1 Create Lane A from an explicit starting branch
Open a new chat in Worktree mode rather than Local mode. OpenAI’s environment-modes documentation, accessed 5 October 2026, describes Worktree mode as isolating changes in a Git worktree. Select the verified starting branch explicitly and attach the prepared local environment. Do not rely on whichever branch happens to be visible in another window.
Before confirming creation, inspect Local with:
git status --short
git branch --show-current
git rev-parse HEAD
Compare the output with the preflight record. If there are new unstaged changes, stop: official guidance says selecting a branch with unstaged local changes applies those changes to the new worktree. Classify, commit, stash or deliberately remove them according to project policy before continuing. The worktree is not a way to escape a dirty starting checkout.
A Codex-managed worktree starts in detached HEAD by default. That is appropriate during investigation because it avoids occupying a persistent branch prematurely. Use a named branch only when the change is ready to persist or publish. Git permits a named branch to be checked out in only one worktree at a time: a branch cannot be checked out in both Local and its worktree.
For Lane A, select the agreed base—for example, the project’s current development branch—without inventing a bug branch during initial diagnosis. The decision rule is simple: remain detached while the root cause or viability is uncertain; create a descriptive named branch only when there is reviewed work worth retaining. If the wrong base was selected, abandon that lane before meaningful edits and recreate it from the correct base rather than layering compensating changes onto the wrong history.
A human should verify inside Lane A’s terminal that git rev-parse HEAD matches the recorded base commit and that git status --short contains no unexpected inherited edits before giving Codex the implementation prompt.
4.2 Create Lane B independently, not as a continuation
Open a second new Worktree chat and repeat the same base and environment selection. Do not continue Lane A’s chat, duplicate its conversational history or ask one chat to coordinate both fixes. OpenAI’s best-practices guidance, accessed 5 October 2026, says to keep one chat per coherent unit of work. Separate chats preserve distinct goals, evidence and stop conditions.
Lane B may start from the same commit as Lane A because each begins as a separate checkout in detached state. This does not authorise both worktrees to occupy the same named branch later. Reserve different branch names for durable results, such as an illustrative fix/parser-empty-token for Lane A and fix/cache-expiry-boundary for Lane B, subject to the repository’s naming policy.
After creation, compare these values in both terminals:
git rev-parse HEAD
git status --short
git worktree list
The commit should match the intended starting point, while the paths should identify different checkouts. Treat this as identity verification, not a functional test. If the status output differs unexpectedly, pause both lanes and determine whether local changes were inherited, setup generated tracked files, or one environment step behaved non-deterministically.
The decision rule is that both lanes may proceed only when their initial tracked state is explainable and the proposed fixes remain independent. If setup reveals that both must edit the same generated artefact or lockfile, close the parallel path and sequence them. A human should compare the initial status and base commit for both chats before allowing either to modify source files.
4.3 Give Lane A a bounded bug-fix card
A bounded prompt names the goal, relevant evidence, exclusions and an observable finish condition. It should not include secrets, private credentials or untrusted issue text that has not been reviewed. Give references to repository paths and sanitised diagnostics instead of pasting uncontrolled data wholesale.
An illustrative Lane A prompt is:
Goal: fix the single reproducible defect where an empty token is handled incorrectly.
Context: the focused failing test is [repository test path]; inspect [relevant parser paths].
Use only the sanitised reproduction supplied in the repository.
Constraints:
- make the smallest behaviour-preserving diff;
- do not refactor adjacent parser code;
- do not edit the lockfile, generated artefacts, migrations or Lane B's cache module;
- do not use network access;
- stop and report if the root cause requires a shared schema, generated file or excluded module.
Done when:
- the focused reproduction test passes;
- the relevant type-check or build action completes;
- report changed files, root cause, exact commands and material output;
- report failed, skipped or unavailable checks and remaining uncertainty;
- state whether the result appears ready for human review.
This is an example prompt, not a promise of a correct fix. Replace placeholders with verified repository facts. “Smallest diff” means the least change needed to satisfy the named behaviour while retaining necessary tests; it does not mean suppressing an essential validation change.
If Codex reports that the defect actually originates in a shared migration, lockfile or the module assigned to Lane B, enforce the stop condition. Do not broaden the prompt mid-flight into a refactor. The human decision is either to reclassify the investigations as dependent and sequence them, or to redefine one lane after confirming that the change surfaces do not overlap.
Before accepting Lane A’s claimed completion, the developer should read its changed-file list, inspect git diff, rerun the focused reproduction in that chat’s terminal and check that the output corresponds to the requested behaviour. A successful command does not by itself accept the lane.
4.4 Give Lane B its own constraints and test evidence
Lane B should use the same prompt structure but different defect evidence, relevant paths and forbidden overlap. Do not refer vaguely to “the other bug”; name Lane B’s own expected behaviour so its conclusion remains independently reviewable.
An illustrative card is:
Goal: fix the single cache-expiry boundary defect described by the focused repository test.
Context: begin with [cache test path] and [cache implementation path].
The expected boundary behaviour is [reviewed requirement].
Constraints:
- minimal diff; no opportunistic cleanup or API redesign;
- do not edit parser files, shared migrations, generated artefacts or the lockfile;
- do not start or mutate an external cache service;
- stop and report if the reproduction depends on shared persistent state.
Done when:
- the focused expiry test demonstrates the expected boundary behaviour;
- the agreed build or type-check runs;
- report the root cause, changed files, commands, material output and uncertainty;
- identify every skipped or unavailable check;
- state whether the result appears ready for human review.
The trade-off is between speed and evidential separation. Reusing Lane A’s conclusions could be faster, but it risks importing assumptions and expanding both chats’ scope. Keep Lane B anchored to its own reproduction. Shared coding conventions may come from repository instructions, but defect-specific reasoning should remain in the relevant chat.
If the test requires a live external cache with shared state, stop rather than treating checkout isolation as service isolation. Either provide an approved per-lane disposable instance or sequence the runtime validation. A human must inspect the service requirement, permission request and cleanup implications before allowing access.
4.5 Enforce a no-opportunistic-refactor checkpoint
After each lane proposes a plan but before broad edits, ask it to list the expected files and explain why each is required. Compare that list with the allowed change surface in the bug card. A test file plus one implementation file may be plausible for a bounded defect; a formatting sweep, module rename or dependency upgrade is a scope warning, not automatic evidence of thoroughness.
Use this checkpoint procedure in each chat:
- Request the suspected root cause and planned file list.
- Reject unrelated cleanup, renaming, abstraction work and style-only changes.
- Require a stop-and-report response for any shared-risk area named in the prompt.
- After implementation, compare the actual changed files with the proposed list.
- Run
git diff --statand then inspect the full diff rather than judging only by file count.
For example, if Lane A discovers duplicated parser logic, it may note that technical debt without consolidating it. The current job is the reproducible empty-token defect, not parser architecture. Refactoring can become a separately planned task after the fix is accepted.
The decision rule is to permit an extra file only when it is necessary to test or implement the named behaviour and the lane explains that necessity. If accepting it creates overlap with Lane B, pause both and sequence the work. Human review is required because neither a concise plan nor a green test proves that unrelated behaviour was preserved.
4.6 Establish the evidence each lane must return
Require both chats to produce comparable evidence without asking them to approve their own work. Each lane should return its base commit, current status, changed-file list, root-cause explanation, exact validation commands, material outputs, skipped checks and remaining uncertainty. Sample output headings could be “Base”, “Files changed”, “Reproduction”, “Broader checks” and “Open risks”; these are examples, not fixed interface labels.
Where useful, run /review against the appropriate base or uncommitted changes. OpenAI’s code-review documentation, accessed 5 October 2026, says local review reports findings without changing the working tree and advises checking the referenced diff and tests against expected behaviour. Its scope can include all repository changes, not only lines authored by Codex, so first confirm that the review is looking at the intended lane and diff.
Before merging, compare the diffs from each worktree and confirm that every bug fix stays inside its intended scope. For a larger process that pairs tested changes with a human checkpoint, see Tested Code and Human Review, which shows where review fits after code has been written and checked.
If review identifies a possible defect, verify it against the latest code and requirement before editing. If it reports no findings, do not treat that as approval. The developer remains responsible for checking the focused reproduction, inspecting the complete diff and deciding whether the result should later receive a branch, commit or pull request.
The final launch gate is human: confirm that Lane A and Lane B still address different behaviours, have different allowed file surfaces, use the intended base, and have no unexplained initial changes. Only then let both investigations continue concurrently. If either lane crosses a shared-risk boundary, suspend parallel execution rather than assuming worktrees will resolve the resulting conflict.
5. Worktree mechanics that prevent branch mistakes
A Codex-managed worktree is a separate checkout of the repository, with its own files but shared Git metadata. That distinction permits two independent chats to edit different checkouts without either chat writing into the main working directory. It does not create a second repository, remove Git’s branch-occupancy rules, or isolate external services. As of 5 October 2026, OpenAI’s Worktrees documentation states that Git allows a branch to be checked out in only one place at a time.
5.1 Confirm that both lanes started from the intended commit
Before either lane edits files, use its chat-scoped terminal to record the checkout state. Run the following commands separately in Lane A and Lane B:
git status --short --branch
git rev-parse --show-toplevel
git rev-parse HEAD
git branch --show-current
The commands answer different questions. git status --short --branch exposes modifications and the current branch presentation; git rev-parse --show-toplevel confirms which checkout the terminal addresses; git rev-parse HEAD records the exact starting commit; and git branch --show-current reports a named branch only if one is checked out. An empty result from the final command can be expected because a new Codex-managed worktree starts in detached HEAD by default.
For example, suppose both fixes were meant to start from the current main commit. The two lanes should record the same commit identifier even though their top-level directories differ. If one identifier differs, stop before implementation. Check whether the wrong base branch was selected or whether the branch moved between chat creation times. Do not conceal the mismatch by resetting a lane after work has begun; first preserve any useful diff, then recreate or deliberately rebase the affected work under human control.
Decision rule: continue in parallel only when each lane has the intended base commit and no unexplained starting modifications. A different directory is correct; a different base commit is acceptable only when it was explicitly planned. The human reviewer should compare both recorded commit identifiers with the intended base before accepting subsequent test evidence.
5.2 Treat detached work as an investigation state, not a defect
A detached checkout points directly at a commit instead of attaching HEAD to a named branch. This is useful while a bounded investigation may still be abandoned: Lane A can reproduce its parsing fault and Lane B can investigate its unrelated cache-expiry fault without prematurely occupying two branch names. File edits and diffs still work normally. The important trade-off is durability: uncommitted work in a managed checkout should not be treated as the durable record of a finished fix.
Keep the lane detached while the root cause remains uncertain. Inspect its state with:
git status --short
git diff --stat
git diff --name-only
As a scoped example, Lane A may be permitted to edit src/parser/token.ts and its focused test. If git diff --name-only also lists a generated client, a lockfile or a migration, pause the lane. Ask why that file changed and determine whether it creates a shared change surface with Lane B. Revert only after inspecting the diff; an unexpected file may contain relevant evidence, while an automatic blanket reset could destroy it.
Decision rule: remain detached while the lane is exploratory, disposable or awaiting proof of the root cause. Create a named branch only when the change is sufficiently coherent to retain, commit, publish or hand to another developer. Before that transition, a human should inspect the file list and verify that the lane still represents one bounded defect rather than an opportunistic refactor.
5.3 Use “Create branch here” only for work worth retaining
When a lane has a minimal, reviewable change, use the documented Create branch here flow in that worktree rather than improvising branch movement from another checkout. Choose a descriptive name that follows the repository’s convention, such as fix/parser-empty-token for Lane A or fix/cache-expiry-boundary for Lane B. These names are illustrative, not Codex-generated guarantees.
Immediately verify the result in that lane:
git status --short --branch
git branch --show-current
git worktree list
The first two commands should identify the new branch and any remaining modifications. git worktree list provides the broader occupancy view: it helps the reviewer see which checkout owns each branch. If the expected name does not appear, or the branch is shown against a different path, do not commit until the discrepancy is understood. Check for a naming collision, an earlier branch with the same name, or a command run in the wrong terminal.
Creating a branch too early increases coordination overhead because that branch becomes occupied by the worktree. Creating it too late leaves a completed change dependent on a managed checkout’s continued retention. Decision rule: branch after the lane has a plausible minimal fix and reproducible evidence, but before treating the work as durable or ready to publish. The human should confirm the branch points to the expected base and that its diff contains only the intended fix.
5.4 Respect the single-checkout rule in Local
Once Lane A owns fix/parser-empty-token, that branch cannot be checked out in both the managed worktree and Local. Attempting to force the same branch into the main checkout is not a parallel-work technique; it violates the occupancy model that protects the two paths from ambiguous branch state.
Before changing anything in Local, run:
git status --short --branch
git worktree list
If Local is on main and the Lane A worktree owns fix/parser-empty-token, leave that ownership intact unless Lane A is the one fix that must move into the foreground. Do not use force options, manually remove worktree metadata, or detach checkouts merely to silence an occupancy error. Such actions can obscure where the authoritative working files are.
For example, suppose Lane A needs the developer’s usual IDE and existing development server, while Lane B can finish entirely in its scoped terminal. Select Lane A for Handoff and leave Lane B in its worktree. The practical trade-off is that only one lane receives the foreground Local environment; choosing both would recreate the branch and runtime contention the workflow is meant to control.
Decision rule: if a branch is already attached to a worktree, either continue there or use Handoff for that lane. Never try to check out the same named branch manually in Local. After the transition, the human should run git worktree list and git status --short --branch again to verify that branch ownership and the visible diff match the intended destination.
5.5 Understand what Handoff changes—and what it does not
Handoff is the documented Codex flow for moving a chat and its Git state safely between Local and its worktree. It is appropriate when one completed lane needs the normal editor, an already configured local server or direct collaboration. It is not a merge, rebase or conflict-resolution facility, and it must not be used as a way to duplicate one checked-out branch.
- Choose the single lane that needs to become foreground work.
- Stop commands still running in that lane, especially watchers or development servers.
- Inspect Local with
git status --short --branch; resolve or deliberately preserve unrelated Local changes before proceeding. - In the selected chat, invoke Handoff to Local using the documented desktop flow.
- After completion, confirm the destination path, branch, commit and modified-file list.
- Leave the other lane in its separate worktree and continue validating it there.
A sample verification sequence after handing Lane A to Local is:
git rev-parse --show-toplevel
git status --short --branch
git rev-parse HEAD
git diff --name-only
git worktree list
If Local had uncommitted edits, if Handoff reports that it cannot complete, or if the post-handoff file list is not Lane A’s expected list, stop. Do not resolve the situation by deleting a checkout or forcing branch operations. Preserve the displayed state, identify which path owns the work and compare the diff with the lane’s pre-handoff record. The human must verify that the chat, files and branch moved as a coherent lane before editing resumes.
Decision rule: use Handoff only when foreground Local tooling is materially necessary. Otherwise, keeping a finished lane in its worktree until review avoids needless branch movement. The cost of leaving it isolated is an extra checkout; the cost of an unnecessary handoff is more state to verify and a greater chance of confusing Local changes with lane changes.
5.6 Stop parallel work when isolation ends at the filesystem boundary
Worktrees separate tracked checkout files, but two commands can still contend for a port, database, cache, container name, credential, test account or remote service. They can also generate changes to the same lockfile or generated output independently. OpenAI makes a narrower separation claim for scheduled tasks—its Automations documentation says worktrees keep scheduled-task changes separate from unfinished local work—but that filesystem separation does not establish runtime independence. This guide’s two supervised lanes require the same caution.
Before running either fix, list its external effects: ports, writable data stores, queues, buckets, browser sessions and service accounts. For example, if both focused tests reset the same development database, do not run them concurrently merely because their source files differ. Give each lane an explicitly isolated resource where the project supports that safely, or sequence the tests.
If investigation reveals that both fixes touch a shared module, migration, generated artefact, lockfile or externally stateful service, pause one lane. Record its current diff and test status, then finish and integrate the other first. Decision rule: parallelism is justified only while both the code change surfaces and consequential runtime effects remain independent. A human—not an automated green status—must approve any claim that separate runtime resources are safe to use.
Repository-level instructions can make these stop conditions explicit, including forbidden paths, approved checks and commands that require human approval. Write those boundaries without placing credentials or untrusted content in agent instructions.
These repository-instruction boundaries remain this article’s own control; for a separate multi-agent control over files and handoff material moving between isolated lanes, see Codex Artifact and Multi-Agent Isolation Playbook, which treats files, logs, summaries and handoff notes as controlled objects requiring separate authorisation, rather than guidance for writing instructions or storing credentials.
6. Validate independently in each chat-scoped terminal
Each lane needs its own evidence trail. As of 5 October 2026, OpenAI’s Integrated terminal documentation says: “Each chat in the ChatGPT desktop app includes a terminal scoped to its current project or worktree.” Scope helps prevent running Lane A’s Git commands in Lane B’s checkout, but it does not prove that a command is correct, complete or successful.
6.1 Establish terminal identity before every consequential command
At the start of a validation session—and again after Handoff—print enough context to identify the lane:
pwd
git rev-parse --show-toplevel
git status --short --branch
git rev-parse HEAD
Label the captured output as Lane A or Lane B in the chat. Avoid pasting secrets, access tokens, private customer data or untrusted issue content into prompts. If a log contains sensitive or attacker-controlled material, retain it outside the prompt and provide only the minimal, sanitised excerpt needed to explain the defect.
For example, Lane B’s expected top-level path should be its own worktree until Handoff. If the path is the Local checkout or Lane A’s directory, do not run a formatter, package installer or test. Close the mistaken terminal context and return to the correct chat. The human reviewer should compare the reported path and branch with git worktree list before accepting any command output as lane-specific evidence.
Decision rule: no consequential command runs until checkout identity is explicit. Repeating four inexpensive identity commands is preferable to attributing a valid test result to the wrong diff.
6.2 Reproduce the defect before changing code
Run the smallest deterministic reproduction named in the lane’s “done when” condition. A focused automated test is preferable where one exists; otherwise use a bounded command that exposes the incorrect behaviour without modifying shared state. Record the command, relevant output and exit status. Do not describe a defect as reproduced merely because a code path looks suspicious.
An illustrative Lane A procedure might be:
npm test -- path/to/parser-empty-token.test.ts --runInBand
printf 'exit=%s\n' "$?"
The command is only an example; use the repository’s documented package manager and test syntax. If the test suite requires network access, a credential or a shared database, stop and assess that requirement rather than broadening permissions automatically. OpenAI’s Agent approvals and security guidance says local network access is off by default and that monitoring does not replace sandboxing, permissions or review of the result.
If reproduction fails for an environmental reason—such as a missing dependency, unavailable fixture or occupied port—classify it as could not run, not as a passing test or a disproved bug. Capture the error without exposing secrets, check whether the approved setup action was completed, and request only the narrow permission or resource genuinely required.
Decision rule: edit only after observing the defect or documenting why faithful reproduction is unavailable. The trade-off is time versus confidence: a quick speculative patch may be smaller initially, but it gives the reviewer no reliable before-and-after comparison. A human should verify that the reproduction actually represents the reported user-visible fault.
6.3 Make the minimal fix and audit the changed-file boundary
Ask each lane to address its single root cause without refactoring neighbouring code. After the edit, inspect both summary and content:
git status --short
git diff --stat
git diff --check
git diff -- path/to/expected-file path/to/expected-test
git diff --check can expose whitespace errors, but a clean result does not establish correctness. Read the actual patch. Compare every changed file with the lane’s permitted scope. For Lane A, a parser implementation and one focused test may be expected; a lockfile, migration or broad formatting change is a reason to pause. For Lane B, the cache-expiry module and its test may be acceptable, while edits to Lane A’s parser indicate that the independence assumption has failed.
If an unexpected generated file appears, determine whether the approved build produced it and whether the repository requires it to be committed. Do not automatically retain or discard it. If both lanes need the same generated artefact, stop parallel work and sequence generation after one fix is integrated.
Decision rule: accept the implementation boundary only when every changed hunk is necessary for the stated defect or its focused regression test. Smaller is not automatically better if it omits required behaviour, but unrelated cleanup belongs in another task. The human reviewer should read the complete patch rather than relying on the changed-file summary.
6.4 Re-run the focused check, then add proportionate coverage
First run the exact reproduction command again. Only after it passes should the lane run the smallest relevant broader check, such as the containing test file, package-level type check or targeted linter. Do not substitute a successful build for a failing behavioural test: compilation, linting and reproduction answer different questions.
A sample evidence sequence is:
# Exact regression check
npm test -- path/to/parser-empty-token.test.ts --runInBand
# Relevant static check, if defined by the repository
npm run typecheck -- --project path/to/package
These are illustrative commands, not guaranteed interfaces. Report each command separately as passed, failed or could not run. Include the relevant failure text and exit status rather than smoothing mixed results into “validation completed”. If the focused test passes but the type check fails in an untouched module, record that distinction and ask a human whether the failure is pre-existing; do not assert that it is unrelated without comparing against the base.
Decision rule: a lane is a handoff candidate when the original reproduction now passes, the relevant broader checks have passed or have a clearly documented limitation, and no unexpected files remain. The trade-off is breadth against signal: running an enormous flaky suite can obscure a bounded result, while running only one assertion may miss an immediate integration break. The human chooses the proportionate breadth based on repository risk.
6.5 Capture evidence in a reviewable lane report
Require each chat to finish with a structured report, not a generic success statement. A suitable example format is:
Lane:
Base commit:
Current branch or detached state:
Defect reproduced with:
Root cause:
Changed files:
Focused check:
Broader check:
Checks not run and reason:
External resources used:
Remaining uncertainty:
Safe to consider for Local handoff: yes/no, with reason
This is a reporting example, not a product-generated guarantee. The report must correspond to terminal output and the current diff. If the lane changed after its tests ran, invalidate the old evidence and rerun the relevant checks. If Handoff occurred, repeat checkout identity and inspect the Local diff before relying on the pre-handoff report.
For additional review, the local /review flow can examine uncommitted changes or compare against a base branch without changing the working tree. However, its scope can include all repository changes, including developer edits. OpenAI’s Code review guidance therefore instructs reviewers to check the referenced diff and tests against expected behaviour. Review findings are evidence to verify, not approval, a merge action or acceptance of the fix.
If the review mentions a file outside the lane’s expected list, first determine whether the wrong repository state was reviewed. Do not ask the review to approve the lane or assume every finding concerns Codex-authored lines. Decision rule: retain only findings that a human can reproduce against the latest intended diff. The final commit or pull-request decision remains with the developer after inspecting the patch, confirming the reported commands and resolving material uncertainty.
6.6 Compare the two evidence trails before allowing either lane forward
Place the reports side by side and compare base commit, changed files, checks and external resources. The lanes remain independent only if their evidence supports that conclusion. A shared test command is not itself a collision, but a shared mutable database, lockfile rewrite or generated output is. Likewise, two clean diffs can still be unsafe if both tests changed the same external account.
If both lanes are valid but only one requires Local tooling, hand off that one and leave the other isolated. If both require the same server or stateful fixture, sequence them. If either report omits a failed or unavailable check, return it to the lane for correction rather than filling the gap by inference.
Final control rule for this stage: proceed only when a human can identify which checkout produced each diff, which exact commands support it, what did not run, and why the two fixes remain independent. Successful setup, tests and automated review are supporting evidence; none replaces inspection of the correct diff or the human decision to commit and open a pull request.
7. Bring one lane forward safely
Keep one fix in its isolated worktree while bringing only the lane that needs foreground tools into Local. Typical reasons include inspecting the change in the usual IDE, using an existing development server, or collaborating through tooling already attached to the main project directory. This is a change of working context, not acceptance of the fix: the human reviewer still needs to verify the diff, reproduction and tests.
7.1 Choose the lane that genuinely needs Local
Do not hand off both fixes merely because their investigations are complete. Select one foreground lane by asking which change needs a facility tied to the Local checkout. Lane A might require browser testing against an existing development server, while Lane B can remain in its chat-scoped terminal for a focused unit test. OpenAI’s Integrated terminal documentation, accessed 5 October 2026, states that each chat in the ChatGPT desktop app has a terminal scoped to its current project or worktree. That makes remaining in the worktree a valid choice when no Local-only facility is required.
- Read each lane’s changed-file summary and validation report.
- Confirm that neither lane has expanded into a shared module, migration, generated artefact, lockfile or externally stateful service.
- Choose the one lane that needs foreground inspection.
- Pause commands in that lane, including watchers, builds and test processes that might continue writing files.
- Leave the other lane in its worktree and avoid changing its Git state while the handoff is in progress.
Example: Lane A changes the parser’s empty-token path and needs an existing development server for an authorised end-to-end input check; Lane B changes the separate cache-expiry module and can use its focused unit test within the worktree. Hand off Lane A and retain Lane B in its worktree. This is only an illustrative decision: terminal scope, test commands and server behaviour depend on the repository.
Decision rule: use Handoff only when Local provides a concrete inspection or collaboration advantage. Otherwise, review and retain the fix in its isolated worktree.
If either lane now touches a shared-risk area, do not hand off and continue both concurrently. Stop one lane, inspect both diffs and sequence the work. A human must confirm that the selected lane is still bounded and that no active process will alter its files during transition.
7.2 Check Local and worktree Git state before Handoff
A Codex-managed worktree begins in detached HEAD by default. A named branch becomes appropriate when the change is worth retaining, sharing or publishing, but Git allows a branch to be checked out in only one worktree at a time. In practical terms, the branch cannot be checked out in both the managed worktree and Local. OpenAI’s Worktrees documentation, accessed 5 October 2026, presents Handoff as the supported way to move the chat and code safely between those contexts.
Before initiating it, inspect both sides rather than assuming that “Local” means clean:
git status --short
git branch --show-current
git rev-parse --show-toplevel
git rev-parse HEAD
Run the commands in the lane’s scoped terminal, then repeat the relevant checks in Local. Record the repository root, current commit and branch state. If the lane has already been placed on a named branch, note where that branch is currently checked out. If Local contains unrelated modifications, stop and classify them before moving anything.
Example: suppose Lane A is ready to persist as fix/parser-empty-token, while Local is on the project’s development branch. Confirm that fix/parser-empty-token belongs only to Lane A’s worktree and that Local has no edits that would be confused with the incoming fix. Do not manually check out fix/parser-empty-token in Local while it remains occupied by the worktree.
Decision rule: proceed only when the source and destination states are understood and Local has no unexplained changes. If Local is dirty, either preserve those changes through the repository’s normal deliberate workflow or finish reviewing them first.
If Git reports that a branch is already checked out elsewhere, treat that message as a guardrail. Do not bypass it with force options or ad hoc manipulation of worktree metadata. A human should identify the occupying checkout and decide whether to use Handoff, leave the lane where it is, or sequence the work.
7.3 Use Handoff instead of creating a duplicate checkout
Start Handoff from the completed lane in the desktop app and select Local as the destination. Follow the displayed transition rather than attempting to reproduce its Git operations manually. As of 5 October 2026, OpenAI’s current Worktrees documentation describes Handoff as moving a chat and its code between Local and Worktree while handling the required Git operations. OpenAI’s What’s new page also describes chat Handoff as moving a chat and its Git state; remote-host features are not required for this local two-lane procedure.
- Stop any long-running command in the source chat.
- Confirm that the destination is the matching local project, not another clone with a similar directory name.
- Initiate Handoff to Local through the desktop workflow.
- Wait for the transition to complete before issuing Git commands elsewhere.
- Reopen the moved chat and verify its current project context.
- Leave the second lane untouched in its own worktree.
Handoff is not a merge, rebase or conflict-resolution engine. It does not justify checking out the same branch twice, and it does not prove that the incoming change is correct. It changes where the chat and associated Git work are being continued. The distinction matters because conflict resolution combines histories or file changes, whereas Handoff addresses safe context movement and branch occupancy.
Example: after Lane A reaches its bounded “done when” condition, move its chat to Local so the existing development server can be used for manual inspection. Lane B remains isolated and can still be reviewed in its worktree. Do not start a second Local checkout of Lane A’s branch as a substitute.
Decision rule: if the goal is to continue the same chat and code in Local, use Handoff. If the goal is to integrate two independently developed changes, stop: that is a separate Git review and integration decision.
If the destination does not match the expected repository, branch or commit, stop before editing. Capture git status, git branch --show-current and git rev-parse HEAD, then reconcile them with the pre-handoff record. A human must confirm that the intended lane, rather than the still-isolated fix, arrived in Local.
7.4 Re-establish identity and runtime safety in Local
After Handoff, verify both Git identity and runtime identity. Worktrees isolate repository files, but they do not necessarily isolate ports, databases, caches, credentials or external services. A development server still running from the old Local checkout might reload against unexpected files, while two test processes could target the same external state.
git status --short
git branch --show-current
git rev-parse --show-toplevel
git diff --stat
git diff --name-only
Then inspect the project-specific runtime before restarting anything. Check which configuration the server will read, which port it will bind and whether the focused test writes to shared state. Do not place credentials, tokens, production records or other untrusted data into the Codex prompt. If access outside the workspace or local network access is requested, evaluate the command and grant only what the specific verification requires. OpenAI’s Agent approvals and security documentation, accessed 5 October 2026, says network access is off by default and distinguishes sandbox permissions from approval policy.
Example: Lane A’s parser fix may need the existing server, but its reproduction should use a non-sensitive test account and an approved local or test endpoint. Lane B must not simultaneously mutate the same test database merely because its files remain in another checkout.
Decision rule: resume the Local server only after the repository root, changed-file list and target environment all match the intended lane. If verification requires broader network or filesystem access, approve it only when the command, destination and data exposure are understood.
If the server loads stale output, stop it, clear only repository-approved disposable artefacts and run the documented setup or build action. Do not delete caches indiscriminately. A human should reproduce the original defect, exercise the corrected behaviour and inspect relevant logs rather than treating a successful server start as acceptance.
7.5 Keep the retained lane isolated
The lane left behind remains a separate review unit. Do not use its branch for the Local handoff, copy its files into Lane A, or run a repository-wide formatter that rewrites both change surfaces. If the retained lane discovers an overlap—such as both fixes needing the same shared lockfile—pause it until the foreground fix has been reviewed and delivered.
Example: Lane B reports that its fix would regenerate the same client artefact now touched by Lane A. That discovery invalidates the parallel assumption. Pause Lane B, record its current diff and status, and sequence it after Lane A’s fix has been delivered and the base branch updated, rather than presenting the two diffs as independent.
Decision rule: maintain parallelism only while changed files and external effects remain separable.
Stop Lane B’s commands, record its status and do not discard its work merely to simplify Lane A. A human must compare both name-only diffs and decide whether Lane B can continue from its original base or must later be reconciled with the delivered Lane A change.
8. Human review, delivery and cleanup
Review the foreground lane as a single proposed bug fix. A setup script, passing command or Codex review is evidence, not acceptance. Delivery requires a human to select the correct comparison, inspect the code and tests, decide what to stage, and determine whether the change is suitable for a commit and pull request (PR)A proposed set of repository changes submitted for review before integration. Open glossary entry.
8.1 Select the review scope before invoking /review
Choose between two materially different scopes: review uncommitted changes when assessing the current working tree, or compare against the intended base branch when assessing everything the proposed branch would deliver. OpenAI’s Code review documentation, accessed 5 October 2026, says local /review can use either scope and reports findings without changing the working tree. Its scope can include developer edits and other repository changes, not only lines produced by Codex.
- Run
git status --shortand identify staged, unstaged and untracked entries. - Use
git difffor unstaged content andgit diff --cachedfor staged content. - Use a base comparison when the question is “What would this branch deliver?”
- Use an uncommitted review when the question is “Are these current edits ready to stage?”
- Invoke
/reviewwith that scope and retain the findings as review leads, not verdicts.
Example: if Lane A contains only unstaged implementation and test changes, review uncommitted changes first. After deliberate staging, inspect the staged diff separately. If the branch also contains an earlier commit, compare the complete branch with its intended base so that the earlier commit is not omitted.
Decision rule: select the scope that matches the impending action. Use staged content for a commit decision and a base-branch comparison for a PR decision.
If the review includes unrelated files, stop and find their source instead of asking the reviewer to ignore them. The human should verify every finding against the latest diff, expected behaviour and actual test evidence. /review does not approve, merge or automatically post a PR review.
8.2 Resolve findings with line comments and direct inspection
Use inline or line-level comments to tie a concern to the exact code that prompted it. This differs from a general chat request: a line comment preserves location and makes it easier to determine whether the issue belongs to the bounded fix. Ask for an explanation or a minimal correction, but inspect the resulting patch rather than assuming the comment was resolved correctly.
Example: on a newly added boundary condition, comment: “Example review request: explain whether an empty value reaches this branch, and change only this condition and its focused test if the case is unhandled.” This is a suggested review method, not a guarantee that a change is necessary or sufficient.
Decision rule: act on a finding only when it is reproducible, relevant to the stated defect and proportionate to the fix. Defer unrelated improvements rather than expanding the lane into refactoring.
If a finding contradicts repository behaviour or an existing test, inspect the referenced code and rerun the smallest discriminating check. Record a finding as rejected or deferred with a reason. A human must decide whether the code, the test or the review observation reflects the intended behaviour.
8.3 Stage and revert at the smallest defensible boundary
The review pane can support stage, unstage and revert operations at whole-diff, file and hunk levels, according to OpenAI’s Code review documentation as accessed 5 October 2026. Choose the smallest level that preserves a coherent change. Staging a whole file is convenient but can include formatting or debug edits; hunk staging is more precise but can accidentally separate code from its required test.
git diff --check
git diff
git add --patch
git diff --cached
git status --short
Example: a file contains the intended validation change plus a temporary log statement. Stage the validation hunk and its focused test, leave the log unstaged, then revert the log only after confirming it has no diagnostic value that must be retained. Never use a broad revert command without first reading the affected paths.
Decision rule: stage only content required to explain and verify the bounded fix. Revert content only when its purpose is understood and it is not part of another person’s or lane’s work.
If hunk staging produces an incomplete or invalid state, unstage and reconstruct the patch by file or by a safer set of related hunks. Run the focused verification against the staged intent where the project permits, and have a human read the full cached diff before committing.
8.4 Re-run the acceptance evidence in the delivered context
Validation performed before Handoff should be repeated where context can affect the result. Run the original reproduction, the focused regression test and only the proportionate lint, type-check or build commands defined for the lane. Report skipped, unavailable or failed checks explicitly. Do not convert an unavailable dependency or denied network request into a passing result.
Example: an illustrative sequence might run the single regression test, then a package-level type-check. If the repository’s build requires network access that is not justified for this review, record the build as not run and explain why; do not relax permissions merely to produce a green summary.
Decision rule: commit only when the defect’s acceptance condition has been observed in the relevant context and remaining uncertainty is acceptable to the human reviewer.
If Local produces a different result from the worktree, investigate configuration, generated files, runtime state and commit identity before changing the implementation. The human should verify command output and the user-visible behaviour independently of Codex’s summary.
For a separate, fictional documentation-repair example that uses source checks, a bounded local diff and human sign-off, see Repair One Authentication Troubleshooting Documentation Error; it illustrates review discipline, not this guide’s two-worktree pull-request workflow.
8.5 Commit, push and open the PR deliberately
Create or retain a named branch only when the change is ready to persist or publish. OpenAI’s Developer settings documentation, accessed 5 October 2026, says Git settings can standardise branch naming and control whether Codex uses force pushes. Repository policy should determine the convention; do not assume that a generated name or force-push choice is safe everywhere.
- Confirm the current branch and staged diff.
- Write a commit message describing the defect and correction, not the use of Codex.
- Commit only the reviewed staged content.
- Push to the intended remote under the repository’s normal policy.
- Open a PR whose description states the reproduction, root cause, changed boundary, verification and remaining uncertainty.
- Inspect the remote PR diff to confirm that it matches the local delivery decision.
Example: a suitable PR description might state that an empty token bypassed one parser branch, identify the focused regression test and disclose that a broader integration suite was not run. This sample structure is illustrative; use the project’s required template and do not claim checks that were not performed.
Decision rule: push only after the cached diff and resulting commit have been inspected. Use force push only where repository policy and the specific branch history permit it.
If the remote diff contains unexpected commits or files, do not merge. Recheck the base branch, commit range and remote target. A human reviewer must make the commit, push and PR decisions and must review consequential behaviour before release.
8.6 Retain durable work before cleanup
Cleanup manages disk use; it is not version control. OpenAI’s Worktrees documentation, accessed 5 October 2026, says managed worktrees retain the most recent 15 by default, allows that limit to be configured or automatic deletion disabled, and documents a snapshot before deletion with restoration offered through the associated chat. Treat that snapshot as recovery support, not as the durable home of an accepted fix.
Before archiving or allowing cleanup, confirm that valuable work exists in a named branch, commit or PR. For the retained second lane, decide whether it is ready to persist, should remain an active worktree, or should be abandoned after deliberate inspection. Do not delete it merely because Lane A has shipped.
Example: Lane A has a reviewed commit and PR, so its managed worktree no longer needs to serve as the sole copy. Lane B remains paused, with its diff and status recorded, pending sequencing against Lane A; retain its worktree until a human deliberately preserves its work in an approved branch, commit or patch, or discards it after review.
Decision rule: clean a managed worktree only when it contains no sole copy of work worth keeping. Raise retention or disable automatic deletion only when the additional storage is justified; do not treat retention as a substitute for commits.
If cleanup removes a worktree whose chat still represents needed work, use only the documented snapshot restoration path and verify the restored status and diff. Do not assume every ignored file, runtime artefact or external service state will return. A human must inspect restored content before continuing or publishing it.
8.7 Close the two-lane exercise with an explicit ledger
Finish by recording the disposition of both fixes. This prevents the completed Local lane from obscuring an unfinished worktree and distinguishes code delivery from workspace housekeeping.
- Lane A: branch, commit or PR reference; reviewed diff scope; tests actually run; remaining uncertainty; worktree retained or eligible for cleanup.
- Lane B: current commit and branch state; changed paths; tests actually run; whether parallel independence still holds; next human action.
- Local: current branch, clean or intentionally dirty status, and any development server still running.
- Shared resources: ports, databases, generated artefacts or external services that must be reset or left intact.
Example: record Lane A as delivered and eligible for cleanup only after checking the remote diff. Record Lane B as retained and paused if it now depends on Lane A. This ledger is an example process output, not a Codex-generated guarantee.
Decision rule: the exercise is complete only when every checkout and shared runtime has an identified owner and next action.
If the ledger cannot explain an untracked file, active process or branch, do not clean up. Inspect it first. Scheduled-task worktrees are a separate feature: OpenAI’s Automations documentation, accessed 5 October 2026, notes that they can separate scheduled changes from unfinished local work, but unattended automation is not a replacement for this supervised delivery and review procedure.
Access 40,000+ AI Prompts for ChatGPT, Claude & Codex — Free!
Subscribe to get instant access to our complete Notion Prompt Library — the largest curated collection of prompts for ChatGPT, Claude, OpenAI Codex, and other leading AI models. Optimized for real-world workflows across coding, research, content creation, and business.
Useful Links
- OpenAI Codex worktrees documentation
- OpenAI Codex local-environment documentation
- OpenAI Codex Local, Worktree and Cloud environment modes
- OpenAI Codex prompting and workflow best practices
- OpenAI Codex code-review documentation
- OpenAI Codex agent approvals and security guidance
- OpenAI Codex integrated-terminal documentation
- OpenAI Codex automations and scheduled-task documentation
- OpenAI Codex developer settings documentation
- OpenAI Codex release chronology and feature updates
- OpenAI Help Centre guide to ChatGPT Work and Codex
