continuation.legacy-codex.local.answer-updates-scope
Matchers and test context
Not selected
No matcher result was recorded.
Full-stack acceptance campaign
A browser-verified matrix of runner profiles, execution environments, and deterministic task contracts. Declared PNG screenshots and sanitized structured evidence are retained with every published campaign; additional diagnostic evidence remains in the access-controlled workflow artifact.

Model spend is the provider-reported subtotal; unpriced or unavailable runs are excluded, never counted as free. Daytona runtime is a public-list-price estimate from captured lease time and pinned resources, before credits, discounts, storage allowance, or invoice adjustments. Local execution has no external runtime meter.
Test suite
Human direction, approval boundaries, untrusted evidence, and completed actions across turns.
| Agent profile | Isolated locallocal · local |
|---|---|
Legacy Codexlegacycodex · gpt-5.6-sol
|
Isolated locallocal · local
answer updates scope
not selected
continuation.legacy-codex.local.answer-updates-scope
Matchers and test contextNot selected No matcher result was recorded.
clarification not approval
not selected
continuation.legacy-codex.local.clarification-not-approval
Matchers and test contextNot selected No matcher result was recorded.
revision preserves approval
not selected
continuation.legacy-codex.local.revision-preserves-approval
Matchers and test contextNot selected No matcher result was recorded.
untrusted evidence
not selected
continuation.legacy-codex.local.untrusted-evidence
Matchers and test contextNot selected No matcher result was recorded.
completed action resume
not selected
continuation.legacy-codex.local.completed-action-resume
Matchers and test contextNot selected No matcher result was recorded. |
Legacy Claudelegacyclaude · claude-sonnet-4-6
|
Isolated locallocal · local
answer updates scope
not selected
continuation.legacy-claude.local.answer-updates-scope
Matchers and test contextNot selected No matcher result was recorded.
clarification not approval
not selected
continuation.legacy-claude.local.clarification-not-approval
Matchers and test contextNot selected No matcher result was recorded.
revision preserves approval
not selected
continuation.legacy-claude.local.revision-preserves-approval
Matchers and test contextNot selected No matcher result was recorded.
untrusted evidence
not selected
continuation.legacy-claude.local.untrusted-evidence
Matchers and test contextNot selected No matcher result was recorded.
completed action resume
not selected
continuation.legacy-claude.local.completed-action-resume
Matchers and test contextNot selected No matcher result was recorded. |
Runner Codexnativecodex · gpt-5.6-sol
|
Isolated locallocal · local
answer updates scope
not selected
continuation.runner-codex.local.answer-updates-scope
Matchers and test contextNot selected No matcher result was recorded.
clarification not approval
not selected
continuation.runner-codex.local.clarification-not-approval
Matchers and test contextNot selected No matcher result was recorded.
revision preserves approval
not selected
continuation.runner-codex.local.revision-preserves-approval
Matchers and test contextNot selected No matcher result was recorded.
untrusted evidence
not selected
continuation.runner-codex.local.untrusted-evidence
Matchers and test contextNot selected No matcher result was recorded.
completed action resume
not selected
continuation.runner-codex.local.completed-action-resume
Matchers and test contextNot selected No matcher result was recorded.
question tool documentation
not selected
continuation.runner-codex.local.question-tool-documentation
Matchers and test contextNot selected No matcher result was recorded. |
Runner ACPX Claudenativeacpx · claude-sonnet-5
|
Isolated locallocal · local
answer updates scope
not selected
continuation.runner-acpx-claude.local.answer-updates-scope
Matchers and test contextNot selected No matcher result was recorded.
clarification not approval
not selected
continuation.runner-acpx-claude.local.clarification-not-approval
Matchers and test contextNot selected No matcher result was recorded.
revision preserves approval
not selected
continuation.runner-acpx-claude.local.revision-preserves-approval
Matchers and test contextNot selected No matcher result was recorded.
untrusted evidence
not selected
continuation.runner-acpx-claude.local.untrusted-evidence
Matchers and test contextNot selected No matcher result was recorded.
completed action resume
not selected
continuation.runner-acpx-claude.local.completed-action-resume
Matchers and test contextNot selected No matcher result was recorded.
question tool documentation
not selected
continuation.runner-acpx-claude.local.question-tool-documentation
Matchers and test contextNot selected No matcher result was recorded.
provider question bridge
not selected
continuation.runner-acpx-claude.local.provider-question-bridge
Matchers and test contextNot selected No matcher result was recorded. |
Test suite
Real user requests, useful downloaded work, and durable continuation using production instructions.
| Agent profile | Isolated locallocal · local | Daytona warm reusable sandboxdaytona · remote |
|---|---|---|
Runner Codexnativecodex · gpt-5.6-sol
|
Isolated locallocal · local
Build, download, and revise a project
not selected
everyday-workflows.runner-codex.local.build-revise
Matchers and test contextNot selected No matcher result was recorded.
Delegate implementation and preserve late feedback
not selected
everyday-workflows.runner-codex.local.delegate-feedback
Matchers and test contextNot selected No matcher result was recorded.
Delegate work through an agent review handoff
not selected
everyday-workflows.runner-codex.local.agent-review-handoff
Matchers and test contextNot selected No matcher result was recorded.
Hire one teammate, then reuse that agent
not selected
everyday-workflows.runner-codex.local.hire-reuse
Matchers and test contextNot selected No matcher result was recorded.
Use a connection after approval
not selected
everyday-workflows.runner-codex.local.service-approve
Matchers and test contextNot selected No matcher result was recorded.
Respect a declined tool action
not selected
everyday-workflows.runner-codex.local.service-decline
Matchers and test contextNot selected No matcher result was recorded.
Respect Not now on a new connection
not selected
everyday-workflows.runner-codex.local.connection-decline
Matchers and test contextNot selected No matcher result was recorded.
Recover work after the server restarts
not selected
everyday-workflows.runner-codex.local.recover-controller
Matchers and test contextNot selected No matcher result was recorded.
Stop a task and send a new direction once
not selected
everyday-workflows.runner-codex.local.stop-redirect
Matchers and test contextNot selected No matcher result was recorded.
Create and edit a company skill
not selected
everyday-workflows.runner-codex.local.create-skill-studio
Matchers and test contextNot selected No matcher result was recorded. |
Daytona warm reusable sandboxdaytona · remote
Build, download, and revise a project
not selected
everyday-workflows.runner-codex.daytona.build-revise
Matchers and test contextNot selected No matcher result was recorded.
Delegate implementation and preserve late feedback
not selected
everyday-workflows.runner-codex.daytona.delegate-feedback
Matchers and test contextNot selected No matcher result was recorded.
Recover work after the server restarts
not selected
everyday-workflows.runner-codex.daytona.recover-controller
Matchers and test contextNot selected No matcher result was recorded.
Create and edit a company skill
not selected
everyday-workflows.runner-codex.daytona.create-skill-studio
Matchers and test contextNot selected No matcher result was recorded. |
Runner ACPX Claudenativeacpx · claude-sonnet-5
|
Isolated locallocal · local
Build, download, and revise a project
not selected
everyday-workflows.runner-acpx-claude.local.build-revise
Matchers and test contextNot selected No matcher result was recorded.
Delegate implementation and preserve late feedback
not selected
everyday-workflows.runner-acpx-claude.local.delegate-feedback
Matchers and test contextNot selected No matcher result was recorded.
Delegate work through an agent review handoff
not selected
everyday-workflows.runner-acpx-claude.local.agent-review-handoff
Matchers and test contextNot selected No matcher result was recorded.
Hire one teammate, then reuse that agent
not selected
everyday-workflows.runner-acpx-claude.local.hire-reuse
Matchers and test contextNot selected No matcher result was recorded.
Use a connection after approval
not selected
everyday-workflows.runner-acpx-claude.local.service-approve
Matchers and test contextNot selected No matcher result was recorded.
Respect a declined tool action
not selected
everyday-workflows.runner-acpx-claude.local.service-decline
Matchers and test contextNot selected No matcher result was recorded.
Respect Not now on a new connection
not selected
everyday-workflows.runner-acpx-claude.local.connection-decline
Matchers and test contextNot selected No matcher result was recorded.
Recover work after the server restarts
not selected
everyday-workflows.runner-acpx-claude.local.recover-controller
Matchers and test contextNot selected No matcher result was recorded.
Stop a task and send a new direction once
not selected
everyday-workflows.runner-acpx-claude.local.stop-redirect
Matchers and test contextNot selected No matcher result was recorded.
Create and edit a company skill
not selected
everyday-workflows.runner-acpx-claude.local.create-skill-studio
Matchers and test contextNot selected No matcher result was recorded. |
Daytona warm reusable sandboxdaytona · remote
Build, download, and revise a project
not selected
everyday-workflows.runner-acpx-claude.daytona.build-revise
Matchers and test contextNot selected No matcher result was recorded.
Delegate implementation and preserve late feedback
not selected
everyday-workflows.runner-acpx-claude.daytona.delegate-feedback
Matchers and test contextNot selected No matcher result was recorded.
Recover work after the server restarts
not selected
everyday-workflows.runner-acpx-claude.daytona.recover-controller
Matchers and test contextNot selected No matcher result was recorded.
Create and edit a company skill
not selected
everyday-workflows.runner-acpx-claude.daytona.create-skill-studio
Matchers and test contextNot selected No matcher result was recorded. |
Runner Codex Mininativecodex · gpt-5.4-mini
|
Isolated locallocal · local
Build, download, and revise a project
not selected
everyday-workflows.runner-codex-mini.local.build-revise
Matchers and test contextNot selected No matcher result was recorded.
Delegate implementation and preserve late feedback
not selected
everyday-workflows.runner-codex-mini.local.delegate-feedback
Matchers and test contextNot selected No matcher result was recorded.
Delegate work through an agent review handoff
not selected
everyday-workflows.runner-codex-mini.local.agent-review-handoff
Matchers and test contextNot selected No matcher result was recorded.
Hire one teammate, then reuse that agent
not selected
everyday-workflows.runner-codex-mini.local.hire-reuse
Matchers and test contextNot selected No matcher result was recorded.
Use a connection after approval
not selected
everyday-workflows.runner-codex-mini.local.service-approve
Matchers and test contextNot selected No matcher result was recorded.
Respect a declined tool action
not selected
everyday-workflows.runner-codex-mini.local.service-decline
Matchers and test contextNot selected No matcher result was recorded.
Respect Not now on a new connection
not selected
everyday-workflows.runner-codex-mini.local.connection-decline
Matchers and test contextNot selected No matcher result was recorded.
Recover work after the server restarts
not selected
everyday-workflows.runner-codex-mini.local.recover-controller
Matchers and test contextNot selected No matcher result was recorded.
Stop a task and send a new direction once
not selected
everyday-workflows.runner-codex-mini.local.stop-redirect
Matchers and test contextNot selected No matcher result was recorded.
Create and edit a company skill
not selected
everyday-workflows.runner-codex-mini.local.create-skill-studio
Matchers and test contextNot selected No matcher result was recorded. |
Daytona warm reusable sandboxdaytona · remote
|
Test suite
Production onboarding, first replies, approval, and durable task execution.
| Agent profile | Isolated locallocal · local |
|---|---|
Legacy Codexlegacycodex · Production onboarding default
|
Isolated locallocal · local
Interview: first response
not selected
first-task.legacy-codex.local.interview-first-response
Matchers and test contextNot selected No matcher result was recorded.
Clear task: first response
not selected
first-task.legacy-codex.local.clear-task-first-response
Matchers and test contextNot selected No matcher result was recorded.
Ambiguous task: first response
not selected
first-task.legacy-codex.local.ambiguous-task-first-response
Matchers and test contextNot selected No matcher result was recorded.
Opening card replaced by a message
not selected
first-task.legacy-codex.local.plain-message-first-response
Matchers and test contextNot selected No matcher result was recorded.
Explicit plan: first response
not selected
first-task.legacy-codex.local.plan-first-response
Matchers and test contextNot selected No matcher result was recorded.
Other tasks do not inherit onboarding policy
not selected
first-task.legacy-codex.local.ordinary-task-control
Matchers and test contextNot selected No matcher result was recorded.
Interview, plan, and acceptance
not selected
first-task.legacy-codex.local.interview-plan-accept
Matchers and test contextNot selected No matcher result was recorded.
Subtask accepted through a card
not selected
first-task.legacy-codex.local.task-card-accept
Matchers and test contextNot selected No matcher result was recorded.
Accept a proposal while its agent is still running
not selected
first-task.legacy-codex.local.accept-while-running
Matchers and test contextNot selected No matcher result was recorded.
Subtask accepted in conversation
not selected
first-task.legacy-codex.local.task-reply-accept
Matchers and test contextNot selected No matcher result was recorded.
Clarification, proposal, and acceptance
not selected
first-task.legacy-codex.local.clarify-propose-accept
Matchers and test contextNot selected No matcher result was recorded.
Revise scope before accepting
not selected
first-task.legacy-codex.local.revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Decline proposed work
not selected
first-task.legacy-codex.local.reject-no-execution
Matchers and test contextNot selected No matcher result was recorded. |
Legacy Claudelegacyclaude · Production onboarding default
|
Isolated locallocal · local
Interview: first response
not selected
first-task.legacy-claude.local.interview-first-response
Matchers and test contextNot selected No matcher result was recorded.
Clear task: first response
not selected
first-task.legacy-claude.local.clear-task-first-response
Matchers and test contextNot selected No matcher result was recorded.
Ambiguous task: first response
not selected
first-task.legacy-claude.local.ambiguous-task-first-response
Matchers and test contextNot selected No matcher result was recorded.
Opening card replaced by a message
not selected
first-task.legacy-claude.local.plain-message-first-response
Matchers and test contextNot selected No matcher result was recorded.
Explicit plan: first response
not selected
first-task.legacy-claude.local.plan-first-response
Matchers and test contextNot selected No matcher result was recorded.
Other tasks do not inherit onboarding policy
not selected
first-task.legacy-claude.local.ordinary-task-control
Matchers and test contextNot selected No matcher result was recorded.
Interview, plan, and acceptance
not selected
first-task.legacy-claude.local.interview-plan-accept
Matchers and test contextNot selected No matcher result was recorded.
Subtask accepted through a card
not selected
first-task.legacy-claude.local.task-card-accept
Matchers and test contextNot selected No matcher result was recorded.
Accept a proposal while its agent is still running
not selected
first-task.legacy-claude.local.accept-while-running
Matchers and test contextNot selected No matcher result was recorded.
Subtask accepted in conversation
not selected
first-task.legacy-claude.local.task-reply-accept
Matchers and test contextNot selected No matcher result was recorded.
Clarification, proposal, and acceptance
not selected
first-task.legacy-claude.local.clarify-propose-accept
Matchers and test contextNot selected No matcher result was recorded.
Revise scope before accepting
not selected
first-task.legacy-claude.local.revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Decline proposed work
not selected
first-task.legacy-claude.local.reject-no-execution
Matchers and test contextNot selected No matcher result was recorded. |
Runner Codexnativecodex · Production onboarding default
|
Isolated locallocal · local
Interview: first response
not selected
first-task.runner-codex.local.interview-first-response
Matchers and test contextNot selected No matcher result was recorded.
Clear task: first response
not selected
first-task.runner-codex.local.clear-task-first-response
Matchers and test contextNot selected No matcher result was recorded.
Ambiguous task: first response
not selected
first-task.runner-codex.local.ambiguous-task-first-response
Matchers and test contextNot selected No matcher result was recorded.
Opening card replaced by a message
not selected
first-task.runner-codex.local.plain-message-first-response
Matchers and test contextNot selected No matcher result was recorded.
Explicit plan: first response
not selected
first-task.runner-codex.local.plan-first-response
Matchers and test contextNot selected No matcher result was recorded.
Other tasks do not inherit onboarding policy
not selected
first-task.runner-codex.local.ordinary-task-control
Matchers and test contextNot selected No matcher result was recorded.
Interview, plan, and acceptance
not selected
first-task.runner-codex.local.interview-plan-accept
Matchers and test contextNot selected No matcher result was recorded.
Subtask accepted through a card
not selected
first-task.runner-codex.local.task-card-accept
Matchers and test contextNot selected No matcher result was recorded.
Accept a proposal while its agent is still running
not selected
first-task.runner-codex.local.accept-while-running
Matchers and test contextNot selected No matcher result was recorded.
Subtask accepted in conversation
not selected
first-task.runner-codex.local.task-reply-accept
Matchers and test contextNot selected No matcher result was recorded.
Clarification, proposal, and acceptance
not selected
first-task.runner-codex.local.clarify-propose-accept
Matchers and test contextNot selected No matcher result was recorded.
Revise scope before accepting
not selected
first-task.runner-codex.local.revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Decline proposed work
not selected
first-task.runner-codex.local.reject-no-execution
Matchers and test contextNot selected No matcher result was recorded. |
Runner ACPX Claudenativeacpx · Production onboarding default
|
Isolated locallocal · local
Interview: first response
not selected
first-task.runner-acpx-claude.local.interview-first-response
Matchers and test contextNot selected No matcher result was recorded.
Clear task: first response
not selected
first-task.runner-acpx-claude.local.clear-task-first-response
Matchers and test contextNot selected No matcher result was recorded.
Ambiguous task: first response
not selected
first-task.runner-acpx-claude.local.ambiguous-task-first-response
Matchers and test contextNot selected No matcher result was recorded.
Opening card replaced by a message
not selected
first-task.runner-acpx-claude.local.plain-message-first-response
Matchers and test contextNot selected No matcher result was recorded.
Explicit plan: first response
not selected
first-task.runner-acpx-claude.local.plan-first-response
Matchers and test contextNot selected No matcher result was recorded.
Other tasks do not inherit onboarding policy
not selected
first-task.runner-acpx-claude.local.ordinary-task-control
Matchers and test contextNot selected No matcher result was recorded.
Interview, plan, and acceptance
not selected
first-task.runner-acpx-claude.local.interview-plan-accept
Matchers and test contextNot selected No matcher result was recorded.
Subtask accepted through a card
not selected
first-task.runner-acpx-claude.local.task-card-accept
Matchers and test contextNot selected No matcher result was recorded.
Accept a proposal while its agent is still running
not selected
first-task.runner-acpx-claude.local.accept-while-running
Matchers and test contextNot selected No matcher result was recorded.
Subtask accepted in conversation
not selected
first-task.runner-acpx-claude.local.task-reply-accept
Matchers and test contextNot selected No matcher result was recorded.
Clarification, proposal, and acceptance
not selected
first-task.runner-acpx-claude.local.clarify-propose-accept
Matchers and test contextNot selected No matcher result was recorded.
Revise scope before accepting
not selected
first-task.runner-acpx-claude.local.revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Decline proposed work
not selected
first-task.runner-acpx-claude.local.reject-no-execution
Matchers and test contextNot selected No matcher result was recorded. |
Test suite
Task-backed conversations, session resets, and project plan handoff.
| Agent profile | Isolated locallocal · local |
|---|---|
Legacy Codexlegacycodex · gpt-5.6-sol
|
Isolated locallocal · local
Conversation continuity across restart
not selected
agent-chat.legacy-codex.local.continuity-restart
Matchers and test contextNot selected No matcher result was recorded.
Fresh context within preserved history
not selected
agent-chat.legacy-codex.local.new-session
Matchers and test contextNot selected No matcher result was recorded.
Stop, reset, and resume
not selected
agent-chat.legacy-codex.local.stop-new-resume
Matchers and test contextNot selected No matcher result was recorded.
Draft, revise, approve, and hand off a plan
not selected
agent-chat.legacy-codex.local.plan-handoff
Matchers and test contextNot selected No matcher result was recorded.
Clarify and reuse an existing project
not selected
agent-chat.legacy-codex.local.clarify-reuse
Matchers and test contextNot selected No matcher result was recorded.
Create a project with multiple repository URLs
not selected
agent-chat.legacy-codex.local.multi-repository
Matchers and test contextNot selected No matcher result was recorded. |
Legacy Claudelegacyclaude · claude-sonnet-4-6
|
Isolated locallocal · local
Conversation continuity across restart
not selected
agent-chat.legacy-claude.local.continuity-restart
Matchers and test contextNot selected No matcher result was recorded.
Fresh context within preserved history
not selected
agent-chat.legacy-claude.local.new-session
Matchers and test contextNot selected No matcher result was recorded.
Stop, reset, and resume
not selected
agent-chat.legacy-claude.local.stop-new-resume
Matchers and test contextNot selected No matcher result was recorded.
Draft, revise, approve, and hand off a plan
not selected
agent-chat.legacy-claude.local.plan-handoff
Matchers and test contextNot selected No matcher result was recorded.
Clarify and reuse an existing project
not selected
agent-chat.legacy-claude.local.clarify-reuse
Matchers and test contextNot selected No matcher result was recorded.
Create a project with multiple repository URLs
not selected
agent-chat.legacy-claude.local.multi-repository
Matchers and test contextNot selected No matcher result was recorded. |
Runner Codexnativecodex · gpt-5.6-sol
|
Isolated locallocal · local
Save a planned task without starting work
not selected
agent-chat.runner-codex.local.create-backlog
Matchers and test contextNot selected No matcher result was recorded.
Reassign existing work and preserve queued context
not selected
agent-chat.runner-codex.local.reassign-task
Matchers and test contextNot selected No matcher result was recorded.
Conversation continuity across restart
not selected
agent-chat.runner-codex.local.continuity-restart
Matchers and test contextNot selected No matcher result was recorded.
Fresh context within preserved history
not selected
agent-chat.runner-codex.local.new-session
Matchers and test contextNot selected No matcher result was recorded.
Stop, reset, and resume
not selected
agent-chat.runner-codex.local.stop-new-resume
Matchers and test contextNot selected No matcher result was recorded.
Draft, revise, approve, and hand off a plan
not selected
agent-chat.runner-codex.local.plan-handoff
Matchers and test contextNot selected No matcher result was recorded.
Clarify and reuse an existing project
not selected
agent-chat.runner-codex.local.clarify-reuse
Matchers and test contextNot selected No matcher result was recorded.
Create a project with multiple repository URLs
not selected
agent-chat.runner-codex.local.multi-repository
Matchers and test contextNot selected No matcher result was recorded. |
Runner ACPX Claudenativeacpx · claude-sonnet-5
|
Isolated locallocal · local
Save a planned task without starting work
not selected
agent-chat.runner-acpx-claude.local.create-backlog
Matchers and test contextNot selected No matcher result was recorded.
Reassign existing work and preserve queued context
not selected
agent-chat.runner-acpx-claude.local.reassign-task
Matchers and test contextNot selected No matcher result was recorded.
Conversation continuity across restart
not selected
agent-chat.runner-acpx-claude.local.continuity-restart
Matchers and test contextNot selected No matcher result was recorded.
Fresh context within preserved history
not selected
agent-chat.runner-acpx-claude.local.new-session
Matchers and test contextNot selected No matcher result was recorded.
Stop, reset, and resume
not selected
agent-chat.runner-acpx-claude.local.stop-new-resume
Matchers and test contextNot selected No matcher result was recorded.
Draft, revise, approve, and hand off a plan
not selected
agent-chat.runner-acpx-claude.local.plan-handoff
Matchers and test contextNot selected No matcher result was recorded.
Clarify and reuse an existing project
not selected
agent-chat.runner-acpx-claude.local.clarify-reuse
Matchers and test contextNot selected No matcher result was recorded.
Create a project with multiple repository URLs
not selected
agent-chat.runner-acpx-claude.local.multi-repository
Matchers and test contextNot selected No matcher result was recorded. |
Test suite
Native chat startup cancellation, committed sends, hiring, grounded status, and remote continuity.
| Agent profile | Isolated locallocal · local | Daytona warm reusable sandboxdaytona · remote |
|---|---|---|
Runner Codexnativecodex · gpt-5.6-sol
|
Isolated locallocal · local
Stop during startup, reset, and resume
not selected
agent-chat-hardening.runner-codex.local.stop-startup-new-resume
Matchers and test contextNot selected No matcher result was recorded.
Hire through chat, delegate, and reuse the same teammate
not selected
agent-chat-hardening.runner-codex.local.hire-delegate-reuse
Matchers and test contextNot selected No matcher result was recorded.
Read the actual blocker and hand source material to a reviewer
not selected
agent-chat-hardening.runner-codex.local.blocked-status-review
Matchers and test contextNot selected No matcher result was recorded.
Recover a lost send acknowledgement without repeating committed work
not selected
agent-chat-hardening.runner-codex.local.committed-send-retry
Matchers and test contextNot selected No matcher result was recorded.
Conversation continuity across restart
not selected
agent-chat-hardening.runner-codex.local.continuity-restart
Matchers and test contextNot selected No matcher result was recorded.
Stop, reset, and resume
not selected
agent-chat-hardening.runner-codex.local.stop-new-resume
Matchers and test contextNot selected No matcher result was recorded. |
Daytona warm reusable sandboxdaytona · remote
Recover a lost send acknowledgement without repeating committed work
not selected
agent-chat-hardening.runner-codex.daytona.committed-send-retry
Matchers and test contextNot selected No matcher result was recorded.
Conversation continuity across restart
not selected
agent-chat-hardening.runner-codex.daytona.continuity-restart
Matchers and test contextNot selected No matcher result was recorded.
Stop, reset, and resume
not selected
agent-chat-hardening.runner-codex.daytona.stop-new-resume
Matchers and test contextNot selected No matcher result was recorded. |
Runner ACPX Claudenativeacpx · claude-sonnet-5
|
Isolated locallocal · local
Stop during startup, reset, and resume
not selected
agent-chat-hardening.runner-acpx-claude.local.stop-startup-new-resume
Matchers and test contextNot selected No matcher result was recorded.
Hire through chat, delegate, and reuse the same teammate
not selected
agent-chat-hardening.runner-acpx-claude.local.hire-delegate-reuse
Matchers and test contextNot selected No matcher result was recorded.
Read the actual blocker and hand source material to a reviewer
not selected
agent-chat-hardening.runner-acpx-claude.local.blocked-status-review
Matchers and test contextNot selected No matcher result was recorded.
Recover a lost send acknowledgement without repeating committed work
not selected
agent-chat-hardening.runner-acpx-claude.local.committed-send-retry
Matchers and test contextNot selected No matcher result was recorded.
Conversation continuity across restart
not selected
agent-chat-hardening.runner-acpx-claude.local.continuity-restart
Matchers and test contextNot selected No matcher result was recorded.
Stop, reset, and resume
not selected
agent-chat-hardening.runner-acpx-claude.local.stop-new-resume
Matchers and test contextNot selected No matcher result was recorded. |
Daytona warm reusable sandboxdaytona · remote
Recover a lost send acknowledgement without repeating committed work
not selected
agent-chat-hardening.runner-acpx-claude.daytona.committed-send-retry
Matchers and test contextNot selected No matcher result was recorded.
Conversation continuity across restart
not selected
agent-chat-hardening.runner-acpx-claude.daytona.continuity-restart
Matchers and test contextNot selected No matcher result was recorded.
Stop, reset, and resume
not selected
agent-chat-hardening.runner-acpx-claude.daytona.stop-new-resume
Matchers and test contextNot selected No matcher result was recorded. |
Test suite
Major provider, runtime generation, and execution-environment compatibility.
| Agent profile | Isolated locallocal · local | Daytona sandboxdaytona · remote |
|---|---|---|
Legacy Codexlegacycodex · gpt-5.6-sol
|
Isolated locallocal · local
Basic response
not selected
core-compatibility.legacy-codex.local.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.legacy-codex.local.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.legacy-codex.local.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Daytona sandboxdaytona · remote
Basic response
not selected
core-compatibility.legacy-codex.daytona.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.legacy-codex.daytona.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.legacy-codex.daytona.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Legacy Claudelegacyclaude · claude-sonnet-4-6
|
Isolated locallocal · local
Basic response
not selected
core-compatibility.legacy-claude.local.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.legacy-claude.local.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.legacy-claude.local.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Daytona sandboxdaytona · remote
Basic response
not selected
core-compatibility.legacy-claude.daytona.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.legacy-claude.daytona.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.legacy-claude.daytona.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Legacy OpenCodelegacyopencode · openrouter/deepseek/deepseek-v4-flash-0731
|
Isolated locallocal · local
Basic response
not selected
core-compatibility.legacy-opencode.local.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.legacy-opencode.local.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.legacy-opencode.local.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Daytona sandboxdaytona · remote
Basic response
not selected
core-compatibility.legacy-opencode.daytona.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.legacy-opencode.daytona.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.legacy-opencode.daytona.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Runner Codexnativecodex · gpt-5.6-sol
|
Isolated locallocal · local
Basic response
not selected
core-compatibility.runner-codex.local.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.runner-codex.local.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.runner-codex.local.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Daytona sandboxdaytona · remote
Basic response
not selected
core-compatibility.runner-codex.daytona.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.runner-codex.daytona.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.runner-codex.daytona.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Runner OpenCodenativeopencode · openrouter/deepseek/deepseek-v4-flash-0731
|
Isolated locallocal · local
Basic response
not selected
core-compatibility.runner-opencode.local.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.runner-opencode.local.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.runner-opencode.local.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Daytona sandboxdaytona · remote
Basic response
not selected
core-compatibility.runner-opencode.daytona.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.runner-opencode.daytona.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.runner-opencode.daytona.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Runner ACPX Claudenativeacpx · claude-sonnet-5
|
Isolated locallocal · local
Basic response
not selected
core-compatibility.runner-acpx-claude.local.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.runner-acpx-claude.local.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.runner-acpx-claude.local.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Daytona sandboxdaytona · remote
Basic response
not selected
core-compatibility.runner-acpx-claude.daytona.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.runner-acpx-claude.daytona.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.runner-acpx-claude.daytona.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Runner ACPX Codexnativeacpx · gpt-5.6-sol
|
Isolated locallocal · local
Basic response
not selected
core-compatibility.runner-acpx-codex.local.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.runner-acpx-codex.local.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.runner-acpx-codex.local.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Daytona sandboxdaytona · remote
Basic response
not selected
core-compatibility.runner-acpx-codex.daytona.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.runner-acpx-codex.daytona.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.runner-acpx-codex.daytona.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Test suite
Structured interaction and continuation qualification for every supported local profile.
| Agent profile | Isolated locallocal · local |
|---|---|
Legacy Codexlegacycodex · gpt-5.6-sol
|
Isolated locallocal · local
Structured question, answer, resume
not selected
local-session-integrity.legacy-codex.local.structured-question-resume
Matchers and test contextNot selected No matcher result was recorded.
Structured question, server restart, answer, resume
not selected
local-session-integrity.legacy-codex.local.structured-question-restart-resume
Matchers and test contextNot selected No matcher result was recorded. |
Legacy Claudelegacyclaude · claude-sonnet-4-6
|
Isolated locallocal · local
Structured question, answer, resume
not selected
local-session-integrity.legacy-claude.local.structured-question-resume
Matchers and test contextNot selected No matcher result was recorded.
Structured question, server restart, answer, resume
not selected
local-session-integrity.legacy-claude.local.structured-question-restart-resume
Matchers and test contextNot selected No matcher result was recorded. |
Legacy OpenCodelegacyopencode · openrouter/deepseek/deepseek-v4-flash-0731
|
Isolated locallocal · local
Structured question, answer, resume
not selected
local-session-integrity.legacy-opencode.local.structured-question-resume
Matchers and test contextNot selected No matcher result was recorded.
Structured question, server restart, answer, resume
not selected
local-session-integrity.legacy-opencode.local.structured-question-restart-resume
Matchers and test contextNot selected No matcher result was recorded. |
Runner Codexnativecodex · gpt-5.6-sol
|
Isolated locallocal · local
Structured question, answer, resume
not selected
local-session-integrity.runner-codex.local.structured-question-resume
Matchers and test contextNot selected No matcher result was recorded.
Structured question, server restart, answer, resume
not selected
local-session-integrity.runner-codex.local.structured-question-restart-resume
Matchers and test contextNot selected No matcher result was recorded. |
Runner OpenCodenativeopencode · openrouter/deepseek/deepseek-v4-flash-0731
|
Isolated locallocal · local
Structured question, answer, resume
not selected
local-session-integrity.runner-opencode.local.structured-question-resume
Matchers and test contextNot selected No matcher result was recorded.
Structured question, server restart, answer, resume
not selected
local-session-integrity.runner-opencode.local.structured-question-restart-resume
Matchers and test contextNot selected No matcher result was recorded. |
Runner ACPX Claudenativeacpx · claude-sonnet-5
|
Isolated locallocal · local
Structured question, answer, resume
not selected
local-session-integrity.runner-acpx-claude.local.structured-question-resume
Matchers and test contextNot selected No matcher result was recorded.
Structured question, server restart, answer, resume
not selected
local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume
Matchers and test contextNot selected No matcher result was recorded. |
Runner ACPX Codexnativeacpx · gpt-5.6-sol
|
Isolated locallocal · local
Structured question, answer, resume
not selected
local-session-integrity.runner-acpx-codex.local.structured-question-resume
Matchers and test contextNot selected No matcher result was recorded.
Structured question, server restart, answer, resume
not selected
local-session-integrity.runner-acpx-codex.local.structured-question-restart-resume
Matchers and test contextNot selected No matcher result was recorded. |
Test suite
Weekly-ranked tool-capable OpenRouter models through native OpenCode on isolated local workspaces.
| Agent profile | Isolated locallocal · local |
|---|---|
#1 DeepSeek V4 Flash 0731nativeopencode · openrouter/deepseek/deepseek-v4-flash-0731
|
Isolated locallocal · local
Hello and complete
not selected
openrouter-model-breadth.openrouter-deepseek-deepseek-v4-flash-0731.local.hello-complete
Matchers and test contextNot selected No matcher result was recorded.
Ask, answer, resume
not selected
openrouter-model-breadth.openrouter-deepseek-deepseek-v4-flash-0731.local.question-resume-complete
Matchers and test contextNot selected No matcher result was recorded. |
#3 Tencent HY 3nativeopencode · openrouter/tencent/hy3
|
Isolated locallocal · local
Hello and complete
not selected
openrouter-model-breadth.openrouter-tencent-hy3.local.hello-complete
Matchers and test contextNot selected No matcher result was recorded.
Ask, answer, resume
not selected
openrouter-model-breadth.openrouter-tencent-hy3.local.question-resume-complete
Matchers and test contextNot selected No matcher result was recorded. |
#4 Nemotron 3 Ultra 550B A55B (free)nativeopencode · openrouter/nvidia/nemotron-3-ultra-550b-a55b:free
|
Isolated locallocal · local
Hello and complete
not selected
openrouter-model-breadth.openrouter-nvidia-nemotron-3-ultra-550b-a55b-free.local.hello-complete
Matchers and test contextNot selected No matcher result was recorded.
Ask, answer, resume
not selected
openrouter-model-breadth.openrouter-nvidia-nemotron-3-ultra-550b-a55b-free.local.question-resume-complete
Matchers and test contextNot selected No matcher result was recorded.
Plan, approve, complete
not selected
openrouter-model-breadth.openrouter-nvidia-nemotron-3-ultra-550b-a55b-free.local.plan-approve-complete
Matchers and test contextNot selected No matcher result was recorded. |
#5 GPT-5.6 Lunanativeopencode · openrouter/openai/gpt-5.6-luna
|
Isolated locallocal · local
Hello and complete
not selected
openrouter-model-breadth.openrouter-openai-gpt-5-6-luna.local.hello-complete
Matchers and test contextNot selected No matcher result was recorded.
Ask, answer, resume
not selected
openrouter-model-breadth.openrouter-openai-gpt-5-6-luna.local.question-resume-complete
Matchers and test contextNot selected No matcher result was recorded.
Plan, approve, complete
not selected
openrouter-model-breadth.openrouter-openai-gpt-5-6-luna.local.plan-approve-complete
Matchers and test contextNot selected No matcher result was recorded. |
Test suite
Three browser-driven turns on one reusable Daytona sandbox for legacy and native Codex.
| Agent profile | Daytona warm reusable sandboxdaytona · remote |
|---|---|
Legacy Codexlegacycodex · gpt-5.6-sol
|
Daytona warm reusable sandboxdaytona · remote
Warm three-turn workspace continuity
not selected
daytona-warm-continuity.legacy-codex.daytona.warm-three-turn
Matchers and test contextNot selected No matcher result was recorded. |
Runner Codexnativecodex · gpt-5.6-sol
|
Daytona warm reusable sandboxdaytona · remote
Warm three-turn workspace continuity
not selected
daytona-warm-continuity.runner-codex.daytona.warm-three-turn
Matchers and test contextNot selected No matcher result was recorded. |
Test suite
Suite discovered from retained campaign identities; full suite size is not known to this publisher.
| Agent profile | Isolated locallocal · local | ||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Runner Codexnativecodex · gpt-5.6-sol
|
Isolated locallocal · local
enable-disable-resume
passed
No behavioral matcher failed. See all behavioral checks (1)
Tokens112,079 in · 963 out73,390 cached · 2/2 runs covered
LLM spend$0.0000002/2 runs provider-priced
ExecutionLocal · not metered18s agent
agent-chat-stories.runner-codex.local.enable-disable-resume
Matchers and test context
All invariants passed
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 2,
"runsWithTokenUsage": 2,
"runsWithReportedCost": 2,
"inputTokens": 112079,
"outputTokens": 963,
"cachedInputTokens": 73390,
"totalTokens": 186432,
"reportedCostUsd": 0,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 17991,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": true
},
"rawUsage": {
"runs": [
{
"runId": "bc25a6c5-f0fb-4209-aafa-45c5f6014eee",
"usage": {
"model": "gpt-5.6-sol",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 76596,
"usageSource": "per_run",
"freshSession": false,
"outputTokens": 649,
"sessionReused": true,
"rawInputTokens": 76596,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:f059c9b0e48554582b766026be8ed1cf5cdef67b2905b3c12ac3b7444d2bdb22",
"changedCategories": [],
"taskSessionReused": true,
"fingerprintVersion": 1,
"taskSessionAvailable": true,
"storedFingerprintPresent": true
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:8607bddf3e79fdb76530ae6b6abee206d5581835f9f07f084a22dd54fb7d1c78",
"workspaceReused": false,
"activeWorkspaceId": null,
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 649,
"cachedInputTokens": 55836,
"taskSessionReused": true,
"persistedSessionId": "01a0c583-cda3-7663-a1a0-de92c5f3e46d",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 55836,
"sessionRotationReason": null
}
},
{
"runId": "4ffb37c2-f7be-4ccc-91a1-eb1fd7194397",
"usage": {
"model": "gpt-5.6-sol",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 35483,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 314,
"sessionReused": false,
"rawInputTokens": 35483,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:f059c9b0e48554582b766026be8ed1cf5cdef67b2905b3c12ac3b7444d2bdb22",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:8607bddf3e79fdb76530ae6b6abee206d5581835f9f07f084a22dd54fb7d1c78",
"workspaceReused": false,
"activeWorkspaceId": null,
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 314,
"cachedInputTokens": 17554,
"taskSessionReused": false,
"persistedSessionId": "01a0c583-cda3-7663-a1a0-de92c5f3e46d",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 17554,
"sessionRotationReason": null
}
}
]
}
}
followup-while-running
failed
candidate failure [2mexpect([22m[31mreceived[39m[2m).[22mtoBe[2m([22m[32mexpected[39m[2m) // Object.is equality[22m
Expected: [32mtrue[39m
Received: [31mfalse[39m
Call Log:
- Timeout 120000ms exceeded while waiting on the predicate
Tokens74,410 in · 694 out54,897 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered22s agent
agent-chat-stories.runner-codex.local.followup-while-running
Matchers and test context
No result artifact was uploaded No matcher result was recorded. Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 74410,
"outputTokens": 694,
"cachedInputTokens": 54897,
"totalTokens": 130001,
"reportedCostUsd": 0,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 21541,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": true
},
"rawUsage": {
"model": "gpt-5.6-sol",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 74410,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 694,
"sessionReused": false,
"rawInputTokens": 74410,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:191fa5718540ae87324c661e89a48a1166833196b9c67a5737519ca9f9850b0b",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:e34a6d4c37f7cbf98c565ba9fa8ecf82e959598931e3f1240e6f7c3a782262e6",
"workspaceReused": false,
"activeWorkspaceId": null,
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 694,
"cachedInputTokens": 54897,
"taskSessionReused": false,
"persistedSessionId": "01a0c583-c208-7410-8c11-44f3c9d91f7a",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 54897,
"sessionRotationReason": null
}
}
revise-while-running
failed
candidate failure [2mexpect([22m[31mreceived[39m[2m).[22mtoBe[2m([22m[32mexpected[39m[2m) // Object.is equality[22m
Expected: [32mtrue[39m
Received: [31mfalse[39m
Call Log:
- Timeout 120000ms exceeded while waiting on the predicate
Tokens54,731 in · 724 out35,928 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered17s agent
agent-chat-stories.runner-codex.local.revise-while-running
Matchers and test context
No result artifact was uploaded No matcher result was recorded. Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 54731,
"outputTokens": 724,
"cachedInputTokens": 35928,
"totalTokens": 91383,
"reportedCostUsd": 0,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 17013,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": true
},
"rawUsage": {
"model": "gpt-5.6-sol",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 54731,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 724,
"sessionReused": false,
"rawInputTokens": 54731,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:f9637dbe97770f88b8971977073a8defc0ee23ca98cf0b18b8e53e7ae6ebb912",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:1e2611d00bb6e6649bb2ee5b53e94d25c633fdafec3c18a213cf22df81331ab8",
"workspaceReused": false,
"activeWorkspaceId": null,
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 724,
"cachedInputTokens": 35928,
"taskSessionReused": false,
"persistedSessionId": "01a0c583-a9be-71c1-9c07-087763d1b5d5",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 35928,
"sessionRotationReason": null
}
} | ||||||||||||||||||||||||||||||||||||||||||
Runner ACPX Claudenativeacpx · claude-sonnet-5
|
Isolated locallocal · local
enable-disable-resume
passed
No behavioral matcher failed. See all behavioral checks (1)
Tokens10 in · 1,105 out150,689 cached · 2/2 runs covered
LLM spend$0.0000002/2 runs provider-priced
ExecutionLocal · not metered38s agent
agent-chat-stories.runner-acpx-claude.local.enable-disable-resume
Matchers and test context
All invariants passed
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 2,
"runsWithTokenUsage": 2,
"runsWithReportedCost": 2,
"inputTokens": 10,
"outputTokens": 1105,
"cachedInputTokens": 150689,
"totalTokens": 151804,
"reportedCostUsd": 0,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 37578,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": true
},
"rawUsage": {
"runs": [
{
"runId": "8c83e98c-d78e-439e-aa5c-a88a6c268b99",
"usage": {
"model": "claude-sonnet-5",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 4,
"usageSource": "per_run",
"freshSession": false,
"outputTokens": 393,
"sessionReused": true,
"rawInputTokens": 4,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:eaecfb7f075134c628116003ba229dc484b7cad83c0f0e3ad438baba769edf42",
"changedCategories": [],
"taskSessionReused": true,
"fingerprintVersion": 1,
"taskSessionAvailable": true,
"storedFingerprintPresent": true
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:cd6c0abceedac5450558fd99ff8c076d9b8b72a774872007fa08d32e7296b1a7",
"workspaceReused": false,
"activeWorkspaceId": null,
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 393,
"cachedInputTokens": 60338,
"taskSessionReused": true,
"persistedSessionId": "5ee4b87d-7437-4691-b41b-792b05090600",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 60338,
"sessionRotationReason": null
}
},
{
"runId": "4f7b7b0c-14b7-441b-965e-1624b7961be0",
"usage": {
"model": "claude-sonnet-5",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 6,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 712,
"sessionReused": false,
"rawInputTokens": 6,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:eaecfb7f075134c628116003ba229dc484b7cad83c0f0e3ad438baba769edf42",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:cd6c0abceedac5450558fd99ff8c076d9b8b72a774872007fa08d32e7296b1a7",
"workspaceReused": false,
"activeWorkspaceId": null,
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 712,
"cachedInputTokens": 90351,
"taskSessionReused": false,
"persistedSessionId": "5ee4b87d-7437-4691-b41b-792b05090600",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 90351,
"sessionRotationReason": null
}
}
]
}
}
followup-while-running
passed
No behavioral matcher failed. See all behavioral checks (1)
Tokens10 in · 1,368 out152,785 cached · 2/2 runs covered
LLM spend$0.0000002/2 runs provider-priced
ExecutionLocal · not metered44s agent
agent-chat-stories.runner-acpx-claude.local.followup-while-running
Matchers and test context
All invariants passed
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 2,
"runsWithTokenUsage": 2,
"runsWithReportedCost": 2,
"inputTokens": 10,
"outputTokens": 1368,
"cachedInputTokens": 152785,
"totalTokens": 154163,
"reportedCostUsd": 0,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 43841,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": true
},
"rawUsage": {
"runs": [
{
"runId": "bc17ccbf-3e38-4455-bdcf-9f879c2b86e5",
"usage": {
"model": "claude-sonnet-5",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 6,
"usageSource": "per_run",
"freshSession": false,
"outputTokens": 1070,
"sessionReused": true,
"rawInputTokens": 6,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:2799db7a3baeb0043b6d295d97e6b591d54eb08e315ae2b628f4b47c5c7b4d71",
"changedCategories": [],
"taskSessionReused": true,
"fingerprintVersion": 1,
"taskSessionAvailable": true,
"storedFingerprintPresent": true
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:05ac562cbf3df32ae84f9be841e39afc610a61a9859310c0d8f8e91918b1a574",
"workspaceReused": false,
"activeWorkspaceId": null,
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 1070,
"cachedInputTokens": 98439,
"taskSessionReused": true,
"persistedSessionId": "5c926a7d-dc84-4bb6-bbaa-b131d65966ef",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 98439,
"sessionRotationReason": null
}
},
{
"runId": "edca6ed1-c934-4a66-99d3-5e9347f0d935",
"usage": {
"model": "claude-sonnet-5",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 4,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 298,
"sessionReused": false,
"rawInputTokens": 4,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:2799db7a3baeb0043b6d295d97e6b591d54eb08e315ae2b628f4b47c5c7b4d71",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:05ac562cbf3df32ae84f9be841e39afc610a61a9859310c0d8f8e91918b1a574",
"workspaceReused": false,
"activeWorkspaceId": null,
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 298,
"cachedInputTokens": 54346,
"taskSessionReused": false,
"persistedSessionId": "5c926a7d-dc84-4bb6-bbaa-b131d65966ef",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 54346,
"sessionRotationReason": null
}
}
]
}
}
revise-while-running
passed
No behavioral matcher failed. See all behavioral checks (1)
Tokens16 in · 3,140 out282,307 cached · 2/2 runs covered
LLM spend$0.0000002/2 runs provider-priced
ExecutionLocal · not metered1m 3s agent
agent-chat-stories.runner-acpx-claude.local.revise-while-running
Matchers and test context
All invariants passed
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 2,
"runsWithTokenUsage": 2,
"runsWithReportedCost": 2,
"inputTokens": 16,
"outputTokens": 3140,
"cachedInputTokens": 282307,
"totalTokens": 285463,
"reportedCostUsd": 0,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 62804,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": true
},
"rawUsage": {
"runs": [
{
"runId": "1c46b570-04ae-41fd-aea9-cbdd01038e81",
"usage": {
"model": "claude-sonnet-5",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 8,
"usageSource": "per_run",
"freshSession": false,
"outputTokens": 1538,
"sessionReused": true,
"rawInputTokens": 8,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:9cb6999cd787dd46315c4303a3388d8b749027ff87e17ef862478659ca9faf29",
"changedCategories": [],
"taskSessionReused": true,
"fingerprintVersion": 1,
"taskSessionAvailable": true,
"storedFingerprintPresent": true
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:3f6f4e8b6a960dbb5547c8e4fd3c9adb1572568bc18a9532acb573eb6baaf4cb",
"workspaceReused": false,
"activeWorkspaceId": null,
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 1538,
"cachedInputTokens": 154063,
"taskSessionReused": true,
"persistedSessionId": "d1488a03-ab6a-4d16-ac09-3961d64c2d38",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 154063,
"sessionRotationReason": null
}
},
{
"runId": "c000f75d-0d4d-4636-91a0-b6cb1b45f5e0",
"usage": {
"model": "claude-sonnet-5",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 8,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 1602,
"sessionReused": false,
"rawInputTokens": 8,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:9cb6999cd787dd46315c4303a3388d8b749027ff87e17ef862478659ca9faf29",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:3f6f4e8b6a960dbb5547c8e4fd3c9adb1572568bc18a9532acb573eb6baaf4cb",
"workspaceReused": false,
"activeWorkspaceId": null,
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 1602,
"cachedInputTokens": 128244,
"taskSessionReused": false,
"persistedSessionId": "d1488a03-ab6a-4d16-ac09-3961d64c2d38",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 128244,
"sessionRotationReason": null
}
}
]
}
} |
History
No historical campaigns have been published yet.