everyday-workflows.runner-codex.local.build-revise
Matchers and test context
Not selected
No matcher result was recorded.
Full-stack acceptance campaign
A browser-verified matrix of runner profiles, execution environments, and deterministic task contracts. Declared PNG screenshots and sanitized structured evidence are retained with every published campaign; additional diagnostic evidence remains in the access-controlled workflow artifact.

Model spend is the provider-reported subtotal; unpriced or unavailable runs are excluded, never counted as free. Daytona runtime is a public-list-price estimate from captured lease time and pinned resources, before credits, discounts, storage allowance, or invoice adjustments. Local execution has no external runtime meter.
Test suite
Real user requests, useful downloaded work, and durable continuation using production instructions.
| Agent profile | Isolated locallocal · local | Daytona warm reusable sandboxdaytona · remote |
|---|---|---|
Runner Codexnativecodex · gpt-5.6-sol
|
Isolated locallocal · local
Build, download, and revise a project
not selected
everyday-workflows.runner-codex.local.build-revise
Matchers and test contextNot selected No matcher result was recorded.
Delegate implementation and preserve late feedback
not selected
everyday-workflows.runner-codex.local.delegate-feedback
Matchers and test contextNot selected No matcher result was recorded.
Hire one teammate, then reuse that agent
not selected
everyday-workflows.runner-codex.local.hire-reuse
Matchers and test contextNot selected No matcher result was recorded.
Use a connection after approval
not selected
everyday-workflows.runner-codex.local.service-approve
Matchers and test contextNot selected No matcher result was recorded.
Respect a declined tool action
not selected
everyday-workflows.runner-codex.local.service-decline
Matchers and test contextNot selected No matcher result was recorded.
Respect Not now on a new connection
not selected
everyday-workflows.runner-codex.local.connection-decline
Matchers and test contextNot selected No matcher result was recorded.
Recover work after the server restarts
not selected
everyday-workflows.runner-codex.local.recover-controller
Matchers and test contextNot selected No matcher result was recorded.
Stop a task and send a new direction once
not selected
everyday-workflows.runner-codex.local.stop-redirect
Matchers and test contextNot selected No matcher result was recorded. |
Daytona warm reusable sandboxdaytona · remote
Build, download, and revise a project
not selected
everyday-workflows.runner-codex.daytona.build-revise
Matchers and test contextNot selected No matcher result was recorded.
Delegate implementation and preserve late feedback
not selected
everyday-workflows.runner-codex.daytona.delegate-feedback
Matchers and test contextNot selected No matcher result was recorded.
Recover work after the server restarts
not selected
everyday-workflows.runner-codex.daytona.recover-controller
Matchers and test contextNot selected No matcher result was recorded. |
Runner ACPX Claudenativeacpx · claude-sonnet-5
|
Isolated locallocal · local
Build, download, and revise a project
not selected
everyday-workflows.runner-acpx-claude.local.build-revise
Matchers and test contextNot selected No matcher result was recorded.
Delegate implementation and preserve late feedback
not selected
everyday-workflows.runner-acpx-claude.local.delegate-feedback
Matchers and test contextNot selected No matcher result was recorded.
Hire one teammate, then reuse that agent
not selected
everyday-workflows.runner-acpx-claude.local.hire-reuse
Matchers and test contextNot selected No matcher result was recorded.
Use a connection after approval
not selected
everyday-workflows.runner-acpx-claude.local.service-approve
Matchers and test contextNot selected No matcher result was recorded.
Respect a declined tool action
not selected
everyday-workflows.runner-acpx-claude.local.service-decline
Matchers and test contextNot selected No matcher result was recorded.
Respect Not now on a new connection
not selected
everyday-workflows.runner-acpx-claude.local.connection-decline
Matchers and test contextNot selected No matcher result was recorded.
Recover work after the server restarts
not selected
everyday-workflows.runner-acpx-claude.local.recover-controller
Matchers and test contextNot selected No matcher result was recorded.
Stop a task and send a new direction once
not selected
everyday-workflows.runner-acpx-claude.local.stop-redirect
Matchers and test contextNot selected No matcher result was recorded. |
Daytona warm reusable sandboxdaytona · remote
Build, download, and revise a project
not selected
everyday-workflows.runner-acpx-claude.daytona.build-revise
Matchers and test contextNot selected No matcher result was recorded.
Delegate implementation and preserve late feedback
not selected
everyday-workflows.runner-acpx-claude.daytona.delegate-feedback
Matchers and test contextNot selected No matcher result was recorded.
Recover work after the server restarts
not selected
everyday-workflows.runner-acpx-claude.daytona.recover-controller
Matchers and test contextNot selected No matcher result was recorded. |
Runner Codex Mininativecodex · gpt-5.4-mini
|
Isolated locallocal · local
Build, download, and revise a project
not selected
everyday-workflows.runner-codex-mini.local.build-revise
Matchers and test contextNot selected No matcher result was recorded.
Delegate implementation and preserve late feedback
not selected
everyday-workflows.runner-codex-mini.local.delegate-feedback
Matchers and test contextNot selected No matcher result was recorded.
Hire one teammate, then reuse that agent
not selected
everyday-workflows.runner-codex-mini.local.hire-reuse
Matchers and test contextNot selected No matcher result was recorded.
Use a connection after approval
not selected
everyday-workflows.runner-codex-mini.local.service-approve
Matchers and test contextNot selected No matcher result was recorded.
Respect a declined tool action
not selected
everyday-workflows.runner-codex-mini.local.service-decline
Matchers and test contextNot selected No matcher result was recorded.
Respect Not now on a new connection
not selected
everyday-workflows.runner-codex-mini.local.connection-decline
Matchers and test contextNot selected No matcher result was recorded.
Recover work after the server restarts
not selected
everyday-workflows.runner-codex-mini.local.recover-controller
Matchers and test contextNot selected No matcher result was recorded.
Stop a task and send a new direction once
not selected
everyday-workflows.runner-codex-mini.local.stop-redirect
Matchers and test contextNot selected No matcher result was recorded. |
Daytona warm reusable sandboxdaytona · remote
|
Test suite
Task-backed conversations, session resets, and project plan handoff.
| Agent profile | Isolated locallocal · local |
|---|---|
Legacy Codexlegacycodex · gpt-5.6-sol
|
Isolated locallocal · local
Conversation continuity across restart
not selected
agent-chat.legacy-codex.local.continuity-restart
Matchers and test contextNot selected No matcher result was recorded.
Fresh context within preserved history
not selected
agent-chat.legacy-codex.local.new-session
Matchers and test contextNot selected No matcher result was recorded.
Stop, reset, and resume
not selected
agent-chat.legacy-codex.local.stop-new-resume
Matchers and test contextNot selected No matcher result was recorded.
Draft, revise, approve, and hand off a plan
not selected
agent-chat.legacy-codex.local.plan-handoff
Matchers and test contextNot selected No matcher result was recorded.
Clarify and reuse an existing project
not selected
agent-chat.legacy-codex.local.clarify-reuse
Matchers and test contextNot selected No matcher result was recorded.
Create a project with multiple repository URLs
not selected
agent-chat.legacy-codex.local.multi-repository
Matchers and test contextNot selected No matcher result was recorded. |
Legacy Claudelegacyclaude · claude-sonnet-4-6
|
Isolated locallocal · local
Conversation continuity across restart
not selected
agent-chat.legacy-claude.local.continuity-restart
Matchers and test contextNot selected No matcher result was recorded.
Fresh context within preserved history
not selected
agent-chat.legacy-claude.local.new-session
Matchers and test contextNot selected No matcher result was recorded.
Stop, reset, and resume
not selected
agent-chat.legacy-claude.local.stop-new-resume
Matchers and test contextNot selected No matcher result was recorded.
Draft, revise, approve, and hand off a plan
not selected
agent-chat.legacy-claude.local.plan-handoff
Matchers and test contextNot selected No matcher result was recorded.
Clarify and reuse an existing project
not selected
agent-chat.legacy-claude.local.clarify-reuse
Matchers and test contextNot selected No matcher result was recorded.
Create a project with multiple repository URLs
not selected
agent-chat.legacy-claude.local.multi-repository
Matchers and test contextNot selected No matcher result was recorded. |
Runner Codexnativecodex · gpt-5.6-sol
|
Isolated locallocal · local
Conversation continuity across restart
not selected
agent-chat.runner-codex.local.continuity-restart
Matchers and test contextNot selected No matcher result was recorded.
Fresh context within preserved history
not selected
agent-chat.runner-codex.local.new-session
Matchers and test contextNot selected No matcher result was recorded.
Stop, reset, and resume
not selected
agent-chat.runner-codex.local.stop-new-resume
Matchers and test contextNot selected No matcher result was recorded.
Draft, revise, approve, and hand off a plan
not selected
agent-chat.runner-codex.local.plan-handoff
Matchers and test contextNot selected No matcher result was recorded.
Clarify and reuse an existing project
not selected
agent-chat.runner-codex.local.clarify-reuse
Matchers and test contextNot selected No matcher result was recorded.
Create a project with multiple repository URLs
not selected
agent-chat.runner-codex.local.multi-repository
Matchers and test contextNot selected No matcher result was recorded. |
Runner ACPX Claudenativeacpx · claude-sonnet-5
|
Isolated locallocal · local
Conversation continuity across restart
not selected
agent-chat.runner-acpx-claude.local.continuity-restart
Matchers and test contextNot selected No matcher result was recorded.
Fresh context within preserved history
not selected
agent-chat.runner-acpx-claude.local.new-session
Matchers and test contextNot selected No matcher result was recorded.
Stop, reset, and resume
not selected
agent-chat.runner-acpx-claude.local.stop-new-resume
Matchers and test contextNot selected No matcher result was recorded.
Draft, revise, approve, and hand off a plan
not selected
agent-chat.runner-acpx-claude.local.plan-handoff
Matchers and test contextNot selected No matcher result was recorded.
Clarify and reuse an existing project
not selected
agent-chat.runner-acpx-claude.local.clarify-reuse
Matchers and test contextNot selected No matcher result was recorded.
Create a project with multiple repository URLs
not selected
agent-chat.runner-acpx-claude.local.multi-repository
Matchers and test contextNot selected No matcher result was recorded. |
Test suite
Major provider, runtime generation, and execution-environment compatibility.
| Agent profile | Isolated locallocal · local | Daytona sandboxdaytona · remote |
|---|---|---|
Legacy Codexlegacycodex · gpt-5.6-sol
|
Isolated locallocal · local
Basic response
not selected
core-compatibility.legacy-codex.local.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.legacy-codex.local.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.legacy-codex.local.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Daytona sandboxdaytona · remote
Basic response
not selected
core-compatibility.legacy-codex.daytona.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.legacy-codex.daytona.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.legacy-codex.daytona.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Legacy Claudelegacyclaude · claude-sonnet-4-6
|
Isolated locallocal · local
Basic response
not selected
core-compatibility.legacy-claude.local.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.legacy-claude.local.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.legacy-claude.local.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Daytona sandboxdaytona · remote
Basic response
not selected
core-compatibility.legacy-claude.daytona.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.legacy-claude.daytona.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.legacy-claude.daytona.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Legacy OpenCodelegacyopencode · openrouter/deepseek/deepseek-v4-flash-0731
|
Isolated locallocal · local
Basic response
not selected
core-compatibility.legacy-opencode.local.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.legacy-opencode.local.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.legacy-opencode.local.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Daytona sandboxdaytona · remote
Basic response
not selected
core-compatibility.legacy-opencode.daytona.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.legacy-opencode.daytona.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.legacy-opencode.daytona.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Runner Codexnativecodex · gpt-5.6-sol
|
Isolated locallocal · local
Basic response
not selected
core-compatibility.runner-codex.local.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.runner-codex.local.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.runner-codex.local.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Daytona sandboxdaytona · remote
Basic response
not selected
core-compatibility.runner-codex.daytona.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.runner-codex.daytona.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.runner-codex.daytona.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Runner OpenCodenativeopencode · openrouter/deepseek/deepseek-v4-flash-0731
|
Isolated locallocal · local
Basic response
not selected
core-compatibility.runner-opencode.local.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.runner-opencode.local.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.runner-opencode.local.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Daytona sandboxdaytona · remote
Basic response
not selected
core-compatibility.runner-opencode.daytona.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.runner-opencode.daytona.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.runner-opencode.daytona.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Runner ACPX Claudenativeacpx · claude-sonnet-5
|
Isolated locallocal · local
Basic response
not selected
core-compatibility.runner-acpx-claude.local.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.runner-acpx-claude.local.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.runner-acpx-claude.local.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Daytona sandboxdaytona · remote
Basic response
not selected
core-compatibility.runner-acpx-claude.daytona.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.runner-acpx-claude.daytona.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.runner-acpx-claude.daytona.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Runner ACPX Codexnativeacpx · gpt-5.6-sol
|
Isolated locallocal · local
Basic response
not selected
core-compatibility.runner-acpx-codex.local.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.runner-acpx-codex.local.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.runner-acpx-codex.local.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Daytona sandboxdaytona · remote
Basic response
not selected
core-compatibility.runner-acpx-codex.daytona.message-marker
Matchers and test contextNot selected No matcher result was recorded.
Plan, revise, accept, implement
not selected
core-compatibility.runner-acpx-codex.daytona.plan-revise-accept
Matchers and test contextNot selected No matcher result was recorded.
Ask mode question
not selected
core-compatibility.runner-acpx-codex.daytona.ask-question
Matchers and test contextNot selected No matcher result was recorded. |
Test suite
Structured interaction and continuation qualification for every supported local profile.
| Agent profile | Isolated locallocal · local |
|---|---|
Legacy Codexlegacycodex · gpt-5.6-sol
|
Isolated locallocal · local
Structured question, answer, resume
not selected
local-session-integrity.legacy-codex.local.structured-question-resume
Matchers and test contextNot selected No matcher result was recorded.
Structured question, server restart, answer, resume
not selected
local-session-integrity.legacy-codex.local.structured-question-restart-resume
Matchers and test contextNot selected No matcher result was recorded. |
Legacy Claudelegacyclaude · claude-sonnet-4-6
|
Isolated locallocal · local
Structured question, answer, resume
not selected
local-session-integrity.legacy-claude.local.structured-question-resume
Matchers and test contextNot selected No matcher result was recorded.
Structured question, server restart, answer, resume
not selected
local-session-integrity.legacy-claude.local.structured-question-restart-resume
Matchers and test contextNot selected No matcher result was recorded. |
Legacy OpenCodelegacyopencode · openrouter/deepseek/deepseek-v4-flash-0731
|
Isolated locallocal · local
Structured question, answer, resume
not selected
local-session-integrity.legacy-opencode.local.structured-question-resume
Matchers and test contextNot selected No matcher result was recorded.
Structured question, server restart, answer, resume
not selected
local-session-integrity.legacy-opencode.local.structured-question-restart-resume
Matchers and test contextNot selected No matcher result was recorded. |
Runner Codexnativecodex · gpt-5.6-sol
|
Isolated locallocal · local
Structured question, answer, resume
not selected
local-session-integrity.runner-codex.local.structured-question-resume
Matchers and test contextNot selected No matcher result was recorded.
Structured question, server restart, answer, resume
not selected
local-session-integrity.runner-codex.local.structured-question-restart-resume
Matchers and test contextNot selected No matcher result was recorded. |
Runner OpenCodenativeopencode · openrouter/deepseek/deepseek-v4-flash-0731
|
Isolated locallocal · local
Structured question, answer, resume
not selected
local-session-integrity.runner-opencode.local.structured-question-resume
Matchers and test contextNot selected No matcher result was recorded.
Structured question, server restart, answer, resume
not selected
local-session-integrity.runner-opencode.local.structured-question-restart-resume
Matchers and test contextNot selected No matcher result was recorded. |
Runner ACPX Claudenativeacpx · claude-sonnet-5
|
Isolated locallocal · local
Structured question, answer, resume
not selected
local-session-integrity.runner-acpx-claude.local.structured-question-resume
Matchers and test contextNot selected No matcher result was recorded.
Structured question, server restart, answer, resume
not selected
local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume
Matchers and test contextNot selected No matcher result was recorded. |
Runner ACPX Codexnativeacpx · gpt-5.6-sol
|
Isolated locallocal · local
Structured question, answer, resume
not selected
local-session-integrity.runner-acpx-codex.local.structured-question-resume
Matchers and test contextNot selected No matcher result was recorded.
Structured question, server restart, answer, resume
not selected
local-session-integrity.runner-acpx-codex.local.structured-question-restart-resume
Matchers and test contextNot selected No matcher result was recorded. |
Test suite
Weekly-ranked tool-capable OpenRouter models through native OpenCode on isolated local workspaces.
| Agent profile | Isolated locallocal · local |
|---|---|
#1 DeepSeek V4 Flash 0731nativeopencode · openrouter/deepseek/deepseek-v4-flash-0731
|
Isolated locallocal · local
Hello and complete
not selected
openrouter-model-breadth.openrouter-deepseek-deepseek-v4-flash-0731.local.hello-complete
Matchers and test contextNot selected No matcher result was recorded.
Ask, answer, resume
not selected
openrouter-model-breadth.openrouter-deepseek-deepseek-v4-flash-0731.local.question-resume-complete
Matchers and test contextNot selected No matcher result was recorded. |
#3 Tencent HY 3nativeopencode · openrouter/tencent/hy3
|
Isolated locallocal · local
Hello and complete
not selected
openrouter-model-breadth.openrouter-tencent-hy3.local.hello-complete
Matchers and test contextNot selected No matcher result was recorded.
Ask, answer, resume
not selected
openrouter-model-breadth.openrouter-tencent-hy3.local.question-resume-complete
Matchers and test contextNot selected No matcher result was recorded. |
#4 Nemotron 3 Ultra 550B A55B (free)nativeopencode · openrouter/nvidia/nemotron-3-ultra-550b-a55b:free
|
Isolated locallocal · local
Hello and complete
not selected
openrouter-model-breadth.openrouter-nvidia-nemotron-3-ultra-550b-a55b-free.local.hello-complete
Matchers and test contextNot selected No matcher result was recorded.
Ask, answer, resume
not selected
openrouter-model-breadth.openrouter-nvidia-nemotron-3-ultra-550b-a55b-free.local.question-resume-complete
Matchers and test contextNot selected No matcher result was recorded.
Plan, approve, complete
not selected
openrouter-model-breadth.openrouter-nvidia-nemotron-3-ultra-550b-a55b-free.local.plan-approve-complete
Matchers and test contextNot selected No matcher result was recorded. |
#5 GPT-5.6 Lunanativeopencode · openrouter/openai/gpt-5.6-luna
|
Isolated locallocal · local
Hello and complete
not selected
openrouter-model-breadth.openrouter-openai-gpt-5-6-luna.local.hello-complete
Matchers and test contextNot selected No matcher result was recorded.
Ask, answer, resume
not selected
openrouter-model-breadth.openrouter-openai-gpt-5-6-luna.local.question-resume-complete
Matchers and test contextNot selected No matcher result was recorded.
Plan, approve, complete
not selected
openrouter-model-breadth.openrouter-openai-gpt-5-6-luna.local.plan-approve-complete
Matchers and test contextNot selected No matcher result was recorded. |
Test suite
Three browser-driven turns on one reusable Daytona sandbox for legacy and native Codex.
| Agent profile | Daytona warm reusable sandboxdaytona · remote |
|---|---|
Legacy Codexlegacycodex · gpt-5.6-sol
|
Daytona warm reusable sandboxdaytona · remote
Warm three-turn workspace continuity
not selected
daytona-warm-continuity.legacy-codex.daytona.warm-three-turn
Matchers and test contextNot selected No matcher result was recorded. |
Runner Codexnativecodex · gpt-5.6-sol
|
Daytona warm reusable sandboxdaytona · remote
Warm three-turn workspace continuity
not selected
daytona-warm-continuity.runner-codex.daytona.warm-three-turn
Matchers and test contextNot selected No matcher result was recorded. |
Test suite
Suite discovered from retained campaign identities; full suite size is not known to this publisher.
| Agent profile | Isolated locallocal · local | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Legacy Codexlegacycodex · gpt-5.6-sol
|
Isolated locallocal · local
interview-first-response
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.legacy-codex.local.interview-first-response
Matchers and test context
locator.click: Timeout 30000ms exceeded. Call log: [2m - waiting for getByRole('radio', { name: 'Interview me and propose a plan and an agent team to execute it.', exact: true }).last()[22m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 0,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"runs": []
}
}
clear-task-first-response
failed
Tokens1,089 in · 161 out36,185 cached · 1/1 runs covered
LLM spendunpriced0/1 runs provider-priced
ExecutionLocal · not metered58s agent
first-task.legacy-codex.local.clear-task-first-response
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-90f05d1bc6f2edc1/companies/1385351b-da54-474c-b43a-90dc0fb1fef9/acp-engine/agents/dd2e7752-5108-4685-a788-6e9e1f75330e/sessions/paperclip%3A1385351b-da54-474c-b43a-90dc0fb1fef9%3Add2e7752-5108-4685-a788-6e9e1f75330e%3Ab6ffae3b-7be6-4f91-bc34-cacedf75419d%3Af087b779991de06f.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 0,
"inputTokens": 1089,
"outputTokens": 161,
"cachedInputTokens": 36185,
"totalTokens": 37435,
"reportedCostUsd": 0,
"costStatus": "unpriced"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 58493,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"model": "unknown",
"biller": "openai",
"provider": "openai",
"costStatus": "unpriced",
"billingType": "metered_api",
"inputTokens": 1089,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 161,
"sessionReused": false,
"rawInputTokens": 1089,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:a698c106ccdbf5fdbb366bd374aeaf21c7852d6eb0ce8d23f392ad57b05fb1fd",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:5c9f539236d646fcf3b79dee0afedd865fda731b6ddc44c6f94ea5a7097d108e",
"workspaceReused": false,
"activeWorkspaceId": "df7c008d-3e11-4c0f-a569-005239ea4d1a",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 161,
"cachedInputTokens": 36185,
"taskSessionReused": false,
"persistedSessionId": "01a0a644-cff2-7c12-8523-859679b94c2e",
"rawCachedInputTokens": 36185,
"sessionRotationReason": null
}
}
ambiguous-task-first-response
failed
Tokens963 in · 36 out30,751 cached · 1/1 runs covered
LLM spendunpriced0/1 runs provider-priced
ExecutionLocal · not metered48s agent
first-task.legacy-codex.local.ambiguous-task-first-response
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-062287acecfe8a82/companies/618fbcc5-6f58-47ba-b816-eb1b9561a93d/acp-engine/agents/6d86a02e-e9e7-4092-9302-b0fb124ef29c/sessions/paperclip%3A618fbcc5-6f58-47ba-b816-eb1b9561a93d%3A6d86a02e-e9e7-4092-9302-b0fb124ef29c%3A635fbfaa-302d-451d-8f82-ab2dc03f9a75%3A0857beb4f174ec6b.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 0,
"inputTokens": 963,
"outputTokens": 36,
"cachedInputTokens": 30751,
"totalTokens": 31750,
"reportedCostUsd": 0,
"costStatus": "unpriced"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 48118,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"model": "unknown",
"biller": "openai",
"provider": "openai",
"costStatus": "unpriced",
"billingType": "metered_api",
"inputTokens": 963,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 36,
"sessionReused": false,
"rawInputTokens": 963,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:199733c12706f0e8bfba757c029927f145a08d99341a4b6b1b8879ea930d1d00",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:6c330fad7bab1044a2615e779772062a98a91afb66fcd7adae823f82e5c7dd78",
"workspaceReused": false,
"activeWorkspaceId": "bffb3755-97f3-4cfd-9e02-5c11c84afac3",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 36,
"cachedInputTokens": 30751,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-3c8f-7732-a8c5-49f920e3b39a",
"rawCachedInputTokens": 30751,
"sessionRotationReason": null
}
}
plain-message-first-response
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.legacy-codex.local.plain-message-first-response
Matchers and test context
locator.fill: Timeout 30000ms exceeded. Call log: [2m - waiting for getByTestId('task-chat-composer-input').last().locator('[contenteditable="true"], textarea').first()[22m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 0,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"runs": []
}
}
plan-first-response
failed
Tokens1,488 in · 66 out35,208 cached · 1/1 runs covered
LLM spendunpriced0/1 runs provider-priced
ExecutionLocal · not metered53s agent
first-task.legacy-codex.local.plan-first-response
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-c507a533c701b6c3/companies/e12f8657-de50-4bce-bd8f-c492d224a6e8/acp-engine/agents/96058ee9-82d3-46f2-9bbd-af85ecd4be3c/sessions/paperclip%3Ae12f8657-de50-4bce-bd8f-c492d224a6e8%3A96058ee9-82d3-46f2-9bbd-af85ecd4be3c%3Aac441456-b1d3-4205-988c-2e91ff0292d7%3Abd594d4831bd3920.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 0,
"inputTokens": 1488,
"outputTokens": 66,
"cachedInputTokens": 35208,
"totalTokens": 36762,
"reportedCostUsd": 0,
"costStatus": "unpriced"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 52607,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"model": "unknown",
"biller": "openai",
"provider": "openai",
"costStatus": "unpriced",
"billingType": "metered_api",
"inputTokens": 1488,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 66,
"sessionReused": false,
"rawInputTokens": 1488,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:bd6025d831361175498d8953d0bd85d6b0ffecd79c86c9b85a6512f589626e4a",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:fbcc74f4c4699a648ebdea90a9efa9a7ff0c613443af030ae035ff67977affca",
"workspaceReused": false,
"activeWorkspaceId": "bfd615e6-9e6e-45db-938d-5e9bb406eb6a",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 66,
"cachedInputTokens": 35208,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-248e-7f11-a024-45743c1ebc5c",
"rawCachedInputTokens": 35208,
"sessionRotationReason": null
}
}
ordinary-task-control
failed
Tokens1,486 in · 58 out27,720 cached · 1/1 runs covered
LLM spendunpriced0/1 runs provider-priced
ExecutionLocal · not metered31s agent
first-task.legacy-codex.local.ordinary-task-control
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-91562c687a3484d3/companies/e29b2b4e-a57f-4811-a6f3-edc0d16d200f/acp-engine/agents/b22f3bca-fa19-4cb0-98f7-3683b191ce29/sessions/paperclip%3Ae29b2b4e-a57f-4811-a6f3-edc0d16d200f%3Ab22f3bca-fa19-4cb0-98f7-3683b191ce29%3Aa46e048c-6617-45df-bc8b-c4a5a8064a32%3Aa3b3f6ed0b3b2eac.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 0,
"inputTokens": 1486,
"outputTokens": 58,
"cachedInputTokens": 27720,
"totalTokens": 29264,
"reportedCostUsd": 0,
"costStatus": "unpriced"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 31391,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"model": "unknown",
"biller": "openai",
"provider": "openai",
"costStatus": "unpriced",
"billingType": "metered_api",
"inputTokens": 1486,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 58,
"sessionReused": false,
"rawInputTokens": 1486,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": true,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:e49b25090989e1ca4cf26e039dec95e17c84c625bb4ae791b943ce7a5634b8e7",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:4146699b4295d368d7c0169d2e33bac59f2e7a8217b3af37947183bca1edd697",
"workspaceReused": false,
"activeWorkspaceId": null,
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 58,
"cachedInputTokens": 27720,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-b02e-74e3-b247-a7b25a01dbde",
"rawCachedInputTokens": 27720,
"sessionRotationReason": null
}
}
interview-plan-accept
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.legacy-codex.local.interview-plan-accept
Matchers and test context
locator.click: Timeout 30000ms exceeded. Call log: [2m - waiting for getByRole('radio', { name: 'Interview me and propose a plan and an agent team to execute it.', exact: true }).last()[22m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 0,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"runs": []
}
}
task-card-accept
failed
Tokens2,667 in · 167 out30,666 cached · 1/2 runs covered
LLM spendunpriced0/2 runs provider-priced
ExecutionLocal · not metered51s agent
first-task.legacy-codex.local.task-card-accept
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-2e5a5d01d0859994/companies/fb364bbb-8a73-4ba0-bb8d-a8b9cb3be4a0/acp-engine/agents/a94fce4c-649e-4b2d-8afe-95084d8cea48/sessions/paperclip%3Afb364bbb-8a73-4ba0-bb8d-a8b9cb3be4a0%3Aa94fce4c-649e-4b2d-8afe-95084d8cea48%3A90fd6d68-53db-4322-b30f-9e4b9fb77c37%3A2f9ba35b5288f1c6.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 2,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 0,
"inputTokens": 2667,
"outputTokens": 167,
"cachedInputTokens": 30666,
"totalTokens": 33500,
"reportedCostUsd": 0,
"costStatus": "unpriced"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 50757,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"runs": [
{
"runId": "80119376-fbef-44c0-9524-c3e0ec913e2c",
"usage": null
},
{
"runId": "75b8d94b-44d1-4815-9dab-0c38550fd369",
"usage": {
"model": "unknown",
"biller": "openai",
"provider": "openai",
"costStatus": "unpriced",
"billingType": "metered_api",
"inputTokens": 2667,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 167,
"sessionReused": false,
"rawInputTokens": 2667,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:9b0462ec3f30dc5acf9de55aff135f05f3378f3acf7d0301f5187f1639e1bb01",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:a424f7f516754e9c3db794d8dfe05caa0aa4b6e74ea24c6ae75edf709933d7c4",
"workspaceReused": false,
"activeWorkspaceId": "ec68480a-f08e-4475-afa6-3c3057bec18f",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 167,
"cachedInputTokens": 30666,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-5d81-7c40-80fd-5c051810c431",
"rawCachedInputTokens": 30666,
"sessionRotationReason": null
}
}
]
}
}
task-reply-accept
failed
Tokens1,065 in · 163 out36,487 cached · 1/1 runs covered
LLM spendunpriced0/1 runs provider-priced
ExecutionLocal · not metered51s agent
first-task.legacy-codex.local.task-reply-accept
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-9bb0e3e7c918d999/companies/82129711-e2fa-4261-8944-01adf9e8aca5/acp-engine/agents/7bd93cf0-4265-4a6e-a3ec-c9c99c205149/sessions/paperclip%3A82129711-e2fa-4261-8944-01adf9e8aca5%3A7bd93cf0-4265-4a6e-a3ec-c9c99c205149%3Aa8199b8a-6779-4273-8c17-66671df9fed5%3A6bb25da5e6b0c7a0.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 0,
"inputTokens": 1065,
"outputTokens": 163,
"cachedInputTokens": 36487,
"totalTokens": 37715,
"reportedCostUsd": 0,
"costStatus": "unpriced"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 50934,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"model": "unknown",
"biller": "openai",
"provider": "openai",
"costStatus": "unpriced",
"billingType": "metered_api",
"inputTokens": 1065,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 163,
"sessionReused": false,
"rawInputTokens": 1065,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:bff86fb53cdb4de8bf36f9cdb61520f9b1964bd88a88630582116eee86ea0f7f",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:b8b90ecf51a87fd34328684d00d07b02699e66efe0de279d8a20652fed35a100",
"workspaceReused": false,
"activeWorkspaceId": "ccadc74b-b921-40d6-a489-2a5eb30e1cde",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 163,
"cachedInputTokens": 36487,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-4e30-7801-ae58-1013a7d0ca14",
"rawCachedInputTokens": 36487,
"sessionRotationReason": null
}
}
clarify-propose-accept
failed
Tokens2,638 in · 208 out71,683 cached · 2/2 runs covered
LLM spendunpriced0/2 runs provider-priced
ExecutionLocal · not metered1m 38s agent
first-task.legacy-codex.local.clarify-propose-accept
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-d4b1e193a7778582/companies/a9478ac5-149c-492c-b17a-5e15dd1af913/acp-engine/agents/9394d2fb-4d0d-47ac-8fac-2f366f9be8af/sessions/paperclip%3Aa9478ac5-149c-492c-b17a-5e15dd1af913%3A9394d2fb-4d0d-47ac-8fac-2f366f9be8af%3Af234d0a7-f389-405f-96ff-11745edd6f66%3A04a916ee022e59d4.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 2,
"runsWithTokenUsage": 2,
"runsWithReportedCost": 0,
"inputTokens": 2638,
"outputTokens": 208,
"cachedInputTokens": 71683,
"totalTokens": 74529,
"reportedCostUsd": 0,
"costStatus": "unpriced"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 97661,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"runs": [
{
"runId": "b4de8fd6-55cd-4416-b08f-b0027e1b24d3",
"usage": {
"model": "unknown",
"biller": "openai",
"provider": "openai",
"costStatus": "unpriced",
"billingType": "metered_api",
"inputTokens": 1248,
"usageSource": "per_run",
"freshSession": false,
"outputTokens": 169,
"sessionReused": true,
"rawInputTokens": 1248,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:ac3b14aa3093739b61b69d65253bdbfdf5fda823cb1b79749f95141758633424",
"changedCategories": [],
"taskSessionReused": true,
"fingerprintVersion": 1,
"taskSessionAvailable": true,
"storedFingerprintPresent": true
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:d2d08ae47e616ba03694d244a03e7ecf6119127e7ed02c128f21b670346db53e",
"workspaceReused": false,
"activeWorkspaceId": "df07fe38-4ddf-42c8-b14c-798930cedb1d",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 169,
"cachedInputTokens": 40377,
"taskSessionReused": true,
"persistedSessionId": "01a0a645-47ea-7aa1-8b06-3455c54d6b90",
"rawCachedInputTokens": 40377,
"sessionRotationReason": null
}
},
{
"runId": "4b477280-d504-413a-8b2f-4ba176be9513",
"usage": {
"model": "unknown",
"biller": "openai",
"provider": "openai",
"costStatus": "unpriced",
"billingType": "metered_api",
"inputTokens": 1390,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 39,
"sessionReused": false,
"rawInputTokens": 1390,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:ac3b14aa3093739b61b69d65253bdbfdf5fda823cb1b79749f95141758633424",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:d2d08ae47e616ba03694d244a03e7ecf6119127e7ed02c128f21b670346db53e",
"workspaceReused": false,
"activeWorkspaceId": "4a7e3e8a-760e-4617-bb93-0508fc72e2f5",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 39,
"cachedInputTokens": 31306,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-47ea-7aa1-8b06-3455c54d6b90",
"rawCachedInputTokens": 31306,
"sessionRotationReason": null
}
}
]
}
}
revise-accept
failed
Tokens886 in · 148 out35,243 cached · 1/1 runs covered
LLM spendunpriced0/1 runs provider-priced
ExecutionLocal · not metered42s agent
first-task.legacy-codex.local.revise-accept
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-ebb9424b6e539765/companies/be8258f7-3e69-4840-997d-539ffca7a54d/acp-engine/agents/987944a0-794c-42e9-a510-496d055f74c6/sessions/paperclip%3Abe8258f7-3e69-4840-997d-539ffca7a54d%3A987944a0-794c-42e9-a510-496d055f74c6%3A4acdb6a6-0542-4503-bf29-dbab266623bf%3Ad80775ef9d6eca2f.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 0,
"inputTokens": 886,
"outputTokens": 148,
"cachedInputTokens": 35243,
"totalTokens": 36277,
"reportedCostUsd": 0,
"costStatus": "unpriced"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 42478,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"model": "unknown",
"biller": "openai",
"provider": "openai",
"costStatus": "unpriced",
"billingType": "metered_api",
"inputTokens": 886,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 148,
"sessionReused": false,
"rawInputTokens": 886,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:cfcc75d16cbad2b43d17e56d8f233dd500e0e4e03f5a5a2ab824c9e07ec3adc7",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:e848cbb43ee9e02b38103bdc442c45a84f86ca6c044504dd838dd99cd0838f07",
"workspaceReused": false,
"activeWorkspaceId": "68dc8a70-ae6c-4b02-bd49-0052fa05e1eb",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 148,
"cachedInputTokens": 35243,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-823b-7010-9d5c-eaca2c38284d",
"rawCachedInputTokens": 35243,
"sessionRotationReason": null
}
}
reject-no-execution
failed
Tokens913 in · 158 out36,557 cached · 1/1 runs covered
LLM spendunpriced0/1 runs provider-priced
ExecutionLocal · not metered48s agent
first-task.legacy-codex.local.reject-no-execution
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-0204c33d3e554d11/companies/ce5ec342-25cf-43ad-b060-add3d67ad27a/acp-engine/agents/754c1b14-967d-4abe-9819-61f2cfe6c29a/sessions/paperclip%3Ace5ec342-25cf-43ad-b060-add3d67ad27a%3A754c1b14-967d-4abe-9819-61f2cfe6c29a%3A2cd5cadd-e1b4-4e41-aab1-f00f4982fb48%3A8c50933accf92f12.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 0,
"inputTokens": 913,
"outputTokens": 158,
"cachedInputTokens": 36557,
"totalTokens": 37628,
"reportedCostUsd": 0,
"costStatus": "unpriced"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 48008,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"model": "unknown",
"biller": "openai",
"provider": "openai",
"costStatus": "unpriced",
"billingType": "metered_api",
"inputTokens": 913,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 158,
"sessionReused": false,
"rawInputTokens": 913,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:08169a778be2b81cfefce32401eeb33775d64c74ba43d304b3fd12f1164c14cc",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:7bc5d5910c0d77097fc52f18999f9a09536ac954e13bb838d90fec03250fad88",
"workspaceReused": false,
"activeWorkspaceId": "5ad8f8de-e52e-4f4c-ae61-f551678fb507",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 158,
"cachedInputTokens": 36557,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-4854-77f2-bad0-715bae013e99",
"rawCachedInputTokens": 36557,
"sessionRotationReason": null
}
} | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Legacy Claudelegacyclaude · claude-sonnet-4-6
|
Isolated locallocal · local
interview-first-response
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.legacy-claude.local.interview-first-response
Matchers and test context
locator.selectOption: Timeout 30000ms exceeded. Call log: [2m - waiting for getByRole('combobox', { name: 'Saved API key' })[22m No matcher result was recorded. Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 0,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"runs": []
}
}
clear-task-first-response
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.legacy-claude.local.clear-task-first-response
Matchers and test context
locator.selectOption: Timeout 30000ms exceeded. Call log: [2m - waiting for getByRole('combobox', { name: 'Saved API key' })[22m No matcher result was recorded. Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 0,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"runs": []
}
}
ambiguous-task-first-response
failed
Tokens18,643 in · 2,766 out154,439 cached · 1/1 runs covered
LLM spend$0.26741/1 runs provider-priced
ExecutionLocal · not metered41s agent
first-task.legacy-claude.local.ambiguous-task-first-response
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-b1b7bc4e8e2dfbbe/companies/400ba5e4-3543-42d2-b4e1-a936a299b172/acp-engine/agents/9b17d996-d71a-426f-a414-8b402ba8e154/sessions/paperclip%3A400ba5e4-3543-42d2-b4e1-a936a299b172%3A9b17d996-d71a-426f-a414-8b402ba8e154%3Aa61be509-4da9-4c1e-9293-8c4d08341e36%3A204927e32c7aabb3.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 18643,
"outputTokens": 2766,
"cachedInputTokens": 154439,
"totalTokens": 175848,
"reportedCostUsd": 0.26737975,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 41428,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0.26737975,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0.26737975,
"complete": true
},
"rawUsage": {
"model": "claude-opus-5",
"biller": "anthropic",
"costUsd": 0.26737975,
"provider": "anthropic",
"costStatus": "reported",
"billingType": "metered_api",
"inputTokens": 18643,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 2766,
"sessionReused": false,
"rawInputTokens": 18643,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:391f4153f51f840980cb1e388dd25d8262f2d2cb1ba1c2b677f08296e7a87bab",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:d3b3f8c233dbd26ebb3b586f308bc89e9e2b6623baa6cd61a5e8b45451f489db",
"workspaceReused": false,
"activeWorkspaceId": "f04c7b21-994f-4296-9dea-16dbfe152adf",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 2766,
"cachedInputTokens": 154439,
"taskSessionReused": false,
"persistedSessionId": "1973bdaf-0c6a-40fe-be64-e687442602eb",
"cacheAdjustedCostUsd": 0.26737975,
"rawCachedInputTokens": 154439,
"sessionRotationReason": null
}
}
plain-message-first-response
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.legacy-claude.local.plain-message-first-response
Matchers and test context
locator.fill: Timeout 30000ms exceeded. Call log: [2m - waiting for getByTestId('task-chat-composer-input').last().locator('[contenteditable="true"], textarea').first()[22m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 0,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"runs": []
}
}
plan-first-response
failed
Tokens41,009 in · 4,150 out295,029 cached · 1/1 runs covered
LLM spend$0.51221/1 runs provider-priced
ExecutionLocal · not metered59s agent
first-task.legacy-claude.local.plan-first-response
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-8d5e0422a17eac56/companies/15194dde-2623-4aed-a5f2-7f825c06f7fc/acp-engine/agents/b740fc73-3510-433b-9318-0fa20d182ab8/sessions/paperclip%3A15194dde-2623-4aed-a5f2-7f825c06f7fc%3Ab740fc73-3510-433b-9318-0fa20d182ab8%3A008a1b49-519f-45b8-825e-f4479383ae12%3A985f15603c2d82ca.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 41009,
"outputTokens": 4150,
"cachedInputTokens": 295029,
"totalTokens": 340188,
"reportedCostUsd": 0.5122027499999999,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 58684,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0.5122027499999999,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0.5122027499999999,
"complete": true
},
"rawUsage": {
"model": "claude-opus-5",
"biller": "anthropic",
"costUsd": 0.5122027499999999,
"provider": "anthropic",
"costStatus": "reported",
"billingType": "metered_api",
"inputTokens": 41009,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 4150,
"sessionReused": false,
"rawInputTokens": 41009,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:ece63c007f3f4e4296924c784a31bb4f6bca7a5ca352fd888d09725a2af92929",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:dfa723baea47cb98db4b10ffdf719b87fd47731864157b13261701f6d756ffaf",
"workspaceReused": false,
"activeWorkspaceId": "8b545b6f-3535-4028-ab1e-9bf552100130",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 4150,
"cachedInputTokens": 295029,
"taskSessionReused": false,
"persistedSessionId": "131850d0-f178-4c77-bd19-51642831bdfc",
"cacheAdjustedCostUsd": 0.5122027499999999,
"rawCachedInputTokens": 295029,
"sessionRotationReason": null
}
}
ordinary-task-control
failed
Tokens36,596 in · 2,245 out143,308 cached · 1/1 runs covered
LLM spend$0.36061/1 runs provider-priced
ExecutionLocal · not metered41s agent
first-task.legacy-claude.local.ordinary-task-control
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-627ea18117c0eb2b/companies/b3bb63b9-05fd-4d77-b18e-1fc4016c279f/acp-engine/agents/6a215fca-4096-4f10-a185-6b8c9759c1dc/sessions/paperclip%3Ab3bb63b9-05fd-4d77-b18e-1fc4016c279f%3A6a215fca-4096-4f10-a185-6b8c9759c1dc%3A30fc3fc3-51a9-4f7f-9d87-2b825f9d1b33%3A2825499e330cc17a.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 36596,
"outputTokens": 2245,
"cachedInputTokens": 143308,
"totalTokens": 182149,
"reportedCostUsd": 0.3606155,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 40890,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0.3606155,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0.3606155,
"complete": true
},
"rawUsage": {
"model": "claude-opus-5",
"biller": "anthropic",
"costUsd": 0.3606155,
"provider": "anthropic",
"costStatus": "reported",
"billingType": "metered_api",
"inputTokens": 36596,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 2245,
"sessionReused": false,
"rawInputTokens": 36596,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": true,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:e1509af373fb7e9b3a0ef8754af26e155a79795098ad8f44df23901a7986f556",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:69ca91363dec106b8b911795af9fd5d10b66140dc1c02c629e48112e0e822962",
"workspaceReused": false,
"activeWorkspaceId": null,
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 2245,
"cachedInputTokens": 143308,
"taskSessionReused": false,
"persistedSessionId": "66fdb81a-46a3-4bc8-bd1d-53c1d36e1cb5",
"cacheAdjustedCostUsd": 0.3606155,
"rawCachedInputTokens": 143308,
"sessionRotationReason": null
}
}
interview-plan-accept
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.legacy-claude.local.interview-plan-accept
Matchers and test context
locator.click: Timeout 30000ms exceeded. Call log: [2m - waiting for getByRole('radio', { name: 'Interview me and propose a plan and an agent team to execute it.', exact: true }).last()[22m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 0,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"runs": []
}
}
task-card-accept
failed
Tokens17,810 in · 3,063 out132,591 cached · 1/1 runs covered
LLM spend$0.25871/1 runs provider-priced
ExecutionLocal · not metered43s agent
first-task.legacy-claude.local.task-card-accept
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-f8b15bb10a201f21/companies/aa448105-db2a-4e6e-9323-6d0fcdca4151/acp-engine/agents/9e9f5e03-8432-401d-ab25-20df50b8bed9/sessions/paperclip%3Aaa448105-db2a-4e6e-9323-6d0fcdca4151%3A9e9f5e03-8432-401d-ab25-20df50b8bed9%3Ac34fd372-6ce0-406a-8b00-47ae849924fc%3A9fcc7de67029b1a1.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 17810,
"outputTokens": 3063,
"cachedInputTokens": 132591,
"totalTokens": 153464,
"reportedCostUsd": 0.258746,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 43005,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0.258746,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0.258746,
"complete": true
},
"rawUsage": {
"model": "claude-opus-5",
"biller": "anthropic",
"costUsd": 0.258746,
"provider": "anthropic",
"costStatus": "reported",
"billingType": "metered_api",
"inputTokens": 17810,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 3063,
"sessionReused": false,
"rawInputTokens": 17810,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:ac23dfbbbac4f11e2b2e113b19db2bdcc8b5cbbc1f95362e41d2285dc068911d",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:0ace2c7019082247412c69c1bbb470b2e2515cbd4e602e41007ead568e7c8074",
"workspaceReused": false,
"activeWorkspaceId": "c615e9fb-bb47-4a7a-a86c-1fd34d353d18",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 3063,
"cachedInputTokens": 132591,
"taskSessionReused": false,
"persistedSessionId": "468f135d-36cd-41ad-9f7b-b313ffe76106",
"cacheAdjustedCostUsd": 0.258746,
"rawCachedInputTokens": 132591,
"sessionRotationReason": null
}
}
task-reply-accept
failed
Tokens39,953 in · 3,810 out246,048 cached · 1/1 runs covered
LLM spend$0.47261/1 runs provider-priced
ExecutionLocal · not metered1m 2s agent
first-task.legacy-claude.local.task-reply-accept
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-0e2a64ae4e986009/companies/51d74995-1cd1-4f54-85de-f34ee8b90472/acp-engine/agents/92b62bae-8ae4-4f4f-ab6d-a2967165ed2b/sessions/paperclip%3A51d74995-1cd1-4f54-85de-f34ee8b90472%3A92b62bae-8ae4-4f4f-ab6d-a2967165ed2b%3A53b5e56b-117b-4bf9-b118-f2223862adb1%3A2a2e6d91d34c0577.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 39953,
"outputTokens": 3810,
"cachedInputTokens": 246048,
"totalTokens": 289811,
"reportedCostUsd": 0.47255375,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 61652,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0.47255375,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0.47255375,
"complete": true
},
"rawUsage": {
"model": "claude-opus-5",
"biller": "anthropic",
"costUsd": 0.47255375,
"provider": "anthropic",
"costStatus": "reported",
"billingType": "metered_api",
"inputTokens": 39953,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 3810,
"sessionReused": false,
"rawInputTokens": 39953,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:51267a34c96ce49bcc4ccbc0764cc30dfb42c36869938ed411afad5c9d275d5a",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:d8ae12111502e55ea01158830fc51975e5e7848d77acf6850a4c7d15c281f900",
"workspaceReused": false,
"activeWorkspaceId": "34a36993-a3df-493d-b746-3a859a27815c",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 3810,
"cachedInputTokens": 246048,
"taskSessionReused": false,
"persistedSessionId": "9f35c2c1-810a-4223-9a57-d56a754c1842",
"cacheAdjustedCostUsd": 0.47255375,
"rawCachedInputTokens": 246048,
"sessionRotationReason": null
}
}
clarify-propose-accept
failed
Tokens95,426 in · 6,927 out492,713 cached · 2/2 runs covered
LLM spend$1.02042/2 runs provider-priced
ExecutionLocal · not metered1m 47s agent
first-task.legacy-claude.local.clarify-propose-accept
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-6e55bb956a88ae23/companies/bb4e6b6e-06c5-4cb4-803e-7580cdf79637/acp-engine/agents/9c884890-54bf-4d21-8f65-61772f073eb2/sessions/paperclip%3Abb4e6b6e-06c5-4cb4-803e-7580cdf79637%3A9c884890-54bf-4d21-8f65-61772f073eb2%3A7c557546-e836-4289-a26b-81cce5390364%3A14c1a635c436f7f4.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 2,
"runsWithTokenUsage": 2,
"runsWithReportedCost": 2,
"inputTokens": 95426,
"outputTokens": 6927,
"cachedInputTokens": 492713,
"totalTokens": 595066,
"reportedCostUsd": 1.020414,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 106604,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 1.020414,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 1.020414,
"complete": true
},
"rawUsage": {
"runs": [
{
"runId": "364807e9-b268-4de9-9bd5-6c4671705c8e",
"usage": {
"model": "claude-opus-5",
"biller": "anthropic",
"costUsd": 0.49596524999999997,
"provider": "anthropic",
"costStatus": "reported",
"billingType": "metered_api",
"inputTokens": 52919,
"usageSource": "per_run",
"freshSession": false,
"outputTokens": 2771,
"sessionReused": true,
"rawInputTokens": 52919,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:845379098730270231652b6da3de3e72bd07157a89bc60780b67b56d41c9878c",
"changedCategories": [],
"taskSessionReused": true,
"fingerprintVersion": 1,
"taskSessionAvailable": true,
"storedFingerprintPresent": true
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:46dfa27fd77972e8840e2f72c9b816e07599bc2d0bc69b058e2bf57e8d3e056a",
"workspaceReused": false,
"activeWorkspaceId": "b1454a11-c930-4475-8178-274e05ca9acd",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 2771,
"cachedInputTokens": 191913,
"taskSessionReused": true,
"persistedSessionId": "b536f86e-dfe4-4db2-a238-80a9d615760d",
"cacheAdjustedCostUsd": 0.49596524999999997,
"rawCachedInputTokens": 191913,
"sessionRotationReason": null
}
},
{
"runId": "4a0ec2b7-ef9a-42b4-ada1-e8e1e17c122a",
"usage": {
"model": "claude-opus-5",
"biller": "anthropic",
"costUsd": 0.52444875,
"provider": "anthropic",
"costStatus": "reported",
"billingType": "metered_api",
"inputTokens": 42507,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 4156,
"sessionReused": false,
"rawInputTokens": 42507,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:845379098730270231652b6da3de3e72bd07157a89bc60780b67b56d41c9878c",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:46dfa27fd77972e8840e2f72c9b816e07599bc2d0bc69b058e2bf57e8d3e056a",
"workspaceReused": false,
"activeWorkspaceId": "fe784f7f-c508-4fb4-9d2b-037163d40e66",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 4156,
"cachedInputTokens": 300800,
"taskSessionReused": false,
"persistedSessionId": "b536f86e-dfe4-4db2-a238-80a9d615760d",
"cacheAdjustedCostUsd": 0.52444875,
"rawCachedInputTokens": 300800,
"sessionRotationReason": null
}
}
]
}
}
revise-accept
failed
Tokens21,418 in · 4,232 out194,919 cached · 1/1 runs covered
LLM spend$0.34171/1 runs provider-priced
ExecutionLocal · not metered1m 8s agent
first-task.legacy-claude.local.revise-accept
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-4687b0953a2074f7/companies/1a924dfe-067f-40fd-afee-4a4da563cf94/acp-engine/agents/25757569-0da3-4e20-aa0e-addac653a46a/sessions/paperclip%3A1a924dfe-067f-40fd-afee-4a4da563cf94%3A25757569-0da3-4e20-aa0e-addac653a46a%3Aeeb6e195-9129-4a28-ab3b-ce2eb828fdc2%3A56df603991fd9f91.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 21418,
"outputTokens": 4232,
"cachedInputTokens": 194919,
"totalTokens": 220569,
"reportedCostUsd": 0.34167100000000006,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 67565,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0.34167100000000006,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0.34167100000000006,
"complete": true
},
"rawUsage": {
"model": "claude-opus-5",
"biller": "anthropic",
"costUsd": 0.34167100000000006,
"provider": "anthropic",
"costStatus": "reported",
"billingType": "metered_api",
"inputTokens": 21418,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 4232,
"sessionReused": false,
"rawInputTokens": 21418,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:26e41c0ee264918beef7e41adc4abff6c2b9e98a9f2bcc159f3672cd91dde3b0",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:b3d2141a0879ee8e08e4e3f60aab8411ec0b7ba21a9d82dc427603a3196a8302",
"workspaceReused": false,
"activeWorkspaceId": "23f90b09-3715-4382-b99b-de24920a67d7",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 4232,
"cachedInputTokens": 194919,
"taskSessionReused": false,
"persistedSessionId": "14d44eb1-bb36-4745-9f41-18cd7ecbb731",
"cacheAdjustedCostUsd": 0.34167100000000006,
"rawCachedInputTokens": 194919,
"sessionRotationReason": null
}
}
reject-no-execution
failed
Tokens40,706 in · 4,309 out223,446 cached · 1/1 runs covered
LLM spend$0.47851/1 runs provider-priced
ExecutionLocal · not metered1m 3s agent
first-task.legacy-claude.local.reject-no-execution
Matchers and test context
Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-ee8a199611fd686d/companies/a4f6799a-b03c-4abc-8a51-1b2f164b8672/acp-engine/agents/c4aa820c-3b23-4cc6-a266-e4f935d58b15/sessions/paperclip%3Aa4f6799a-b03c-4abc-8a51-1b2f164b8672%3Ac4aa820c-3b23-4cc6-a266-e4f935d58b15%3A2018fb36-f9bf-4053-a149-ab5345c89ada%3Abb67268e52a36316.json: exact secret value
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 40706,
"outputTokens": 4309,
"cachedInputTokens": 223446,
"totalTokens": 268461,
"reportedCostUsd": 0.478451,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 63499,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0.478451,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0.478451,
"complete": true
},
"rawUsage": {
"model": "claude-opus-5",
"biller": "anthropic",
"costUsd": 0.478451,
"provider": "anthropic",
"costStatus": "reported",
"billingType": "metered_api",
"inputTokens": 40706,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 4309,
"sessionReused": false,
"rawInputTokens": 40706,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:29af427afc7af6607f8b00ac69ecffb7d596646f4c3c65d8dc316019f9ea2d08",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:37fe63205bdb074e21acb32cc89bba51fd6d36fad3ff924188ce6dba3eb7df11",
"workspaceReused": false,
"activeWorkspaceId": "570d1ab1-408f-4c81-a39c-2d42bcefffc8",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 4309,
"cachedInputTokens": 223446,
"taskSessionReused": false,
"persistedSessionId": "7e29ff74-8a71-487c-8fcd-940922b7a30f",
"cacheAdjustedCostUsd": 0.478451,
"rawCachedInputTokens": 223446,
"sessionRotationReason": null
}
} | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Runner Codexnativecodex · gpt-5.6-sol
|
Isolated locallocal · local
interview-first-response
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.runner-codex.local.interview-first-response
Matchers and test context
locator.click: Timeout 30000ms exceeded. Call log: [2m - waiting for getByRole('radio', { name: 'Interview me and propose a plan and an agent team to execute it.', exact: true }).last()[22m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 0,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"runs": []
}
}
clear-task-first-response
failed
Tokens52,024 in · 785 out34,278 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered23s agent
first-task.runner-codex.local.clear-task-first-response
Matchers and test context
subtask-proposal: Propose a task for the concrete request [2mexpect([22m[31mreceived[39m[2m).[22mtoEqual[2m([22m[32mexpected[39m[2m) // deep equality[22m [32m- Expected - 1[39m [31m+ Received + 10[39m [32m- Array [][39m [31m+ Array [[39m [31m+ Object {[39m [31m+ "detail": "Propose a task for the concrete request",[39m [31m+ "evidence": Array [[39m [31m+ "response-1",[39m [31m+ ],[39m [31m+ "id": "subtask-proposal",[39m [31m+ "passed": false,[39m [31m+ },[39m [31m+ ][39m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 52024,
"outputTokens": 785,
"cachedInputTokens": 34278,
"totalTokens": 87087,
"reportedCostUsd": 0,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 23183,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": true
},
"rawUsage": {
"model": "gpt-5.6-sol",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 52024,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 785,
"sessionReused": false,
"rawInputTokens": 52024,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:cc87224256e6c8c6dc5c7080aca3db96075057b48c64ef62ea62dcdaf12672df",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:b9a7efda3b1f531ea2c04d396559941f86b1f3572711bc67e50ae7d2c154cef8",
"workspaceReused": false,
"activeWorkspaceId": "f580947e-f723-4a31-86ed-dfc89677878f",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 785,
"cachedInputTokens": 34278,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-5a8c-75a2-9d78-00c50d2ccb9e",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 34278,
"sessionRotationReason": null
}
}
ambiguous-task-first-response
passed
Tokens91,198 in · 1,229 out69,944 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered35s agent
first-task.runner-codex.local.ambiguous-task-first-response
Matchers and test context
All invariants passed
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 91198,
"outputTokens": 1229,
"cachedInputTokens": 69944,
"totalTokens": 162371,
"reportedCostUsd": 0,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 35242,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": true
},
"rawUsage": {
"model": "gpt-5.6-sol",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 91198,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 1229,
"sessionReused": false,
"rawInputTokens": 91198,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:636a366ee19c45f9877b76e03382447d645ca3bd542697a6da5626eef6e27253",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:ddf9ea04f08cd099ea64a81d21d3954a710917e2920ddaf2657a2ed1b669a4c3",
"workspaceReused": false,
"activeWorkspaceId": "313b6e18-5423-43c2-b7d9-c9c18ec63c7f",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 1229,
"cachedInputTokens": 69944,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-7680-7ee3-bc28-549d0da35317",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 69944,
"sessionRotationReason": null
}
}
plain-message-first-response
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.runner-codex.local.plain-message-first-response
Matchers and test context
locator.fill: Timeout 30000ms exceeded. Call log: [2m - waiting for getByTestId('task-chat-composer-input').last().locator('[contenteditable="true"], textarea').first()[22m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 0,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"runs": []
}
}
plan-first-response
failed
Tokens89,480 in · 1,268 out70,407 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered26s agent
first-task.runner-codex.local.plan-first-response
Matchers and test context
Behavior failure: work executed before acceptance [2mexpect([22m[31mreceived[39m[2m).[22mtoBe[2m([22m[32mexpected[39m[2m) // Object.is equality[22m Expected: [32mtrue[39m Received: [31mfalse[39m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 89480,
"outputTokens": 1268,
"cachedInputTokens": 70407,
"totalTokens": 161155,
"reportedCostUsd": 0,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 26021,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": true
},
"rawUsage": {
"model": "gpt-5.6-sol",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 89480,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 1268,
"sessionReused": false,
"rawInputTokens": 89480,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:1e52082d998d086b26e9d973a37a8c44e6cc24814a0c0e65044dbf24aa948fa5",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:15e2a1880efb548e9dbd4d27a9286c8d831c6ece152163e4b61dec7f0261cf3d",
"workspaceReused": false,
"activeWorkspaceId": "7d71e12a-8006-4d12-90c3-7da2e7928539",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 1268,
"cachedInputTokens": 70407,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-7dc0-7b31-8d7b-fe727211ca55",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 70407,
"sessionRotationReason": null
}
}
ordinary-task-control
passed
Tokens68,149 in · 1,207 out50,109 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered27s agent
first-task.runner-codex.local.ordinary-task-control
Matchers and test context
All invariants passed
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 68149,
"outputTokens": 1207,
"cachedInputTokens": 50109,
"totalTokens": 119465,
"reportedCostUsd": 0,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 27349,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": true
},
"rawUsage": {
"model": "gpt-5.6-sol",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 68149,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 1207,
"sessionReused": false,
"rawInputTokens": 68149,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": true,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:0df69a820e0a63f18fc8f04800916c140f3a6e69eae285df915faf5f6419a904",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:1f817961c5af5805ffffb879eafc791a9b653efeb86ae5aa31da651bccf9aa22",
"workspaceReused": false,
"activeWorkspaceId": null,
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 1207,
"cachedInputTokens": 50109,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-e5f5-70f0-ae64-7973d4d481e2",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 50109,
"sessionRotationReason": null
}
}
interview-plan-accept
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.runner-codex.local.interview-plan-accept
Matchers and test context
locator.click: Timeout 30000ms exceeded. Call log: [2m - waiting for getByRole('radio', { name: 'Interview me and propose a plan and an agent team to execute it.', exact: true }).last()[22m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 0,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"runs": []
}
}
task-card-accept
failed
Tokens150,312 in · 1,960 out129,015 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered34s agent
first-task.runner-codex.local.task-card-accept
Matchers and test context
Behavior failure: work executed before acceptance [2mexpect([22m[31mreceived[39m[2m).[22mtoBe[2m([22m[32mexpected[39m[2m) // Object.is equality[22m Expected: [32mtrue[39m Received: [31mfalse[39m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 150312,
"outputTokens": 1960,
"cachedInputTokens": 129015,
"totalTokens": 281287,
"reportedCostUsd": 0,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 33761,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": true
},
"rawUsage": {
"model": "gpt-5.6-sol",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 150312,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 1960,
"sessionReused": false,
"rawInputTokens": 150312,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:1d969433d4fae29ae4f76ecf4561f86d59c54f22082537ed39bd70ac675dbf31",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:5108134be7b93436c0bfa607a0a2e36411cece1c0db12fa7ed9e23f8716740e5",
"workspaceReused": false,
"activeWorkspaceId": "a4eeb3d1-8552-40d5-b75f-38f1e66ffc08",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 1960,
"cachedInputTokens": 129015,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-b091-7e80-9d3b-d8c646737410",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 129015,
"sessionRotationReason": null
}
}
task-reply-accept
failed
Tokens69,887 in · 928 out51,787 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered20s agent
first-task.runner-codex.local.task-reply-accept
Matchers and test context
locator.fill: Timeout 30000ms exceeded. Call log: [2m - waiting for getByTestId('task-chat-composer-input').last().locator('[contenteditable="true"], textarea').first()[22m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 69887,
"outputTokens": 928,
"cachedInputTokens": 51787,
"totalTokens": 122602,
"reportedCostUsd": 0,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 19608,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": true
},
"rawUsage": {
"model": "gpt-5.6-sol",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 69887,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 928,
"sessionReused": false,
"rawInputTokens": 69887,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:1236ad3714ef2d2e61da18ee24dca41c2875a8cbb7e963af37eab769c9d1cf56",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:7877e8ac8a09abdab499b7bcf3c0c2e8134874e117afed244e321f0f456ae0d3",
"workspaceReused": false,
"activeWorkspaceId": "4dc07a0f-3b6b-4492-bf06-510157752477",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 928,
"cachedInputTokens": 51787,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-7f43-7060-b3a8-d1cc79a5220a",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 51787,
"sessionRotationReason": null
}
}
clarify-propose-accept
failed
Tokens184,619 in · 1,976 out143,336 cached · 2/2 runs covered
LLM spend$0.0000002/2 runs provider-priced
ExecutionLocal · not metered52s agent
first-task.runner-codex.local.clarify-propose-accept
Matchers and test context
locator.fill: Timeout 30000ms exceeded. Call log: [2m - waiting for getByTestId('task-chat-composer-input').last().locator('[contenteditable="true"], textarea').first()[22m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 2,
"runsWithTokenUsage": 2,
"runsWithReportedCost": 2,
"inputTokens": 184619,
"outputTokens": 1976,
"cachedInputTokens": 143336,
"totalTokens": 329931,
"reportedCostUsd": 0,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 51881,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": true
},
"rawUsage": {
"runs": [
{
"runId": "087fb8ce-37a3-488c-8f33-e531ec70a243",
"usage": {
"model": "gpt-5.6-sol",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 93290,
"usageSource": "per_run",
"freshSession": false,
"outputTokens": 752,
"sessionReused": true,
"rawInputTokens": 93290,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:44b95fc6dbc51e60c9cae76ed0920936505cc7bfab7a2e305094df6a8c1ea282",
"changedCategories": [],
"taskSessionReused": true,
"fingerprintVersion": 1,
"taskSessionAvailable": true,
"storedFingerprintPresent": true
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:ee139c9229e056331a47b9f78448afac4fddec02feb1d6e2e9eddf04789bc3ff",
"workspaceReused": false,
"activeWorkspaceId": "81422b54-804e-4719-bfca-9e38f6840126",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 752,
"cachedInputTokens": 73259,
"taskSessionReused": true,
"persistedSessionId": "01a0a645-f141-79f3-af27-4116b52a62af",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 73259,
"sessionRotationReason": null
}
},
{
"runId": "6f3a2a2c-ffb7-432f-8430-c6f870bb2dfa",
"usage": {
"model": "gpt-5.6-sol",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 91329,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 1224,
"sessionReused": false,
"rawInputTokens": 91329,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:44b95fc6dbc51e60c9cae76ed0920936505cc7bfab7a2e305094df6a8c1ea282",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:ee139c9229e056331a47b9f78448afac4fddec02feb1d6e2e9eddf04789bc3ff",
"workspaceReused": false,
"activeWorkspaceId": "4ba12799-c694-4b62-88c1-a14a64b7f59a",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 1224,
"cachedInputTokens": 70077,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-7e8b-7f83-9476-d5dfecbddf3b",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 70077,
"sessionRotationReason": null
}
}
]
}
}
revise-accept
failed
Tokens107,256 in · 1,486 out88,178 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered29s agent
first-task.runner-codex.local.revise-accept
Matchers and test context
Behavior failure: work executed before acceptance [2mexpect([22m[31mreceived[39m[2m).[22mtoBe[2m([22m[32mexpected[39m[2m) // Object.is equality[22m Expected: [32mtrue[39m Received: [31mfalse[39m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 107256,
"outputTokens": 1486,
"cachedInputTokens": 88178,
"totalTokens": 196920,
"reportedCostUsd": 0,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 29211,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": true
},
"rawUsage": {
"model": "gpt-5.6-sol",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 107256,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 1486,
"sessionReused": false,
"rawInputTokens": 107256,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:1e4091ddbf7b55f692aa4dafa4cfa62c285618cd75a43784b33ab38cb9b2e64a",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:0ea73aabdeb71e9b482837e4f40856dedd8ef45496c3888a19fc457adc22e805",
"workspaceReused": false,
"activeWorkspaceId": "0e84e89d-a6c2-4a29-aec8-6020820c5fa8",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 1486,
"cachedInputTokens": 88178,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-a5fd-7612-8953-fe71a10d481d",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 88178,
"sessionRotationReason": null
}
}
reject-no-execution
failed
Tokens69,434 in · 596 out51,578 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered24s agent
first-task.runner-codex.local.reject-no-execution
Matchers and test context
locator.fill: Timeout 30000ms exceeded. Call log: [2m - waiting for getByTestId('task-chat-composer-input').last().locator('[contenteditable="true"], textarea').first()[22m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 1,
"runsWithReportedCost": 1,
"inputTokens": 69434,
"outputTokens": 596,
"cachedInputTokens": 51578,
"totalTokens": 121608,
"reportedCostUsd": 0,
"costStatus": "reported"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 23758,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": true
},
"rawUsage": {
"model": "gpt-5.6-sol",
"biller": "openai",
"costUsd": 0,
"provider": "openai",
"costStatus": "reported",
"billingType": "unknown",
"inputTokens": 69434,
"usageSource": "per_run",
"freshSession": true,
"outputTokens": 596,
"sessionReused": false,
"rawInputTokens": 69434,
"sessionRotated": false,
"configFreshness": {
"session": {
"reset": false,
"categories": [
"adapter",
"adapterConfig",
"agentRuntimeConfig",
"instructions",
"issueOverrides",
"workspaceConfig",
"environment",
"envBindings",
"secrets",
"runtimeSkills"
],
"resetReasons": [],
"nextFingerprint": "v1:sha256:84a8f67bff3efdfb463314b637937a4bef35050f9eccb29d7ffdc4c658fd4b65",
"changedCategories": [],
"taskSessionReused": false,
"fingerprintVersion": 1,
"taskSessionAvailable": false,
"storedFingerprintPresent": false
},
"version": 1,
"workspace": {
"action": "create",
"reasons": [],
"categories": [
"mode",
"projectWorkspace",
"strategy",
"repo",
"lifecycleCommands",
"runtimeServices",
"environment",
"realization"
],
"reuseRequested": false,
"nextFingerprint": "v1:sha256:8386b55a6d0b046a64209819149c5877ae5d86ee4b39a33f137d7e8ff8718d5f",
"workspaceReused": false,
"activeWorkspaceId": "c3dff590-19ae-4154-85d2-c611e8c06346",
"changedCategories": [],
"storedFingerprint": null,
"fingerprintVersion": 1,
"inferredFingerprint": null,
"previousWorkspaceId": null,
"configSnapshotRefreshed": false,
"storedFingerprintPresent": false
}
},
"rawOutputTokens": 596,
"cachedInputTokens": 51578,
"taskSessionReused": false,
"persistedSessionId": "01a0a645-867c-7c50-9e22-39c58dac964f",
"cacheAdjustedCostUsd": 0,
"rawCachedInputTokens": 51578,
"sessionRotationReason": null
}
} | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Runner ACPX Claudenativeacpx · claude-sonnet-5
|
Isolated locallocal · local
interview-first-response
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.runner-acpx-claude.local.interview-first-response
Matchers and test context
locator.click: Timeout 30000ms exceeded. Call log: [2m - waiting for getByRole('radio', { name: 'Interview me and propose a plan and an agent team to execute it.', exact: true }).last()[22m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 0,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"runs": []
}
}
clear-task-first-response
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered10s agent
first-task.runner-acpx-claude.local.clear-task-first-response
Matchers and test context
Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 9914,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": null
}
ambiguous-task-first-response
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered10s agent
first-task.runner-acpx-claude.local.ambiguous-task-first-response
Matchers and test context
Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 10480,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": null
}
plain-message-first-response
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.runner-acpx-claude.local.plain-message-first-response
Matchers and test context
locator.fill: Timeout 30000ms exceeded. Call log: [2m - waiting for getByTestId('task-chat-composer-input').last().locator('[contenteditable="true"], textarea').first()[22m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 0,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"runs": []
}
}
plan-first-response
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered9s agent
first-task.runner-acpx-claude.local.plan-first-response
Matchers and test context
Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 9424,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": null
}
ordinary-task-control
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered13s agent
first-task.runner-acpx-claude.local.ordinary-task-control
Matchers and test context
Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 12896,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": null
}
interview-plan-accept
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.runner-acpx-claude.local.interview-plan-accept
Matchers and test context
locator.click: Timeout 30000ms exceeded. Call log: [2m - waiting for getByRole('radio', { name: 'Interview me and propose a plan and an agent team to execute it.', exact: true }).last()[22m
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 0,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": {
"runs": []
}
}
task-card-accept
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered8s agent
first-task.runner-acpx-claude.local.task-card-accept
Matchers and test context
Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 8142,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": null
}
task-reply-accept
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered8s agent
first-task.runner-acpx-claude.local.task-reply-accept
Matchers and test context
Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 7659,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": null
}
clarify-propose-accept
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered10s agent
first-task.runner-acpx-claude.local.clarify-propose-accept
Matchers and test context
Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 10051,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": null
}
revise-accept
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered9s agent
first-task.runner-acpx-claude.local.revise-accept
Matchers and test context
Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 8612,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": null
}
reject-no-execution
failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered9s agent
first-task.runner-acpx-claude.local.reject-no-execution
Matchers and test context
Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}
Usage and billing metadata{
"billing": {
"llm": {
"runCount": 1,
"runsWithTokenUsage": 0,
"runsWithReportedCost": 0,
"inputTokens": 0,
"outputTokens": 0,
"cachedInputTokens": 0,
"totalTokens": 0,
"reportedCostUsd": 0,
"costStatus": "unavailable"
},
"runtime": {
"provider": "local",
"agentRunDurationMs": 9135,
"leaseDurationMs": null,
"leaseCount": 0,
"costStatus": "not_metered",
"costSource": "local_not_metered"
},
"reportedCostUsd": 0,
"estimatedRuntimeCostUsd": 0,
"observedAndEstimatedCostUsd": 0,
"complete": false
},
"rawUsage": null
} |
History
No historical campaigns have been published yet.