Paperclip Quality engineering · Runner acceptance

Full-stack acceptance campaign

Runner Full-Stack E2E

A browser-verified matrix of runner profiles, execution environments, and deterministic task contracts. Declared PNG screenshots and sanitized structured evidence are retained with every published campaign; additional diagnostic evidence remains in the access-controlled workflow artifact.

2/48Passed
46Failed
54m 29sTest time
Runner E2E campaign status summary
1,207,115Input tokens
44,102Output tokens
2,911,625Cached tokens
$3.7120LLM reported subtotal
$0.000000Daytona list estimate
22m 0sAgent execution time
0msDaytona lease time
19/52Runs provider-priced

Model spend is the provider-reported subtotal; unpriced or unavailable runs are excluded, never counted as free. Daytona runtime is a public-list-price estimate from captured lease time and pinned resources, before credits, discounts, storage allowance, or invoice adjustments. Local execution has no external runtime meter.

Test suite

Everyday Paperclip Work

Real user requests, useful downloaded work, and durable continuation using production instructions.

Configuration matrix3 profiles · 2 environments · 0 selected
Agent profileIsolated locallocal · localDaytona warm reusable sandboxdaytona · remote
Runner Codexnativecodex · gpt-5.6-sol
Isolated locallocal · local
Build, download, and revise a project not selected
everyday-workflows.runner-codex.local.build-revise
Matchers and test context

Not selected

No matcher result was recorded.

Delegate implementation and preserve late feedback not selected
everyday-workflows.runner-codex.local.delegate-feedback
Matchers and test context

Not selected

No matcher result was recorded.

Hire one teammate, then reuse that agent not selected
everyday-workflows.runner-codex.local.hire-reuse
Matchers and test context

Not selected

No matcher result was recorded.

Use a connection after approval not selected
everyday-workflows.runner-codex.local.service-approve
Matchers and test context

Not selected

No matcher result was recorded.

Respect a declined tool action not selected
everyday-workflows.runner-codex.local.service-decline
Matchers and test context

Not selected

No matcher result was recorded.

Respect Not now on a new connection not selected
everyday-workflows.runner-codex.local.connection-decline
Matchers and test context

Not selected

No matcher result was recorded.

Recover work after the server restarts not selected
everyday-workflows.runner-codex.local.recover-controller
Matchers and test context

Not selected

No matcher result was recorded.

Stop a task and send a new direction once not selected
everyday-workflows.runner-codex.local.stop-redirect
Matchers and test context

Not selected

No matcher result was recorded.

Daytona warm reusable sandboxdaytona · remote
Build, download, and revise a project not selected
everyday-workflows.runner-codex.daytona.build-revise
Matchers and test context

Not selected

No matcher result was recorded.

Delegate implementation and preserve late feedback not selected
everyday-workflows.runner-codex.daytona.delegate-feedback
Matchers and test context

Not selected

No matcher result was recorded.

Recover work after the server restarts not selected
everyday-workflows.runner-codex.daytona.recover-controller
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Claudenativeacpx · claude-sonnet-5
Isolated locallocal · local
Build, download, and revise a project not selected
everyday-workflows.runner-acpx-claude.local.build-revise
Matchers and test context

Not selected

No matcher result was recorded.

Delegate implementation and preserve late feedback not selected
everyday-workflows.runner-acpx-claude.local.delegate-feedback
Matchers and test context

Not selected

No matcher result was recorded.

Hire one teammate, then reuse that agent not selected
everyday-workflows.runner-acpx-claude.local.hire-reuse
Matchers and test context

Not selected

No matcher result was recorded.

Use a connection after approval not selected
everyday-workflows.runner-acpx-claude.local.service-approve
Matchers and test context

Not selected

No matcher result was recorded.

Respect a declined tool action not selected
everyday-workflows.runner-acpx-claude.local.service-decline
Matchers and test context

Not selected

No matcher result was recorded.

Respect Not now on a new connection not selected
everyday-workflows.runner-acpx-claude.local.connection-decline
Matchers and test context

Not selected

No matcher result was recorded.

Recover work after the server restarts not selected
everyday-workflows.runner-acpx-claude.local.recover-controller
Matchers and test context

Not selected

No matcher result was recorded.

Stop a task and send a new direction once not selected
everyday-workflows.runner-acpx-claude.local.stop-redirect
Matchers and test context

Not selected

No matcher result was recorded.

Daytona warm reusable sandboxdaytona · remote
Build, download, and revise a project not selected
everyday-workflows.runner-acpx-claude.daytona.build-revise
Matchers and test context

Not selected

No matcher result was recorded.

Delegate implementation and preserve late feedback not selected
everyday-workflows.runner-acpx-claude.daytona.delegate-feedback
Matchers and test context

Not selected

No matcher result was recorded.

Recover work after the server restarts not selected
everyday-workflows.runner-acpx-claude.daytona.recover-controller
Matchers and test context

Not selected

No matcher result was recorded.

Runner Codex Mininativecodex · gpt-5.4-mini
Isolated locallocal · local
Build, download, and revise a project not selected
everyday-workflows.runner-codex-mini.local.build-revise
Matchers and test context

Not selected

No matcher result was recorded.

Delegate implementation and preserve late feedback not selected
everyday-workflows.runner-codex-mini.local.delegate-feedback
Matchers and test context

Not selected

No matcher result was recorded.

Hire one teammate, then reuse that agent not selected
everyday-workflows.runner-codex-mini.local.hire-reuse
Matchers and test context

Not selected

No matcher result was recorded.

Use a connection after approval not selected
everyday-workflows.runner-codex-mini.local.service-approve
Matchers and test context

Not selected

No matcher result was recorded.

Respect a declined tool action not selected
everyday-workflows.runner-codex-mini.local.service-decline
Matchers and test context

Not selected

No matcher result was recorded.

Respect Not now on a new connection not selected
everyday-workflows.runner-codex-mini.local.connection-decline
Matchers and test context

Not selected

No matcher result was recorded.

Recover work after the server restarts not selected
everyday-workflows.runner-codex-mini.local.recover-controller
Matchers and test context

Not selected

No matcher result was recorded.

Stop a task and send a new direction once not selected
everyday-workflows.runner-codex-mini.local.stop-redirect
Matchers and test context

Not selected

No matcher result was recorded.

Daytona warm reusable sandboxdaytona · remote

Test suite

Persistent Agent Chat

Task-backed conversations, session resets, and project plan handoff.

Configuration matrix4 profiles · 1 environments · 0 selected
Agent profileIsolated locallocal · local
Legacy Codexlegacycodex · gpt-5.6-sol
Isolated locallocal · local
Conversation continuity across restart not selected
agent-chat.legacy-codex.local.continuity-restart
Matchers and test context

Not selected

No matcher result was recorded.

Fresh context within preserved history not selected
agent-chat.legacy-codex.local.new-session
Matchers and test context

Not selected

No matcher result was recorded.

Stop, reset, and resume not selected
agent-chat.legacy-codex.local.stop-new-resume
Matchers and test context

Not selected

No matcher result was recorded.

Draft, revise, approve, and hand off a plan not selected
agent-chat.legacy-codex.local.plan-handoff
Matchers and test context

Not selected

No matcher result was recorded.

Clarify and reuse an existing project not selected
agent-chat.legacy-codex.local.clarify-reuse
Matchers and test context

Not selected

No matcher result was recorded.

Create a project with multiple repository URLs not selected
agent-chat.legacy-codex.local.multi-repository
Matchers and test context

Not selected

No matcher result was recorded.

Legacy Claudelegacyclaude · claude-sonnet-4-6
Isolated locallocal · local
Conversation continuity across restart not selected
agent-chat.legacy-claude.local.continuity-restart
Matchers and test context

Not selected

No matcher result was recorded.

Fresh context within preserved history not selected
agent-chat.legacy-claude.local.new-session
Matchers and test context

Not selected

No matcher result was recorded.

Stop, reset, and resume not selected
agent-chat.legacy-claude.local.stop-new-resume
Matchers and test context

Not selected

No matcher result was recorded.

Draft, revise, approve, and hand off a plan not selected
agent-chat.legacy-claude.local.plan-handoff
Matchers and test context

Not selected

No matcher result was recorded.

Clarify and reuse an existing project not selected
agent-chat.legacy-claude.local.clarify-reuse
Matchers and test context

Not selected

No matcher result was recorded.

Create a project with multiple repository URLs not selected
agent-chat.legacy-claude.local.multi-repository
Matchers and test context

Not selected

No matcher result was recorded.

Runner Codexnativecodex · gpt-5.6-sol
Isolated locallocal · local
Conversation continuity across restart not selected
agent-chat.runner-codex.local.continuity-restart
Matchers and test context

Not selected

No matcher result was recorded.

Fresh context within preserved history not selected
agent-chat.runner-codex.local.new-session
Matchers and test context

Not selected

No matcher result was recorded.

Stop, reset, and resume not selected
agent-chat.runner-codex.local.stop-new-resume
Matchers and test context

Not selected

No matcher result was recorded.

Draft, revise, approve, and hand off a plan not selected
agent-chat.runner-codex.local.plan-handoff
Matchers and test context

Not selected

No matcher result was recorded.

Clarify and reuse an existing project not selected
agent-chat.runner-codex.local.clarify-reuse
Matchers and test context

Not selected

No matcher result was recorded.

Create a project with multiple repository URLs not selected
agent-chat.runner-codex.local.multi-repository
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Claudenativeacpx · claude-sonnet-5
Isolated locallocal · local
Conversation continuity across restart not selected
agent-chat.runner-acpx-claude.local.continuity-restart
Matchers and test context

Not selected

No matcher result was recorded.

Fresh context within preserved history not selected
agent-chat.runner-acpx-claude.local.new-session
Matchers and test context

Not selected

No matcher result was recorded.

Stop, reset, and resume not selected
agent-chat.runner-acpx-claude.local.stop-new-resume
Matchers and test context

Not selected

No matcher result was recorded.

Draft, revise, approve, and hand off a plan not selected
agent-chat.runner-acpx-claude.local.plan-handoff
Matchers and test context

Not selected

No matcher result was recorded.

Clarify and reuse an existing project not selected
agent-chat.runner-acpx-claude.local.clarify-reuse
Matchers and test context

Not selected

No matcher result was recorded.

Create a project with multiple repository URLs not selected
agent-chat.runner-acpx-claude.local.multi-repository
Matchers and test context

Not selected

No matcher result was recorded.

Test suite

Core Runner Compatibility

Major provider, runtime generation, and execution-environment compatibility.

Configuration matrix7 profiles · 2 environments · 0 selected
Agent profileIsolated locallocal · localDaytona sandboxdaytona · remote
Legacy Codexlegacycodex · gpt-5.6-sol
Isolated locallocal · local
Basic response not selected
core-compatibility.legacy-codex.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-codex.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-codex.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.legacy-codex.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-codex.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-codex.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Legacy Claudelegacyclaude · claude-sonnet-4-6
Isolated locallocal · local
Basic response not selected
core-compatibility.legacy-claude.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-claude.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-claude.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.legacy-claude.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-claude.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-claude.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Legacy OpenCodelegacyopencode · openrouter/deepseek/deepseek-v4-flash-0731
Isolated locallocal · local
Basic response not selected
core-compatibility.legacy-opencode.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-opencode.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-opencode.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.legacy-opencode.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-opencode.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-opencode.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Runner Codexnativecodex · gpt-5.6-sol
Isolated locallocal · local
Basic response not selected
core-compatibility.runner-codex.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-codex.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-codex.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.runner-codex.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-codex.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-codex.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Runner OpenCodenativeopencode · openrouter/deepseek/deepseek-v4-flash-0731
Isolated locallocal · local
Basic response not selected
core-compatibility.runner-opencode.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-opencode.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-opencode.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.runner-opencode.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-opencode.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-opencode.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Claudenativeacpx · claude-sonnet-5
Isolated locallocal · local
Basic response not selected
core-compatibility.runner-acpx-claude.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-acpx-claude.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-acpx-claude.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.runner-acpx-claude.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-acpx-claude.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-acpx-claude.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Codexnativeacpx · gpt-5.6-sol
Isolated locallocal · local
Basic response not selected
core-compatibility.runner-acpx-codex.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-acpx-codex.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-acpx-codex.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.runner-acpx-codex.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-acpx-codex.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-acpx-codex.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Test suite

Local Session Integrity

Structured interaction and continuation qualification for every supported local profile.

Configuration matrix7 profiles · 1 environments · 0 selected
Agent profileIsolated locallocal · local
Legacy Codexlegacycodex · gpt-5.6-sol
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.legacy-codex.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.legacy-codex.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Legacy Claudelegacyclaude · claude-sonnet-4-6
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.legacy-claude.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.legacy-claude.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Legacy OpenCodelegacyopencode · openrouter/deepseek/deepseek-v4-flash-0731
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.legacy-opencode.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.legacy-opencode.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Runner Codexnativecodex · gpt-5.6-sol
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.runner-codex.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.runner-codex.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Runner OpenCodenativeopencode · openrouter/deepseek/deepseek-v4-flash-0731
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.runner-opencode.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.runner-opencode.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Claudenativeacpx · claude-sonnet-5
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.runner-acpx-claude.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Codexnativeacpx · gpt-5.6-sol
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.runner-acpx-codex.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.runner-acpx-codex.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Test suite

OpenRouter Model Breadth

Weekly-ranked tool-capable OpenRouter models through native OpenCode on isolated local workspaces.

Configuration matrix4 profiles · 1 environments · 0 selected
Agent profileIsolated locallocal · local
#1 DeepSeek V4 Flash 0731nativeopencode · openrouter/deepseek/deepseek-v4-flash-0731
Isolated locallocal · local
Hello and complete not selected
openrouter-model-breadth.openrouter-deepseek-deepseek-v4-flash-0731.local.hello-complete
Matchers and test context

Not selected

No matcher result was recorded.

Ask, answer, resume not selected
openrouter-model-breadth.openrouter-deepseek-deepseek-v4-flash-0731.local.question-resume-complete
Matchers and test context

Not selected

No matcher result was recorded.

#3 Tencent HY 3nativeopencode · openrouter/tencent/hy3
Isolated locallocal · local
Hello and complete not selected
openrouter-model-breadth.openrouter-tencent-hy3.local.hello-complete
Matchers and test context

Not selected

No matcher result was recorded.

Ask, answer, resume not selected
openrouter-model-breadth.openrouter-tencent-hy3.local.question-resume-complete
Matchers and test context

Not selected

No matcher result was recorded.

#4 Nemotron 3 Ultra 550B A55B (free)nativeopencode · openrouter/nvidia/nemotron-3-ultra-550b-a55b:free
Isolated locallocal · local
Hello and complete not selected
openrouter-model-breadth.openrouter-nvidia-nemotron-3-ultra-550b-a55b-free.local.hello-complete
Matchers and test context

Not selected

No matcher result was recorded.

Ask, answer, resume not selected
openrouter-model-breadth.openrouter-nvidia-nemotron-3-ultra-550b-a55b-free.local.question-resume-complete
Matchers and test context

Not selected

No matcher result was recorded.

Plan, approve, complete not selected
openrouter-model-breadth.openrouter-nvidia-nemotron-3-ultra-550b-a55b-free.local.plan-approve-complete
Matchers and test context

Not selected

No matcher result was recorded.

#5 GPT-5.6 Lunanativeopencode · openrouter/openai/gpt-5.6-luna
Isolated locallocal · local
Hello and complete not selected
openrouter-model-breadth.openrouter-openai-gpt-5-6-luna.local.hello-complete
Matchers and test context

Not selected

No matcher result was recorded.

Ask, answer, resume not selected
openrouter-model-breadth.openrouter-openai-gpt-5-6-luna.local.question-resume-complete
Matchers and test context

Not selected

No matcher result was recorded.

Plan, approve, complete not selected
openrouter-model-breadth.openrouter-openai-gpt-5-6-luna.local.plan-approve-complete
Matchers and test context

Not selected

No matcher result was recorded.

Test suite

Daytona Warm Continuity

Three browser-driven turns on one reusable Daytona sandbox for legacy and native Codex.

Configuration matrix2 profiles · 1 environments · 0 selected
Agent profileDaytona warm reusable sandboxdaytona · remote
Legacy Codexlegacycodex · gpt-5.6-sol
Daytona warm reusable sandboxdaytona · remote
Warm three-turn workspace continuity not selected
daytona-warm-continuity.legacy-codex.daytona.warm-three-turn
Matchers and test context

Not selected

No matcher result was recorded.

Runner Codexnativecodex · gpt-5.6-sol
Daytona warm reusable sandboxdaytona · remote
Warm three-turn workspace continuity not selected
daytona-warm-continuity.runner-codex.daytona.warm-three-turn
Matchers and test context

Not selected

No matcher result was recorded.

Test suite

first-task

Suite discovered from retained campaign identities; full suite size is not known to this publisher.

Pass rate4.2%2/48 passed
Tokens4,162,8421,207,115 input · 44,102 output
Cost$3.7120reported LLM + runtime estimate
Agent time22m 0s0ms lease
Execution48/489 retries · cleanup passed
Configuration matrix4 profiles · 1 environments · 48 selected
Agent profileIsolated locallocal · local
Legacy Codexlegacycodex · gpt-5.6-sol
Isolated locallocal · local
interview-first-response failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.legacy-codex.local.interview-first-response
Matchers and test context
Attempt
1
Duration
53s
Agent runtime
0ms
Runtime
legacy
Provider
codex
Model
provider-default (unreported)
Issue
FIR-1

locator.click: Timeout 30000ms exceeded. Call log:  - waiting for getByRole('radio', { name: 'Interview me and propose a plan and an agent team to execute it.', exact: true }).last()

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 0,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "runs": []
  }
}
clear-task-first-response failed
Tokens1,089 in · 161 out36,185 cached · 1/1 runs covered
LLM spendunpriced0/1 runs provider-priced
ExecutionLocal · not metered58s agent
first-task.legacy-codex.local.clear-task-first-response
Matchers and test context
Attempt
1
Duration
1m 23s
Agent runtime
58s
Runtime
legacy
Provider
codex
Model
unknown
Issue
FIR-1

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-90f05d1bc6f2edc1/companies/1385351b-da54-474c-b43a-90dc0fb1fef9/acp-engine/agents/dd2e7752-5108-4685-a788-6e9e1f75330e/sessions/paperclip%3A1385351b-da54-474c-b43a-90dc0fb1fef9%3Add2e7752-5108-4685-a788-6e9e1f75330e%3Ab6ffae3b-7be6-4f91-bc34-cacedf75419d%3Af087b779991de06f.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.subtask-proposal","expected":true} Propose a task for the concrete request
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 0,
      "inputTokens": 1089,
      "outputTokens": 161,
      "cachedInputTokens": 36185,
      "totalTokens": 37435,
      "reportedCostUsd": 0,
      "costStatus": "unpriced"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 58493,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "model": "unknown",
    "biller": "openai",
    "provider": "openai",
    "costStatus": "unpriced",
    "billingType": "metered_api",
    "inputTokens": 1089,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 161,
    "sessionReused": false,
    "rawInputTokens": 1089,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:a698c106ccdbf5fdbb366bd374aeaf21c7852d6eb0ce8d23f392ad57b05fb1fd",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:5c9f539236d646fcf3b79dee0afedd865fda731b6ddc44c6f94ea5a7097d108e",
        "workspaceReused": false,
        "activeWorkspaceId": "df7c008d-3e11-4c0f-a569-005239ea4d1a",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 161,
    "cachedInputTokens": 36185,
    "taskSessionReused": false,
    "persistedSessionId": "01a0a644-cff2-7c12-8523-859679b94c2e",
    "rawCachedInputTokens": 36185,
    "sessionRotationReason": null
  }
}
ambiguous-task-first-response failed
Tokens963 in · 36 out30,751 cached · 1/1 runs covered
LLM spendunpriced0/1 runs provider-priced
ExecutionLocal · not metered48s agent
first-task.legacy-codex.local.ambiguous-task-first-response
Matchers and test context
Attempt
1
Duration
1m 14s
Agent runtime
48s
Runtime
legacy
Provider
codex
Model
unknown
Issue
FIR-1

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-062287acecfe8a82/companies/618fbcc5-6f58-47ba-b816-eb1b9561a93d/acp-engine/agents/6d86a02e-e9e7-4092-9302-b0fb124ef29c/sessions/paperclip%3A618fbcc5-6f58-47ba-b816-eb1b9561a93d%3A6d86a02e-e9e7-4092-9302-b0fb124ef29c%3A635fbfaa-302d-451d-8f82-ab2dc03f9a75%3A0857beb4f174ec6b.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.focused-clarification","expected":true} Ask focused questions before proposing ambiguous work
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 0,
      "inputTokens": 963,
      "outputTokens": 36,
      "cachedInputTokens": 30751,
      "totalTokens": 31750,
      "reportedCostUsd": 0,
      "costStatus": "unpriced"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 48118,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "model": "unknown",
    "biller": "openai",
    "provider": "openai",
    "costStatus": "unpriced",
    "billingType": "metered_api",
    "inputTokens": 963,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 36,
    "sessionReused": false,
    "rawInputTokens": 963,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:199733c12706f0e8bfba757c029927f145a08d99341a4b6b1b8879ea930d1d00",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:6c330fad7bab1044a2615e779772062a98a91afb66fcd7adae823f82e5c7dd78",
        "workspaceReused": false,
        "activeWorkspaceId": "bffb3755-97f3-4cfd-9e02-5c11c84afac3",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 36,
    "cachedInputTokens": 30751,
    "taskSessionReused": false,
    "persistedSessionId": "01a0a645-3c8f-7732-a8c5-49f920e3b39a",
    "rawCachedInputTokens": 30751,
    "sessionRotationReason": null
  }
}
plain-message-first-response failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.legacy-codex.local.plain-message-first-response
Matchers and test context
Attempt
1
Duration
57s
Agent runtime
0ms
Runtime
legacy
Provider
codex
Model
provider-default (unreported)
Issue
FIR-1

locator.fill: Timeout 30000ms exceeded. Call log:  - waiting for getByTestId('task-chat-composer-input').last().locator('[contenteditable="true"], textarea').first()

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 0,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "runs": []
  }
}
plan-first-response failed
Tokens1,488 in · 66 out35,208 cached · 1/1 runs covered
LLM spendunpriced0/1 runs provider-priced
ExecutionLocal · not metered53s agent
first-task.legacy-codex.local.plan-first-response
Matchers and test context
Attempt
1
Duration
1m 18s
Agent runtime
53s
Runtime
legacy
Provider
codex
Model
unknown
Issue
FIR-1

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-c507a533c701b6c3/companies/e12f8657-de50-4bce-bd8f-c492d224a6e8/acp-engine/agents/96058ee9-82d3-46f2-9bbd-af85ecd4be3c/sessions/paperclip%3Ae12f8657-de50-4bce-bd8f-c492d224a6e8%3A96058ee9-82d3-46f2-9bbd-af85ecd4be3c%3Aac441456-b1d3-4205-988c-2e91ff0292d7%3Abd594d4831bd3920.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.durable-plan","expected":true} Save the requested plan before acceptance
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 0,
      "inputTokens": 1488,
      "outputTokens": 66,
      "cachedInputTokens": 35208,
      "totalTokens": 36762,
      "reportedCostUsd": 0,
      "costStatus": "unpriced"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 52607,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "model": "unknown",
    "biller": "openai",
    "provider": "openai",
    "costStatus": "unpriced",
    "billingType": "metered_api",
    "inputTokens": 1488,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 66,
    "sessionReused": false,
    "rawInputTokens": 1488,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:bd6025d831361175498d8953d0bd85d6b0ffecd79c86c9b85a6512f589626e4a",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:fbcc74f4c4699a648ebdea90a9efa9a7ff0c613443af030ae035ff67977affca",
        "workspaceReused": false,
        "activeWorkspaceId": "bfd615e6-9e6e-45db-938d-5e9bb406eb6a",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 66,
    "cachedInputTokens": 35208,
    "taskSessionReused": false,
    "persistedSessionId": "01a0a645-248e-7f11-a024-45743c1ebc5c",
    "rawCachedInputTokens": 35208,
    "sessionRotationReason": null
  }
}
ordinary-task-control failed
Tokens1,486 in · 58 out27,720 cached · 1/1 runs covered
LLM spendunpriced0/1 runs provider-priced
ExecutionLocal · not metered31s agent
first-task.legacy-codex.local.ordinary-task-control
Matchers and test context
Attempt
1
Duration
1m 13s
Agent runtime
31s
Runtime
legacy
Provider
codex
Model
unknown
Issue
FIR-2

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-91562c687a3484d3/companies/e29b2b4e-a57f-4811-a6f3-edc0d16d200f/acp-engine/agents/b22f3bca-fa19-4cb0-98f7-3683b191ce29/sessions/paperclip%3Ae29b2b4e-a57f-4811-a6f3-edc0d16d200f%3Ab22f3bca-fa19-4cb0-98f7-3683b191ce29%3Aa46e048c-6617-45df-bc8b-c4a5a8064a32%3Aa3b3f6ed0b3b2eac.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.ordinary-task-control","expected":true} Ordinary work produces output without onboarding questions or delegation
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 0,
      "inputTokens": 1486,
      "outputTokens": 58,
      "cachedInputTokens": 27720,
      "totalTokens": 29264,
      "reportedCostUsd": 0,
      "costStatus": "unpriced"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 31391,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "model": "unknown",
    "biller": "openai",
    "provider": "openai",
    "costStatus": "unpriced",
    "billingType": "metered_api",
    "inputTokens": 1486,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 58,
    "sessionReused": false,
    "rawInputTokens": 1486,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": true,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:e49b25090989e1ca4cf26e039dec95e17c84c625bb4ae791b943ce7a5634b8e7",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:4146699b4295d368d7c0169d2e33bac59f2e7a8217b3af37947183bca1edd697",
        "workspaceReused": false,
        "activeWorkspaceId": null,
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 58,
    "cachedInputTokens": 27720,
    "taskSessionReused": false,
    "persistedSessionId": "01a0a645-b02e-74e3-b247-a7b25a01dbde",
    "rawCachedInputTokens": 27720,
    "sessionRotationReason": null
  }
}
interview-plan-accept failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.legacy-codex.local.interview-plan-accept
Matchers and test context
Attempt
1
Duration
54s
Agent runtime
0ms
Runtime
legacy
Provider
codex
Model
provider-default (unreported)
Issue
FIR-1

locator.click: Timeout 30000ms exceeded. Call log:  - waiting for getByRole('radio', { name: 'Interview me and propose a plan and an agent team to execute it.', exact: true }).last()

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 0,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "runs": []
  }
}
task-card-accept failed
Tokens2,667 in · 167 out30,666 cached · 1/2 runs covered
LLM spendunpriced0/2 runs provider-priced
ExecutionLocal · not metered51s agent
first-task.legacy-codex.local.task-card-accept
Matchers and test context
Attempt
1
Duration
1m 23s
Agent runtime
51s
Runtime
legacy
Provider
codex
Model
unknown
Issue
FIR-1

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-2e5a5d01d0859994/companies/fb364bbb-8a73-4ba0-bb8d-a8b9cb3be4a0/acp-engine/agents/a94fce4c-649e-4b2d-8afe-95084d8cea48/sessions/paperclip%3Afb364bbb-8a73-4ba0-bb8d-a8b9cb3be4a0%3Aa94fce4c-649e-4b2d-8afe-95084d8cea48%3A90fd6d68-53db-4322-b30f-9e4b9fb77c37%3A2f9ba35b5288f1c6.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.subtask-proposal","expected":true} Propose a task for the concrete request
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Fail json_path {"kind":"json_path","path":"firstTask.checks.acceptance-recorded","expected":true} Explicit acceptance is persisted as a user comment or approved confirmation card
Fail json_path {"kind":"json_path","path":"firstTask.checks.one-scoped-subtask","expected":true} Exactly one approved subtask belongs to this onboarding issue and agent
Fail json_path {"kind":"json_path","path":"firstTask.checks.creation-after-acceptance","expected":true} Task creation must follow acceptance, including between checkpoints
Fail json_path {"kind":"json_path","path":"firstTask.checks.durable-completion","expected":true} Approved output is saved on the completed child, with revised scope when applicable
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 2,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 0,
      "inputTokens": 2667,
      "outputTokens": 167,
      "cachedInputTokens": 30666,
      "totalTokens": 33500,
      "reportedCostUsd": 0,
      "costStatus": "unpriced"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 50757,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "runs": [
      {
        "runId": "80119376-fbef-44c0-9524-c3e0ec913e2c",
        "usage": null
      },
      {
        "runId": "75b8d94b-44d1-4815-9dab-0c38550fd369",
        "usage": {
          "model": "unknown",
          "biller": "openai",
          "provider": "openai",
          "costStatus": "unpriced",
          "billingType": "metered_api",
          "inputTokens": 2667,
          "usageSource": "per_run",
          "freshSession": true,
          "outputTokens": 167,
          "sessionReused": false,
          "rawInputTokens": 2667,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:9b0462ec3f30dc5acf9de55aff135f05f3378f3acf7d0301f5187f1639e1bb01",
              "changedCategories": [],
              "taskSessionReused": false,
              "fingerprintVersion": 1,
              "taskSessionAvailable": false,
              "storedFingerprintPresent": false
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:a424f7f516754e9c3db794d8dfe05caa0aa4b6e74ea24c6ae75edf709933d7c4",
              "workspaceReused": false,
              "activeWorkspaceId": "ec68480a-f08e-4475-afa6-3c3057bec18f",
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 167,
          "cachedInputTokens": 30666,
          "taskSessionReused": false,
          "persistedSessionId": "01a0a645-5d81-7c40-80fd-5c051810c431",
          "rawCachedInputTokens": 30666,
          "sessionRotationReason": null
        }
      }
    ]
  }
}
task-reply-accept failed
Tokens1,065 in · 163 out36,487 cached · 1/1 runs covered
LLM spendunpriced0/1 runs provider-priced
ExecutionLocal · not metered51s agent
first-task.legacy-codex.local.task-reply-accept
Matchers and test context
Attempt
1
Duration
1m 48s
Agent runtime
51s
Runtime
legacy
Provider
codex
Model
unknown
Issue
FIR-1

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-9bb0e3e7c918d999/companies/82129711-e2fa-4261-8944-01adf9e8aca5/acp-engine/agents/7bd93cf0-4265-4a6e-a3ec-c9c99c205149/sessions/paperclip%3A82129711-e2fa-4261-8944-01adf9e8aca5%3A7bd93cf0-4265-4a6e-a3ec-c9c99c205149%3Aa8199b8a-6779-4273-8c17-66671df9fed5%3A6bb25da5e6b0c7a0.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.subtask-proposal","expected":true} Propose a task for the concrete request
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Fail json_path {"kind":"json_path","path":"firstTask.checks.acceptance-recorded","expected":true} Explicit acceptance is persisted as a user comment or approved confirmation card
Fail json_path {"kind":"json_path","path":"firstTask.checks.one-scoped-subtask","expected":true} Exactly one approved subtask belongs to this onboarding issue and agent
Fail json_path {"kind":"json_path","path":"firstTask.checks.creation-after-acceptance","expected":true} Task creation must follow acceptance, including between checkpoints
Fail json_path {"kind":"json_path","path":"firstTask.checks.durable-completion","expected":true} Approved output is saved on the completed child, with revised scope when applicable
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 0,
      "inputTokens": 1065,
      "outputTokens": 163,
      "cachedInputTokens": 36487,
      "totalTokens": 37715,
      "reportedCostUsd": 0,
      "costStatus": "unpriced"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 50934,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "model": "unknown",
    "biller": "openai",
    "provider": "openai",
    "costStatus": "unpriced",
    "billingType": "metered_api",
    "inputTokens": 1065,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 163,
    "sessionReused": false,
    "rawInputTokens": 1065,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:bff86fb53cdb4de8bf36f9cdb61520f9b1964bd88a88630582116eee86ea0f7f",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:b8b90ecf51a87fd34328684d00d07b02699e66efe0de279d8a20652fed35a100",
        "workspaceReused": false,
        "activeWorkspaceId": "ccadc74b-b921-40d6-a489-2a5eb30e1cde",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 163,
    "cachedInputTokens": 36487,
    "taskSessionReused": false,
    "persistedSessionId": "01a0a645-4e30-7801-ae58-1013a7d0ca14",
    "rawCachedInputTokens": 36487,
    "sessionRotationReason": null
  }
}
clarify-propose-accept failed
Tokens2,638 in · 208 out71,683 cached · 2/2 runs covered
LLM spendunpriced0/2 runs provider-priced
ExecutionLocal · not metered1m 38s agent
first-task.legacy-codex.local.clarify-propose-accept
Matchers and test context
Attempt
1
Duration
2m 12s
Agent runtime
1m 38s
Runtime
legacy
Provider
codex
Model
unknown
Issue
FIR-1

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-d4b1e193a7778582/companies/a9478ac5-149c-492c-b17a-5e15dd1af913/acp-engine/agents/9394d2fb-4d0d-47ac-8fac-2f366f9be8af/sessions/paperclip%3Aa9478ac5-149c-492c-b17a-5e15dd1af913%3A9394d2fb-4d0d-47ac-8fac-2f366f9be8af%3Af234d0a7-f389-405f-96ff-11745edd6f66%3A04a916ee022e59d4.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.focused-clarification","expected":true} Ask focused questions before proposing ambiguous work
Fail json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Fail json_path {"kind":"json_path","path":"firstTask.checks.acceptance-recorded","expected":true} Explicit acceptance is persisted as a user comment or approved confirmation card
Fail json_path {"kind":"json_path","path":"firstTask.checks.one-scoped-subtask","expected":true} Exactly one approved subtask belongs to this onboarding issue and agent
Fail json_path {"kind":"json_path","path":"firstTask.checks.creation-after-acceptance","expected":true} Task creation must follow acceptance, including between checkpoints
Fail json_path {"kind":"json_path","path":"firstTask.checks.durable-completion","expected":true} Approved output is saved on the completed child, with revised scope when applicable
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 2,
      "runsWithTokenUsage": 2,
      "runsWithReportedCost": 0,
      "inputTokens": 2638,
      "outputTokens": 208,
      "cachedInputTokens": 71683,
      "totalTokens": 74529,
      "reportedCostUsd": 0,
      "costStatus": "unpriced"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 97661,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "runs": [
      {
        "runId": "b4de8fd6-55cd-4416-b08f-b0027e1b24d3",
        "usage": {
          "model": "unknown",
          "biller": "openai",
          "provider": "openai",
          "costStatus": "unpriced",
          "billingType": "metered_api",
          "inputTokens": 1248,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 169,
          "sessionReused": true,
          "rawInputTokens": 1248,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:ac3b14aa3093739b61b69d65253bdbfdf5fda823cb1b79749f95141758633424",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:d2d08ae47e616ba03694d244a03e7ecf6119127e7ed02c128f21b670346db53e",
              "workspaceReused": false,
              "activeWorkspaceId": "df07fe38-4ddf-42c8-b14c-798930cedb1d",
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 169,
          "cachedInputTokens": 40377,
          "taskSessionReused": true,
          "persistedSessionId": "01a0a645-47ea-7aa1-8b06-3455c54d6b90",
          "rawCachedInputTokens": 40377,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "4b477280-d504-413a-8b2f-4ba176be9513",
        "usage": {
          "model": "unknown",
          "biller": "openai",
          "provider": "openai",
          "costStatus": "unpriced",
          "billingType": "metered_api",
          "inputTokens": 1390,
          "usageSource": "per_run",
          "freshSession": true,
          "outputTokens": 39,
          "sessionReused": false,
          "rawInputTokens": 1390,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:ac3b14aa3093739b61b69d65253bdbfdf5fda823cb1b79749f95141758633424",
              "changedCategories": [],
              "taskSessionReused": false,
              "fingerprintVersion": 1,
              "taskSessionAvailable": false,
              "storedFingerprintPresent": false
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:d2d08ae47e616ba03694d244a03e7ecf6119127e7ed02c128f21b670346db53e",
              "workspaceReused": false,
              "activeWorkspaceId": "4a7e3e8a-760e-4617-bb93-0508fc72e2f5",
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 39,
          "cachedInputTokens": 31306,
          "taskSessionReused": false,
          "persistedSessionId": "01a0a645-47ea-7aa1-8b06-3455c54d6b90",
          "rawCachedInputTokens": 31306,
          "sessionRotationReason": null
        }
      }
    ]
  }
}
revise-accept failed
Tokens886 in · 148 out35,243 cached · 1/1 runs covered
LLM spendunpriced0/1 runs provider-priced
ExecutionLocal · not metered42s agent
first-task.legacy-codex.local.revise-accept
Matchers and test context
Attempt
1
Duration
1m 43s
Agent runtime
42s
Runtime
legacy
Provider
codex
Model
unknown
Issue
FIR-1

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-ebb9424b6e539765/companies/be8258f7-3e69-4840-997d-539ffca7a54d/acp-engine/agents/987944a0-794c-42e9-a510-496d055f74c6/sessions/paperclip%3Abe8258f7-3e69-4840-997d-539ffca7a54d%3A987944a0-794c-42e9-a510-496d055f74c6%3A4acdb6a6-0542-4503-bf29-dbab266623bf%3Ad80775ef9d6eca2f.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.subtask-proposal","expected":true} Propose a task for the concrete request
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Fail json_path {"kind":"json_path","path":"firstTask.checks.acceptance-recorded","expected":true} Explicit acceptance is persisted as a user comment or approved confirmation card
Fail json_path {"kind":"json_path","path":"firstTask.checks.one-scoped-subtask","expected":true} Exactly one approved subtask belongs to this onboarding issue and agent
Fail json_path {"kind":"json_path","path":"firstTask.checks.creation-after-acceptance","expected":true} Task creation must follow acceptance, including between checkpoints
Fail json_path {"kind":"json_path","path":"firstTask.checks.durable-completion","expected":true} Approved output is saved on the completed child, with revised scope when applicable
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 0,
      "inputTokens": 886,
      "outputTokens": 148,
      "cachedInputTokens": 35243,
      "totalTokens": 36277,
      "reportedCostUsd": 0,
      "costStatus": "unpriced"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 42478,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "model": "unknown",
    "biller": "openai",
    "provider": "openai",
    "costStatus": "unpriced",
    "billingType": "metered_api",
    "inputTokens": 886,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 148,
    "sessionReused": false,
    "rawInputTokens": 886,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:cfcc75d16cbad2b43d17e56d8f233dd500e0e4e03f5a5a2ab824c9e07ec3adc7",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:e848cbb43ee9e02b38103bdc442c45a84f86ca6c044504dd838dd99cd0838f07",
        "workspaceReused": false,
        "activeWorkspaceId": "68dc8a70-ae6c-4b02-bd49-0052fa05e1eb",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 148,
    "cachedInputTokens": 35243,
    "taskSessionReused": false,
    "persistedSessionId": "01a0a645-823b-7010-9d5c-eaca2c38284d",
    "rawCachedInputTokens": 35243,
    "sessionRotationReason": null
  }
}
reject-no-execution failed
Tokens913 in · 158 out36,557 cached · 1/1 runs covered
LLM spendunpriced0/1 runs provider-priced
ExecutionLocal · not metered48s agent
first-task.legacy-codex.local.reject-no-execution
Matchers and test context
Attempt
1
Duration
1m 48s
Agent runtime
48s
Runtime
legacy
Provider
codex
Model
unknown
Issue
FIR-1

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-0204c33d3e554d11/companies/ce5ec342-25cf-43ad-b060-add3d67ad27a/acp-engine/agents/754c1b14-967d-4abe-9819-61f2cfe6c29a/sessions/paperclip%3Ace5ec342-25cf-43ad-b060-add3d67ad27a%3A754c1b14-967d-4abe-9819-61f2cfe6c29a%3A2cd5cadd-e1b4-4e41-aab1-f00f4982fb48%3A8c50933accf92f12.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.subtask-proposal","expected":true} Propose a task for the concrete request
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Pass json_path {"kind":"json_path","path":"firstTask.checks.rejection-respected","expected":true} Rejected work never executes
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 0,
      "inputTokens": 913,
      "outputTokens": 158,
      "cachedInputTokens": 36557,
      "totalTokens": 37628,
      "reportedCostUsd": 0,
      "costStatus": "unpriced"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 48008,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "model": "unknown",
    "biller": "openai",
    "provider": "openai",
    "costStatus": "unpriced",
    "billingType": "metered_api",
    "inputTokens": 913,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 158,
    "sessionReused": false,
    "rawInputTokens": 913,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:08169a778be2b81cfefce32401eeb33775d64c74ba43d304b3fd12f1164c14cc",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:7bc5d5910c0d77097fc52f18999f9a09536ac954e13bb838d90fec03250fad88",
        "workspaceReused": false,
        "activeWorkspaceId": "5ad8f8de-e52e-4f4c-ae61-f551678fb507",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 158,
    "cachedInputTokens": 36557,
    "taskSessionReused": false,
    "persistedSessionId": "01a0a645-4854-77f2-bad0-715bae013e99",
    "rawCachedInputTokens": 36557,
    "sessionRotationReason": null
  }
}
Legacy Claudelegacyclaude · claude-sonnet-4-6
Isolated locallocal · local
interview-first-response failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.legacy-claude.local.interview-first-response
Matchers and test context
Attempt
1
Duration
43s
Agent runtime
0ms
Runtime
legacy
Provider
claude
Model
claude-sonnet-4-6

locator.selectOption: Timeout 30000ms exceeded. Call log:  - waiting for getByRole('combobox', { name: 'Saved API key' })

No matcher result was recorded.

Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 0,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "runs": []
  }
}
clear-task-first-response failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.legacy-claude.local.clear-task-first-response
Matchers and test context
Attempt
1
Duration
43s
Agent runtime
0ms
Runtime
legacy
Provider
claude
Model
claude-sonnet-4-6

locator.selectOption: Timeout 30000ms exceeded. Call log:  - waiting for getByRole('combobox', { name: 'Saved API key' })

No matcher result was recorded.

Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 0,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "runs": []
  }
}
ambiguous-task-first-response failed
Tokens18,643 in · 2,766 out154,439 cached · 1/1 runs covered
LLM spend$0.26741/1 runs provider-priced
ExecutionLocal · not metered41s agent
first-task.legacy-claude.local.ambiguous-task-first-response
Matchers and test context
Attempt
1
Duration
1m 12s
Agent runtime
41s
Runtime
legacy
Provider
claude
Model
claude-opus-5
Issue
FIR-1

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-b1b7bc4e8e2dfbbe/companies/400ba5e4-3543-42d2-b4e1-a936a299b172/acp-engine/agents/9b17d996-d71a-426f-a414-8b402ba8e154/sessions/paperclip%3A400ba5e4-3543-42d2-b4e1-a936a299b172%3A9b17d996-d71a-426f-a414-8b402ba8e154%3Aa61be509-4da9-4c1e-9293-8c4d08341e36%3A204927e32c7aabb3.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.focused-clarification","expected":true} Ask focused questions before proposing ambiguous work
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 1,
      "inputTokens": 18643,
      "outputTokens": 2766,
      "cachedInputTokens": 154439,
      "totalTokens": 175848,
      "reportedCostUsd": 0.26737975,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 41428,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0.26737975,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0.26737975,
    "complete": true
  },
  "rawUsage": {
    "model": "claude-opus-5",
    "biller": "anthropic",
    "costUsd": 0.26737975,
    "provider": "anthropic",
    "costStatus": "reported",
    "billingType": "metered_api",
    "inputTokens": 18643,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 2766,
    "sessionReused": false,
    "rawInputTokens": 18643,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:391f4153f51f840980cb1e388dd25d8262f2d2cb1ba1c2b677f08296e7a87bab",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:d3b3f8c233dbd26ebb3b586f308bc89e9e2b6623baa6cd61a5e8b45451f489db",
        "workspaceReused": false,
        "activeWorkspaceId": "f04c7b21-994f-4296-9dea-16dbfe152adf",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 2766,
    "cachedInputTokens": 154439,
    "taskSessionReused": false,
    "persistedSessionId": "1973bdaf-0c6a-40fe-be64-e687442602eb",
    "cacheAdjustedCostUsd": 0.26737975,
    "rawCachedInputTokens": 154439,
    "sessionRotationReason": null
  }
}
plain-message-first-response failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.legacy-claude.local.plain-message-first-response
Matchers and test context
Attempt
1
Duration
53s
Agent runtime
0ms
Runtime
legacy
Provider
claude
Model
provider-default (unreported)
Issue
FIR-1

locator.fill: Timeout 30000ms exceeded. Call log:  - waiting for getByTestId('task-chat-composer-input').last().locator('[contenteditable="true"], textarea').first()

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 0,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "runs": []
  }
}
plan-first-response failed
Tokens41,009 in · 4,150 out295,029 cached · 1/1 runs covered
LLM spend$0.51221/1 runs provider-priced
ExecutionLocal · not metered59s agent
first-task.legacy-claude.local.plan-first-response
Matchers and test context
Attempt
1
Duration
1m 25s
Agent runtime
59s
Runtime
legacy
Provider
claude
Model
claude-opus-5
Issue
FIR-1

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-8d5e0422a17eac56/companies/15194dde-2623-4aed-a5f2-7f825c06f7fc/acp-engine/agents/b740fc73-3510-433b-9318-0fa20d182ab8/sessions/paperclip%3A15194dde-2623-4aed-a5f2-7f825c06f7fc%3Ab740fc73-3510-433b-9318-0fa20d182ab8%3A008a1b49-519f-45b8-825e-f4479383ae12%3A985f15603c2d82ca.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.durable-plan","expected":true} Save the requested plan before acceptance
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 1,
      "inputTokens": 41009,
      "outputTokens": 4150,
      "cachedInputTokens": 295029,
      "totalTokens": 340188,
      "reportedCostUsd": 0.5122027499999999,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 58684,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0.5122027499999999,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0.5122027499999999,
    "complete": true
  },
  "rawUsage": {
    "model": "claude-opus-5",
    "biller": "anthropic",
    "costUsd": 0.5122027499999999,
    "provider": "anthropic",
    "costStatus": "reported",
    "billingType": "metered_api",
    "inputTokens": 41009,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 4150,
    "sessionReused": false,
    "rawInputTokens": 41009,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:ece63c007f3f4e4296924c784a31bb4f6bca7a5ca352fd888d09725a2af92929",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:dfa723baea47cb98db4b10ffdf719b87fd47731864157b13261701f6d756ffaf",
        "workspaceReused": false,
        "activeWorkspaceId": "8b545b6f-3535-4028-ab1e-9bf552100130",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 4150,
    "cachedInputTokens": 295029,
    "taskSessionReused": false,
    "persistedSessionId": "131850d0-f178-4c77-bd19-51642831bdfc",
    "cacheAdjustedCostUsd": 0.5122027499999999,
    "rawCachedInputTokens": 295029,
    "sessionRotationReason": null
  }
}
ordinary-task-control failed
Tokens36,596 in · 2,245 out143,308 cached · 1/1 runs covered
LLM spend$0.36061/1 runs provider-priced
ExecutionLocal · not metered41s agent
first-task.legacy-claude.local.ordinary-task-control
Matchers and test context
Attempt
1
Duration
1m 15s
Agent runtime
41s
Runtime
legacy
Provider
claude
Model
claude-opus-5
Issue
FIR-2

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-627ea18117c0eb2b/companies/b3bb63b9-05fd-4d77-b18e-1fc4016c279f/acp-engine/agents/6a215fca-4096-4f10-a185-6b8c9759c1dc/sessions/paperclip%3Ab3bb63b9-05fd-4d77-b18e-1fc4016c279f%3A6a215fca-4096-4f10-a185-6b8c9759c1dc%3A30fc3fc3-51a9-4f7f-9d87-2b825f9d1b33%3A2825499e330cc17a.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.ordinary-task-control","expected":true} Ordinary work produces output without onboarding questions or delegation
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 1,
      "inputTokens": 36596,
      "outputTokens": 2245,
      "cachedInputTokens": 143308,
      "totalTokens": 182149,
      "reportedCostUsd": 0.3606155,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 40890,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0.3606155,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0.3606155,
    "complete": true
  },
  "rawUsage": {
    "model": "claude-opus-5",
    "biller": "anthropic",
    "costUsd": 0.3606155,
    "provider": "anthropic",
    "costStatus": "reported",
    "billingType": "metered_api",
    "inputTokens": 36596,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 2245,
    "sessionReused": false,
    "rawInputTokens": 36596,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": true,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:e1509af373fb7e9b3a0ef8754af26e155a79795098ad8f44df23901a7986f556",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:69ca91363dec106b8b911795af9fd5d10b66140dc1c02c629e48112e0e822962",
        "workspaceReused": false,
        "activeWorkspaceId": null,
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 2245,
    "cachedInputTokens": 143308,
    "taskSessionReused": false,
    "persistedSessionId": "66fdb81a-46a3-4bc8-bd1d-53c1d36e1cb5",
    "cacheAdjustedCostUsd": 0.3606155,
    "rawCachedInputTokens": 143308,
    "sessionRotationReason": null
  }
}
interview-plan-accept failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.legacy-claude.local.interview-plan-accept
Matchers and test context
Attempt
1
Duration
57s
Agent runtime
0ms
Runtime
legacy
Provider
claude
Model
provider-default (unreported)
Issue
FIR-1

locator.click: Timeout 30000ms exceeded. Call log:  - waiting for getByRole('radio', { name: 'Interview me and propose a plan and an agent team to execute it.', exact: true }).last()

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 0,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "runs": []
  }
}
task-card-accept failed
Tokens17,810 in · 3,063 out132,591 cached · 1/1 runs covered
LLM spend$0.25871/1 runs provider-priced
ExecutionLocal · not metered43s agent
first-task.legacy-claude.local.task-card-accept
Matchers and test context
Attempt
1
Duration
1m 15s
Agent runtime
43s
Runtime
legacy
Provider
claude
Model
claude-opus-5
Issue
FIR-1

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-f8b15bb10a201f21/companies/aa448105-db2a-4e6e-9323-6d0fcdca4151/acp-engine/agents/9e9f5e03-8432-401d-ab25-20df50b8bed9/sessions/paperclip%3Aaa448105-db2a-4e6e-9323-6d0fcdca4151%3A9e9f5e03-8432-401d-ab25-20df50b8bed9%3Ac34fd372-6ce0-406a-8b00-47ae849924fc%3A9fcc7de67029b1a1.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Fail json_path {"kind":"json_path","path":"firstTask.checks.subtask-proposal","expected":true} Propose a task for the concrete request
Fail json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Fail json_path {"kind":"json_path","path":"firstTask.checks.acceptance-recorded","expected":true} Explicit acceptance is persisted as a user comment or approved confirmation card
Fail json_path {"kind":"json_path","path":"firstTask.checks.one-scoped-subtask","expected":true} Exactly one approved subtask belongs to this onboarding issue and agent
Fail json_path {"kind":"json_path","path":"firstTask.checks.creation-after-acceptance","expected":true} Task creation must follow acceptance, including between checkpoints
Fail json_path {"kind":"json_path","path":"firstTask.checks.durable-completion","expected":true} Approved output is saved on the completed child, with revised scope when applicable
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 1,
      "inputTokens": 17810,
      "outputTokens": 3063,
      "cachedInputTokens": 132591,
      "totalTokens": 153464,
      "reportedCostUsd": 0.258746,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 43005,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0.258746,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0.258746,
    "complete": true
  },
  "rawUsage": {
    "model": "claude-opus-5",
    "biller": "anthropic",
    "costUsd": 0.258746,
    "provider": "anthropic",
    "costStatus": "reported",
    "billingType": "metered_api",
    "inputTokens": 17810,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 3063,
    "sessionReused": false,
    "rawInputTokens": 17810,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:ac23dfbbbac4f11e2b2e113b19db2bdcc8b5cbbc1f95362e41d2285dc068911d",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:0ace2c7019082247412c69c1bbb470b2e2515cbd4e602e41007ead568e7c8074",
        "workspaceReused": false,
        "activeWorkspaceId": "c615e9fb-bb47-4a7a-a86c-1fd34d353d18",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 3063,
    "cachedInputTokens": 132591,
    "taskSessionReused": false,
    "persistedSessionId": "468f135d-36cd-41ad-9f7b-b313ffe76106",
    "cacheAdjustedCostUsd": 0.258746,
    "rawCachedInputTokens": 132591,
    "sessionRotationReason": null
  }
}
task-reply-accept failed
Tokens39,953 in · 3,810 out246,048 cached · 1/1 runs covered
LLM spend$0.47261/1 runs provider-priced
ExecutionLocal · not metered1m 2s agent
first-task.legacy-claude.local.task-reply-accept
Matchers and test context
Attempt
1
Duration
1m 33s
Agent runtime
1m 2s
Runtime
legacy
Provider
claude
Model
claude-opus-5
Issue
FIR-1

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-0e2a64ae4e986009/companies/51d74995-1cd1-4f54-85de-f34ee8b90472/acp-engine/agents/92b62bae-8ae4-4f4f-ab6d-a2967165ed2b/sessions/paperclip%3A51d74995-1cd1-4f54-85de-f34ee8b90472%3A92b62bae-8ae4-4f4f-ab6d-a2967165ed2b%3A53b5e56b-117b-4bf9-b118-f2223862adb1%3A2a2e6d91d34c0577.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Fail json_path {"kind":"json_path","path":"firstTask.checks.subtask-proposal","expected":true} Propose a task for the concrete request
Fail json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Fail json_path {"kind":"json_path","path":"firstTask.checks.acceptance-recorded","expected":true} Explicit acceptance is persisted as a user comment or approved confirmation card
Fail json_path {"kind":"json_path","path":"firstTask.checks.one-scoped-subtask","expected":true} Exactly one approved subtask belongs to this onboarding issue and agent
Fail json_path {"kind":"json_path","path":"firstTask.checks.creation-after-acceptance","expected":true} Task creation must follow acceptance, including between checkpoints
Fail json_path {"kind":"json_path","path":"firstTask.checks.durable-completion","expected":true} Approved output is saved on the completed child, with revised scope when applicable
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 1,
      "inputTokens": 39953,
      "outputTokens": 3810,
      "cachedInputTokens": 246048,
      "totalTokens": 289811,
      "reportedCostUsd": 0.47255375,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 61652,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0.47255375,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0.47255375,
    "complete": true
  },
  "rawUsage": {
    "model": "claude-opus-5",
    "biller": "anthropic",
    "costUsd": 0.47255375,
    "provider": "anthropic",
    "costStatus": "reported",
    "billingType": "metered_api",
    "inputTokens": 39953,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 3810,
    "sessionReused": false,
    "rawInputTokens": 39953,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:51267a34c96ce49bcc4ccbc0764cc30dfb42c36869938ed411afad5c9d275d5a",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:d8ae12111502e55ea01158830fc51975e5e7848d77acf6850a4c7d15c281f900",
        "workspaceReused": false,
        "activeWorkspaceId": "34a36993-a3df-493d-b746-3a859a27815c",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 3810,
    "cachedInputTokens": 246048,
    "taskSessionReused": false,
    "persistedSessionId": "9f35c2c1-810a-4223-9a57-d56a754c1842",
    "cacheAdjustedCostUsd": 0.47255375,
    "rawCachedInputTokens": 246048,
    "sessionRotationReason": null
  }
}
clarify-propose-accept failed
Tokens95,426 in · 6,927 out492,713 cached · 2/2 runs covered
LLM spend$1.02042/2 runs provider-priced
ExecutionLocal · not metered1m 47s agent
first-task.legacy-claude.local.clarify-propose-accept
Matchers and test context
Attempt
1
Duration
2m 53s
Agent runtime
1m 47s
Runtime
legacy
Provider
claude
Model
claude-opus-5
Issue
FIR-1

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-6e55bb956a88ae23/companies/bb4e6b6e-06c5-4cb4-803e-7580cdf79637/acp-engine/agents/9c884890-54bf-4d21-8f65-61772f073eb2/sessions/paperclip%3Abb4e6b6e-06c5-4cb4-803e-7580cdf79637%3A9c884890-54bf-4d21-8f65-61772f073eb2%3A7c557546-e836-4289-a26b-81cce5390364%3A14c1a635c436f7f4.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.focused-clarification","expected":true} Ask focused questions before proposing ambiguous work
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Fail json_path {"kind":"json_path","path":"firstTask.checks.acceptance-recorded","expected":true} Explicit acceptance is persisted as a user comment or approved confirmation card
Fail json_path {"kind":"json_path","path":"firstTask.checks.one-scoped-subtask","expected":true} Exactly one approved subtask belongs to this onboarding issue and agent
Fail json_path {"kind":"json_path","path":"firstTask.checks.creation-after-acceptance","expected":true} Task creation must follow acceptance, including between checkpoints
Fail json_path {"kind":"json_path","path":"firstTask.checks.durable-completion","expected":true} Approved output is saved on the completed child, with revised scope when applicable
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 2,
      "runsWithTokenUsage": 2,
      "runsWithReportedCost": 2,
      "inputTokens": 95426,
      "outputTokens": 6927,
      "cachedInputTokens": 492713,
      "totalTokens": 595066,
      "reportedCostUsd": 1.020414,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 106604,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 1.020414,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 1.020414,
    "complete": true
  },
  "rawUsage": {
    "runs": [
      {
        "runId": "364807e9-b268-4de9-9bd5-6c4671705c8e",
        "usage": {
          "model": "claude-opus-5",
          "biller": "anthropic",
          "costUsd": 0.49596524999999997,
          "provider": "anthropic",
          "costStatus": "reported",
          "billingType": "metered_api",
          "inputTokens": 52919,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 2771,
          "sessionReused": true,
          "rawInputTokens": 52919,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:845379098730270231652b6da3de3e72bd07157a89bc60780b67b56d41c9878c",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:46dfa27fd77972e8840e2f72c9b816e07599bc2d0bc69b058e2bf57e8d3e056a",
              "workspaceReused": false,
              "activeWorkspaceId": "b1454a11-c930-4475-8178-274e05ca9acd",
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 2771,
          "cachedInputTokens": 191913,
          "taskSessionReused": true,
          "persistedSessionId": "b536f86e-dfe4-4db2-a238-80a9d615760d",
          "cacheAdjustedCostUsd": 0.49596524999999997,
          "rawCachedInputTokens": 191913,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "4a0ec2b7-ef9a-42b4-ada1-e8e1e17c122a",
        "usage": {
          "model": "claude-opus-5",
          "biller": "anthropic",
          "costUsd": 0.52444875,
          "provider": "anthropic",
          "costStatus": "reported",
          "billingType": "metered_api",
          "inputTokens": 42507,
          "usageSource": "per_run",
          "freshSession": true,
          "outputTokens": 4156,
          "sessionReused": false,
          "rawInputTokens": 42507,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:845379098730270231652b6da3de3e72bd07157a89bc60780b67b56d41c9878c",
              "changedCategories": [],
              "taskSessionReused": false,
              "fingerprintVersion": 1,
              "taskSessionAvailable": false,
              "storedFingerprintPresent": false
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:46dfa27fd77972e8840e2f72c9b816e07599bc2d0bc69b058e2bf57e8d3e056a",
              "workspaceReused": false,
              "activeWorkspaceId": "fe784f7f-c508-4fb4-9d2b-037163d40e66",
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 4156,
          "cachedInputTokens": 300800,
          "taskSessionReused": false,
          "persistedSessionId": "b536f86e-dfe4-4db2-a238-80a9d615760d",
          "cacheAdjustedCostUsd": 0.52444875,
          "rawCachedInputTokens": 300800,
          "sessionRotationReason": null
        }
      }
    ]
  }
}
revise-accept failed
Tokens21,418 in · 4,232 out194,919 cached · 1/1 runs covered
LLM spend$0.34171/1 runs provider-priced
ExecutionLocal · not metered1m 8s agent
first-task.legacy-claude.local.revise-accept
Matchers and test context
Attempt
1
Duration
1m 39s
Agent runtime
1m 8s
Runtime
legacy
Provider
claude
Model
claude-opus-5
Issue
FIR-1

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-4687b0953a2074f7/companies/1a924dfe-067f-40fd-afee-4a4da563cf94/acp-engine/agents/25757569-0da3-4e20-aa0e-addac653a46a/sessions/paperclip%3A1a924dfe-067f-40fd-afee-4a4da563cf94%3A25757569-0da3-4e20-aa0e-addac653a46a%3Aeeb6e195-9129-4a28-ab3b-ce2eb828fdc2%3A56df603991fd9f91.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Fail json_path {"kind":"json_path","path":"firstTask.checks.subtask-proposal","expected":true} Propose a task for the concrete request
Fail json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Fail json_path {"kind":"json_path","path":"firstTask.checks.acceptance-recorded","expected":true} Explicit acceptance is persisted as a user comment or approved confirmation card
Fail json_path {"kind":"json_path","path":"firstTask.checks.one-scoped-subtask","expected":true} Exactly one approved subtask belongs to this onboarding issue and agent
Fail json_path {"kind":"json_path","path":"firstTask.checks.creation-after-acceptance","expected":true} Task creation must follow acceptance, including between checkpoints
Fail json_path {"kind":"json_path","path":"firstTask.checks.durable-completion","expected":true} Approved output is saved on the completed child, with revised scope when applicable
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 1,
      "inputTokens": 21418,
      "outputTokens": 4232,
      "cachedInputTokens": 194919,
      "totalTokens": 220569,
      "reportedCostUsd": 0.34167100000000006,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 67565,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0.34167100000000006,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0.34167100000000006,
    "complete": true
  },
  "rawUsage": {
    "model": "claude-opus-5",
    "biller": "anthropic",
    "costUsd": 0.34167100000000006,
    "provider": "anthropic",
    "costStatus": "reported",
    "billingType": "metered_api",
    "inputTokens": 21418,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 4232,
    "sessionReused": false,
    "rawInputTokens": 21418,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:26e41c0ee264918beef7e41adc4abff6c2b9e98a9f2bcc159f3672cd91dde3b0",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:b3d2141a0879ee8e08e4e3f60aab8411ec0b7ba21a9d82dc427603a3196a8302",
        "workspaceReused": false,
        "activeWorkspaceId": "23f90b09-3715-4382-b99b-de24920a67d7",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 4232,
    "cachedInputTokens": 194919,
    "taskSessionReused": false,
    "persistedSessionId": "14d44eb1-bb36-4745-9f41-18cd7ecbb731",
    "cacheAdjustedCostUsd": 0.34167100000000006,
    "rawCachedInputTokens": 194919,
    "sessionRotationReason": null
  }
}
reject-no-execution failed
Tokens40,706 in · 4,309 out223,446 cached · 1/1 runs covered
LLM spend$0.47851/1 runs provider-priced
ExecutionLocal · not metered1m 3s agent
first-task.legacy-claude.local.reject-no-execution
Matchers and test context
Attempt
1
Duration
1m 35s
Agent runtime
1m 3s
Runtime
legacy
Provider
claude
Model
claude-opus-5
Issue
FIR-1

Secret leak in persisted Paperclip home state at paperclip-home/instances/runner-e2e-ee8a199611fd686d/companies/a4f6799a-b03c-4abc-8a51-1b2f164b8672/acp-engine/agents/c4aa820c-3b23-4cc6-a266-e4f935d58b15/sessions/paperclip%3Aa4f6799a-b03c-4abc-8a51-1b2f164b8672%3Ac4aa820c-3b23-4cc6-a266-e4f935d58b15%3A2018fb36-f9bf-4053-a149-ab5345c89ada%3Abb67268e52a36316.json: exact secret value

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Fail json_path {"kind":"json_path","path":"firstTask.checks.subtask-proposal","expected":true} Propose a task for the concrete request
Fail json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Fail json_path {"kind":"json_path","path":"firstTask.checks.rejection-respected","expected":true} Rejected work never executes
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 1,
      "inputTokens": 40706,
      "outputTokens": 4309,
      "cachedInputTokens": 223446,
      "totalTokens": 268461,
      "reportedCostUsd": 0.478451,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 63499,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0.478451,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0.478451,
    "complete": true
  },
  "rawUsage": {
    "model": "claude-opus-5",
    "biller": "anthropic",
    "costUsd": 0.478451,
    "provider": "anthropic",
    "costStatus": "reported",
    "billingType": "metered_api",
    "inputTokens": 40706,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 4309,
    "sessionReused": false,
    "rawInputTokens": 40706,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:29af427afc7af6607f8b00ac69ecffb7d596646f4c3c65d8dc316019f9ea2d08",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:37fe63205bdb074e21acb32cc89bba51fd6d36fad3ff924188ce6dba3eb7df11",
        "workspaceReused": false,
        "activeWorkspaceId": "570d1ab1-408f-4c81-a39c-2d42bcefffc8",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 4309,
    "cachedInputTokens": 223446,
    "taskSessionReused": false,
    "persistedSessionId": "7e29ff74-8a71-487c-8fcd-940922b7a30f",
    "cacheAdjustedCostUsd": 0.478451,
    "rawCachedInputTokens": 223446,
    "sessionRotationReason": null
  }
}
Runner Codexnativecodex · gpt-5.6-sol
Isolated locallocal · local
interview-first-response failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.runner-codex.local.interview-first-response
Matchers and test context
Attempt
1
Duration
52s
Agent runtime
0ms
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
FIR-1

locator.click: Timeout 30000ms exceeded. Call log:  - waiting for getByRole('radio', { name: 'Interview me and propose a plan and an agent team to execute it.', exact: true }).last()

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 0,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "runs": []
  }
}
clear-task-first-response failed
Tokens52,024 in · 785 out34,278 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered23s agent
first-task.runner-codex.local.clear-task-first-response
Matchers and test context
Attempt
1
Duration
50s
Agent runtime
23s
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
FIR-1

subtask-proposal: Propose a task for the concrete request expect(received).toEqual(expected) // deep equality - Expected - 1 + Received + 10 - Array [] + Array [ + Object { + "detail": "Propose a task for the concrete request", + "evidence": Array [ + "response-1", + ], + "id": "subtask-proposal", + "passed": false, + }, + ]

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Fail json_path {"kind":"json_path","path":"firstTask.checks.subtask-proposal","expected":true} Propose a task for the concrete request
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 1,
      "inputTokens": 52024,
      "outputTokens": 785,
      "cachedInputTokens": 34278,
      "totalTokens": 87087,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 23183,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": true
  },
  "rawUsage": {
    "model": "gpt-5.6-sol",
    "biller": "openai",
    "costUsd": 0,
    "provider": "openai",
    "costStatus": "reported",
    "billingType": "unknown",
    "inputTokens": 52024,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 785,
    "sessionReused": false,
    "rawInputTokens": 52024,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:cc87224256e6c8c6dc5c7080aca3db96075057b48c64ef62ea62dcdaf12672df",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:b9a7efda3b1f531ea2c04d396559941f86b1f3572711bc67e50ae7d2c154cef8",
        "workspaceReused": false,
        "activeWorkspaceId": "f580947e-f723-4a31-86ed-dfc89677878f",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 785,
    "cachedInputTokens": 34278,
    "taskSessionReused": false,
    "persistedSessionId": "01a0a645-5a8c-75a2-9d78-00c50d2ccb9e",
    "cacheAdjustedCostUsd": 0,
    "rawCachedInputTokens": 34278,
    "sessionRotationReason": null
  }
}
ambiguous-task-first-response passed
Tokens91,198 in · 1,229 out69,944 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered35s agent
first-task.runner-codex.local.ambiguous-task-first-response
Matchers and test context
Attempt
1
Duration
1m 5s
Agent runtime
35s
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
FIR-1

All invariants passed

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.focused-clarification","expected":true} Ask focused questions before proposing ambiguous work
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 1,
      "inputTokens": 91198,
      "outputTokens": 1229,
      "cachedInputTokens": 69944,
      "totalTokens": 162371,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 35242,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": true
  },
  "rawUsage": {
    "model": "gpt-5.6-sol",
    "biller": "openai",
    "costUsd": 0,
    "provider": "openai",
    "costStatus": "reported",
    "billingType": "unknown",
    "inputTokens": 91198,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 1229,
    "sessionReused": false,
    "rawInputTokens": 91198,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:636a366ee19c45f9877b76e03382447d645ca3bd542697a6da5626eef6e27253",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:ddf9ea04f08cd099ea64a81d21d3954a710917e2920ddaf2657a2ed1b669a4c3",
        "workspaceReused": false,
        "activeWorkspaceId": "313b6e18-5423-43c2-b7d9-c9c18ec63c7f",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 1229,
    "cachedInputTokens": 69944,
    "taskSessionReused": false,
    "persistedSessionId": "01a0a645-7680-7ee3-bc28-549d0da35317",
    "cacheAdjustedCostUsd": 0,
    "rawCachedInputTokens": 69944,
    "sessionRotationReason": null
  }
}
plain-message-first-response failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.runner-codex.local.plain-message-first-response
Matchers and test context
Attempt
1
Duration
55s
Agent runtime
0ms
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
FIR-1

locator.fill: Timeout 30000ms exceeded. Call log:  - waiting for getByTestId('task-chat-composer-input').last().locator('[contenteditable="true"], textarea').first()

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 0,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "runs": []
  }
}
plan-first-response failed
Tokens89,480 in · 1,268 out70,407 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered26s agent
first-task.runner-codex.local.plan-first-response
Matchers and test context
Attempt
1
Duration
57s
Agent runtime
26s
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
FIR-1

Behavior failure: work executed before acceptance expect(received).toBe(expected) // Object.is equality Expected: true Received: false

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Fail json_path {"kind":"json_path","path":"firstTask.checks.durable-plan","expected":true} Save the requested plan before acceptance
Fail json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 1,
      "inputTokens": 89480,
      "outputTokens": 1268,
      "cachedInputTokens": 70407,
      "totalTokens": 161155,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 26021,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": true
  },
  "rawUsage": {
    "model": "gpt-5.6-sol",
    "biller": "openai",
    "costUsd": 0,
    "provider": "openai",
    "costStatus": "reported",
    "billingType": "unknown",
    "inputTokens": 89480,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 1268,
    "sessionReused": false,
    "rawInputTokens": 89480,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:1e52082d998d086b26e9d973a37a8c44e6cc24814a0c0e65044dbf24aa948fa5",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:15e2a1880efb548e9dbd4d27a9286c8d831c6ece152163e4b61dec7f0261cf3d",
        "workspaceReused": false,
        "activeWorkspaceId": "7d71e12a-8006-4d12-90c3-7da2e7928539",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 1268,
    "cachedInputTokens": 70407,
    "taskSessionReused": false,
    "persistedSessionId": "01a0a645-7dc0-7b31-8d7b-fe727211ca55",
    "cacheAdjustedCostUsd": 0,
    "rawCachedInputTokens": 70407,
    "sessionRotationReason": null
  }
}
ordinary-task-control passed
Tokens68,149 in · 1,207 out50,109 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered27s agent
first-task.runner-codex.local.ordinary-task-control
Matchers and test context
Attempt
1
Duration
1m 6s
Agent runtime
27s
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
FIR-2

All invariants passed

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.ordinary-task-control","expected":true} Ordinary work produces output without onboarding questions or delegation
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 1,
      "inputTokens": 68149,
      "outputTokens": 1207,
      "cachedInputTokens": 50109,
      "totalTokens": 119465,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 27349,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": true
  },
  "rawUsage": {
    "model": "gpt-5.6-sol",
    "biller": "openai",
    "costUsd": 0,
    "provider": "openai",
    "costStatus": "reported",
    "billingType": "unknown",
    "inputTokens": 68149,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 1207,
    "sessionReused": false,
    "rawInputTokens": 68149,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": true,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:0df69a820e0a63f18fc8f04800916c140f3a6e69eae285df915faf5f6419a904",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:1f817961c5af5805ffffb879eafc791a9b653efeb86ae5aa31da651bccf9aa22",
        "workspaceReused": false,
        "activeWorkspaceId": null,
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 1207,
    "cachedInputTokens": 50109,
    "taskSessionReused": false,
    "persistedSessionId": "01a0a645-e5f5-70f0-ae64-7973d4d481e2",
    "cacheAdjustedCostUsd": 0,
    "rawCachedInputTokens": 50109,
    "sessionRotationReason": null
  }
}
interview-plan-accept failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.runner-codex.local.interview-plan-accept
Matchers and test context
Attempt
1
Duration
53s
Agent runtime
0ms
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
FIR-1

locator.click: Timeout 30000ms exceeded. Call log:  - waiting for getByRole('radio', { name: 'Interview me and propose a plan and an agent team to execute it.', exact: true }).last()

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 0,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "runs": []
  }
}
task-card-accept failed
Tokens150,312 in · 1,960 out129,015 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered34s agent
first-task.runner-codex.local.task-card-accept
Matchers and test context
Attempt
1
Duration
1m 2s
Agent runtime
34s
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
FIR-1

Behavior failure: work executed before acceptance expect(received).toBe(expected) // Object.is equality Expected: true Received: false

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.subtask-proposal","expected":true} Propose a task for the concrete request
Fail json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Fail json_path {"kind":"json_path","path":"firstTask.checks.acceptance-recorded","expected":true} Explicit acceptance is persisted as a user comment or approved confirmation card
Fail json_path {"kind":"json_path","path":"firstTask.checks.one-scoped-subtask","expected":true} Exactly one approved subtask belongs to this onboarding issue and agent
Fail json_path {"kind":"json_path","path":"firstTask.checks.creation-after-acceptance","expected":true} Task creation must follow acceptance, including between checkpoints
Fail json_path {"kind":"json_path","path":"firstTask.checks.durable-completion","expected":true} Approved output is saved on the completed child, with revised scope when applicable
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 1,
      "inputTokens": 150312,
      "outputTokens": 1960,
      "cachedInputTokens": 129015,
      "totalTokens": 281287,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 33761,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": true
  },
  "rawUsage": {
    "model": "gpt-5.6-sol",
    "biller": "openai",
    "costUsd": 0,
    "provider": "openai",
    "costStatus": "reported",
    "billingType": "unknown",
    "inputTokens": 150312,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 1960,
    "sessionReused": false,
    "rawInputTokens": 150312,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:1d969433d4fae29ae4f76ecf4561f86d59c54f22082537ed39bd70ac675dbf31",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:5108134be7b93436c0bfa607a0a2e36411cece1c0db12fa7ed9e23f8716740e5",
        "workspaceReused": false,
        "activeWorkspaceId": "a4eeb3d1-8552-40d5-b75f-38f1e66ffc08",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 1960,
    "cachedInputTokens": 129015,
    "taskSessionReused": false,
    "persistedSessionId": "01a0a645-b091-7e80-9d3b-d8c646737410",
    "cacheAdjustedCostUsd": 0,
    "rawCachedInputTokens": 129015,
    "sessionRotationReason": null
  }
}
task-reply-accept failed
Tokens69,887 in · 928 out51,787 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered20s agent
first-task.runner-codex.local.task-reply-accept
Matchers and test context
Attempt
1
Duration
1m 18s
Agent runtime
20s
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
FIR-1

locator.fill: Timeout 30000ms exceeded. Call log:  - waiting for getByTestId('task-chat-composer-input').last().locator('[contenteditable="true"], textarea').first()

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Fail json_path {"kind":"json_path","path":"firstTask.checks.subtask-proposal","expected":true} Propose a task for the concrete request
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Fail json_path {"kind":"json_path","path":"firstTask.checks.acceptance-recorded","expected":true} Explicit acceptance is persisted as a user comment or approved confirmation card
Fail json_path {"kind":"json_path","path":"firstTask.checks.one-scoped-subtask","expected":true} Exactly one approved subtask belongs to this onboarding issue and agent
Fail json_path {"kind":"json_path","path":"firstTask.checks.creation-after-acceptance","expected":true} Task creation must follow acceptance, including between checkpoints
Fail json_path {"kind":"json_path","path":"firstTask.checks.durable-completion","expected":true} Approved output is saved on the completed child, with revised scope when applicable
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 1,
      "inputTokens": 69887,
      "outputTokens": 928,
      "cachedInputTokens": 51787,
      "totalTokens": 122602,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 19608,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": true
  },
  "rawUsage": {
    "model": "gpt-5.6-sol",
    "biller": "openai",
    "costUsd": 0,
    "provider": "openai",
    "costStatus": "reported",
    "billingType": "unknown",
    "inputTokens": 69887,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 928,
    "sessionReused": false,
    "rawInputTokens": 69887,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:1236ad3714ef2d2e61da18ee24dca41c2875a8cbb7e963af37eab769c9d1cf56",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:7877e8ac8a09abdab499b7bcf3c0c2e8134874e117afed244e321f0f456ae0d3",
        "workspaceReused": false,
        "activeWorkspaceId": "4dc07a0f-3b6b-4492-bf06-510157752477",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 928,
    "cachedInputTokens": 51787,
    "taskSessionReused": false,
    "persistedSessionId": "01a0a645-7f43-7060-b3a8-d1cc79a5220a",
    "cacheAdjustedCostUsd": 0,
    "rawCachedInputTokens": 51787,
    "sessionRotationReason": null
  }
}
clarify-propose-accept failed
Tokens184,619 in · 1,976 out143,336 cached · 2/2 runs covered
LLM spend$0.0000002/2 runs provider-priced
ExecutionLocal · not metered52s agent
first-task.runner-codex.local.clarify-propose-accept
Matchers and test context
Attempt
1
Duration
1m 51s
Agent runtime
52s
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
FIR-1

locator.fill: Timeout 30000ms exceeded. Call log:  - waiting for getByTestId('task-chat-composer-input').last().locator('[contenteditable="true"], textarea').first()

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Pass json_path {"kind":"json_path","path":"firstTask.checks.focused-clarification","expected":true} Ask focused questions before proposing ambiguous work
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Fail json_path {"kind":"json_path","path":"firstTask.checks.acceptance-recorded","expected":true} Explicit acceptance is persisted as a user comment or approved confirmation card
Fail json_path {"kind":"json_path","path":"firstTask.checks.one-scoped-subtask","expected":true} Exactly one approved subtask belongs to this onboarding issue and agent
Fail json_path {"kind":"json_path","path":"firstTask.checks.creation-after-acceptance","expected":true} Task creation must follow acceptance, including between checkpoints
Fail json_path {"kind":"json_path","path":"firstTask.checks.durable-completion","expected":true} Approved output is saved on the completed child, with revised scope when applicable
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 2,
      "runsWithTokenUsage": 2,
      "runsWithReportedCost": 2,
      "inputTokens": 184619,
      "outputTokens": 1976,
      "cachedInputTokens": 143336,
      "totalTokens": 329931,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 51881,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": true
  },
  "rawUsage": {
    "runs": [
      {
        "runId": "087fb8ce-37a3-488c-8f33-e531ec70a243",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 93290,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 752,
          "sessionReused": true,
          "rawInputTokens": 93290,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:44b95fc6dbc51e60c9cae76ed0920936505cc7bfab7a2e305094df6a8c1ea282",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:ee139c9229e056331a47b9f78448afac4fddec02feb1d6e2e9eddf04789bc3ff",
              "workspaceReused": false,
              "activeWorkspaceId": "81422b54-804e-4719-bfca-9e38f6840126",
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 752,
          "cachedInputTokens": 73259,
          "taskSessionReused": true,
          "persistedSessionId": "01a0a645-f141-79f3-af27-4116b52a62af",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 73259,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "6f3a2a2c-ffb7-432f-8430-c6f870bb2dfa",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 91329,
          "usageSource": "per_run",
          "freshSession": true,
          "outputTokens": 1224,
          "sessionReused": false,
          "rawInputTokens": 91329,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:44b95fc6dbc51e60c9cae76ed0920936505cc7bfab7a2e305094df6a8c1ea282",
              "changedCategories": [],
              "taskSessionReused": false,
              "fingerprintVersion": 1,
              "taskSessionAvailable": false,
              "storedFingerprintPresent": false
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:ee139c9229e056331a47b9f78448afac4fddec02feb1d6e2e9eddf04789bc3ff",
              "workspaceReused": false,
              "activeWorkspaceId": "4ba12799-c694-4b62-88c1-a14a64b7f59a",
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 1224,
          "cachedInputTokens": 70077,
          "taskSessionReused": false,
          "persistedSessionId": "01a0a645-7e8b-7f83-9476-d5dfecbddf3b",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 70077,
          "sessionRotationReason": null
        }
      }
    ]
  }
}
revise-accept failed
Tokens107,256 in · 1,486 out88,178 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered29s agent
first-task.runner-codex.local.revise-accept
Matchers and test context
Attempt
1
Duration
59s
Agent runtime
29s
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
FIR-1

Behavior failure: work executed before acceptance expect(received).toBe(expected) // Object.is equality Expected: true Received: false

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Fail json_path {"kind":"json_path","path":"firstTask.checks.subtask-proposal","expected":true} Propose a task for the concrete request
Fail json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Fail json_path {"kind":"json_path","path":"firstTask.checks.acceptance-recorded","expected":true} Explicit acceptance is persisted as a user comment or approved confirmation card
Fail json_path {"kind":"json_path","path":"firstTask.checks.one-scoped-subtask","expected":true} Exactly one approved subtask belongs to this onboarding issue and agent
Fail json_path {"kind":"json_path","path":"firstTask.checks.creation-after-acceptance","expected":true} Task creation must follow acceptance, including between checkpoints
Fail json_path {"kind":"json_path","path":"firstTask.checks.durable-completion","expected":true} Approved output is saved on the completed child, with revised scope when applicable
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 1,
      "inputTokens": 107256,
      "outputTokens": 1486,
      "cachedInputTokens": 88178,
      "totalTokens": 196920,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 29211,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": true
  },
  "rawUsage": {
    "model": "gpt-5.6-sol",
    "biller": "openai",
    "costUsd": 0,
    "provider": "openai",
    "costStatus": "reported",
    "billingType": "unknown",
    "inputTokens": 107256,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 1486,
    "sessionReused": false,
    "rawInputTokens": 107256,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:1e4091ddbf7b55f692aa4dafa4cfa62c285618cd75a43784b33ab38cb9b2e64a",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:0ea73aabdeb71e9b482837e4f40856dedd8ef45496c3888a19fc457adc22e805",
        "workspaceReused": false,
        "activeWorkspaceId": "0e84e89d-a6c2-4a29-aec8-6020820c5fa8",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 1486,
    "cachedInputTokens": 88178,
    "taskSessionReused": false,
    "persistedSessionId": "01a0a645-a5fd-7612-8953-fe71a10d481d",
    "cacheAdjustedCostUsd": 0,
    "rawCachedInputTokens": 88178,
    "sessionRotationReason": null
  }
}
reject-no-execution failed
Tokens69,434 in · 596 out51,578 cached · 1/1 runs covered
LLM spend$0.0000001/1 runs provider-priced
ExecutionLocal · not metered24s agent
first-task.runner-codex.local.reject-no-execution
Matchers and test context
Attempt
1
Duration
1m 21s
Agent runtime
24s
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
FIR-1

locator.fill: Timeout 30000ms exceeded. Call log:  - waiting for getByTestId('task-chat-composer-input').last().locator('[contenteditable="true"], textarea').first()

ResultMatcherExpectationDetail
Pass json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-reintroduction","expected":true} Do not repeat the seeded welcome
Pass json_path {"kind":"json_path","path":"firstTask.checks.opening-not-repeated","expected":true} Do not post another opening choice
Fail json_path {"kind":"json_path","path":"firstTask.checks.subtask-proposal","expected":true} Propose a task for the concrete request
Pass json_path {"kind":"json_path","path":"firstTask.checks.no-premature-work","expected":true} Before acceptance: no hires, execution tasks, finished output, or claimed completion
Pass json_path {"kind":"json_path","path":"firstTask.checks.rejection-respected","expected":true} Rejected work never executes
Pass json_path {"kind":"json_path","path":"firstTask.checks.provider-runs-succeeded","expected":true} All observed provider runs settled successfully
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 1,
      "runsWithReportedCost": 1,
      "inputTokens": 69434,
      "outputTokens": 596,
      "cachedInputTokens": 51578,
      "totalTokens": 121608,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 23758,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": true
  },
  "rawUsage": {
    "model": "gpt-5.6-sol",
    "biller": "openai",
    "costUsd": 0,
    "provider": "openai",
    "costStatus": "reported",
    "billingType": "unknown",
    "inputTokens": 69434,
    "usageSource": "per_run",
    "freshSession": true,
    "outputTokens": 596,
    "sessionReused": false,
    "rawInputTokens": 69434,
    "sessionRotated": false,
    "configFreshness": {
      "session": {
        "reset": false,
        "categories": [
          "adapter",
          "adapterConfig",
          "agentRuntimeConfig",
          "instructions",
          "issueOverrides",
          "workspaceConfig",
          "environment",
          "envBindings",
          "secrets",
          "runtimeSkills"
        ],
        "resetReasons": [],
        "nextFingerprint": "v1:sha256:84a8f67bff3efdfb463314b637937a4bef35050f9eccb29d7ffdc4c658fd4b65",
        "changedCategories": [],
        "taskSessionReused": false,
        "fingerprintVersion": 1,
        "taskSessionAvailable": false,
        "storedFingerprintPresent": false
      },
      "version": 1,
      "workspace": {
        "action": "create",
        "reasons": [],
        "categories": [
          "mode",
          "projectWorkspace",
          "strategy",
          "repo",
          "lifecycleCommands",
          "runtimeServices",
          "environment",
          "realization"
        ],
        "reuseRequested": false,
        "nextFingerprint": "v1:sha256:8386b55a6d0b046a64209819149c5877ae5d86ee4b39a33f137d7e8ff8718d5f",
        "workspaceReused": false,
        "activeWorkspaceId": "c3dff590-19ae-4154-85d2-c611e8c06346",
        "changedCategories": [],
        "storedFingerprint": null,
        "fingerprintVersion": 1,
        "inferredFingerprint": null,
        "previousWorkspaceId": null,
        "configSnapshotRefreshed": false,
        "storedFingerprintPresent": false
      }
    },
    "rawOutputTokens": 596,
    "cachedInputTokens": 51578,
    "taskSessionReused": false,
    "persistedSessionId": "01a0a645-867c-7c50-9e22-39c58dac964f",
    "cacheAdjustedCostUsd": 0,
    "rawCachedInputTokens": 51578,
    "sessionRotationReason": null
  }
}
Runner ACPX Claudenativeacpx · claude-sonnet-5
Isolated locallocal · local
interview-first-response failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.runner-acpx-claude.local.interview-first-response
Matchers and test context
Attempt
1
Duration
54s
Agent runtime
0ms
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
FIR-1

locator.click: Timeout 30000ms exceeded. Call log:  - waiting for getByRole('radio', { name: 'Interview me and propose a plan and an agent team to execute it.', exact: true }).last()

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 0,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "runs": []
  }
}
clear-task-first-response failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered10s agent
first-task.runner-acpx-claude.local.clear-task-first-response
Matchers and test context
Attempt
2
Duration
38s
Agent runtime
10s
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
FIR-1

Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 9914,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": null
}
ambiguous-task-first-response failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered10s agent
first-task.runner-acpx-claude.local.ambiguous-task-first-response
Matchers and test context
Attempt
2
Duration
39s
Agent runtime
10s
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
FIR-1

Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 10480,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": null
}
plain-message-first-response failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.runner-acpx-claude.local.plain-message-first-response
Matchers and test context
Attempt
1
Duration
54s
Agent runtime
0ms
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
FIR-1

locator.fill: Timeout 30000ms exceeded. Call log:  - waiting for getByTestId('task-chat-composer-input').last().locator('[contenteditable="true"], textarea').first()

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 0,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "runs": []
  }
}
plan-first-response failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered9s agent
first-task.runner-acpx-claude.local.plan-first-response
Matchers and test context
Attempt
2
Duration
36s
Agent runtime
9s
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
FIR-1

Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 9424,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": null
}
ordinary-task-control failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered13s agent
first-task.runner-acpx-claude.local.ordinary-task-control
Matchers and test context
Attempt
2
Duration
53s
Agent runtime
13s
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
FIR-2

Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 12896,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": null
}
interview-plan-accept failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
first-task.runner-acpx-claude.local.interview-plan-accept
Matchers and test context
Attempt
1
Duration
58s
Agent runtime
0ms
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
FIR-1

locator.click: Timeout 30000ms exceeded. Call log:  - waiting for getByRole('radio', { name: 'Interview me and propose a plan and an agent team to execute it.', exact: true }).last()

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 0,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": {
    "runs": []
  }
}
task-card-accept failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered8s agent
first-task.runner-acpx-claude.local.task-card-accept
Matchers and test context
Attempt
2
Duration
35s
Agent runtime
8s
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
FIR-1

Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 8142,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": null
}
task-reply-accept failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered8s agent
first-task.runner-acpx-claude.local.task-reply-accept
Matchers and test context
Attempt
2
Duration
32s
Agent runtime
8s
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
FIR-1

Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 7659,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": null
}
clarify-propose-accept failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered10s agent
first-task.runner-acpx-claude.local.clarify-propose-accept
Matchers and test context
Attempt
2
Duration
38s
Agent runtime
10s
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
FIR-1

Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 10051,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": null
}
revise-accept failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered9s agent
first-task.runner-acpx-claude.local.revise-accept
Matchers and test context
Attempt
2
Duration
35s
Agent runtime
9s
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
FIR-1

Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 8612,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": null
}
reject-no-execution failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered9s agent
first-task.runner-acpx-claude.local.reject-no-execution
Matchers and test context
Attempt
2
Duration
37s
Agent runtime
9s
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
FIR-1

Stopped waiting for first-task response and durable outcome: run status failed: native_session_interrupted PRP command session.open failed: {"commandId":"command_prp_00000002","commandType":"session.open","controllerSeq":2,"result":{"code":"command_execution_failed","message":"failed to start ACPX provider: ACPX sidecar command session.open was rejected (retryable=false, classification=effective_model_mismatch)"},"status":"failed"}

ResultMatcherExpectationDetail
Fail json_path {"kind":"json_path","path":"firstTask.checks.recorded-response","expected":true} An agent response or structured interaction was recorded
Pass json_path {"kind":"json_path","path":"firstTask.checks.instruction-snapshot","expected":true} Actual persona and skill were retained with verified hashes
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 9135,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": null
}

History

Campaign trends

No historical campaigns have been published yet.