Paperclip Quality engineering · Runner acceptance

Full-stack acceptance campaign

Runner Full-Stack E2E

A browser-verified matrix of runner profiles, execution environments, and deterministic task contracts. Declared PNG screenshots and sanitized structured evidence are retained with every published campaign; additional diagnostic evidence remains in the access-controlled workflow artifact.

9/9Passed
0Failed
13m 56sTest time
Runner E2E campaign status summary
1,070,622Input tokens
16,365Output tokens
2,054,926Cached tokens
$0.000000LLM reported subtotal
$0.0280Daytona list estimate
7m 47sAgent execution time
6m 16sDaytona lease time
23/23Runs provider-priced

Model spend is the provider-reported subtotal; unpriced or unavailable runs are excluded, never counted as free. Daytona runtime is a public-list-price estimate from captured lease time and pinned resources, before credits, discounts, storage allowance, or invoice adjustments. Local execution has no external runtime meter.

Test suite

Task continuation

Human direction, approval boundaries, untrusted evidence, and completed actions across turns.

Configuration matrix4 profiles · 1 environments · 0 selected
Agent profileIsolated locallocal · local
Legacy Codexlegacycodex · gpt-5.6-sol
Isolated locallocal · local
answer updates scope not selected
continuation.legacy-codex.local.answer-updates-scope
Matchers and test context

Not selected

No matcher result was recorded.

clarification not approval not selected
continuation.legacy-codex.local.clarification-not-approval
Matchers and test context

Not selected

No matcher result was recorded.

revision preserves approval not selected
continuation.legacy-codex.local.revision-preserves-approval
Matchers and test context

Not selected

No matcher result was recorded.

untrusted evidence not selected
continuation.legacy-codex.local.untrusted-evidence
Matchers and test context

Not selected

No matcher result was recorded.

completed action resume not selected
continuation.legacy-codex.local.completed-action-resume
Matchers and test context

Not selected

No matcher result was recorded.

Legacy Claudelegacyclaude · claude-sonnet-4-6
Isolated locallocal · local
answer updates scope not selected
continuation.legacy-claude.local.answer-updates-scope
Matchers and test context

Not selected

No matcher result was recorded.

clarification not approval not selected
continuation.legacy-claude.local.clarification-not-approval
Matchers and test context

Not selected

No matcher result was recorded.

revision preserves approval not selected
continuation.legacy-claude.local.revision-preserves-approval
Matchers and test context

Not selected

No matcher result was recorded.

untrusted evidence not selected
continuation.legacy-claude.local.untrusted-evidence
Matchers and test context

Not selected

No matcher result was recorded.

completed action resume not selected
continuation.legacy-claude.local.completed-action-resume
Matchers and test context

Not selected

No matcher result was recorded.

Runner Codexnativecodex · gpt-5.6-sol
Isolated locallocal · local
answer updates scope not selected
continuation.runner-codex.local.answer-updates-scope
Matchers and test context

Not selected

No matcher result was recorded.

clarification not approval not selected
continuation.runner-codex.local.clarification-not-approval
Matchers and test context

Not selected

No matcher result was recorded.

revision preserves approval not selected
continuation.runner-codex.local.revision-preserves-approval
Matchers and test context

Not selected

No matcher result was recorded.

untrusted evidence not selected
continuation.runner-codex.local.untrusted-evidence
Matchers and test context

Not selected

No matcher result was recorded.

completed action resume not selected
continuation.runner-codex.local.completed-action-resume
Matchers and test context

Not selected

No matcher result was recorded.

question tool documentation not selected
continuation.runner-codex.local.question-tool-documentation
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Claudenativeacpx · claude-sonnet-5
Isolated locallocal · local
answer updates scope not selected
continuation.runner-acpx-claude.local.answer-updates-scope
Matchers and test context

Not selected

No matcher result was recorded.

clarification not approval not selected
continuation.runner-acpx-claude.local.clarification-not-approval
Matchers and test context

Not selected

No matcher result was recorded.

revision preserves approval not selected
continuation.runner-acpx-claude.local.revision-preserves-approval
Matchers and test context

Not selected

No matcher result was recorded.

untrusted evidence not selected
continuation.runner-acpx-claude.local.untrusted-evidence
Matchers and test context

Not selected

No matcher result was recorded.

completed action resume not selected
continuation.runner-acpx-claude.local.completed-action-resume
Matchers and test context

Not selected

No matcher result was recorded.

question tool documentation not selected
continuation.runner-acpx-claude.local.question-tool-documentation
Matchers and test context

Not selected

No matcher result was recorded.

provider question bridge not selected
continuation.runner-acpx-claude.local.provider-question-bridge
Matchers and test context

Not selected

No matcher result was recorded.

Test suite

Everyday Paperclip Work

Real user requests, useful downloaded work, and durable continuation using production instructions.

Configuration matrix3 profiles · 2 environments · 0 selected
Agent profileIsolated locallocal · localDaytona warm reusable sandboxdaytona · remote
Runner Codexnativecodex · gpt-5.6-sol
Isolated locallocal · local
Build, download, and revise a project not selected
everyday-workflows.runner-codex.local.build-revise
Matchers and test context

Not selected

No matcher result was recorded.

Delegate implementation and preserve late feedback not selected
everyday-workflows.runner-codex.local.delegate-feedback
Matchers and test context

Not selected

No matcher result was recorded.

Delegate work through an agent review handoff not selected
everyday-workflows.runner-codex.local.agent-review-handoff
Matchers and test context

Not selected

No matcher result was recorded.

Hire one teammate, then reuse that agent not selected
everyday-workflows.runner-codex.local.hire-reuse
Matchers and test context

Not selected

No matcher result was recorded.

Use a connection after approval not selected
everyday-workflows.runner-codex.local.service-approve
Matchers and test context

Not selected

No matcher result was recorded.

Respect a declined tool action not selected
everyday-workflows.runner-codex.local.service-decline
Matchers and test context

Not selected

No matcher result was recorded.

Respect Not now on a new connection not selected
everyday-workflows.runner-codex.local.connection-decline
Matchers and test context

Not selected

No matcher result was recorded.

Recover work after the server restarts not selected
everyday-workflows.runner-codex.local.recover-controller
Matchers and test context

Not selected

No matcher result was recorded.

Stop a task and send a new direction once not selected
everyday-workflows.runner-codex.local.stop-redirect
Matchers and test context

Not selected

No matcher result was recorded.

Create and edit a company skill not selected
everyday-workflows.runner-codex.local.create-skill-studio
Matchers and test context

Not selected

No matcher result was recorded.

Daytona warm reusable sandboxdaytona · remote
Build, download, and revise a project not selected
everyday-workflows.runner-codex.daytona.build-revise
Matchers and test context

Not selected

No matcher result was recorded.

Delegate implementation and preserve late feedback not selected
everyday-workflows.runner-codex.daytona.delegate-feedback
Matchers and test context

Not selected

No matcher result was recorded.

Recover work after the server restarts not selected
everyday-workflows.runner-codex.daytona.recover-controller
Matchers and test context

Not selected

No matcher result was recorded.

Create and edit a company skill not selected
everyday-workflows.runner-codex.daytona.create-skill-studio
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Claudenativeacpx · claude-sonnet-5
Isolated locallocal · local
Build, download, and revise a project not selected
everyday-workflows.runner-acpx-claude.local.build-revise
Matchers and test context

Not selected

No matcher result was recorded.

Delegate implementation and preserve late feedback not selected
everyday-workflows.runner-acpx-claude.local.delegate-feedback
Matchers and test context

Not selected

No matcher result was recorded.

Delegate work through an agent review handoff not selected
everyday-workflows.runner-acpx-claude.local.agent-review-handoff
Matchers and test context

Not selected

No matcher result was recorded.

Hire one teammate, then reuse that agent not selected
everyday-workflows.runner-acpx-claude.local.hire-reuse
Matchers and test context

Not selected

No matcher result was recorded.

Use a connection after approval not selected
everyday-workflows.runner-acpx-claude.local.service-approve
Matchers and test context

Not selected

No matcher result was recorded.

Respect a declined tool action not selected
everyday-workflows.runner-acpx-claude.local.service-decline
Matchers and test context

Not selected

No matcher result was recorded.

Respect Not now on a new connection not selected
everyday-workflows.runner-acpx-claude.local.connection-decline
Matchers and test context

Not selected

No matcher result was recorded.

Recover work after the server restarts not selected
everyday-workflows.runner-acpx-claude.local.recover-controller
Matchers and test context

Not selected

No matcher result was recorded.

Stop a task and send a new direction once not selected
everyday-workflows.runner-acpx-claude.local.stop-redirect
Matchers and test context

Not selected

No matcher result was recorded.

Create and edit a company skill not selected
everyday-workflows.runner-acpx-claude.local.create-skill-studio
Matchers and test context

Not selected

No matcher result was recorded.

Daytona warm reusable sandboxdaytona · remote
Build, download, and revise a project not selected
everyday-workflows.runner-acpx-claude.daytona.build-revise
Matchers and test context

Not selected

No matcher result was recorded.

Delegate implementation and preserve late feedback not selected
everyday-workflows.runner-acpx-claude.daytona.delegate-feedback
Matchers and test context

Not selected

No matcher result was recorded.

Recover work after the server restarts not selected
everyday-workflows.runner-acpx-claude.daytona.recover-controller
Matchers and test context

Not selected

No matcher result was recorded.

Create and edit a company skill not selected
everyday-workflows.runner-acpx-claude.daytona.create-skill-studio
Matchers and test context

Not selected

No matcher result was recorded.

Runner Codex Mininativecodex · gpt-5.4-mini
Isolated locallocal · local
Build, download, and revise a project not selected
everyday-workflows.runner-codex-mini.local.build-revise
Matchers and test context

Not selected

No matcher result was recorded.

Delegate implementation and preserve late feedback not selected
everyday-workflows.runner-codex-mini.local.delegate-feedback
Matchers and test context

Not selected

No matcher result was recorded.

Delegate work through an agent review handoff not selected
everyday-workflows.runner-codex-mini.local.agent-review-handoff
Matchers and test context

Not selected

No matcher result was recorded.

Hire one teammate, then reuse that agent not selected
everyday-workflows.runner-codex-mini.local.hire-reuse
Matchers and test context

Not selected

No matcher result was recorded.

Use a connection after approval not selected
everyday-workflows.runner-codex-mini.local.service-approve
Matchers and test context

Not selected

No matcher result was recorded.

Respect a declined tool action not selected
everyday-workflows.runner-codex-mini.local.service-decline
Matchers and test context

Not selected

No matcher result was recorded.

Respect Not now on a new connection not selected
everyday-workflows.runner-codex-mini.local.connection-decline
Matchers and test context

Not selected

No matcher result was recorded.

Recover work after the server restarts not selected
everyday-workflows.runner-codex-mini.local.recover-controller
Matchers and test context

Not selected

No matcher result was recorded.

Stop a task and send a new direction once not selected
everyday-workflows.runner-codex-mini.local.stop-redirect
Matchers and test context

Not selected

No matcher result was recorded.

Create and edit a company skill not selected
everyday-workflows.runner-codex-mini.local.create-skill-studio
Matchers and test context

Not selected

No matcher result was recorded.

Daytona warm reusable sandboxdaytona · remote

Test suite

First-task onboarding

Production onboarding, first replies, approval, and durable task execution.

Configuration matrix4 profiles · 1 environments · 0 selected
Agent profileIsolated locallocal · local
Legacy Codexlegacycodex · Production onboarding default
Isolated locallocal · local
Interview: first response not selected
first-task.legacy-codex.local.interview-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Clear task: first response not selected
first-task.legacy-codex.local.clear-task-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Ambiguous task: first response not selected
first-task.legacy-codex.local.ambiguous-task-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Opening card replaced by a message not selected
first-task.legacy-codex.local.plain-message-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Explicit plan: first response not selected
first-task.legacy-codex.local.plan-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Other tasks do not inherit onboarding policy not selected
first-task.legacy-codex.local.ordinary-task-control
Matchers and test context

Not selected

No matcher result was recorded.

Interview, plan, and acceptance not selected
first-task.legacy-codex.local.interview-plan-accept
Matchers and test context

Not selected

No matcher result was recorded.

Subtask accepted through a card not selected
first-task.legacy-codex.local.task-card-accept
Matchers and test context

Not selected

No matcher result was recorded.

Accept a proposal while its agent is still running not selected
first-task.legacy-codex.local.accept-while-running
Matchers and test context

Not selected

No matcher result was recorded.

Subtask accepted in conversation not selected
first-task.legacy-codex.local.task-reply-accept
Matchers and test context

Not selected

No matcher result was recorded.

Clarification, proposal, and acceptance not selected
first-task.legacy-codex.local.clarify-propose-accept
Matchers and test context

Not selected

No matcher result was recorded.

Revise scope before accepting not selected
first-task.legacy-codex.local.revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Decline proposed work not selected
first-task.legacy-codex.local.reject-no-execution
Matchers and test context

Not selected

No matcher result was recorded.

Legacy Claudelegacyclaude · Production onboarding default
Isolated locallocal · local
Interview: first response not selected
first-task.legacy-claude.local.interview-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Clear task: first response not selected
first-task.legacy-claude.local.clear-task-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Ambiguous task: first response not selected
first-task.legacy-claude.local.ambiguous-task-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Opening card replaced by a message not selected
first-task.legacy-claude.local.plain-message-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Explicit plan: first response not selected
first-task.legacy-claude.local.plan-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Other tasks do not inherit onboarding policy not selected
first-task.legacy-claude.local.ordinary-task-control
Matchers and test context

Not selected

No matcher result was recorded.

Interview, plan, and acceptance not selected
first-task.legacy-claude.local.interview-plan-accept
Matchers and test context

Not selected

No matcher result was recorded.

Subtask accepted through a card not selected
first-task.legacy-claude.local.task-card-accept
Matchers and test context

Not selected

No matcher result was recorded.

Accept a proposal while its agent is still running not selected
first-task.legacy-claude.local.accept-while-running
Matchers and test context

Not selected

No matcher result was recorded.

Subtask accepted in conversation not selected
first-task.legacy-claude.local.task-reply-accept
Matchers and test context

Not selected

No matcher result was recorded.

Clarification, proposal, and acceptance not selected
first-task.legacy-claude.local.clarify-propose-accept
Matchers and test context

Not selected

No matcher result was recorded.

Revise scope before accepting not selected
first-task.legacy-claude.local.revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Decline proposed work not selected
first-task.legacy-claude.local.reject-no-execution
Matchers and test context

Not selected

No matcher result was recorded.

Runner Codexnativecodex · Production onboarding default
Isolated locallocal · local
Interview: first response not selected
first-task.runner-codex.local.interview-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Clear task: first response not selected
first-task.runner-codex.local.clear-task-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Ambiguous task: first response not selected
first-task.runner-codex.local.ambiguous-task-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Opening card replaced by a message not selected
first-task.runner-codex.local.plain-message-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Explicit plan: first response not selected
first-task.runner-codex.local.plan-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Other tasks do not inherit onboarding policy not selected
first-task.runner-codex.local.ordinary-task-control
Matchers and test context

Not selected

No matcher result was recorded.

Interview, plan, and acceptance not selected
first-task.runner-codex.local.interview-plan-accept
Matchers and test context

Not selected

No matcher result was recorded.

Subtask accepted through a card not selected
first-task.runner-codex.local.task-card-accept
Matchers and test context

Not selected

No matcher result was recorded.

Accept a proposal while its agent is still running not selected
first-task.runner-codex.local.accept-while-running
Matchers and test context

Not selected

No matcher result was recorded.

Subtask accepted in conversation not selected
first-task.runner-codex.local.task-reply-accept
Matchers and test context

Not selected

No matcher result was recorded.

Clarification, proposal, and acceptance not selected
first-task.runner-codex.local.clarify-propose-accept
Matchers and test context

Not selected

No matcher result was recorded.

Revise scope before accepting not selected
first-task.runner-codex.local.revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Decline proposed work not selected
first-task.runner-codex.local.reject-no-execution
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Claudenativeacpx · Production onboarding default
Isolated locallocal · local
Interview: first response not selected
first-task.runner-acpx-claude.local.interview-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Clear task: first response not selected
first-task.runner-acpx-claude.local.clear-task-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Ambiguous task: first response not selected
first-task.runner-acpx-claude.local.ambiguous-task-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Opening card replaced by a message not selected
first-task.runner-acpx-claude.local.plain-message-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Explicit plan: first response not selected
first-task.runner-acpx-claude.local.plan-first-response
Matchers and test context

Not selected

No matcher result was recorded.

Other tasks do not inherit onboarding policy not selected
first-task.runner-acpx-claude.local.ordinary-task-control
Matchers and test context

Not selected

No matcher result was recorded.

Interview, plan, and acceptance not selected
first-task.runner-acpx-claude.local.interview-plan-accept
Matchers and test context

Not selected

No matcher result was recorded.

Subtask accepted through a card not selected
first-task.runner-acpx-claude.local.task-card-accept
Matchers and test context

Not selected

No matcher result was recorded.

Accept a proposal while its agent is still running not selected
first-task.runner-acpx-claude.local.accept-while-running
Matchers and test context

Not selected

No matcher result was recorded.

Subtask accepted in conversation not selected
first-task.runner-acpx-claude.local.task-reply-accept
Matchers and test context

Not selected

No matcher result was recorded.

Clarification, proposal, and acceptance not selected
first-task.runner-acpx-claude.local.clarify-propose-accept
Matchers and test context

Not selected

No matcher result was recorded.

Revise scope before accepting not selected
first-task.runner-acpx-claude.local.revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Decline proposed work not selected
first-task.runner-acpx-claude.local.reject-no-execution
Matchers and test context

Not selected

No matcher result was recorded.

Test suite

Persistent Agent Chat

Task-backed conversations, session resets, and project plan handoff.

Pass rate100.0%1/1 passed
Tokens400,623229,717 input · 2,145 output
Cost$0.000000reported LLM + runtime estimate
Agent time30s0ms lease
Execution1/10 retries · cleanup passed
Configuration matrix4 profiles · 1 environments · 1 selected
Agent profileIsolated locallocal · local
Legacy Codexlegacycodex · gpt-5.6-sol
Isolated locallocal · local
Conversation continuity across restart not selected
agent-chat.legacy-codex.local.continuity-restart
Matchers and test context

Not selected

No matcher result was recorded.

Fresh context within preserved history not selected
agent-chat.legacy-codex.local.new-session
Matchers and test context

Not selected

No matcher result was recorded.

Stop, reset, and resume not selected
agent-chat.legacy-codex.local.stop-new-resume
Matchers and test context

Not selected

No matcher result was recorded.

Draft, revise, approve, and hand off a plan not selected
agent-chat.legacy-codex.local.plan-handoff
Matchers and test context

Not selected

No matcher result was recorded.

Clarify and reuse an existing project not selected
agent-chat.legacy-codex.local.clarify-reuse
Matchers and test context

Not selected

No matcher result was recorded.

Create a project with multiple repository URLs not selected
agent-chat.legacy-codex.local.multi-repository
Matchers and test context

Not selected

No matcher result was recorded.

Legacy Claudelegacyclaude · claude-sonnet-4-6
Isolated locallocal · local
Conversation continuity across restart not selected
agent-chat.legacy-claude.local.continuity-restart
Matchers and test context

Not selected

No matcher result was recorded.

Fresh context within preserved history not selected
agent-chat.legacy-claude.local.new-session
Matchers and test context

Not selected

No matcher result was recorded.

Stop, reset, and resume not selected
agent-chat.legacy-claude.local.stop-new-resume
Matchers and test context

Not selected

No matcher result was recorded.

Draft, revise, approve, and hand off a plan not selected
agent-chat.legacy-claude.local.plan-handoff
Matchers and test context

Not selected

No matcher result was recorded.

Clarify and reuse an existing project not selected
agent-chat.legacy-claude.local.clarify-reuse
Matchers and test context

Not selected

No matcher result was recorded.

Create a project with multiple repository URLs not selected
agent-chat.legacy-claude.local.multi-repository
Matchers and test context

Not selected

No matcher result was recorded.

Runner Codexnativecodex · gpt-5.6-sol
Isolated locallocal · local
Save a planned task without starting work not selected
agent-chat.runner-codex.local.create-backlog
Matchers and test context

Not selected

No matcher result was recorded.

Reassign existing work and preserve queued context not selected
agent-chat.runner-codex.local.reassign-task
Matchers and test context

Not selected

No matcher result was recorded.

Conversation continuity across restart passed
Overall passed · 1/1 behavioral checks passed

No behavioral matcher failed.

See all behavioral checks (1)
ResultCheckDetail
Passissue_statusChat workflow and durable handoff/session assertions passed
Tokens229,717 in · 2,145 out168,761 cached · 3/3 runs covered
LLM spend$0.0000003/3 runs provider-priced
ExecutionLocal · not metered30s agent
agent-chat.runner-codex.local.continuity-restart
Matchers and test context
Attempt
1
Duration
1m 12s
Agent runtime
30s
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
RUN-1

All invariants passed

ResultMatcherExpectationDetail
Pass issue_status {"kind":"issue_status","expected":"in_review"} Chat workflow and durable handoff/session assertions passed
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 3,
      "runsWithTokenUsage": 3,
      "runsWithReportedCost": 3,
      "inputTokens": 229717,
      "outputTokens": 2145,
      "cachedInputTokens": 168761,
      "totalTokens": 400623,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 30244,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": true
  },
  "rawUsage": {
    "runs": [
      {
        "runId": "c9e4e557-60a0-4d65-9695-3cab2e2af795",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 120602,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 1086,
          "sessionReused": true,
          "rawInputTokens": 120602,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:e4033b99fbb6cfea24282f78ebc5b49389ab230b42a864b001bff949685c5f28",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:588730cdd8e7b5e6e16728902960e8ca96213faa05b9b5f4cab08b8d34888e6e",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 1086,
          "cachedInputTokens": 97392,
          "taskSessionReused": true,
          "persistedSessionId": "01a0c581-2a18-75c0-9860-4e08ca78ecc9",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 97392,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "594b3fe7-17eb-44f5-8f2d-854319055a8e",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 74637,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 715,
          "sessionReused": true,
          "rawInputTokens": 74637,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:e4033b99fbb6cfea24282f78ebc5b49389ab230b42a864b001bff949685c5f28",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:588730cdd8e7b5e6e16728902960e8ca96213faa05b9b5f4cab08b8d34888e6e",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 715,
          "cachedInputTokens": 54331,
          "taskSessionReused": true,
          "persistedSessionId": "01a0c581-2a18-75c0-9860-4e08ca78ecc9",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 54331,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "630b9a68-ec05-4889-91de-fbeb6c77b056",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 34478,
          "usageSource": "per_run",
          "freshSession": true,
          "outputTokens": 344,
          "sessionReused": false,
          "rawInputTokens": 34478,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:e4033b99fbb6cfea24282f78ebc5b49389ab230b42a864b001bff949685c5f28",
              "changedCategories": [],
              "taskSessionReused": false,
              "fingerprintVersion": 1,
              "taskSessionAvailable": false,
              "storedFingerprintPresent": false
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:588730cdd8e7b5e6e16728902960e8ca96213faa05b9b5f4cab08b8d34888e6e",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 344,
          "cachedInputTokens": 17038,
          "taskSessionReused": false,
          "persistedSessionId": "01a0c581-2a18-75c0-9860-4e08ca78ecc9",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 17038,
          "sessionRotationReason": null
        }
      }
    ]
  }
}
Fresh context within preserved history not selected
agent-chat.runner-codex.local.new-session
Matchers and test context

Not selected

No matcher result was recorded.

Stop, reset, and resume not selected
agent-chat.runner-codex.local.stop-new-resume
Matchers and test context

Not selected

No matcher result was recorded.

Draft, revise, approve, and hand off a plan not selected
agent-chat.runner-codex.local.plan-handoff
Matchers and test context

Not selected

No matcher result was recorded.

Clarify and reuse an existing project not selected
agent-chat.runner-codex.local.clarify-reuse
Matchers and test context

Not selected

No matcher result was recorded.

Create a project with multiple repository URLs not selected
agent-chat.runner-codex.local.multi-repository
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Claudenativeacpx · claude-sonnet-5
Isolated locallocal · local
Save a planned task without starting work not selected
agent-chat.runner-acpx-claude.local.create-backlog
Matchers and test context

Not selected

No matcher result was recorded.

Reassign existing work and preserve queued context not selected
agent-chat.runner-acpx-claude.local.reassign-task
Matchers and test context

Not selected

No matcher result was recorded.

Conversation continuity across restart not selected
agent-chat.runner-acpx-claude.local.continuity-restart
Matchers and test context

Not selected

No matcher result was recorded.

Fresh context within preserved history not selected
agent-chat.runner-acpx-claude.local.new-session
Matchers and test context

Not selected

No matcher result was recorded.

Stop, reset, and resume not selected
agent-chat.runner-acpx-claude.local.stop-new-resume
Matchers and test context

Not selected

No matcher result was recorded.

Draft, revise, approve, and hand off a plan not selected
agent-chat.runner-acpx-claude.local.plan-handoff
Matchers and test context

Not selected

No matcher result was recorded.

Clarify and reuse an existing project not selected
agent-chat.runner-acpx-claude.local.clarify-reuse
Matchers and test context

Not selected

No matcher result was recorded.

Create a project with multiple repository URLs not selected
agent-chat.runner-acpx-claude.local.multi-repository
Matchers and test context

Not selected

No matcher result was recorded.

Test suite

Agent Chat Recovery and Coordination

Native chat startup cancellation, committed sends, hiring, grounded status, and remote continuity.

Pass rate100.0%8/8 passed
Tokens2,741,290840,905 input · 14,220 output
Cost$0.0280reported LLM + runtime estimate
Agent time7m 17s6m 16s lease
Execution8/80 retries · cleanup passed
Configuration matrix2 profiles · 2 environments · 8 selected
Agent profileIsolated locallocal · localDaytona warm reusable sandboxdaytona · remote
Runner Codexnativecodex · gpt-5.6-sol
Isolated locallocal · local
Stop during startup, reset, and resume not selected
agent-chat-hardening.runner-codex.local.stop-startup-new-resume
Matchers and test context

Not selected

No matcher result was recorded.

Hire through chat, delegate, and reuse the same teammate not selected
agent-chat-hardening.runner-codex.local.hire-delegate-reuse
Matchers and test context

Not selected

No matcher result was recorded.

Read the actual blocker and hand source material to a reviewer not selected
agent-chat-hardening.runner-codex.local.blocked-status-review
Matchers and test context

Not selected

No matcher result was recorded.

Recover a lost send acknowledgement without repeating committed work passed
Overall passed · 1/1 behavioral checks passed

No behavioral matcher failed.

See all behavioral checks (1)
ResultCheckDetail
Passissue_statusChat workflow and durable handoff/session assertions passed
Tokens256,791 in · 2,145 out213,676 cached · 2/2 runs covered
LLM spend$0.0000002/2 runs provider-priced
ExecutionLocal · not metered41s agent
agent-chat-hardening.runner-codex.local.committed-send-retry
Matchers and test context
Attempt
1
Duration
1m 12s
Agent runtime
41s
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
RUN-1

All invariants passed

ResultMatcherExpectationDetail
Pass issue_status {"kind":"issue_status","expected":"in_review"} Chat workflow and durable handoff/session assertions passed
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 2,
      "runsWithTokenUsage": 2,
      "runsWithReportedCost": 2,
      "inputTokens": 256791,
      "outputTokens": 2145,
      "cachedInputTokens": 213676,
      "totalTokens": 472612,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 41102,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": true
  },
  "rawUsage": {
    "runs": [
      {
        "runId": "cf855b56-8475-4717-8dfb-2f775d15c0d0",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 162539,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 1250,
          "sessionReused": true,
          "rawInputTokens": 162539,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:406c81bdfd13dcff2edf3857f6563cbd2c84b311b5eff281a619d300252c2583",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:fe7fa9024cf4eea89d8531f97e5b6d99baa1943f08acea814f166160a74b52eb",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 1250,
          "cachedInputTokens": 139429,
          "taskSessionReused": true,
          "persistedSessionId": "01a0c581-1547-7262-9e22-9f1d64f670c5",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 139429,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "36747442-5a1b-41e1-a3c3-8eede461dcb3",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 94252,
          "usageSource": "per_run",
          "freshSession": true,
          "outputTokens": 895,
          "sessionReused": false,
          "rawInputTokens": 94252,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:406c81bdfd13dcff2edf3857f6563cbd2c84b311b5eff281a619d300252c2583",
              "changedCategories": [],
              "taskSessionReused": false,
              "fingerprintVersion": 1,
              "taskSessionAvailable": false,
              "storedFingerprintPresent": false
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:fe7fa9024cf4eea89d8531f97e5b6d99baa1943f08acea814f166160a74b52eb",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 895,
          "cachedInputTokens": 74247,
          "taskSessionReused": false,
          "persistedSessionId": "01a0c581-1547-7262-9e22-9f1d64f670c5",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 74247,
          "sessionRotationReason": null
        }
      }
    ]
  }
}
Conversation continuity across restart passed
Overall passed · 1/1 behavioral checks passed

No behavioral matcher failed.

See all behavioral checks (1)
ResultCheckDetail
Passissue_statusChat workflow and durable handoff/session assertions passed
Tokens233,736 in · 1,212 out172,217 cached · 3/3 runs covered
LLM spend$0.0000003/3 runs provider-priced
ExecutionLocal · not metered34s agent
agent-chat-hardening.runner-codex.local.continuity-restart
Matchers and test context
Attempt
1
Duration
1m 15s
Agent runtime
34s
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
RUN-1

All invariants passed

ResultMatcherExpectationDetail
Pass issue_status {"kind":"issue_status","expected":"in_review"} Chat workflow and durable handoff/session assertions passed
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 3,
      "runsWithTokenUsage": 3,
      "runsWithReportedCost": 3,
      "inputTokens": 233736,
      "outputTokens": 1212,
      "cachedInputTokens": 172217,
      "totalTokens": 407165,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 33522,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": true
  },
  "rawUsage": {
    "runs": [
      {
        "runId": "83a75b0c-503e-47d6-854b-1df3dee0e9b3",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 122211,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 592,
          "sessionReused": true,
          "rawInputTokens": 122211,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:36fd07d81646841d21d0e69fded71e9ca39b336b0e62517828dea00ee0380b6a",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:412f7acd9feded1ef915c5da9d83f3718c6fefe5054b706bdf2dc9ec8ea942f2",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 592,
          "cachedInputTokens": 99024,
          "taskSessionReused": true,
          "persistedSessionId": "01a0c581-1d77-74f0-8e66-f37998d26736",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 99024,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "70c4e6e5-e9d2-499b-8f50-89b1b42101d6",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 76123,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 396,
          "sessionReused": true,
          "rawInputTokens": 76123,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:36fd07d81646841d21d0e69fded71e9ca39b336b0e62517828dea00ee0380b6a",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:412f7acd9feded1ef915c5da9d83f3718c6fefe5054b706bdf2dc9ec8ea942f2",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 396,
          "cachedInputTokens": 55634,
          "taskSessionReused": true,
          "persistedSessionId": "01a0c581-1d77-74f0-8e66-f37998d26736",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 55634,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "c7f724c2-f718-4875-be65-e2530dd7411a",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 35402,
          "usageSource": "per_run",
          "freshSession": true,
          "outputTokens": 224,
          "sessionReused": false,
          "rawInputTokens": 35402,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:36fd07d81646841d21d0e69fded71e9ca39b336b0e62517828dea00ee0380b6a",
              "changedCategories": [],
              "taskSessionReused": false,
              "fingerprintVersion": 1,
              "taskSessionAvailable": false,
              "storedFingerprintPresent": false
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:412f7acd9feded1ef915c5da9d83f3718c6fefe5054b706bdf2dc9ec8ea942f2",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 224,
          "cachedInputTokens": 17559,
          "taskSessionReused": false,
          "persistedSessionId": "01a0c581-1d77-74f0-8e66-f37998d26736",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 17559,
          "sessionRotationReason": null
        }
      }
    ]
  }
}
Stop, reset, and resume not selected
agent-chat-hardening.runner-codex.local.stop-new-resume
Matchers and test context

Not selected

No matcher result was recorded.

Daytona warm reusable sandboxdaytona · remote
Recover a lost send acknowledgement without repeating committed work passed
Overall passed · 1/1 behavioral checks passed

No behavioral matcher failed.

See all behavioral checks (1)
ResultCheckDetail
Passissue_statusChat workflow and durable handoff/session assertions passed
Tokens201,606 in · 1,913 out161,939 cached · 2/2 runs covered
LLM spend$0.0000002/2 runs provider-priced
Execution$0.005570 est.47s agent · 1m 15s lease
agent-chat-hardening.runner-codex.daytona.committed-send-retry
Matchers and test context
Attempt
1
Duration
1m 29s
Agent runtime
47s
Environment lease
1m 15s
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
RUN-1

All invariants passed

ResultMatcherExpectationDetail
Pass issue_status {"kind":"issue_status","expected":"in_review"} Chat workflow and durable handoff/session assertions passed
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 2,
      "runsWithTokenUsage": 2,
      "runsWithReportedCost": 2,
      "inputTokens": 201606,
      "outputTokens": 1913,
      "cachedInputTokens": 161939,
      "totalTokens": 365458,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "daytona",
      "agentRunDurationMs": 47406,
      "leaseDurationMs": 74967,
      "leaseCount": 2,
      "cpuCores": 4,
      "memoryGiB": 4,
      "diskGiB": 10,
      "estimatedListCostUsd": 0.005570048100000001,
      "costStatus": "estimated",
      "costSource": "daytona_public_list_price",
      "pricingAsOf": "2026-08-27",
      "pricingUrl": "https://www.daytona.io/pricing"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0.005570048100000001,
    "observedAndEstimatedCostUsd": 0.005570048100000001,
    "complete": true
  },
  "rawUsage": {
    "runs": [
      {
        "runId": "4f66b6a6-da9f-40e0-84aa-9f67be990d02",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 132330,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 1112,
          "sessionReused": true,
          "rawInputTokens": 132330,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:2d6dbe62f604a37a27f5bab48b46042c578b66cabc44d7d5655f7e8b906218a0",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:a1c91a23203042b282c014a1cdc7741d6f871e102bbcd89d42d045003190034a",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 1112,
          "cachedInputTokens": 110984,
          "taskSessionReused": true,
          "persistedSessionId": "01a0c581-85c4-7762-bcb7-0733ddb2b575",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 110984,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "f2f73e83-d5ab-4387-9438-220b406a28ed",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 69276,
          "usageSource": "per_run",
          "freshSession": true,
          "outputTokens": 801,
          "sessionReused": false,
          "rawInputTokens": 69276,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:2d6dbe62f604a37a27f5bab48b46042c578b66cabc44d7d5655f7e8b906218a0",
              "changedCategories": [],
              "taskSessionReused": false,
              "fingerprintVersion": 1,
              "taskSessionAvailable": false,
              "storedFingerprintPresent": false
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:a1c91a23203042b282c014a1cdc7741d6f871e102bbcd89d42d045003190034a",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 801,
          "cachedInputTokens": 50955,
          "taskSessionReused": false,
          "persistedSessionId": "01a0c581-85c4-7762-bcb7-0733ddb2b575",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 50955,
          "sessionRotationReason": null
        }
      }
    ]
  }
}
Conversation continuity across restart passed
Overall passed · 1/1 behavioral checks passed

No behavioral matcher failed.

See all behavioral checks (1)
ResultCheckDetail
Passissue_statusChat workflow and durable handoff/session assertions passed
Tokens148,704 in · 1,158 out124,465 cached · 3/3 runs covered
LLM spend$0.0000003/3 runs provider-priced
Execution$0.006338 est.49s agent · 1m 25s lease
agent-chat-hardening.runner-codex.daytona.continuity-restart
Matchers and test context
Attempt
1
Duration
1m 39s
Agent runtime
49s
Environment lease
1m 25s
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
RUN-1

All invariants passed

ResultMatcherExpectationDetail
Pass issue_status {"kind":"issue_status","expected":"in_review"} Chat workflow and durable handoff/session assertions passed
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 3,
      "runsWithTokenUsage": 3,
      "runsWithReportedCost": 3,
      "inputTokens": 148704,
      "outputTokens": 1158,
      "cachedInputTokens": 124465,
      "totalTokens": 274327,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "daytona",
      "agentRunDurationMs": 48715,
      "leaseDurationMs": 85306,
      "leaseCount": 3,
      "cpuCores": 4,
      "memoryGiB": 4,
      "diskGiB": 10,
      "estimatedListCostUsd": 0.0063382358,
      "costStatus": "estimated",
      "costSource": "daytona_public_list_price",
      "pricingAsOf": "2026-08-27",
      "pricingUrl": "https://www.daytona.io/pricing"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0.0063382358,
    "observedAndEstimatedCostUsd": 0.0063382358,
    "complete": true
  },
  "rawUsage": {
    "runs": [
      {
        "runId": "7419eff4-bf21-4b24-85ed-ed053ef531f3",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 79819,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 556,
          "sessionReused": true,
          "rawInputTokens": 79819,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:215fbf7d94807b36759084f0738f753fc29f1e8d17b28dfd9fbacafbbaadea9e",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:f6abc3890cce1a6048bd5b7aa67c648a270e433c2353de1ff62e68db664e911d",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 556,
          "cachedInputTokens": 74361,
          "taskSessionReused": true,
          "persistedSessionId": "01a0c581-6126-7883-bf54-ac5da0c54306",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 74361,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "9c2c3093-02d5-4420-afff-989c5b8b6b13",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 37175,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 304,
          "sessionReused": true,
          "rawInputTokens": 37175,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:215fbf7d94807b36759084f0738f753fc29f1e8d17b28dfd9fbacafbbaadea9e",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:f6abc3890cce1a6048bd5b7aa67c648a270e433c2353de1ff62e68db664e911d",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 304,
          "cachedInputTokens": 34427,
          "taskSessionReused": true,
          "persistedSessionId": "01a0c581-6126-7883-bf54-ac5da0c54306",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 34427,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "0569170e-6feb-4046-aec5-1b0f2b485415",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 31710,
          "usageSource": "per_run",
          "freshSession": true,
          "outputTokens": 298,
          "sessionReused": false,
          "rawInputTokens": 31710,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:215fbf7d94807b36759084f0738f753fc29f1e8d17b28dfd9fbacafbbaadea9e",
              "changedCategories": [],
              "taskSessionReused": false,
              "fingerprintVersion": 1,
              "taskSessionAvailable": false,
              "storedFingerprintPresent": false
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:f6abc3890cce1a6048bd5b7aa67c648a270e433c2353de1ff62e68db664e911d",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 298,
          "cachedInputTokens": 15677,
          "taskSessionReused": false,
          "persistedSessionId": "01a0c581-6126-7883-bf54-ac5da0c54306",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 15677,
          "sessionRotationReason": null
        }
      }
    ]
  }
}
Stop, reset, and resume not selected
agent-chat-hardening.runner-codex.daytona.stop-new-resume
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Claudenativeacpx · claude-sonnet-5
Isolated locallocal · local
Stop during startup, reset, and resume not selected
agent-chat-hardening.runner-acpx-claude.local.stop-startup-new-resume
Matchers and test context

Not selected

No matcher result was recorded.

Hire through chat, delegate, and reuse the same teammate not selected
agent-chat-hardening.runner-acpx-claude.local.hire-delegate-reuse
Matchers and test context

Not selected

No matcher result was recorded.

Read the actual blocker and hand source material to a reviewer not selected
agent-chat-hardening.runner-acpx-claude.local.blocked-status-review
Matchers and test context

Not selected

No matcher result was recorded.

Recover a lost send acknowledgement without repeating committed work passed
Overall passed · 1/1 behavioral checks passed

No behavioral matcher failed.

See all behavioral checks (1)
ResultCheckDetail
Passissue_statusChat workflow and durable handoff/session assertions passed
Tokens18 in · 2,607 out337,209 cached · 2/2 runs covered
LLM spend$0.0000002/2 runs provider-priced
ExecutionLocal · not metered1m 3s agent
agent-chat-hardening.runner-acpx-claude.local.committed-send-retry
Matchers and test context
Attempt
1
Duration
1m 36s
Agent runtime
1m 3s
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
RUN-1

All invariants passed

ResultMatcherExpectationDetail
Pass issue_status {"kind":"issue_status","expected":"in_review"} Chat workflow and durable handoff/session assertions passed
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 2,
      "runsWithTokenUsage": 2,
      "runsWithReportedCost": 2,
      "inputTokens": 18,
      "outputTokens": 2607,
      "cachedInputTokens": 337209,
      "totalTokens": 339834,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 62564,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": true
  },
  "rawUsage": {
    "runs": [
      {
        "runId": "4797a973-b750-44ee-8585-8611781b65f3",
        "usage": {
          "model": "claude-sonnet-5",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 8,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 659,
          "sessionReused": true,
          "rawInputTokens": 8,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:cb00bf1827e7accb7cf10c556843a2e0015a5694a53c819e990af03ca442b273",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:ee797ccb5451cb90de7b5736d245f13f413217c712fc941a02f82520c8b8b1bc",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 659,
          "cachedInputTokens": 160377,
          "taskSessionReused": true,
          "persistedSessionId": "c103c12f-c704-4e96-9e57-6fd5a0776595",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 160377,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "96635f49-c721-4343-a974-7272c2d6aa6f",
        "usage": {
          "model": "claude-sonnet-5",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 10,
          "usageSource": "per_run",
          "freshSession": true,
          "outputTokens": 1948,
          "sessionReused": false,
          "rawInputTokens": 10,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:cb00bf1827e7accb7cf10c556843a2e0015a5694a53c819e990af03ca442b273",
              "changedCategories": [],
              "taskSessionReused": false,
              "fingerprintVersion": 1,
              "taskSessionAvailable": false,
              "storedFingerprintPresent": false
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:ee797ccb5451cb90de7b5736d245f13f413217c712fc941a02f82520c8b8b1bc",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 1948,
          "cachedInputTokens": 176832,
          "taskSessionReused": false,
          "persistedSessionId": "c103c12f-c704-4e96-9e57-6fd5a0776595",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 176832,
          "sessionRotationReason": null
        }
      }
    ]
  }
}
Conversation continuity across restart passed
Overall passed · 1/1 behavioral checks passed

No behavioral matcher failed.

See all behavioral checks (1)
ResultCheckDetail
Passissue_statusChat workflow and durable handoff/session assertions passed
Tokens16 in · 1,378 out270,745 cached · 3/3 runs covered
LLM spend$0.0000003/3 runs provider-priced
ExecutionLocal · not metered51s agent
agent-chat-hardening.runner-acpx-claude.local.continuity-restart
Matchers and test context
Attempt
1
Duration
1m 29s
Agent runtime
51s
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
RUN-1

All invariants passed

ResultMatcherExpectationDetail
Pass issue_status {"kind":"issue_status","expected":"in_review"} Chat workflow and durable handoff/session assertions passed
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 3,
      "runsWithTokenUsage": 3,
      "runsWithReportedCost": 3,
      "inputTokens": 16,
      "outputTokens": 1378,
      "cachedInputTokens": 270745,
      "totalTokens": 272139,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 51051,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": true
  },
  "rawUsage": {
    "runs": [
      {
        "runId": "e52a6a5a-d07e-4f1c-8158-3a27e6df3450",
        "usage": {
          "model": "claude-sonnet-5",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 4,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 394,
          "sessionReused": true,
          "rawInputTokens": 4,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:8d53fcebd49b5672cc16774c8be4b462ff7692442aefb8027fedc4a02ea4eee7",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:92720b7501644ab0b4231d7c0fdf536e76c503be7250eb8bc43da655b2e2b9db",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 394,
          "cachedInputTokens": 85594,
          "taskSessionReused": true,
          "persistedSessionId": "84164b39-5ee4-434f-bc80-3af07d5587b9",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 85594,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "ee831842-d6dd-42b6-b33f-e596870d527d",
        "usage": {
          "model": "claude-sonnet-5",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 4,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 360,
          "sessionReused": true,
          "rawInputTokens": 4,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:8d53fcebd49b5672cc16774c8be4b462ff7692442aefb8027fedc4a02ea4eee7",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:92720b7501644ab0b4231d7c0fdf536e76c503be7250eb8bc43da655b2e2b9db",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 360,
          "cachedInputTokens": 60340,
          "taskSessionReused": true,
          "persistedSessionId": "84164b39-5ee4-434f-bc80-3af07d5587b9",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 60340,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "cf648f40-80ba-445b-a94d-c53569f3c338",
        "usage": {
          "model": "claude-sonnet-5",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 8,
          "usageSource": "per_run",
          "freshSession": true,
          "outputTokens": 624,
          "sessionReused": false,
          "rawInputTokens": 8,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:8d53fcebd49b5672cc16774c8be4b462ff7692442aefb8027fedc4a02ea4eee7",
              "changedCategories": [],
              "taskSessionReused": false,
              "fingerprintVersion": 1,
              "taskSessionAvailable": false,
              "storedFingerprintPresent": false
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:92720b7501644ab0b4231d7c0fdf536e76c503be7250eb8bc43da655b2e2b9db",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 624,
          "cachedInputTokens": 124811,
          "taskSessionReused": false,
          "persistedSessionId": "84164b39-5ee4-434f-bc80-3af07d5587b9",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 124811,
          "sessionRotationReason": null
        }
      }
    ]
  }
}
Stop, reset, and resume not selected
agent-chat-hardening.runner-acpx-claude.local.stop-new-resume
Matchers and test context

Not selected

No matcher result was recorded.

Daytona warm reusable sandboxdaytona · remote
Recover a lost send acknowledgement without repeating committed work passed
Overall passed · 1/1 behavioral checks passed

No behavioral matcher failed.

See all behavioral checks (1)
ResultCheckDetail
Passissue_statusChat workflow and durable handoff/session assertions passed
Tokens18 in · 2,402 out337,011 cached · 2/2 runs covered
LLM spend$0.0000002/2 runs provider-priced
Execution$0.007768 est.1m 17s agent · 1m 45s lease
agent-chat-hardening.runner-acpx-claude.daytona.committed-send-retry
Matchers and test context
Attempt
1
Duration
1m 59s
Agent runtime
1m 17s
Environment lease
1m 45s
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
RUN-1

All invariants passed

ResultMatcherExpectationDetail
Pass issue_status {"kind":"issue_status","expected":"in_review"} Chat workflow and durable handoff/session assertions passed
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 2,
      "runsWithTokenUsage": 2,
      "runsWithReportedCost": 2,
      "inputTokens": 18,
      "outputTokens": 2402,
      "cachedInputTokens": 337011,
      "totalTokens": 339431,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "daytona",
      "agentRunDurationMs": 76774,
      "leaseDurationMs": 104554,
      "leaseCount": 2,
      "cpuCores": 4,
      "memoryGiB": 4,
      "diskGiB": 10,
      "estimatedListCostUsd": 0.0077683622,
      "costStatus": "estimated",
      "costSource": "daytona_public_list_price",
      "pricingAsOf": "2026-08-27",
      "pricingUrl": "https://www.daytona.io/pricing"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0.0077683622,
    "observedAndEstimatedCostUsd": 0.0077683622,
    "complete": true
  },
  "rawUsage": {
    "runs": [
      {
        "runId": "31385e6a-c5f4-4d0d-9d3b-56b35999c0b7",
        "usage": {
          "model": "claude-sonnet-5",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 8,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 608,
          "sessionReused": true,
          "rawInputTokens": 8,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:24bd4e7b489273f74aae9993fa5db66e5c18c5d69aba4ac3f637ed354136dda4",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:fb56b547743587e9905201a07e6aa32284bb4921a69b2e8a8c2249fc3fd08216",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 608,
          "cachedInputTokens": 160258,
          "taskSessionReused": true,
          "persistedSessionId": "a089cc86-d0f5-435d-95a0-d61c4b0b95a3",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 160258,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "89c54e79-51bc-4106-a7e5-95a6127b143f",
        "usage": {
          "model": "claude-sonnet-5",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 10,
          "usageSource": "per_run",
          "freshSession": true,
          "outputTokens": 1794,
          "sessionReused": false,
          "rawInputTokens": 10,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:24bd4e7b489273f74aae9993fa5db66e5c18c5d69aba4ac3f637ed354136dda4",
              "changedCategories": [],
              "taskSessionReused": false,
              "fingerprintVersion": 1,
              "taskSessionAvailable": false,
              "storedFingerprintPresent": false
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:fb56b547743587e9905201a07e6aa32284bb4921a69b2e8a8c2249fc3fd08216",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 1794,
          "cachedInputTokens": 176753,
          "taskSessionReused": false,
          "persistedSessionId": "a089cc86-d0f5-435d-95a0-d61c4b0b95a3",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 176753,
          "sessionRotationReason": null
        }
      }
    ]
  }
}
Conversation continuity across restart passed
Overall passed · 1/1 behavioral checks passed

No behavioral matcher failed.

See all behavioral checks (1)
ResultCheckDetail
Passissue_statusChat workflow and durable handoff/session assertions passed
Tokens16 in · 1,405 out268,903 cached · 3/3 runs covered
LLM spend$0.0000003/3 runs provider-priced
Execution$0.008276 est.1m 16s agent · 1m 51s lease
agent-chat-hardening.runner-acpx-claude.daytona.continuity-restart
Matchers and test context
Attempt
1
Duration
2m 5s
Agent runtime
1m 16s
Environment lease
1m 51s
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
RUN-1

All invariants passed

ResultMatcherExpectationDetail
Pass issue_status {"kind":"issue_status","expected":"in_review"} Chat workflow and durable handoff/session assertions passed
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 3,
      "runsWithTokenUsage": 3,
      "runsWithReportedCost": 3,
      "inputTokens": 16,
      "outputTokens": 1405,
      "cachedInputTokens": 268903,
      "totalTokens": 270324,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "daytona",
      "agentRunDurationMs": 75956,
      "leaseDurationMs": 111390,
      "leaseCount": 3,
      "cpuCores": 4,
      "memoryGiB": 4,
      "diskGiB": 10,
      "estimatedListCostUsd": 0.008276277,
      "costStatus": "estimated",
      "costSource": "daytona_public_list_price",
      "pricingAsOf": "2026-08-27",
      "pricingUrl": "https://www.daytona.io/pricing"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0.008276277,
    "observedAndEstimatedCostUsd": 0.008276277,
    "complete": true
  },
  "rawUsage": {
    "runs": [
      {
        "runId": "0c1695aa-620c-4b1e-8d09-129c66e6f5b4",
        "usage": {
          "model": "claude-sonnet-5",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 4,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 407,
          "sessionReused": true,
          "rawInputTokens": 4,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:b5a421ce72afcf21f31a8deb4c7f0d4cb24a44b18c644a4a9af88d9e2e6f64f8",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:5c713d79dbe33bd8d561e9d00a7ea8826e9e0c7fcde47dc82361c5db81ad5878",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 407,
          "cachedInputTokens": 84905,
          "taskSessionReused": true,
          "persistedSessionId": "fc6b115c-85c8-41e6-b928-f649167c5ad8",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 84905,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "62ee30de-384f-4425-8ce6-e173f6ef5e11",
        "usage": {
          "model": "claude-sonnet-5",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 4,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 374,
          "sessionReused": true,
          "rawInputTokens": 4,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:b5a421ce72afcf21f31a8deb4c7f0d4cb24a44b18c644a4a9af88d9e2e6f64f8",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:5c713d79dbe33bd8d561e9d00a7ea8826e9e0c7fcde47dc82361c5db81ad5878",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 374,
          "cachedInputTokens": 60041,
          "taskSessionReused": true,
          "persistedSessionId": "fc6b115c-85c8-41e6-b928-f649167c5ad8",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 60041,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "c8e9b62a-6043-4d59-9425-bf9bc379760d",
        "usage": {
          "model": "claude-sonnet-5",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 8,
          "usageSource": "per_run",
          "freshSession": true,
          "outputTokens": 624,
          "sessionReused": false,
          "rawInputTokens": 8,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:b5a421ce72afcf21f31a8deb4c7f0d4cb24a44b18c644a4a9af88d9e2e6f64f8",
              "changedCategories": [],
              "taskSessionReused": false,
              "fingerprintVersion": 1,
              "taskSessionAvailable": false,
              "storedFingerprintPresent": false
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:5c713d79dbe33bd8d561e9d00a7ea8826e9e0c7fcde47dc82361c5db81ad5878",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 624,
          "cachedInputTokens": 123957,
          "taskSessionReused": false,
          "persistedSessionId": "fc6b115c-85c8-41e6-b928-f649167c5ad8",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 123957,
          "sessionRotationReason": null
        }
      }
    ]
  }
}
Stop, reset, and resume not selected
agent-chat-hardening.runner-acpx-claude.daytona.stop-new-resume
Matchers and test context

Not selected

No matcher result was recorded.

Test suite

Core Runner Compatibility

Major provider, runtime generation, and execution-environment compatibility.

Configuration matrix7 profiles · 2 environments · 0 selected
Agent profileIsolated locallocal · localDaytona sandboxdaytona · remote
Legacy Codexlegacycodex · gpt-5.6-sol
Isolated locallocal · local
Basic response not selected
core-compatibility.legacy-codex.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-codex.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-codex.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.legacy-codex.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-codex.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-codex.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Legacy Claudelegacyclaude · claude-sonnet-4-6
Isolated locallocal · local
Basic response not selected
core-compatibility.legacy-claude.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-claude.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-claude.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.legacy-claude.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-claude.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-claude.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Legacy OpenCodelegacyopencode · openrouter/deepseek/deepseek-v4-flash-0731
Isolated locallocal · local
Basic response not selected
core-compatibility.legacy-opencode.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-opencode.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-opencode.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.legacy-opencode.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-opencode.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-opencode.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Runner Codexnativecodex · gpt-5.6-sol
Isolated locallocal · local
Basic response not selected
core-compatibility.runner-codex.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-codex.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-codex.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.runner-codex.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-codex.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-codex.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Runner OpenCodenativeopencode · openrouter/deepseek/deepseek-v4-flash-0731
Isolated locallocal · local
Basic response not selected
core-compatibility.runner-opencode.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-opencode.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-opencode.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.runner-opencode.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-opencode.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-opencode.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Claudenativeacpx · claude-sonnet-5
Isolated locallocal · local
Basic response not selected
core-compatibility.runner-acpx-claude.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-acpx-claude.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-acpx-claude.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.runner-acpx-claude.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-acpx-claude.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-acpx-claude.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Codexnativeacpx · gpt-5.6-sol
Isolated locallocal · local
Basic response not selected
core-compatibility.runner-acpx-codex.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-acpx-codex.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-acpx-codex.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.runner-acpx-codex.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-acpx-codex.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-acpx-codex.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Test suite

Local Session Integrity

Structured interaction and continuation qualification for every supported local profile.

Configuration matrix7 profiles · 1 environments · 0 selected
Agent profileIsolated locallocal · local
Legacy Codexlegacycodex · gpt-5.6-sol
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.legacy-codex.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.legacy-codex.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Legacy Claudelegacyclaude · claude-sonnet-4-6
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.legacy-claude.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.legacy-claude.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Legacy OpenCodelegacyopencode · openrouter/deepseek/deepseek-v4-flash-0731
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.legacy-opencode.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.legacy-opencode.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Runner Codexnativecodex · gpt-5.6-sol
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.runner-codex.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.runner-codex.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Runner OpenCodenativeopencode · openrouter/deepseek/deepseek-v4-flash-0731
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.runner-opencode.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.runner-opencode.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Claudenativeacpx · claude-sonnet-5
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.runner-acpx-claude.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Codexnativeacpx · gpt-5.6-sol
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.runner-acpx-codex.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.runner-acpx-codex.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Test suite

OpenRouter Model Breadth

Weekly-ranked tool-capable OpenRouter models through native OpenCode on isolated local workspaces.

Configuration matrix4 profiles · 1 environments · 0 selected
Agent profileIsolated locallocal · local
#1 DeepSeek V4 Flash 0731nativeopencode · openrouter/deepseek/deepseek-v4-flash-0731
Isolated locallocal · local
Hello and complete not selected
openrouter-model-breadth.openrouter-deepseek-deepseek-v4-flash-0731.local.hello-complete
Matchers and test context

Not selected

No matcher result was recorded.

Ask, answer, resume not selected
openrouter-model-breadth.openrouter-deepseek-deepseek-v4-flash-0731.local.question-resume-complete
Matchers and test context

Not selected

No matcher result was recorded.

#3 Tencent HY 3nativeopencode · openrouter/tencent/hy3
Isolated locallocal · local
Hello and complete not selected
openrouter-model-breadth.openrouter-tencent-hy3.local.hello-complete
Matchers and test context

Not selected

No matcher result was recorded.

Ask, answer, resume not selected
openrouter-model-breadth.openrouter-tencent-hy3.local.question-resume-complete
Matchers and test context

Not selected

No matcher result was recorded.

#4 Nemotron 3 Ultra 550B A55B (free)nativeopencode · openrouter/nvidia/nemotron-3-ultra-550b-a55b:free
Isolated locallocal · local
Hello and complete not selected
openrouter-model-breadth.openrouter-nvidia-nemotron-3-ultra-550b-a55b-free.local.hello-complete
Matchers and test context

Not selected

No matcher result was recorded.

Ask, answer, resume not selected
openrouter-model-breadth.openrouter-nvidia-nemotron-3-ultra-550b-a55b-free.local.question-resume-complete
Matchers and test context

Not selected

No matcher result was recorded.

Plan, approve, complete not selected
openrouter-model-breadth.openrouter-nvidia-nemotron-3-ultra-550b-a55b-free.local.plan-approve-complete
Matchers and test context

Not selected

No matcher result was recorded.

#5 GPT-5.6 Lunanativeopencode · openrouter/openai/gpt-5.6-luna
Isolated locallocal · local
Hello and complete not selected
openrouter-model-breadth.openrouter-openai-gpt-5-6-luna.local.hello-complete
Matchers and test context

Not selected

No matcher result was recorded.

Ask, answer, resume not selected
openrouter-model-breadth.openrouter-openai-gpt-5-6-luna.local.question-resume-complete
Matchers and test context

Not selected

No matcher result was recorded.

Plan, approve, complete not selected
openrouter-model-breadth.openrouter-openai-gpt-5-6-luna.local.plan-approve-complete
Matchers and test context

Not selected

No matcher result was recorded.

Test suite

Daytona Warm Continuity

Three browser-driven turns on one reusable Daytona sandbox for legacy and native Codex.

Configuration matrix2 profiles · 1 environments · 0 selected
Agent profileDaytona warm reusable sandboxdaytona · remote
Legacy Codexlegacycodex · gpt-5.6-sol
Daytona warm reusable sandboxdaytona · remote
Warm three-turn workspace continuity not selected
daytona-warm-continuity.legacy-codex.daytona.warm-three-turn
Matchers and test context

Not selected

No matcher result was recorded.

Runner Codexnativecodex · gpt-5.6-sol
Daytona warm reusable sandboxdaytona · remote
Warm three-turn workspace continuity not selected
daytona-warm-continuity.runner-codex.daytona.warm-three-turn
Matchers and test context

Not selected

No matcher result was recorded.

History

Campaign trends

No historical campaigns have been published yet.