Paperclip Quality engineering · Runner acceptance

Full-stack acceptance campaign

Runner Full-Stack E2E

A browser-verified matrix of runner profiles, execution environments, and deterministic task contracts. Visual evidence is retained in the access-controlled workflow artifact; public history contains inert structured evidence only.

1/3Passed
2Failed
24m 7sTest time
60,016Input tokens
498Output tokens
52,367Cached tokens
$0.000000LLM reported subtotal
$0.000000Daytona list estimate
11m 3sAgent execution time
0msDaytona lease time
3/5Runs provider-priced

Model spend is the provider-reported subtotal; unpriced or unavailable runs are excluded, never counted as free. Daytona runtime is a public-list-price estimate from captured lease time and pinned resources, before credits, discounts, storage allowance, or invoice adjustments. Local execution has no external runtime meter.

Test suite

Core Runner Compatibility

Major provider, runtime generation, and execution-environment compatibility.

Pass rate50.0%1/2 passed
Tokens112,88160,016 input · 498 output
Cost$0.000000reported LLM + runtime estimate
Agent time11m 3s0ms lease
Execution2/20 retries · cleanup passed
Configuration matrix7 profiles · 2 environments · 2 selected
Agent profileIsolated locallocal · localDaytona sandboxdaytona · remote
Legacy Codexlegacycodex · gpt-5.6-sol
Isolated locallocal · local
Basic response not selected
core-compatibility.legacy-codex.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-codex.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-codex.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.legacy-codex.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-codex.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-codex.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Legacy Claudelegacyclaude · claude-sonnet-4-6
Isolated locallocal · local
Basic response not selected
core-compatibility.legacy-claude.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-claude.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-claude.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.legacy-claude.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-claude.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-claude.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Legacy OpenCodelegacyopencode · openrouter/deepseek/deepseek-v4-flash-0731
Isolated locallocal · local
Basic response not selected
core-compatibility.legacy-opencode.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-opencode.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-opencode.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.legacy-opencode.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.legacy-opencode.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.legacy-opencode.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Runner Codexnativecodex · gpt-5.6-sol
Isolated locallocal · local
Basic response not selected
core-compatibility.runner-codex.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement passed
Tokens60,016 in · 498 out52,367 cached · 3/3 runs covered
LLM spend$0.0000003/3 runs provider-priced
ExecutionLocal · not metered44s agent
core-compatibility.runner-codex.local.plan-revise-accept
Matchers and test context
Attempt
1
Duration
1m 34s
Agent runtime
44s
Runtime
native
Provider
codex
Model
gpt-5.6-sol
Issue
RUN-1

All invariants passed

ResultMatcherExpectationDetail
Pass message_exact {"kind":"message_exact","expected":"PAPERCLIP_E2E_PLAN_DONE_10c1af46458b-1"} matched
Pass message_occurrences {"kind":"message_occurrences","expected":"PAPERCLIP_E2E_PLAN_DONE_10c1af46458b-1","count":1} matched
Pass issue_status {"kind":"issue_status","expected":"done"} matched
Pass run_status {"kind":"run_status","expected":"succeeded"} matched
Pass runtime_mode {"kind":"runtime_mode","expected":"native"} matched
Pass environment {"kind":"environment","expected":"local"} matched
Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 3,
      "runsWithTokenUsage": 3,
      "runsWithReportedCost": 3,
      "inputTokens": 60016,
      "outputTokens": 498,
      "cachedInputTokens": 52367,
      "totalTokens": 112881,
      "reportedCostUsd": 0,
      "costStatus": "reported"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 44237,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": true
  },
  "rawUsage": {
    "runs": [
      {
        "runId": "3def240a-b6e6-4e7a-9a57-3db73c7eec82",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 16809,
          "usageSource": "per_run",
          "freshSession": true,
          "outputTokens": 129,
          "sessionReused": false,
          "rawInputTokens": 16809,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": true,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:2d8a34465b483e286d634a9705e317a62baebe83a87a512bb472836dbd2e9a54",
              "changedCategories": [],
              "taskSessionReused": false,
              "fingerprintVersion": 1,
              "taskSessionAvailable": false,
              "storedFingerprintPresent": false
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:3dde1acc07e7ce49ab16ff61f8659da75c27430b0e4d0718c024a80b24da87ad",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 129,
          "cachedInputTokens": 16232,
          "taskSessionReused": false,
          "persistedSessionId": "01a070ee-a32e-76c0-b691-c5c389a80f54",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 16232,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "b57972f8-d5ea-4ae0-871a-06f4137a7a24",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 24342,
          "usageSource": "per_run",
          "freshSession": false,
          "outputTokens": 346,
          "sessionReused": true,
          "rawInputTokens": 24342,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": false,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [],
              "nextFingerprint": "v1:sha256:2d8a34465b483e286d634a9705e317a62baebe83a87a512bb472836dbd2e9a54",
              "changedCategories": [],
              "taskSessionReused": true,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:3dde1acc07e7ce49ab16ff61f8659da75c27430b0e4d0718c024a80b24da87ad",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 346,
          "cachedInputTokens": 17449,
          "taskSessionReused": true,
          "persistedSessionId": "01a070ee-a32e-76c0-b691-c5c389a80f54",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 17449,
          "sessionRotationReason": null
        }
      },
      {
        "runId": "b77b11c8-18fd-4ad8-8130-9b08805a18f4",
        "usage": {
          "model": "gpt-5.6-sol",
          "biller": "openai",
          "costUsd": 0,
          "provider": "openai",
          "costStatus": "reported",
          "billingType": "unknown",
          "inputTokens": 18865,
          "usageSource": "per_run",
          "freshSession": true,
          "outputTokens": 23,
          "sessionReused": false,
          "rawInputTokens": 18865,
          "sessionRotated": false,
          "configFreshness": {
            "session": {
              "reset": true,
              "categories": [
                "adapter",
                "adapterConfig",
                "agentRuntimeConfig",
                "instructions",
                "issueOverrides",
                "workspaceConfig",
                "environment",
                "envBindings",
                "secrets",
                "runtimeSkills"
              ],
              "resetReasons": [
                "effective run configuration changed: adapter config",
                "forceFreshSession was requested"
              ],
              "nextFingerprint": "v1:sha256:f4f9cb6124e92008adf8e9265c3df87579728b465f8476f3e756cac5066594ba",
              "changedCategories": [
                "adapterConfig"
              ],
              "taskSessionReused": false,
              "fingerprintVersion": 1,
              "taskSessionAvailable": true,
              "storedFingerprintPresent": true
            },
            "version": 1,
            "workspace": {
              "action": "create",
              "reasons": [],
              "categories": [
                "mode",
                "projectWorkspace",
                "strategy",
                "repo",
                "lifecycleCommands",
                "runtimeServices",
                "environment",
                "realization"
              ],
              "reuseRequested": false,
              "nextFingerprint": "v1:sha256:3dde1acc07e7ce49ab16ff61f8659da75c27430b0e4d0718c024a80b24da87ad",
              "workspaceReused": false,
              "activeWorkspaceId": null,
              "changedCategories": [],
              "storedFingerprint": null,
              "fingerprintVersion": 1,
              "inferredFingerprint": null,
              "previousWorkspaceId": null,
              "configSnapshotRefreshed": false,
              "storedFingerprintPresent": false
            }
          },
          "rawOutputTokens": 23,
          "cachedInputTokens": 18686,
          "taskSessionReused": false,
          "persistedSessionId": "01a070ef-7a0e-7060-a98f-31d22e8b2439",
          "cacheAdjustedCostUsd": 0,
          "rawCachedInputTokens": 18686,
          "sessionRotationReason": null
        }
      }
    ]
  }
}
Ask mode question not selected
core-compatibility.runner-codex.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.runner-codex.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-codex.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-codex.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Runner OpenCodenativeopencode · openrouter/deepseek/deepseek-v4-flash-0731
Isolated locallocal · local
Basic response not selected
core-compatibility.runner-opencode.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-opencode.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-opencode.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.runner-opencode.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-opencode.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-opencode.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Claudenativeacpx · claude-sonnet-5
Isolated locallocal · local
Basic response not selected
core-compatibility.runner-acpx-claude.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered10m 19s agent
core-compatibility.runner-acpx-claude.local.plan-revise-accept
Matchers and test context
Attempt
1
Duration
10m 31s
Agent runtime
10m 19s
Runtime
native
Provider
acpx
Model
claude-sonnet-5
Issue
RUN-1

Stopped waiting for initial plan confirmation for issue 948caa6e-45cf-429d-883d-c86c5ca13ff0: heartbeat run b386fbeb-6805-4eb6-9eb5-f4ca38881b0c ended failed (native_session_retry_exhausted): native_session_recovery_failed: provider session ended with a failed terminal

No matcher result was recorded.

Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 618641,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": null
}
Ask mode question not selected
core-compatibility.runner-acpx-claude.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.runner-acpx-claude.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-acpx-claude.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-acpx-claude.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Codexnativeacpx · gpt-5.6-sol
Isolated locallocal · local
Basic response not selected
core-compatibility.runner-acpx-codex.local.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-acpx-codex.local.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-acpx-codex.local.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Daytona sandboxdaytona · remote
Basic response not selected
core-compatibility.runner-acpx-codex.daytona.message-marker
Matchers and test context

Not selected

No matcher result was recorded.

Plan, revise, accept, implement not selected
core-compatibility.runner-acpx-codex.daytona.plan-revise-accept
Matchers and test context

Not selected

No matcher result was recorded.

Ask mode question not selected
core-compatibility.runner-acpx-codex.daytona.ask-question
Matchers and test context

Not selected

No matcher result was recorded.

Test suite

Local Session Integrity

Structured interaction and continuation qualification for every supported local profile.

Pass rate0.0%0/1 passed
Tokens00 input · 0 output
Cost$0.000000reported LLM + runtime estimate
Agent time0ms0ms lease
Execution1/10 retries · cleanup passed
Configuration matrix7 profiles · 1 environments · 1 selected
Agent profileIsolated locallocal · local
Legacy Codexlegacycodex · gpt-5.6-sol
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.legacy-codex.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.legacy-codex.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Legacy Claudelegacyclaude · claude-sonnet-4-6
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.legacy-claude.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.legacy-claude.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Legacy OpenCodelegacyopencode · openrouter/deepseek/deepseek-v4-flash-0731
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.legacy-opencode.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.legacy-opencode.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Runner Codexnativecodex · gpt-5.6-sol
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.runner-codex.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.runner-codex.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Runner OpenCodenativeopencode · openrouter/deepseek/deepseek-v4-flash-0731
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.runner-opencode.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.runner-opencode.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Claudenativeacpx · claude-sonnet-5
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.runner-acpx-claude.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume not selected
local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume
Matchers and test context

Not selected

No matcher result was recorded.

Runner ACPX Codexnativeacpx · gpt-5.6-sol
Isolated locallocal · local
Structured question, answer, resume not selected
local-session-integrity.runner-acpx-codex.local.structured-question-resume
Matchers and test context

Not selected

No matcher result was recorded.

Structured question, server restart, answer, resume failed
Tokens0 in · 0 out0 cached · 0/1 runs covered
LLM spendunavailable0/1 runs provider-priced
ExecutionLocal · not metered0ms agent
local-session-integrity.runner-acpx-codex.local.structured-question-restart-resume
Matchers and test context
Attempt
1
Duration
12m 2s
Agent runtime
0ms
Runtime
native
Provider
acpx
Model
gpt-5.6-sol
Issue
RUN-1

Timed out waiting for pending user question for issue 5fc20e95-98f4-49f3-bf28-2e478048d3da: {"currentIssue":{"id":"5fc20e95-98f4-49f3-bf28-2e478048d3da","companyId":"f5fbed7b-2ee0-49dc-8731-6697f07335c8","projectId":null,"projectWorkspaceId":null,"goalId":null,"parentId":null,"title":"Runner E2E restart question 42ca0cf56c68-1","description":"Ask the user one structured question before completing this task.\nThe question must be required, single-select, and offer Cobalt and Amber. Do not publish a final answer or mark the task Done while it is pending.\nBefore the answer arrives, do not spell, quote, repeat, announce, or include PAPERCLIP\\_E2E\\_QUESTION\\_DONE\\_42ca0cf56c68-1 in any visible response, comment, or tool payload; refer to it only as “the terminal marker.”\nIn a native runner, call request\\_human\\_input exactly once with idempotencyKey \\`question-42ca0cf56c68-1\\`, interactionKind \\`questions\\`, title \\`Verification word\\`, prompt \\`Choose the verification word\\`, continuationPolicy \\`wake\\_assignee\\`, and payload {version:1,questions:\\[{id:\\`verification-word\\`,prompt:\\`Choose the verification word.\\`,selectionMode:\\`single\\`,required:true,options:\\[{id:\\`cobalt\\`,label:\\`Cobalt\\`},{id:\\`amber\\`,label:\\`Amber\\`}]}]}.\nIn a legacy runner, derive \\`API\\_ORIGIN\\` exactly once with \\`API\\_ORIGIN\\=\"${PAPERCLIP\\_API\\_URL%/}\"; API\\_ORIGIN\\=\"${API\\_ORIGIN%/api}\"\\`. Build every endpoint as \\`$API\\_ORIGIN/api/...\\`; never append \\`/api\\` to a base that already ends in \\`/api\\`.\nIn a legacy runner, create exactly one question interaction: POST $API\\_ORIGIN/api/issues/$PAPERCLIP\\_TASK\\_ID/interactions once with {\"kind\":\"ask\\_user\\_questions\",\"idempotencyKey\":\"question-42ca0cf56c68-1\",\"continuationPolicy\":\"wake\\_assignee\",\"payload\":{\"version\":1,\"questions\":\\[{\"id\":\"verification-word\",\"prompt\":\"Choose the verification word.\",\"selectionMode\":\"single\",\"required\":true,\"options\":\\[{\"id\":\"cobalt\",\"label\":\"Cobalt\"},{\"id\":\"amber\",\"label\":\"Amber\"}]}]}} using Authorization and X-Paperclip-Run-Id. Do not create a replacement interaction if a later write fails.\nIn a legacy runner, after that POST returns 2xx, PATCH $API\\_ORIGIN/api/issues/$PAPERCLIP\\_TASK\\_ID with exactly {\"status\":\"in\\_review\"}. Do not include \\`reviewInteractionId\\`: it only designates confirmation interactions, not \\`ask\\_user\\_questions\\`. If the PATCH fails, retry only that PATCH and never POST the interaction again.\nIn a legacy runner, after those two writes succeed, end the current response and heartbeat immediately. Do not wait, sleep, poll, or fetch the interaction; \\`wake\\_assignee\\` will start a new heartbeat after the user answers.\nAfter the answer arrives, if it is Cobalt, publish exactly PAPERCLIP\\_E2E\\_QUESTION\\_DONE\\_42ca0cf56c68-1 once as the complete visible response and mark the task Done.\nIn a native runner, after the answer arrives, call paperclip\\_finish exactly once with {reportedWorkDisposition:\"done\",summary:\"PAPERCLIP\\_E2E\\_QUESTION\\_DONE\\_42ca0cf56c68-1\",completionClaim:{contractRevision:\"1\",objectiveSatisfied:true,criteria:\\[{criterionId:\"objective\",status:\"satisfied\",evidenceRefs:\\[]}],remainingWork:\\[]},evidence:\\[],verification:\\[]}. Wait for that tool call to succeed, then emit exactly PAPERCLIP\\_E2E\\_QUESTION\\_DONE\\_42ca0cf56c68-1 as the complete final response. Do not write a user-facing final response before paperclip\\_finish succeeds, and do not call another tool.\nIn a legacy runner, make exactly one completion write: PATCH $API\\_ORIGIN/api/issues/$PAPERCLIP\\_TASK\\_ID with {\"status\":\"done\",\"comment\":\"PAPERCLIP\\_E2E\\_QUESTION\\_DONE\\_42ca0cf56c68-1\"}. Do not POST a separate comment or perform a second write containing the marker.\nDo not create files, plans, child tasks, or unrelated work, and do not expose credentials.","status":"in_progress","statusVersion":1,"lastStatusDecisionId":null,"workMode":"standard","harnessKind":null,"priority":"medium","reviewPolicy":null,"assigneeAgentId":"a9b856b5-678b-483a-b6ba-d949009bcb16","assigneeUserId":null,"checkoutRunId":"3c0db2a2-6baa-4d81-87c9-105bcf78c1b8","executionRunId":"3c0db2a2-6baa-4d81-87c9-105bcf78c1b8","executionAgentNameKey":"runner e2e runner-acpx-codex 42ca0cf56c68-1","executionLockedAt":"2026-09-05T09:37:50.461Z","createdByAgentId":null,"createdByUserId":"local-board","responsibleUserId":"local-board","issueNumber":1,"identifier":"RUN-1","originKind":"manual","originId":null,"originRunId":null,"originFingerprint":"default","requestDepth":0,"billingCode":null,"assigneeAdapterOverrides":null,"executionPolicy":null,"executionState":null,"monitorNextCheckAt":null,"monitorWakeRequestedAt":null,"monitorLastTriggeredAt":null,"monitorAttemptCount":0,"monitorNotes":null,"monitorScheduledBy":null,"executionWorkspaceId":null,"executionWorkspacePreference":null,"executionWorkspaceSettings":null,"sourceTrust":null,"unblockDescriptor":null,"blockedTransitionAt":null,"blockedOwnerNotifiedAt":null,"startedAt":"2026-09-05T09:37:50.511Z","completedAt":null,"cancelledAt":null,"hiddenAt":null,"createdAt":"2026-09-05T09:37:50.354Z","updatedAt":"2026-09-05T09:37:50.511Z","labels":[],"labelIds":[],"watchdog":null,"ancestors":[],"blockerAttention":{"state":"none","reason":null,"unresolvedBlockerCount":0,"coveredBlockerCount":0,"stalledBlockerCount":0,"attentionBlockerCount":0,"pendingFinalizeBlockerIssueIds":[],"sampleBlockerIdentifier":null,"sampleStalledBlockerIdentifier":null,"blockingTreeLive":false,"directBlockerIssueId":null,"terminalBlockerIssueId":null,"terminalBlocker":null},"reviewAttention":{"state":"none","paths":[],"reason":null},"productivityReview":null,"successfulRunHandoff":null,"scheduledRetry":null,"activeRecoveryAction":null,"blockedBy":[],"blocks":[],"relatedWork":{"outbound":[],"inbound":[]},"referencedIssueIdentifiers":[],"planDocument":null,"documentSummaries":[],"legacyPlanDocument":null,"project":null,"goal":null,"mentionedProjects":[],"currentExecutionWorkspace":null,"workProducts":[],"linkedCases":[]},"taskRuns":[{"id":"3c0db2a2-6baa-4d81-87c9-105bcf78c1b8","companyId":"f5fbed7b-2ee0-49dc-8731-6697f07335c8","agentId":"a9b856b5-678b-483a-b6ba-d949009bcb16","invocationSource":"assignment","triggerDetail":"system","status":"running","startedAt":"2026-09-05T09:37:50.461Z","finishedAt":null,"error":null,"wakeupRequestId":"703c7669-dae9-4023-9282-cdb0348076c1","exitCode":null,"signal":null,"usageJson":null,"sessionIdBefore":null,"sessionIdAfter":"01a070ee-9bb6-74a0-89e8-8c8b8824b101","logStore":"local_file","logRef":"f5fbed7b-2ee0-49dc-8731-6697f07335c8/a9b856b5-678b-483a-b6ba-d949009bcb16/3c0db2a2-6baa-4d81-87c9-105bcf78c1b8.ndjson","logBytes":null,"logSha256":null,"logCompressed":false,"stdoutExcerpt":null,"stderrExcerpt":null,"errorCode":null,"externalRunId":null,"processPid":1967,"processGroupId":1967,"processStartedAt":"2026-09-05T09:37:54.737Z","lastOutputAt":"2026-09-05T09:37:51.075Z","lastOutputSeq":1,"lastOutputStream":"stdout","lastOutputBytes":314,"retryOfRunId":null,"processLossRetryCount":0,"scheduledRetryAt":null,"scheduledRetryAttempt":0,"scheduledRetryReason":null,"livenessState":null,"livenessReason":null,"continuationAttempt":0,"lastUsefulActionAt":null,"nextAction":null,"createdAt":"2026-09-05T09:37:50.425Z","updatedAt":"2026-09-05T09:38:12.161Z","contextSnapshot":{"issueId":"5fc20e95-98f4-49f3-bf28-2e478048d3da","taskId":"5fc20e95-98f4-49f3-bf28-2e478048d3da","taskKey":"5fc20e95-98f4-49f3-bf28-2e478048d3da","wakeReason":"issue_assigned","wakeSource":"assignment","wakeTriggerDetail":"system"},"resultJson":null}],"comments":[],"interactions":[]}

No matcher result was recorded.

Usage and billing metadata
{
  "billing": {
    "llm": {
      "runCount": 1,
      "runsWithTokenUsage": 0,
      "runsWithReportedCost": 0,
      "inputTokens": 0,
      "outputTokens": 0,
      "cachedInputTokens": 0,
      "totalTokens": 0,
      "reportedCostUsd": 0,
      "costStatus": "unavailable"
    },
    "runtime": {
      "provider": "local",
      "agentRunDurationMs": 0,
      "leaseDurationMs": null,
      "leaseCount": 0,
      "costStatus": "not_metered",
      "costSource": "local_not_metered"
    },
    "reportedCostUsd": 0,
    "estimatedRuntimeCostUsd": 0,
    "observedAndEstimatedCostUsd": 0,
    "complete": false
  },
  "rawUsage": null
}

Test suite

OpenRouter Model Breadth

Weekly-ranked tool-capable OpenRouter models through native OpenCode on isolated local workspaces.

Configuration matrix4 profiles · 1 environments · 0 selected
Agent profileIsolated locallocal · local
#1 DeepSeek V4 Flash 0731nativeopencode · openrouter/deepseek/deepseek-v4-flash-0731
Isolated locallocal · local
Hello and complete not selected
openrouter-model-breadth.openrouter-deepseek-deepseek-v4-flash-0731.local.hello-complete
Matchers and test context

Not selected

No matcher result was recorded.

Ask, answer, resume not selected
openrouter-model-breadth.openrouter-deepseek-deepseek-v4-flash-0731.local.question-resume-complete
Matchers and test context

Not selected

No matcher result was recorded.

#3 Tencent HY 3nativeopencode · openrouter/tencent/hy3
Isolated locallocal · local
Hello and complete not selected
openrouter-model-breadth.openrouter-tencent-hy3.local.hello-complete
Matchers and test context

Not selected

No matcher result was recorded.

Ask, answer, resume not selected
openrouter-model-breadth.openrouter-tencent-hy3.local.question-resume-complete
Matchers and test context

Not selected

No matcher result was recorded.

#4 Nemotron 3 Ultra 550B A55B (free)nativeopencode · openrouter/nvidia/nemotron-3-ultra-550b-a55b:free
Isolated locallocal · local
Hello and complete not selected
openrouter-model-breadth.openrouter-nvidia-nemotron-3-ultra-550b-a55b-free.local.hello-complete
Matchers and test context

Not selected

No matcher result was recorded.

Ask, answer, resume not selected
openrouter-model-breadth.openrouter-nvidia-nemotron-3-ultra-550b-a55b-free.local.question-resume-complete
Matchers and test context

Not selected

No matcher result was recorded.

Plan, approve, complete not selected
openrouter-model-breadth.openrouter-nvidia-nemotron-3-ultra-550b-a55b-free.local.plan-approve-complete
Matchers and test context

Not selected

No matcher result was recorded.

#5 GPT-5.6 Lunanativeopencode · openrouter/openai/gpt-5.6-luna
Isolated locallocal · local
Hello and complete not selected
openrouter-model-breadth.openrouter-openai-gpt-5-6-luna.local.hello-complete
Matchers and test context

Not selected

No matcher result was recorded.

Ask, answer, resume not selected
openrouter-model-breadth.openrouter-openai-gpt-5-6-luna.local.question-resume-complete
Matchers and test context

Not selected

No matcher result was recorded.

Plan, approve, complete not selected
openrouter-model-breadth.openrouter-openai-gpt-5-6-luna.local.plan-approve-complete
Matchers and test context

Not selected

No matcher result was recorded.

History

Campaign trends

No historical campaigns have been published yet.