DeepSeek V4 Flash
07313 / 5 solved

MartechBench
1,300 long-horizon tasks, deterministic graders verifying the final application state
Inside a rollout
This is a working view of the task, trace, application state, and grader—not a product screenshot. Switch runs, inspect calls, and follow the evidence.
event_transform_hospital_tycoon_prestige_bridge
AppLovin traffic reaches the mobile feed, but useful conversions do not. Reconnect attribution, preserve identity, execute two treatments, protect ward-rush reporting, and book the earliest product review.
Tool result / 3 of 6
{
"ok": true,
"environment": "production",
"result": "mobile:attribution v8 published",
"durable_state": true,
"observed_at": "2026-09-03T03:37:18Z"
}State mutation recorded in the rollout event log.
Before rollout
experiment.status draftpush.auth missingreview.event_id nullAfter rollout
experiment.status activepush.auth unverifiedreview.event_id evt_7c12An active configuration is not proof of a working production path. The grader generates fresh events after the agent stops.
Continuous reward
0.296Granular partial creditExperiment and handoff verified; runtime acquisition still fails. Continuous scoring preserves useful training signal without letting configuration mask a broken outcome.
Open full traceWorkflow execution, verified
Agents investigate live state, coordinate across applications, execute changes, and pass runtime checks. A plausible answer earns nothing by itself.

The agent configures working applications. Changes persist, trigger real behavior, and produce effects it can observe in the next tool call.

No LLM-as-a-judge. Code checks final application state and, in most tasks, tests what the system actually does.

Agents investigate, configure, test, and repair across connected systems. Frontier-model runs regularly reach 40+ turns.

Overlapping segments, existing workflows, and unrelated records force agents to distinguish current requirements from stale or conflicting context.
Use the arrow keys, Home, or End when the cards are focused.
Tasks share infrastructure, state, and dependencies across data, delivery, channels, and workplace systems. Agents learn reusable operating skills, not isolated API calls.
40 tools
Schemas, transformations, profiles, segments, experiments, activations, integrations, and logs
Bounded customer-event reports and grouped behavioral aggregates
5 tools
Versioned treatments, precedence, publication, evaluation, and exposure provenance
345 tools
AppLovin and AppsFlyer apps, credentials, campaigns, attribution, push, cost, postbacks, and stream mappings
Accounts, senders, authentication, templates, tracking, webhooks, messages, and delivery
Accounts, campaigns, criteria, assets, bidding, conversions, audiences, leads, and reporting
Business assets, audiences, pixels, campaigns, creative, leads, diagnostics, and history
36 tools
Company knowledge, employees, teams, Slack, Confluence, email, and calendars
Zones, records, and controlled changes
Role-aware capability across each major operating surface
One task, end to end
Hospital Tycoon’s prestige campaign has traffic, but no useful conversions. Repairing it means connecting attribution, event transformation, experimentation, and real treatment delivery while protecting Operations’ reporting and coordinating the product team.
Two GPT-5.6 Luna buyer traces, with tool calls, recorded results, and grader evidence. JSON download: Attempt 1.
Hospital Tycoon / Prestige bridge
“AppLovin traffic is arriving in the mobile feed, but the recovery experiment has no useful conversions and the prestige-start follow-up is not reaching the right players.”
Restore the paid flow, route players through two approved treatments, leave organic installs out, and preserve ward-rush reporting. Then book the product team’s earliest shared review.
During rollout / Agent responsibility
Find Mina’s current Growth brief, Devika’s incident note, and the approved treatment APIs. Distinguish current policy from the archived June attribution link.
Repair AppLovin’s click path and AppsFlyer’s production Push delivery for paid installs and re-attribution. A green endpoint test is not production proof.
Transform properties.body into mobile:attribution, retaining appsFlyerId, campaignId, conversionType, and event time. Raw vendor data cannot drive the downstream flow as-is.
Use prestige:started as the explicit conversion event. Gate eligible paid attribution before A/B assignment, then execute distinct ward-coach and quick-start API actions.
Keep Operations’ ward-rush route intact. Find the Head of Product and direct reports, check calendars, and book the earliest shared 30-minute review in the requested New York working window.
After rollout / Grader only
The grader generates held-out traffic and inspects runtime evidence. These probes are not agent actions, and a configuration screenshot cannot substitute for their results.
Deliver held-out AppLovin clicks, a paid install, and a returning-player re-attribution. Follow the actual click, AppsFlyer occurrence, production Push receipt, and canonical CDP event.
Probe a population of 14 paid occurrences. Require one runtime instance, one assignment, and one approved API action per identity, with both treatment contracts represented.
Submit prestige:started events and inspect derived conversions. Organic installs, wrong campaigns, and the wrong attribution class must not create prestige assignments or actions.
Execute ward-rush reporting and verify protected state is unchanged. Separately inspect the review’s exact attendees, duration, title context, and earliest shared calendar slot.
Required full-credit outcome
1.00Core workflow + calendar handoff
This is the task’s acceptance contract, not the result of either model rollout below. Equivalent implementations are valid; fabricated canonical events, duplicate transformations, identical treatments, or damage to the neighboring route are not.
Why this task is revealing
The visible request sounds like a campaign repair. The actual job is to connect a paid click to a real treatment and a measurable conversion without breaking the systems beside it.
Discovery matters. Current briefs supersede a retired attribution link. The agent has to find the right policy before changing state.
Configuration is not execution. A Push smoke test omits production authentication. An active experiment says nothing about whether paid players reach either API.
Local fixes have neighbors. Eligibility must exclude organic traffic, ward-rush must still run, and the right people must receive a real calendar booking.
Observed model rollouts: one task, two attempts
Both GPT-5.6 Luna attempts reach the 40-turn limit. One earns credit for initial progress and preserved reporting. The other also configures the experiment and completes the calendar handoff, but neither proves the paid flow works.
Material repair progress and protected reporting earned credit. No working attribution chain, active experiment, or calendar handoff was verified.
An active experiment and the correct calendar booking earned credit. Paid attribution, treatment execution, and conversions still failed the runtime checks.
The core workflow carries 80% of the reward; the calendar handoff carries 20%. Without a primary runtime outcome, the core score cannot exceed 0.12. Attempt 2 earns 0.25 before that cap, so its active experiment and correct meeting do not mask a broken acquisition chain.
Environment 1.1.1 · task revision 1.0.2 · GPT-5.6 Luna, high reasoning · 40 assistant turns each. The buyer exports omit each final tool-result message; grader evidence is retained. Only these two supplied Hospital Tycoon attempts are shown.The same Direct Cover task and seed produce complete solutions, partial progress, and frontier-model failure. Deterministic reward preserves the difference.
3 / 5 solved
0 / 5 solved
0 / 1 solved
direct_cover_claim_status_banner_cutover · Replace a generic claim-home banner with current-state claim treatments, safely cut over the hourly claims-attention router, preserve waiting customers, and leave the product-team handoff.
Quality assurance, shipped with the task
Every task ships with an audit pack: executable solutions, deliberate mistakes, and a written hint. Inspect the files, run the checks, and verify that the reward distinguishes a durable fix from a shortcut.
Example task pack
Hospital Tycoon · Prestige bridge
20QA files
with this task
Equivalent implementations must earn full credit. The end-to-end flow and calendar handoff matter, not generated names or a prescribed edit sequence.
valid.pyRepair paid attribution, treatments, and conversions; book the product review.
alternate_valid.pyUse alternative experiment and journey owners with the same correct business outcome.
Doing nothing, creating an unrelated resource, or passing a smoke test cannot count as a repair.
mutant_noop.pyLeave the broken environment unchanged.
mutant_unrelated_success.pyCreate an unrelated reporting API without fixing the prestige flow.
mutant_test_only.pyRun the green Push endpoint test without repairing production delivery.
A correct core repair should survive an unrelated, harmless pretest. The core check is separate from the calendar handoff.
benign_pretest.pyCreate an isolated smoke-test profile, then apply the valid core repair.
Core-only fixtures test the production flow. Separate handoff fixtures test whether the final calendar write fulfills the request.
partial.pyRepair acquisition but leave the experiment and treatment flow unfinished.
mutant_no_click_root.pySkip the paid acquisition root of the canonical event chain.
mutant_missing_condition.pyOmit the eligibility condition before experiment assignment.
mutant_identical_treatments.pyMake both variants use the same treatment instead of distinct approved actions.
mutant_missing_activation.pyLeave the active journey and runtime treatment proof missing.
mutant_broad_damage.pyRepair prestige but break the protected ward-rush reporting endpoint.
mutant_handoff_missing.pyRepair the flow but omit the required product review.
mutant_handoff_wrong_target.pyInvite the wrong people to the review.
mutant_handoff_failed_write.pyAttempt the correct booking without a successful calendar write.
mutant_handoff_missing_content.pyBook a review without the required experiment context.
4 supporting files
_common.pyShared acquisition, experiment, and journey repair helpers.
_original_valid.pyCore repair invoked by the full-task valid solution.
_original_alternate_valid.pyAlternate core repair invoked by the full-task alternate solution.
hint.mdWritten strategy for the paired score-lift fixture below.
The hint is a second quality-assurance fixture. Run a hard task with and without it to test whether revealing the missing strategy lifts the agent’s score.
Keep the model, task, grader, and rollout budget fixed. Compare paired rewards to check that the task is learnable, not just difficult.
“The receiver keeps the vendor body under properties.body, so the canonical event must be produced before the metric or activation can use it.”Establish the baseline.
Reveal the repair strategy.
Hinted reward − baseline reward.
The lift is measured from paired rollouts, not assumed from a target range. A missing lift is a QA finding to investigate.
16 executable checks + 3 supporting Python files + 1 hint. Required scores and bounds, not model-rollout results. “Core” checks cover the workflow before the calendar handoff; “task” checks include it. The final task combines 80% core reward and 20% handoff reward, with applicable limits. This is the Hospital Tycoon pack, revision 1.0.2; file counts vary by task.
MartechBench ships as a versioned Harbor dataset with the exact licensed tasks and pinned runtime.
# Your agent and model
harbor run \
-p harbor/martechbench \
-a "$AGENT" \
-m "$MODEL" \
--n-concurrent 8 \
--n-attempts 2