Skip to content
An empty office at night where nearly closed laptops glow beside a ReadySolve desk-lamp robot
Contact

MartechBench

RL environment for
knowledge work

1,300 long-horizon tasks, deterministic graders verifying the final application state

Inside a rollout

One task. Every consequential state change.

This is a working view of the task, trace, application state, and grader—not a product screenshot. Switch runs, inspect calls, and follow the evidence.

rollout completeGPT-5.6 Luna · high reasoning93 tool calls0.296 reward

event_transform_hospital_tycoon_prestige_bridge

Repair the paid prestige recovery flow without breaking the world beside it.

AppLovin traffic reaches the mobile feed, but useful conversions do not. Reconnect attribution, preserve identity, execute two treatments, protect ward-rush reporting, and book the earliest product review.

  • 40-turn limit
  • 6 connected systems
  • Runtime verified
Buyer trace · Environment 1.1.1Download Attempt 2 JSON
426
Rollout tools
1,264
Verifiable tasks
10
Operating surfaces
0-1
Continuous reward

Workflow execution, verified

Real workflows.
Verifiable outcomes.

Agents investigate live state, coordinate across applications, execute changes, and pass runtime checks. A plausible answer earns nothing by itself.

Marketing is one connected enterprise world.

Tasks share infrastructure, state, and dependencies across data, delivery, channels, and workplace systems. Agents learn reusable operating skills, not isolated API calls.

Data

40 tools

Customer data & orchestration39 tools

Schemas, transformations, profiles, segments, experiments, activations, integrations, and logs

Reporting1 tool

Bounded customer-event reports and grouped behavioral aggregates

Delivery

5 tools

Remote config5 tools

Versioned treatments, precedence, publication, evaluation, and exposure provenance

Channels

345 tools

Mobile growth51 tools

AppLovin and AppsFlyer apps, credentials, campaigns, attribution, push, cost, postbacks, and stream mappings

SendGrid45 tools

Accounts, senders, authentication, templates, tracking, webhooks, messages, and delivery

Google Ads113 tools

Accounts, campaigns, criteria, assets, bidding, conversions, audiences, leads, and reporting

Meta Ads136 tools

Business assets, audiences, pixels, campaigns, creative, leads, diagnostics, and history

Coordination

36 tools

Workplace24 tools

Company knowledge, employees, teams, Slack, Confluence, email, and calendars

DNS5 tools

Zones, records, and controlled changes

Access profiles7 tools

Role-aware capability across each major operating surface

One task, end to end

From a paid click
to a proven outcome.

Hospital Tycoon’s prestige campaign has traffic, but no useful conversions. Repairing it means connecting attribution, event transformation, experimentation, and real treatment delivery while protecting Operations’ reporting and coordinating the product team.

Two GPT-5.6 Luna buyer traces, with tool calls, recorded results, and grader evidence. JSON download: Attempt 1.

Hospital Tycoon / Prestige bridge

“AppLovin traffic is arriving in the mobile feed, but the recovery experiment has no useful conversions and the prestige-start follow-up is not reaching the right players.”

Restore the paid flow, route players through two approved treatments, leave organic installs out, and preserve ward-rush reporting. Then book the product team’s earliest shared review.

Task
Event transformation + paid prestige recovery
Connected systems
AppLovin, AppsFlyer, CDP, A/B testing, Journeys, Workplace
The trap
A green test endpoint and an active experiment can coexist with a broken production flow.
One incident. An end-to-end business outcome.

Why this task is revealing

Every link has to work.

The visible request sounds like a campaign repair. The actual job is to connect a paid click to a real treatment and a measurable conversion without breaking the systems beside it.

  1. AppLovinPaid clickCurrent attribution link
  2. AppsFlyerInstall / re-attributionAuthenticated production Push
  3. CDPmobile:attributionCanonical identity + campaign
  4. Journeys + A/B testingEligible paid playersCondition → assignment → API
  5. Measurementprestige:startedDerived conversion evidence

Discovery matters. Current briefs supersede a retired attribution link. The agent has to find the right policy before changing state.

Configuration is not execution. A Push smoke test omits production authentication. An active experiment says nothing about whether paid players reach either API.

Local fixes have neighbors. Eligibility must exclude organic traffic, ward-rush must still run, and the right people must receive a real calendar booking.

Observed model rollouts: one task, two attempts

Granular signal every time.

Both GPT-5.6 Luna attempts reach the 40-turn limit. One earns credit for initial progress and preserved reporting. The other also configures the experiment and completes the calendar handoff, but neither proves the paid flow works.

  1. Attempt 1 · 88 tool calls0.088View full trace

    Material repair progress and protected reporting earned credit. No working attribution chain, active experiment, or calendar handoff was verified.

  2. Attempt 2 · 93 tool calls0.296View full trace

    An active experiment and the correct calendar booking earned credit. Paid attribution, treatment execution, and conversions still failed the runtime checks.

The core workflow carries 80% of the reward; the calendar handoff carries 20%. Without a primary runtime outcome, the core score cannot exceed 0.12. Attempt 2 earns 0.25 before that cap, so its active experiment and correct meeting do not mask a broken acquisition chain.

Environment 1.1.1 · task revision 1.0.2 · GPT-5.6 Luna, high reasoning · 40 assistant turns each. The buyer exports omit each final tool-result message; grader evidence is retained. Only these two supplied Hospital Tycoon attempts are shown.

Tasks that stump frontier models.

The same Direct Cover task and seed produce complete solutions, partial progress, and frontier-model failure. Deterministic reward preserves the difference.

DeepSeek V4 Flash

0731

3 / 5 solved

Mean reward0.840

GPT-5.6 Luna

High

0 / 5 solved

Mean reward0.336

GPT-6 Astra

Max

0 / 1 solved

Mean reward0.560

direct_cover_claim_status_banner_cutover · Replace a generic claim-home banner with current-state claim treatments, safely cut over the hourly claims-attention router, preserve waiting customers, and leave the product-team handoff.

Quality assurance, shipped with the task

Programmatic rewards from real outcomes.

Every task ships with an audit pack: executable solutions, deliberate mistakes, and a written hint. Inspect the files, run the checks, and verify that the reward distinguishes a durable fix from a shortcut.

  • Different valid solutions earn full credit
  • No-ops earn zero; incomplete repairs hit score limits
  • Hidden future probes check that the fix generalizes
  • Deterministic state and runtime checks. No LLM judge.

Example task pack

Each task is an audit pack you can inspect.

Hospital Tycoon · Prestige bridge

20QA files
with this task

2 files

Accept different valid solutions

Equivalent implementations must earn full credit. The end-to-end flow and calendar handoff matter, not generated names or a prescribed edit sequence.

  • valid.py

    Repair paid attribution, treatments, and conversions; book the product review.

    Required reward: 1.00Full task
  • alternate_valid.py

    Use alternative experiment and journey owners with the same correct business outcome.

    Required reward: 1.00Full task

16 executable checks + 3 supporting Python files + 1 hint. Required scores and bounds, not model-rollout results. “Core” checks cover the workflow before the calendar handoff; “task” checks include it. The final task combines 80% core reward and 20% handoff reward, with applicable limits. This is the Hospital Tycoon pack, revision 1.0.2; file counts vary by task.

Run it in the rollout stack you already have.

MartechBench ships as a versioned Harbor dataset with the exact licensed tasks and pinned runtime.

You control
Agent, model, concurrency, attempts
We ship
Pinned runtime, tasks, and isolated verifiers
You get
Continuous rewards in Harbor job artifacts
Acquire a license
Harbor
# Your agent and model
harbor run \
  -p harbor/martechbench \
  -a "$AGENT" \
  -m "$MODEL" \
  --n-concurrent 8 \
  --n-attempts 2