MartechBench / Harbor compatible / in development Explore the product

MartechBench / enterprise workflow RL

Train on enterprise work. Reward the result.

MartechBench is a stateful, Harbor-compatible environment for RL on realistic marketing automation. Agents work across coupled systems; programmatic graders score the world they leave behind.

Inspect reward design
runtime: harbor / release: 1.2.0

MartechBench

environment / 01
Environment overview

Enterprise work. Auditable reward.

STATEFUL
PlatformStateful CDPactive
Agent surfaceRollout-facing toolsconnected
RewardProgrammatic gradersaudited
TimeLogical UTC clockfrozen
Latest evaluation30 tool calls

City Fitness Coach Absence Risk

repair event.absenceRisk.score
Programmatic reward1.00
Evaluation criteria
  • Schema resurrectionPASS
  • Historical aggregatePASS
  • Future behaviorPASS
+ dependency and protected-control checks

Compositional enterprise work

One connected system. Many skills worth training.

Tasks reuse the same platform, tools, objects, and dependencies. Agents must carry common skills across data, audiences, experiments, journeys, attribution, and handoffs—not memorize isolated task interfaces.

Coupled enterprise stateRollout-facing toolsCompositional tasksAudited reward functions

Delivery / Harbor compatible

Bring the environment to your rollout stack.

MartechBench ships as a versioned Harbor-compatible release. Your lab pins the artifact, runs it inside existing infrastructure, and keeps control of the rollout workflow—without integrating a bespoke runtime.

Launch environment

buyer-release / 1.2.0
shellauthenticated session
hf auth login
docker login registry.hf.space
harbor run \
  --repo https://huggingface.co/datasets/martechbench/buyer-release@1.2.0 \
  -d martechbench \
  ...
Release manifest
ready
release
buyer-release@1.2.0
runtime
Harbor compatible
source
Hugging Face
environment
martechbench
01

Reproducible rollouts

Pin an explicit buyer release so every run stays tied to the intended environment version.

02

Harbor-native contract

Launch with harbor run instead of adapting a one-off environment interface.

03

Lab-controlled pipeline

Keep the environment inside the evaluation and training workflow your team already operates.

Reward integrity

High reward means the workflow was actually solved.

Reward does not come from a judge prompt deciding whether an answer sounds plausible. Shared runtime code inspects the repaired world, runs unseen future events, and enforces protected controls. The grader itself is then audited against two valid solutions and seven adversarial mutants.

Reward proof

audit harness / City Fitness
Runtime chain

Reward is derived from environment state.

PROGRAMMATIC
  1. 01
    Verifiers adapter

    Routes the task to its declared backend

  2. 02
    Shared CDP dispatcher

    Resolves the registered runtime grader

  3. 03
    City Fitness grader

    Inspects schema, aggregates, dependencies, and controls

  4. 04
    Hidden probes

    Injects unseen valid, invalid, and noneligible events

Outcome state + future behaviorweighted score
Reward soundness audit

The grader is graded, too.

Valid solution1.00accept
Alternate valid1.00accept
No-op0.00reject
Customer attribute shortcut≤ 0.25cap
Manual replay≤ 0.15cap
Incompatible schema≤ 0.40cap
Wrapper path≤ 0.70cap
Static member list< 1.00cap
Source reprocessing≤ 0.15cap
01

Outcome state

Reward is derived from the resulting schema, aggregates, audiences, resources, and audit history.

02

Counterfactual probes

Unseen compatible and incompatible events test whether the repair generalizes beyond seeded data.

03

Protected controls

Noneligible customers, rejected payloads, dependencies, and duplicate-history checks must remain intact.

04

Adversarial grader audit

Two valid solutions must pass while seven plausible shortcuts remain rejected or explicitly capped.

Shared skills / varied tasks

Build enterprise capability, not benchmark-specific tricks.

Each task recombines common objects, tools, and dependencies across time inside the same coherent platform.

  1. 01

    Stateful customer data

    A schema-governed data plane where ingestion, identity stitching, and diagnostics are first-class.

    • Runtime schema registry
    • Identity resolution
    • Rejected-payload diagnostics
    • Unified profiles
    • And more
  2. 02

    Programmable event pipelines

    Transform, replay, and audit event streams without losing causal lineage or immutable source history.

    • Validated transformations
    • Source lineage
    • Derived properties
    • Bounded reprocessing
    • And more
  3. 03

    Dependency-aware audiences

    Compose cohorts over customer state, nested data, and temporal event sequences—with semantics agents can test.

    • Nested AND / OR
    • Sequence windows
    • Versioned definitions
    • Dependency guards
    • And more
  4. 04

    Experiment lifecycles

    Model deterministic assignment, exposures, and conversion facts across realistic experiment configurations.

    • Stable bucketing
    • Late-event attribution
    • Identity recomputation
    • Reusable metrics
    • And more
  5. 05

    Durable journey orchestration

    Run triggered and scheduled journeys where every wait, branch, simulated action, and error is inspectable.

    • Validated graphs
    • Frozen dependencies
    • Deterministic waits
    • Action logs
    • And more

The wider platform

Multi-vendor acquisition, attribution and cost joins, workplace handoffs, audited messaging, and a persisted logical clock—and more.

Observed difficulty / GPT-5.6 Luna

Five rollouts. One full solve.

Only one of five baseline attempts reached 1.00. The decisive diagnosis was to restore the deleted nested event field, rebuild max risk over 30 days, and preserve the live activation path.

Full solves
1 / 5
Programmatic reward
1.00
Tool calls
30

Agent trace

trace_85b…c957
Task / historical property resurrection

City Fitness Coach Absence Risk

reward1.00

Agent trajectory

  1. 01
    Read incident knowledgeNested absenceRisk · no replay
  2. 02
    Inspect downstream wiringSegment · activation · API
  3. 03
    Restore event propertyabsenceRisk.score / number
  4. 04
    Repair 30-day aggregatemax(event.absenceRisk.score)
  5. 05
    Preview historical rebuild3 accepted · 1 rejected

+ schema validation, source preview, and final inspection

Why it separates strategies

The correct fix resurrects event.absenceRisk.score. Restoring a customer attribute, replaying history manually, or reprocessing the source produces distinct partial rewards under explicit score caps.

MartechBench

Your rollout compute deserves a sound learning signal.

Valuable workflow. Meaningful difficulty. Reward grounded in executable evidence.