AI Training Data Broker

For buyers3 min read

What is RL data? What AI labs mean when they ask for it

Reinforcement learning environments, graders and trajectories explained for business owners, why labs keep asking for them, and what it means for the price of your records.

By Jimmy.


RL is reinforcement learning: a model tries a task, gets scored, and learns from the score. An RL environment is a sandboxed copy of a real tool or workflow where an AI agent practices tasks and gets graded. "RL data" is shorthand for everything that feeds that: the environment, the tasks inside it, the grader that scores each attempt, and the recorded trajectories of attempts. Labs ask for it because the open internet is used up and the next skill they are buying is "doing real work inside real tools," which only exists behind company firewalls.

The four pieces

Term What it is Where it comes from
Environment A reset-able replica of an app or workflow with realistic data and state: a CRM, a ticket queue, a chip-design toolchain, a browser Built by a vendor or lab, often seeded with a real company's records
Task A goal plus a starting state. "Rebook this delayed shipment." "Close this ticket correctly." Mined from real tickets, threads and procedures
Grader (also called verifier or reward) Code or rules that check whether the task was really done, by inspecting the system's state, not by reading the agent's text Written by experts. The hardest and most valuable part
Trajectory The full record of one attempt: every action, state change and intermediate result Produced by running agents in the environment, or reconstructed from real work history

Scale AI's description is the clearest public one: environments "record every action, state change, and intermediate outcome," and verifiers check "for real, measurable changes in the system," not "whether an agent produced plausible text." One buyer's founder put the current ask more bluntly in October 2026: labs want "RL environments with proper graders," and the hot item that month was environments for physical chip design "graded by actual signoff tools."

Why it matters if you are selling records

You are not being asked to build an environment. You are being told that your records are worth more when they show complete, verifiable workflows with outcomes, because that is what gets turned into tasks and graders. Think of the layers:

  1. Raw records. Emails, tickets, docs, CRM rows, code. What a company has.
  2. Workflow data. The same records reconstructed as request, systems used, actions, handoffs, outcome.
  3. Trajectories. Structured sequences of actions toward an outcome.
  4. Environments. A controlled copy where an agent performs those tasks and is scored.

Each step up adds value and work. The seller supplies the seed. micro1 says outright that it uses de-identified company data to build reinforcement learning environments that reflect real business operations, and has committed a billion dollars over twelve months to buying the seed. So when a buyer asks whether your tickets close with a resolution, whether deals are marked won or lost, whether approvals are logged, that is the RL question in plain clothes.

What it pays

Only one source publishes unit prices, and they are for finished environments, not for seed records:

  • A single task with a goal and a verifier: about $200 to $2,000.
  • An interface replica of a website or app: around $20,000.
  • A high-fidelity replica of a complex product: around $300,000.
  • Exclusive: roughly four to five times the non-exclusive price.

Market signals: Anthropic's leadership reportedly discussed spending more than $1 billion on environments in a year. Scale AI has said nearly half of its new data projects involve environments. A vendor directory counts about 40 companies selling them. There are skeptics on the record too: an OpenAI executive said he is "short" on environment startups, and Andrej Karpathy has said he is bullish on environments but bearish on reinforcement learning specifically.

The practical version for a company owner

  • Your support tickets with resolutions are worth more than your email archive.
  • Your chat threads where a decision got argued and changed are worth more than your finished reports.
  • A record that spans ticket, chat and CRM for the same event is worth more than any one of them.
  • Phone-first businesses leave thin records: "ticket created, ticket closed."
  • Industries with few written records (trades, labs, chip design) are scarce and priced that way.

Questions people ask

Do I need to format my data as trajectories before selling?

No. Buyers and vendors do that work and it is where much of their margin sits. Your job is to have the records connected and the outcomes recorded.

What is reward hacking?

An agent finds a way to score well on the grader without doing the task. Buyers test graders hard for it before paying, which is why a grader built from real signoff rules is worth more than one that checks for plausible text.

Is this only for software companies?

No. The industries buyers named in October 2026 were finance, legal, semiconductors, biotech, manufacturing and aerospace, plus "weirdly specific workflow data" from trades like roofing, tax and freight.

Sources

Drafted with AI tools, checked and edited by Jimmy, last reviewed October 10, 2026. Not legal advice.


See what I can source