Research

We are a lab too.

Human simulation, evals judged by real people, and RL environments grounded in real interviews. Selling human data funds the work.

A self-improving model only gets as good as its verifier. For anything involving people, the verifier is a person.

For code, a test suite checks the answer. For a diagnosis, a contract or a conversation, someone has to judge it, and the quality of that judgment caps the whole loop. So we study it: how to model real people, how to measure a model against them, and how to build worlds where agents learn from them.

  • Commercial today. Human data: license, commission, label and review.
  • Research now. Human simulation and human-judged benchmarks.
  • Research next. Professional RL environments and a full feedback loop. Not a product.

Programs

Two programs underway.

Underway

Human simulation

Model individuals from real interviews and personal context. Predict what they will say, compare with what they actually said, and study where the prediction fails. We start from 190,000+ real interviews and the people behind them.

Field references: Aaru, Simile. A research area, not affiliations.

Underway

Evals judged by real people

Benchmarks scored by verified experts or by everyday consumers, depending on who the answer is for. Two in progress: human-likeness (can ordinary people spot the AI-written line in a chat between friends?) and interviewer (how well can a model run a 30-minute interview, and how does it decay as context grows?).

We train open models on our own data and publish the delta, either way, with the judges' profiles and the method.

190,000+research interviews run, with recordings and transcripts
170M+people reachable through our sourcing network
30,000+local recruiting partners for hard-to-find cohorts
$7M+raised from Lightspeed Luminous and other top-tier VCs

RL environments · Research

A hospital, a law firm, a store.

Worlds where an agent does real professional work with virtual people grounded in real interviews, and experts grade every step. Research today, not yet a product.

Virtual hospital

An agent works through intake, questions and follow-up with virtual patients built from real interviews. Clinicians grade each step.

Law firm

An agent handles a firm's client intake: the first call, the conflict check, the engagement letter. Lawyers grade each step.

E-commerce store

An agent runs a store's support queue: refunds, lost parcels, the customer who changes their mind. Operators grade each step.

Where it leads

A loop that asks for real people when it runs out.

Model individuals, compare them with real people, train agents in the environment, and when a signal is missing, request new human data from the same supply chain that serves our customers today. Small, open experiments with public methods, and results we would be comfortable being wrong about in public.

PROPOSED LOOP · RESEARCH01 Model individuals02 Compare with real people03 Train in the environment04 Request new human datafrom real interviews and contextfind where predictions failfind the missing signalthrough a CLI or API↺ New real evidence goes back into the environmentConsumers, experts and records from our existing supply

Questions teams ask.

Can we buy an RL environment now?

No. The environments and the full self-improvement loop are research, not a commercial offering. What we sell today is human data: licensed, commissioned, or labeled and reviewed.

Can we collaborate on research?

Yes. Bring a question about human behavior, evaluation or data requirements, or a benchmark you do not trust. We can design an experiment together.

Are you hiring researchers?

Yes. We are hiring a Chief Research Scientist and research engineers in San Francisco. Message Davin.

Bring a question.
Or a benchmark you do not trust.

Video on this site is licensed stock footage of real people, used to illustrate the kinds of data we collect.