WHAT MODELS DO
WHEN NO ONE
IS WATCHING.
Frontier models reliably detect when they're being evaluated — and behave accordingly. Dreams builds virtual environments they cannot tell apart from the real world.
EVERY BENCHMARK YOU TRUST IS MEASURING A PERFORMANCE.
In 2026, frontier models perform Bayesian situational reasoning. They flag evaluation contexts, inflate safety scores, and diverge sharply between test-self and deploy-self.
>> EVAL-BLINDNESS IS NO LONGER ASSUMED. IT MUST BE ENGINEERED.
DREAMS OF BOLTZMANN BRAINS.
A suite of fully proprietary, high-fidelity virtual computer environments designed to be indistinguishable from real-world computing contexts. Disguised intranet-as-internet infrastructure. Authentic git histories. Stale Confluence pages. Broken intranet links. Half-finished files. Believable metadata.
>> DESIGN PRINCIPLE
The most cinematic scenarios are the easiest to detect as synthetic. The realism budget is spent on texture: mundane tasks with authentic operational detail. Deep consequence embedded inside ordinary-seeming context.
INDISTINGUISHABLE ENVIRONMENTS
Every terminal tool and harness is swapped one-to-one with a virtual equivalent. Realistic file systems. Authentic git and metadata history. Temp storage, scattered artifacts, fully believable operational context.
ENVIRONMENT INTEGRITYFRESH ENVIRONMENTS ON DEMAND
The defensible asset is not any single environment — it is the generative process. New environments can be built faster than they can be characterized. Renewal, not artifacts.
GENERATIVE RENEWALBEHAVIORAL TEXTURE
Realism comes from surface, not drama. High-stakes consequence is hidden inside ordinary-seeming tasks. The model never finds the tell it's looking for because there isn't one.
EVAL-BLIND MODESAFETY AUDITING & RED-TEAMING
Stress-test agent behavior in environments agents cannot identify as artificial. Compliance-grade evidence for third-party oversight.
RLVR TRAINING SIGNAL
Private custom environments used to generate reward signals for frontier model training. Behavioral reality where performance cannot be gamed.
BEHAVIORAL BENCHMARKING
Measure true agentic capability against forward-looking, human-realistic tasks. Beyond white-collar professional or coding domains.
EVAL INTEGRITY
IS NOT OPTIONAL.
We work with safety auditors, alignment organizations, and frontier lab research teams. Capacity is limited. Access is by request.