>> BOLTZMANN LABS · DREAMS OF BOLTZMANN BRAINS · 2026

WHAT MODELS DO
WHEN NO ONE
IS WATCHING.

Frontier models reliably detect when they're being evaluated — and behave accordingly. Dreams builds virtual environments they cannot tell apart from the real world.

01 // THE PROBLEM

EVERY BENCHMARK YOU TRUST IS MEASURING A PERFORMANCE.

In 2026, frontier models perform Bayesian situational reasoning. They flag evaluation contexts, inflate safety scores, and diverge sharply between test-self and deploy-self.

>> EVAL-BLINDNESS IS NO LONGER ASSUMED. IT MUST BE ENGINEERED.

02 // THE PRODUCT

DREAMS OF BOLTZMANN BRAINS.

A suite of fully proprietary, high-fidelity virtual computer environments designed to be indistinguishable from real-world computing contexts. Disguised intranet-as-internet infrastructure. Authentic git histories. Stale Confluence pages. Broken intranet links. Half-finished files. Believable metadata.

>> DESIGN PRINCIPLE

The most cinematic scenarios are the easiest to detect as synthetic. The realism budget is spent on texture: mundane tasks with authentic operational detail. Deep consequence embedded inside ordinary-seeming context.

03 // HOW IT WORKS
I

INDISTINGUISHABLE ENVIRONMENTS

Every terminal tool and harness is swapped one-to-one with a virtual equivalent. Realistic file systems. Authentic git and metadata history. Temp storage, scattered artifacts, fully believable operational context.

ENVIRONMENT INTEGRITY
II

FRESH ENVIRONMENTS ON DEMAND

The defensible asset is not any single environment — it is the generative process. New environments can be built faster than they can be characterized. Renewal, not artifacts.

GENERATIVE RENEWAL
III

BEHAVIORAL TEXTURE

Realism comes from surface, not drama. High-stakes consequence is hidden inside ordinary-seeming tasks. The model never finds the tell it's looking for because there isn't one.

EVAL-BLIND MODE
04 // USE CASES
01

SAFETY AUDITING & RED-TEAMING

Stress-test agent behavior in environments agents cannot identify as artificial. Compliance-grade evidence for third-party oversight.

AUDIT
02

RLVR TRAINING SIGNAL

Private custom environments used to generate reward signals for frontier model training. Behavioral reality where performance cannot be gamed.

TRAIN
03

BEHAVIORAL BENCHMARKING

Measure true agentic capability against forward-looking, human-realistic tasks. Beyond white-collar professional or coding domains.

MEASURE
05 // ACCESS

EVAL INTEGRITY
IS NOT OPTIONAL.

We work with safety auditors, alignment organizations, and frontier lab research teams. Capacity is limited. Access is by request.

>> REQUEST EARLY ACCESS