San Francisco Startup C5R Builds Fully AI-Operating Research Lab In 12 Weeks

Quick Reads
- Startup C5R built Facility-0, a research lab where AI models design and run real experiments across biology, chemistry, and materials science, in just 12 weeks.
- The company released SciUniverse, a benchmark testing whether AI models can turn a scientific goal into a verified real-world result.
- Early results show models understand the science but make basic lab mistakes, like pipetting frozen samples and contaminating DNA.
San Francisco Tech startup C5R has announced the construction of a fully AI-operated research facility metaphase built in just 12 weeks.
In this new facility, artificial intelligence models independently design experiments, control laboratory instruments, and interpret results across biology, chemistry, and materials science, all without humans running the bench work.
“In 12 weeks, we built a research facility that is run entirely by AI.”
– C5R, via X
Inside Facility-0
The facility, called Facility-0, gives models a physical workspace built around an ontology of lab actions, an inventory management system, and a scheduler that coordinates machines and people.
C5R integrated more than 40 scientific instruments by reverse engineering drivers and building custom hardware adapters, turning lab equipment most software was never designed to talk to into something a model can operate directly.
Given a research goal, models explore available inventory, read equipment specs, and design experiments as code, which is then converted into equipment commands or instructions for human technicians. The models then analyze the resulting measurements and decide what to try next.
This environment is uniquely suited for frontier agentic models like OpenAI’s GPT-6 Astra, which are designed specifically for complex multi-step workflows, computer use, and autonomous software manipulation.
Because Astra is engineered to act like a native “human computer operator” interacting with custom software platforms and reading screens directly, it represents the exact class of AI capable of driving the digital backbone of this pioneering AI-operated research facility metaphase.
SciUniverse Benchmark
To measure whether any of this actually works, C5R released SciUniverse, a benchmark evaluating whether models can turn a scientific objective into a verifiable real-world result inside Facility-0 or its digital twin.
Early results point to a consistent gap: even advanced systems with elite theoretical intelligence show a major disconnect when facing physical execution. Models show strong theoretical scientific knowledge but miss details that come from hands-on lab experience.
According to C5R, models pipette frozen samples, contaminate DNA, fail to account for evaporating solvents, and vortex open containers mistakes any trained lab technician would catch instinctively.
C5R argues this reflects a structural gap between coding and science. Software and math give models instant, unambiguous verification loops, since code either runs or it doesn’t. Physical lab work gives feedback that is slow and ambiguous, and a ruined sample doesn’t return an error message.
While a model like GPT-6 Astra has saturated abstract reasoning benchmarks (such as scoring 99.9% on ARC-AGI-3) and achieved perfect scores on pure software execution tasks, SciUniverse exposes the fact that mastering the physical nuances of a wet lab requires an entirely different layer of operational feedback.
The Next Frontier of Science
C5R joins a small group of companies building AI-run lab infrastructure, including Flagship Pioneering’s Lila Sciences and Medra. These pioneers are betting that the next constraint on scientific progress is physical execution, not model intelligence.
By utilizing benchmarks like SciUniverse, the industry can finally measure how effectively frontier models can bridge the gap between digital reasoning and flawless real-world execution within an AI-operated research facility metaphase.





