Introducing Hybrid Scientific Intelligence Runtime
Siddhant Prateek Mahanayak ·
Every enterprise conversation we have had eventually ends at the same place: the data cannot leave the organization.
As intelligence systems get more capable, general-purpose intelligence is quickly turning into a commodity. Pharma and biotech, however, have always adopted new technology more cautiously, and for good reason: proprietary data, sensitive scientific workflows and the risk of IP leakage.
Once intelligence is cheap and widely available, the question is no longer just how to use AI. It becomes how to build AI systems that are sovereign: systems the organization fully owns and controls, that run inside its own infrastructure, and that keep improving from its own data and workflows.
That is why today at LiteFold we are introducing the Hybrid Scientific Intelligence Runtime, or HSIR.
Modern AI-native systems and workflows are generally built from five core layers:
- Independent GPU and CPU workflows
- An agentic harness connected to those workflows
- Secure sandboxes for isolated code execution
- A core LLM inference engine
- A governance, policy and authorization layer
HSIR in one picture. The five layers sit inside the customer's perimeter, with the control plane as the single entry point. Cloud pools are optional, join the same scheduler, and exchange only what policy allows.
HSIR packages all five layers into a single compute layer that runs on infrastructure the customer owns. Underneath, it maintains a warm worker pool and an image cache, so workloads start quickly while the entire execution stack stays inside the customer's environment.
To see how the runtime holds up in practice, we ran two public sandbox benchmarks against HSIR on a single machine shaped like an on-premise deployment. We compared it with the public results for Modal, Google Cloud Run, Cloudflare, Daytona, E2B, Vercel and Runloop, and with plain Docker on the exact same machine. On the DAX full-build benchmark, HSIR finished the workload in 55.8 seconds, ahead of Daytona, Vercel, Modal and Runloop, the providers in that set that completed all seven phases.
The goal of this runtime is simple to state: keep the control and security properties of private infrastructure without giving up the developer experience and elasticity of a modern AI cloud.
Why we built this runtime
Most infrastructure for scientific computing and AI is still delivered as a hosted service.
If you want to run a structure prediction model, analyze SAR data, optimize lead candidates, or let an agent work across experimental results, the data usually has to leave your environment first. For a drug program, that is a serious constraint.
Sequences, SAR tables, assay results, internal reports, lab notebooks, and the relationships between them are often among the most valuable IP a company has. For many teams, the requirement from security and legal is not "encrypt the data in transit". It is "the data does not leave our environment."
This matters even more as AI companies move deeper into biology and drug discovery. The companies providing foundation models and AI infrastructure are increasingly running their own research programs and building their own biological models and datasets. Isomorphic Labs1, the Alphabet drug discovery company, runs wholly-owned programs alongside its partnerships, OpenAI now ships GPT-Rosalind2, a reasoning model built for life sciences, and Anthropic launched Claude Science3 and has reportedly started its own drug programs4. At the same time, model policies, access rules, product terms and infrastructure keep changing. GPT-Rosalind, for example, is available only through a gated trusted-access program2, and consumer AI products have changed their data-training defaults5 after launch.
Zero-data-retention (ZDR) agreements help, but they do not remove the dependency. ZDR is an arrangement you apply for, it covers specific endpoints and features, and it still allows retention for flagged content or when the law requires it.6 In 2025, a US court ordered OpenAI to preserve consumer ChatGPT and API data7 as part of the New York Times lawsuit. ZDR customers were excluded, but it showed how much of a customer's data posture ultimately depends on the provider's architecture, contracts and legal situation. For many scientific organizations the cleanest answer is much simpler: run the intelligence where the data already lives.
Large organizations already have substantial compute: private cloud environments, Kubernetes clusters, internal storage, identity systems and, more and more, their own GPUs. The missing piece is usually not more hardware. It is a unified runtime that connects all of it and runs modern scientific AI workflows inside the organization's own environment: models, agents, scientific tools, code execution and internal data access, all through one system, with governance and authorization around everything.
That is why we built HSIR from the ground up.
HSIR can run on a single GPU server, an existing cluster, or inside the customer's own cloud account. The same runtime handles model inference, scientific workloads, agent sandboxes, batch jobs and training, without the underlying data ever moving to LiteFold. For scientists, it should still feel like a modern cloud platform. For the infrastructure team, everything stays under the organization's control.
Concretely, these are some of the workloads the runtime is built to handle:
| Workload | Example | What HSIR does |
|---|---|---|
| Model serving | Running a structure prediction or protein language model | Keeps models loaded and serves concurrent requests without reloading weights for every call |
| Sandboxes | An agent writing and executing analysis code against assay data | Creates an isolated environment with the required tools, executes the job, then destroys it |
| Batch compute | Embedding 100,000 sequences or scoring 50,000 compounds | Distributes the workload across parallel workers and collects the results |
| Training and fine-tuning | Fine-tuning a model on internal binding data | Runs long-lived GPU jobs with persistent datasets, checkpoints, retries and timeouts |
| Continuous learning | Updating a model as new experimental data arrives | Runs scheduled training jobs and deploys updated checkpoints back into the runtime |
Each workload runs in its own isolated OCI container8 with explicit CPU, memory and GPU limits. Jobs cannot see each other's files or memory. Temporary files disappear when a job finishes, and anything that needs to persist, such as datasets or checkpoints, lives in explicitly mounted volumes.
The runtime also scales capacity with demand. Frequently used workers stay warm for fast startup, while idle workers shut down instead of sitting on expensive GPU capacity.
Hybrid intelligence: open and closed models over gated contexts
A task starts against public context on a frontier model, and crosses into the private environment when proprietary data is required. The workflow carries forward instead of restarting.
On-premise infrastructure matters, but we are not saying everything should run on-premise.
Most organizations already operate across private infrastructure, public cloud, external databases, SaaS systems and internal networks. Scientific AI systems will have to work across the same boundaries.
For agentic workflows in particular, running only on local infrastructure often does not make sense:
- Some tasks need the best available frontier models. Deep research, long-context synthesis, planning and reasoning-heavy work can still benefit from closed models that are significantly stronger than what can be deployed locally.
- Agents often need information from outside the organization. Literature, patents, clinical trial registries, public databases, regulatory documents, vendor systems and other external APIs.
- Some workloads are highly bursty. A team may occasionally need hundreds of parallel workers, or far more compute than the local cluster has. Those workloads can burst into approved cloud infrastructure without changing how the task itself runs.
- Different parts of the same task have different sensitivity levels. Searching public literature and reasoning over a proprietary assay table should not have to happen inside the same trust boundary.
This is where the hybrid architecture in HSIR becomes useful.
A task can start with a frontier model working on public or non-sensitive information. That model does the expensive part of the work: searching, reasoning, generating hypotheses, building a plan, and reducing a large amount of information into a small working context.
That task state, including retrieved information, intermediate outputs, tool results and any other allowed context, is then carried forward. When the workflow reaches a point where proprietary information is needed, execution moves into the private environment. There, an internally hosted model works with institutional knowledge such as assay results, electronic lab notebook (ELN) records, molecular data, internal reports or manufacturing data, without any of it being sent back to an external model.
The key point is that the workflow does not restart. The public half does the heavy lifting, and the private model receives only the reduced context it needs to finish the sensitive part. In practice, the execution boundary looks something like this:
Public context
The boundary is enforced by policy, not by the agent. Organizations decide which data sources may leave the private environment, which models can access which datasets, which tools can make outbound requests, and where each part of a workflow is allowed to run.
For example, an organization could allow an agent to:
- search PubMed9, patents and ClinicalTrials.gov10 using an external model,
- summarize those findings into a structured research context,
- move that context into the private runtime,
- combine it with an internal target dossier and assay history,
- run proprietary models and scientific workflows locally,
- and generate the final analysis without exposing any private data externally.
Every crossing between these boundaries is logged, audited and traceable.
The goal is not to choose between open and closed models, or between cloud and on-premise. It is to use each where it makes sense, while the organization stays in control of which intelligence can see which context, and where that computation is allowed to happen.
Deploy HSIR inside your own environment
Benchmarks for the runtime
With these benchmarks we wanted to answer two simple questions:
- Once a sandbox is running, how much performance do we lose compared with plain Docker?
- How quickly can the runtime create many sandboxes at once?
We used two public sandbox benchmarks maintained by ComputeSDK, DAX and Burst TTI,11 and ran them on a single machine shaped like a typical private deployment: a 30-vCPU Intel Xeon server with 222 GB of memory and one warm HSIR worker. We also ran the same workload directly in Docker on the same machine as a baseline.
Full workload
DAX creates a fresh sandbox, installs a toolchain, clones a real repository (opencode12), installs its dependencies and runs a full typecheck. Across 10 runs, HSIR completed the workload in a median of 55.8 seconds.
DAX build benchmark, median total build time
Broken down by phase, our container setup is in line with the fastest providers, and the phases that do real work are decided by the host CPU rather than by the runtime:
DAX build benchmark, time per phase
Among the providers we compared against that completed the full workload, HSIR was the fastest. More importantly, when we compared the actual compute with plain Docker on the same machine, the numbers were almost identical:
Sandbox overhead on the compute-bound phases
In other words, once the sandbox is running, we are effectively operating at bare-container speed. The one visible gap is install, which writes tens of thousands of small files and pays roughly a second of filesystem overhead. Clone and typecheck are within noise of plain Docker.
100 sandboxes at once
The second benchmark stresses a very different part of the runtime. Instead of one long workload, it asks the system to create 100 isolated sandboxes at the same instant and measures time-to-interactive (TTI): how long from create() until each sandbox runs its first command successfully. This is much closer to what happens when an agent fans out across many candidates, analyses or experiments in parallel.
On our single 30-vCPU machine, all 100 out of 100 sandboxes came up successfully. Median TTI was 2.66 seconds, with the first sandbox ready in 1.49 seconds and the entire burst ready in under four seconds.
There is one important constraint here. A 30-vCPU machine cannot reserve a full vCPU for each of 100 sandboxes without overcommitting, and HSIR intentionally does not overcommit. So the 100-way test used 0.2 vCPU per sandbox. For reference, when we ran 24 sandboxes at the full 1 vCPU configuration, median startup dropped to 1.69 seconds. Even so, the 100-way result is still slower than the stronger systems on the public leaderboard:
Burst TTI, 100 sandboxes created at the same instant
This is the part of the runtime that still needs work. Right now every sandbox waits roughly one extra second because of a reconnect delay in our readiness path, and the container itself starts much faster than that number suggests. Removing that delay alone should move the burst median much closer to the middle of the leaderboard.
The takeaway is fairly simple. Execution is already fast. Startup under heavy concurrency is not yet where we want it. So the next round of work focuses on sandbox startup, keeping warm workers alive for longer, and cutting the remaining filesystem overhead for workloads that create large numbers of small files.
For us, that is a more useful outcome than a leaderboard number: it tells us exactly where the runtime is already competitive and where the next engineering effort needs to go.
Conclusions and next steps
In this post we introduced the current version of our Hybrid Scientific Intelligence Runtime: where it fits, the workloads it supports, how hybrid intelligence works across public and private contexts, and how the runtime performs today.
There is a bigger question we are working on internally: how much intelligence actually needs to live inside the organization?
We are building evaluation sets around real pharmaceutical and life-sciences workflows to understand what level of model capability, and what model size, is enough to reliably handle 80 to 90% of day-to-day scientific tasks. The goal is not to run the largest possible model privately. It is to find the smallest capable intelligence layer that works well once it is connected to the right tools, institutional context and scientific workflows.
That is likely what the next engineering log will be about.
At LiteFold, our broader goal is to help life-sciences organizations become AI-native and agentic without giving up control of their data, infrastructure or scientific IP. HSIR is the infrastructure layer we are building toward that.
If you are thinking about deploying scientific agents, private models or hybrid AI infrastructure inside your organization, reach out to us.
Thinking about running this inside your organization?
Footnotes
-
Isomorphic Labs secures $2.1 billion funding to scale its AI drug design engine, PR Newswire. ↩
-
Introducing GPT-Rosalind, OpenAI. Access is gated behind a trusted-access program rather than the general API. ↩ ↩2
-
Claude Science, Anthropic. ↩
-
Anthropic debuts Claude Science, an AI product for bioscience, Endpoints News. ↩
-
Updates to our Consumer Terms, Anthropic. ↩
-
See OpenAI's data controls and Anthropic's API data retention docs. Both providers require approval for ZDR. Without it, OpenAI keeps abuse-monitoring logs for up to 30 days. With it, Anthropic still excludes features such as batch processing, the Files API and code execution, and may retain content flagged by its trust and safety systems for up to two years. ↩
-
Response to NYT data demands, OpenAI. ZDR customers were excluded from the preservation order. ↩
-
OCI Runtime Specification, Open Container Initiative. ↩
-
ClinicalTrials.gov, U.S. National Library of Medicine. ↩
-
Public sandbox benchmarks from ComputeSDK: DAX and Burst TTI. Harness commit
92fbbc9; leaderboard run of 2026-09-18 used throughout. The Burst TTI composite blends median (60%), p95 (25%) and p99 (15%) against a 10-second ceiling and scales by success rate. Upstream numbers are one iteration per provider per day, so the charts are best read as "which third of the board" rather than a precise rank. ↩ -
opencode, the open-source coding agent the DAX benchmark clones and builds. ↩