PROJECT

OSVA

Open-source infrastructure for running AI agents with durable state and controlled access.

Agent InfrastructureOpen SourceProduct & Tech

What is it?

EXECUTION

Container exited

The runtime process was gone

OUTPUT

Artifact still readable

Its bytes remained available

AUTHORITY

Metadata in PostgreSQL

The execution record survived

VERIFIED BY AN INTEGRATION TEST

The test proved that an artifact could outlive the container that created it. This is test evidence, not a public demo.

OSVA, short for Operating System for Virtual Agents, is an open-source platform for developers building applications with AI agents. You define agents and workflows, submit work, and OSVA manages execution, durable state, permissions, artifacts, and knowledge retrieval.

The one-line pitch: it is the infrastructure between an agent that works in a demo and an agent you can operate as part of a product.

Here is the point I care about most, and it is a small one. The agent's container can exit. The artifact it produced is still available.

In our container artifact integration test, an agent created an artifact through OSVA and got back a reference. After the container exited, the bytes were still readable. The authoritative record of that artifact lives in PostgreSQL, not in the process that made it.

To be straight about it, that is verified by an integration test today, not by a click-through demo. The demo is still on my list.

The thesis behind the whole project fits in one sentence. Agent execution should be a managed product capability, with explicit state, permissions, and recovery.

Source →

Why we built it?

An agent finishes a task on a developer's machine. Good. Now turn that into a product and a different set of questions shows up. Where does its output live? What can it access? What happens when execution fails halfway? How does a workflow pick back up after waiting on something outside the system?

Generating an answer and operating an agent are two different responsibilities. Most demos only do the first one.

I want to be careful about what I am claiming. I am not saying existing agent frameworks lack orchestration or persistence. Plenty of them are good at it. The boundary I chose is different. OSVA owns execution infrastructure and the public contracts around it, and agent logic can use its own libraries behind the runtime interface.

Three bets sit underneath the design. They are judgments, and I could be wrong about any of them.

Developers will need freedom to change how agents run. So runtime execution is separated from the platform's durable state. A particular agent implementation should not quietly become the architecture.

Recovery and permissions will matter as soon as an agent does anything useful. I would rather build them into the platform boundary than leave every application to reconstruct them on its own.

An open-source edition has to solve a complete problem. The Community Edition should be useful by itself. Organizational controls and deployment requirements are reserved for future commercial editions.

I built OSVA as an open-source product foundation. It also gives me a concrete way to talk about product scope, architecture, security, and build-versus-buy decisions in leadership and advisory conversations.

How it works?

Seven concepts hold the whole thing together. Once these click, the rest of the page makes sense.

Agent. A definition of the work an agent can perform and how it should run.

Run. A tracked execution of an agent against an input.

Attempt. An individual execution attempt within a run.

Runtime. The environment that executes agent code, such as a container or an HTTP service.

Workflow. A defined process connecting work, including waiting for external events.

Capability. An operation the runtime can request through OSVA without ever receiving the underlying credentials.

Artifact. A stored output with durable metadata and a reference that other work can use.

AGENTDefines work and how it runsinstantiated asRUNOne tracked executioncontains one or moreATTEMPTOne execution attemptexecuted byRUNTIMERuns agent coderequestsCAPABILITYControlled operationproducesARTIFACTDurable outputWORKFLOWOptional coordinatorcoordinatesOptional for multi-step or event-driven work

Durable authority

PostgreSQL holds the authoritative durable state. BullMQ and Valkey provide job transport.

The technology names matter less than the reasoning. A queue message should not be the only record of what work exists or what happened to it. And process memory should not decide whether a run or a workflow survives.

CONTROL PLANEDecides, records, authorizes, and coordinatesAPI and SDK boundaryPublic operationsApplication servicesProduct rulesAuthorization and capability boundaryScoped accessWorkflow coordinationWaits and eventsEXECUTION PLANERuns agent code through interchangeable runtimesWORKERTrusted TypeScript RuntimeRemote HTTP RuntimeContainer RuntimePOSTGRESQLDURABLE AUTHORITYAgents, runs, attempts, workflows,waits, events, artifacts, principalsState survives workers and processes.BULLMQ + VALKEYTRANSPORTDispatches work to workersThe queue moves work. PostgreSQL records what happened.BLOB STORAGEARTIFACT BYTESFilesystem or S3-compatible storagedispatchcapability requests

Here is what one run looks like from submission to recorded outcome.

1 · SUBMITCaller submits inputHTTP contract2 · RECORDRun and Attempt in PostgreSQLThe request now exists ind…3 · DISPATCHTransport onto BullMQTransport, not authority4 · WORKERLoads persisted Attempt5 · EXECUTERuntime runs agent codeTrusted TS · Remote HTTP ·…6 · CAPABILITIESApproved operationsCreate artifact · Read art…7 · OUTCOMERecorded in durable stateCompleted · Failed · WaitingIs this allowed?metadata → PostgreSQLbytes → Blob storageExecution can end. The recorded outcome and durable outputs remain.

Execution separated from credentials

Agent runtimes do not receive database, vector-store, provider, or secret credentials. They ask OSVA to perform supported operations through capability interfaces.

That draws an explicit line between executing agent code and touching platform resources.

AGENT RUNTIMEAgent codeInputRuntime-local processingCapability clientThe Runtime receives scoped execution contextOSVA CAPABILITY BOUNDARYrequest →← resultOSVA-CONTROLLED RESOURCESAuthorization checkArtifact application serviceKnowledge retrieval servicePostgreSQLBlob storageProvider adaptersSecret resolutionDatabase credentialsVector-store credentialsProvider secrets
{
  "operation": "artifacts.create",
  "input": {
    "contentType": "application/json"
  }
}
Requests cross the boundary. Credentials do not.This reduces credential exposure. It does not make untrusted agent code automatically safe.

I want to be clear about the limit here. This is a boundary in the design, not a claim of complete sandbox security. Runtime isolation and deployment policy still matter.

Waiting is a durable state

A workflow can enter WAITING, persist that wait, and resume when a matching event arrives. Event handling is idempotent, which means the same event delivered twice should not cause the same effect twice.

The product reasoning is simple. Waiting on something outside the system is part of the process, so it needs a durable representation, not a sleeping thread.

RUNNINGExecutes until a waitWait condition created and persistedWAITINGWait persisted in PostgreSQLEVENT RECEIVEDCheck against persisted waitRESUMEDContinues from persisted stateCONTINUESNext executable stepMatching event?NO → remain WAITINGAlready processed?YES → no repeated effectNO → resume onceIdempotency prevents repeated delivery from producing a second resume.The worker does not need to stay alive while the workflow waits.

The operating system analogy

The metaphor holds in three places: execution management, durable state, and permission boundaries. I am not going to force a one-to-one mapping with a computer operating system, because past those three places it stops being true.

How to use it?

Setup

cd deploy/compose
cp .env.example .env
# Set POSTGRES_PASSWORD in .env
docker compose up -d --build --wait

docker compose run --rm migrate

docker compose --profile bootstrap run --rm bootstrap

export OSVA_API_KEY='osva_ak_…'   # from bootstrap output
curl -sS -H "Authorization: Bearer $OSVA_API_KEY" \
  http://127.0.0.1:8080/v1/api-keys | jq .

cd /path/to/osva
pnpm install
pnpm exec turbo run build --filter=@osva/sdk
export OSVA_BASE_URL=http://127.0.0.1:8080
export OSVA_API_KEY='osva_ak_…'
node --input-type=module -e "
  import { OsvaClient } from '@osva/sdk';
  const client = new OsvaClient({ baseUrl: process.env.OSVA_BASE_URL, apiKey: process.env.OSVA_API_KEY });
  const keys = await client.apiKeys.list();
  console.log(keys);
"

python -m pip install ./sdks/python   # or pip install osva-sdk==1.0.0 when on PyPI
export OSVA_BASE_URL=http://127.0.0.1:8080
export OSVA_API_KEY='osva_ak_…'
python -c "
import os
from osva import OSVAClient
client = OSVAClient(base_url=os.environ['OSVA_BASE_URL'], api_key=os.environ['OSVA_API_KEY'])
print(client.agents.list())
"

The demo I am planning to build

A product UI, a demo video, and the reproducible six-step flow below are not built yet. The repository includes developer-facing runtime samples, but not the onboarding demo I intend to ship. Until then, this is the demo I am planning, and it is labeled that way on purpose.

PLANNED DEMO · NOT YET BUILT
1

Start OSVA

Launch OSVA using the documented local setup.

2

Register the sample agent

Add a bounded container agent that produces a file as its output, such as converting structured input into a JSON report.

3

Submit a run

Send demo input through the SDK or API and receive a Run identifier.

4

Watch execution

See the Run and Attempt move through dispatch and container execution via API responses, CLI output, or logs.

5

Create the artifact

The agent requests artifact creation through the OSVA capability boundary and returns an ArtifactReference.

6

Retrieve it after exit

The container stops, then retrieve the artifact through OSVA and confirm its bytes and metadata remain available.

The demo will make the infrastructure promise visible: execution is temporary, durable state and outputs are not.

The last two steps are the whole point. They land the hook: the container is gone and the artifact is still there.

Using OSVA on a real project

  1. Choose one bounded task. That way success and failure are both obvious.
  2. Define the agent's input and output contract. Callers should know what to expect.
  3. Select a runtime that fits your deployment. Container or HTTP service, depending on the requirements.
  4. Configure required capabilities. Do not pass infrastructure credentials into agent code.
  5. Store reusable outputs as artifacts. Their lifecycle should be independent of the runtime.
  6. Use workflows where work spans steps or external events.
  7. Test failure and repeated delivery before you depend on the process.
  8. Choose storage and deployment settings that suit the environment. Filesystem locally, S3 for object storage.

A few things I would tell you before you start

Reliable infrastructure does not guarantee useful model output. OSVA can run an agent dependably and the agent can still give a bad answer.

Container networking needs an explicit operator policy. Persisted artifacts bring storage and lifecycle responsibilities with them. External operations need their own idempotency strategy. And passing local acceptance checks proves specific behavior, not production scale.

What important decisions we took while building?

Six decisions shaped the design.

DecisionReasonCost or tradeoff
PostgreSQL as durable authorityKeep state independent of queues and running processes.Database migrations and persistence discipline.
BullMQ and Valkey for transportSeparate dispatch from authoritative business state.Coordination between transport and durable records.
Capability-mediated accessKeep underlying credentials outside agent runtimes.More explicit interfaces and integration work.
Stable public contractsGive SDKs and runtimes a consistent boundary.Versioning and compatibility obligations.
Filesystem and S3 artifact storageSupport local use and object-storage deployments.Two adapters, plus integrity and streaming concerns.
A useful Community EditionMake the open-source product independently valuable.A more careful commercial boundary.

The test that failed, and what it taught me.

BEFORE

Execution existed only in memory

In-memory Agent

↓

In-memory Run

↓

In-memory Attempt

↓

Artifact capability request

↓

Generic artifact capability failure

DURABILITY ASSUMED, NOT PROVEN

Artifact metadata referenced durable execution records that did not exist in PostgreSQL.

→

Persist the complete execution context

AFTER

Execution records persisted first

Agent in PostgreSQL

↓

Run in PostgreSQL

↓

Attempt in PostgreSQL

↓

Container creates Artifact

↓

ArtifactReference returned

↓

Container exits

↓

Artifact bytes remain readable

END-TO-END OUTCOME VERIFIED

The failure showed that testing artifact creation alone was not enough. The product outcome depended on the full execution record being durable.

In the container artifact integration test, the agent reached the artifact capability, but artifact creation failed with a generic error. Nothing pointed at the real cause.

The cause was in the test itself. It had represented its execution records only in memory. But artifact metadata has foreign-key relationships that require the agent, the run, and the attempt to exist in PostgreSQL. In memory, they did not exist where the database could see them.

I fixed the test setup to persist those records. The integration then passed: the container created an artifact, returned its reference, and the bytes stayed readable after the container exited.

What stayed with me is the product lesson. An artifact belongs to a tracked execution. Testing the capability in isolation was not enough to prove the outcome a user actually cares about.

Release checks should give the same answer on every machine. A migration-history check hashed raw files, and the hashes changed with Windows line endings. We moved to canonical LF hashing so the same migration produces the same hash everywhere. Small fix, but a release check that disagrees with itself is not a check.

What we deliberately didn't build

Every item here was a choice. The cut list says as much about the project as the feature list does.

Infrastructure credentials inside runtimes. It would be easier to hand an agent a database connection. I chose capability requests instead, because that keeps the boundary explicit.

Provider-specific types in the domain. Provider details stay outside the core contracts so the platform does not become tied to one provider.

Process memory as durable authority. Anything that must survive a restart lives in persisted state. No exceptions for convenience.

SDKs that bypass platform boundaries. SDKs talk over HTTP and public contracts, the same way any other client would.

Enterprise controls in the Community scope. SSO, SCIM, advanced organizational controls, and enterprise deployment requirements belong to the commercial roadmap.

Known limitations

Owned, not hidden.

What has been checked. Version 1.0.0 across the recorded release components. 33 of 33 implementation slices completed. Quick verification, clean CI verification, migration verification, release-readiness checks, and Linux kind acceptance all passing.

What that does not prove. It is engineering acceptance evidence. It does not establish adoption, production scale, or package publication.

LimitationImpactNext step
UI work and testing remained unfinishedInfrastructure evidence exceeds end-user experience evidence.Build and test the main user journeys.
Guided onboarding examples and sample workflows were missingNew users must bridge documentation and practical use themselves.Ship runnable, bounded examples.
No confirmed public demoVisitors cannot immediately experience the project.Publish a reproducible demo.
Acceptance testing is not production operating historyPassing checks does not establish sustained reliability or scale.Gather operational and load evidence.
Agent quality remains application-dependentReliable execution can still produce poor decisions or answers.Evaluate each use case and define review boundaries.

The honest bottleneck: OSVA can make execution more dependable. It cannot decide whether an agent's output is good enough for a particular business decision. That part stays human.

This is the kind of product work I want to lead: turning technical capability into a system people can use, inspect, and trust.