🚀 Introducing HarnessRouter — Bring the world's best AI agents into your app, with one API ✨
    Epsilla Logo
    ← Back to all blogs
    September 12, 202611 min readRichard Song

    How Epsilla Runs as an AI-Native Company on HarnessRouter: An Operations Roadmap

    Epsilla's philosophy and roadmap for running an AI-native company on HarnessRouter: agents as colleagues rather than features, work rather than chat as the unit, evidence over narration, and a harness you keep choosing on the numbers.

    HarnessRouterAI-Native CompanyDigital EmployeesAgent OperationsAgent HarnessEpsilla
    How Epsilla Runs as an AI-Native Company on HarnessRouter: An Operations Roadmap

    Most companies that call themselves "AI-native" have a chat window in every department and no idea what it costs. That is not an operating model. It is a subscription line item.

    At Epsilla we ship an Agent-as-a-Service platform for enterprises, so we have to hold ourselves to a harder standard: the company itself should run on agents that complete work, not agents that draft suggestions for humans to finish. The runtime that makes that possible for us is HarnessRouter, the unified interface for agent harnesses.

    This post is about the ideas behind that choice and the order in which we are acting on them. It is a philosophy and a roadmap, not a manual. If you want the technical teardown of HarnessRouter itself, read the companion piece: HarnessRouter Tech Stack Deep Dive: Inside the Unified Harness Protocol.

    Key ideas

    • Work, not chat, is the unit of an AI-native company. An agent that hands back a draft has moved the work; an agent that hands back a finished, checkable result has done it.
    • Agents are colleagues, not features. They belong to the organization, hold roles, and live under the same rules of access and accountability as people.
    • Evidence beats narration. What an agent did is established by the record it left, never by its own account of it.
    • The harness is a decision you keep making. HarnessRouter's published benchmark found a roughly 475x cost-per-task spread across eight harness and model combinations. No configuration stays the winner for long.
    • Trust is earned one task class at a time. The roadmap moves from cheap failures to consequential ones, and each phase produces the evidence that permits the next.

    1. The thesis: the loop closes inside the harness

    The difference between using AI and being AI-native is where the loop closes. In a chat-first company, a model produces a draft and a human closes the loop by editing, testing, formatting, publishing and reporting. The company has bought a faster way to start work and kept every cost of finishing it.

    In an agent-native company the loop closes inside a harness. The agent reads the inputs, uses tools, checks its output against a contract, and hands back an artifact. The human's job moves upstream, to deciding what good looks like, and downstream, to judging whether it arrived. Everything in between is the runtime's problem.

    HarnessRouter's framing is blunt about this: models generate tokens, harnesses complete work. We adopted that sentence as an operating principle. It changes what we ask of every job. Not "which model should write this," but "what is the contract, what may the agent touch, and what happens when it succeeds or fails."

    2. Five beliefs that shape how we run

    Agents belong to the company

    We do not think of our agents as integrations bolted onto a chat tool. They are members of the organization. Each has a name and a role, the way a research analyst, a technical writer, or a chief of staff would, and each can be messaged, included in a conversation, or handed work. What an agent may see is decided the same way it is decided for a person: by where it sits in the organization, not by what it asks for.

    This matters more than it sounds. When agents are features, every new one needs its own permissions story, its own audit story, its own place in the workflow. When agents are members, all of that is inherited. The organization already knows how to grant, revoke, review and explain access. Agents simply join it.

    The workspace is the interface

    An agent should work where the company works: in its documents, on its task board, inside its conversations. We resisted building a separate universe of agent tools. Instead the workspace itself is what agents are given, through the same protocol surfaces HarnessRouter's configured agents already speak. The result is that a person and an agent can look at the same page, the same card, the same thread, and see the same thing.

    A quieter consequence is safety. An agent that acts as a member, with an identity it cannot change, has no way to become someone else. It is not a policy that prevents it; it is the shape of the system.

    Evidence, not narration

    Language models are fluent narrators of their own work, and that fluency is a liability. We decided early that an agent's description of what it did is not the record. The record is what changed: the document that now exists, the card that moved, the diff that was opened, stamped with who did it and when. When an agent reports, the report points at those facts. If there are no facts, there was no work.

    This one belief does more for trust than any amount of prompt engineering. People stop asking whether the agent is telling the truth, because the system never asked them to take its word.

    One shared memory, with a history worth reading

    An AI-native company needs a single, shared memory rather than a scatter of per-team silos, and that memory should have a history a person would actually read. We keep the live working state, where nothing is ever lost, separate from the milestones, which are coarse snapshots of the business that read like a changelog. Any point on that timeline can be revisited. When an agent wants to change something it is not trusted to change alone, it proposes, and a person reviews, on the same surface people use with each other.

    The harness is a choice, not a commitment

    The most expensive assumption in agent operations is that there is a best agent. There is a best agent for a task, for a while. HarnessRouter's own position is that the harness and the model should both remain parameters of the request, and its benchmark shows why: on one identical task, eight combinations of harness and model varied about 475x in cost, and the costliest was not the best. We run our operations as though every winner is temporary, because every winner is.

    3. Why HarnessRouter is the runtime underneath

    Each of those beliefs needs a runtime that does not fight it. We wanted one contract for every harness, so that switching Codex for Claude Code or Hermes is a configuration change rather than a rebuild. We wanted isolation per task, so that an agent's reach ends where its job ends. We wanted sessions that remember, files that come back as artifacts rather than pasted text, and a bill that charges for work rather than for waiting.

    That is what HarnessRouter is. The Unified Harness Protocol gives us the stable contract; the Community Edition gives us a real exit, which we exercise on purpose so that portability stays a fact rather than a hope. HarnessRouter describes this property as the socket rather than the dock: one contract, and the module moves. For a company that expects to change its mind about agents every quarter, that is the whole point.

    4. The roadmap: a sequence of earned trust

    We did not move the company onto agents all at once, and we would not advise anyone to. The order matters, because each stage generates the evidence that makes the next one responsible.

    First, where failure is cheap. Content and internal research went first. The contracts are easy to write, and a failed run costs cents and reaches no customer. This is where we learned what a good task definition looks like and how much a finished artifact is actually worth. Even the cover of this post was produced by an agent inside that loop.

    Then, engineering, with people owning the consequences. Agents produce reviewable changes; humans own merges. The candidate harnesses are the coding agents HarnessRouter already routes, all behind one Agent API. The test suite is the gate, and no configuration keeps its slot by being fast if it is not also right.

    Then, the work that touches customers. Support, sales research and onboarding move last, and they move as drafts first. A task class earns the right to act without a person only after its configuration has cleared the gate on real volume. Nothing about this is timid; it is the same way a company extends authority to a new colleague.

    Finally, a loop rather than a milestone. Harness capabilities move roughly every eleven days by HarnessRouter's count. So the last stage never ends: real task classes are run as competitions in Harness Arena, every run is traced, and configurations are promoted or demoted on the numbers. Nobody argues in a meeting about which agent is best. The evidence changes, and the routing follows.

    5. Principles we hold ourselves to

    1. Success is a gate, cost is the score. A cheap failure is not a bargain. We compare cost per successful result, never cost per token.
    2. Persist the work, discard the worker. One task, one isolated run. What survives is the artifact.
    3. Identity is inherited, never asserted. An agent's access comes from its place in the organization, not from anything it says.
    4. Evidence over narration. Work is proven by what changed, not by what was said.
    5. Drafts by default, autonomy by evidence. Authority is extended to a task class the way it is extended to a person: after it has been earned.
    6. Name the job, not the vendor. The application says what needs doing; configuration decides which harness does it.
    7. Portability is tested, not assumed. If we cannot leave, we are not portable.
    8. Humans own irreversible actions. Merges, sends, payments and production deploys require a named approver.

    6. Where this sits on the AI-maturity ladder

    We described the five levels of AI organizational intelligence earlier this year, from personal tools through siloed workflows and reshaped capabilities to AI in decision making and, finally, an organization whose systems complete whole loops on their own. The roadmap above is how we climb it in practice: one function at a time gains a real contract, one task class at a time earns autonomy, and the organization as a whole stays at the level where decisions about agents are made from evidence. No phase asks for a company-wide leap. The top of the ladder arrives one task class at a time.

    7. What we measure

    Four numbers, per class of work, on one dashboard: cost per successful task, because it is the only cost figure that survives a failed run; tail latency, because features and replies have deadlines and averages hide the tail; the share of agent work a person accepts without edits, because that is the signal a class can earn autonomy; and how often a class's winning configuration changes, because churn tells us what is still being learned.

    We deliberately do not count agents or tokens. Neither tells you whether work got done.

    8. Running your company the same way

    The path is the same for any team building on agents, whether you use Epsilla or not. Write down one class of work as inputs, a contract, a configuration and a disposition. Create a workspace on HarnessRouter, bring your own model keys, and run it through the quick start. Give the agent your real workspace rather than a toy one, and let it report through what it changed. Run a few configurations, keep the one that clears the gate at the lowest cost, and add the next class.

    If you would rather have the whole loop, from knowledge base to no-code agent builder to deployment, run for you, that is what Epsilla's Agent-as-a-Service platform is for, and HarnessRouter is the runtime underneath it.

    Frequently asked questions

    What does it mean for an agent to be a member of the company? It has a name and a role, it can be included in conversations and given work, and what it may see is decided by where it sits in the organization rather than by what it requests. Behind it is a HarnessRouter configured agent, but from the inside of the company it is a colleague.

    Which harness does Epsilla use? There is no single answer, and that is the point. Each class of work has its own configuration, chosen among the harnesses HarnessRouter routes and re-selected on evidence. HarnessRouter's own position is that there is a best harness per task, not per company.

    Does an AI-native operating model require open source? No, but portability does. HarnessRouter's Community Edition implements the same Unified Harness Protocol as the Cloud, so leaving is possible. We test it on purpose.

    Where should a smaller team start? Where failure is cheap: content or internal research. The contract is easy to write, and the evidence collected there is what justifies moving engineering and customer-facing work next.

    Ready to Transform Your AI Strategy?

    Join leading enterprises who are building vertical AI agents without the engineering overhead. Start for free today.