Skip to main content

Custom agent systems

Turn proven bottlenecks into measured agents.

Friday starts with the workflow evidence, defines a narrow agent, and releases changes only when they improve the real task without breaking cost, latency, or control.

Real task setPermissioned toolsMeasured release
Evaluation release / candidate 07Controlled

Evidence set

42 replayable cases
ready

Agent candidate

3 allowed actions
scoped

Release threshold

Quality · latency · cost
measured

Release decision

Ship only measured gains

The workflow becomes the specification

A defined job, not a broad assistant.

Friday begins only after a constraint is visible, frequent, and measurable. The evidence defines what the agent sees, does, and must prove.

  1. 01Observe

    Use the evidence the workflow already produces

    Messages, tickets, documents, code changes, and the decisions that connect them.

  2. 02Act

    Allow only the smallest useful action set

    The agent receives the specific tools required for the job—not general workspace access.

  3. 03Approve

    Pause where consequences begin

    External, irreversible, or high-impact actions wait for an accountable reviewer.

  4. 04Prove

    Define success before the first release

    Quality, corrections, latency, and cost are measured against the current workflow.

Self-improving means controlled release

Every failure makes the test harder to fool.

The agent never rewrites itself in production. Reviewed outcomes become replayable tests, and only a measured improvement is promoted.

  1. 01Capture

    Real tasks, edge cases, reviewed failures, and approved outcomes become the test set.

  2. 02Replay

    Prompts, tools, routes, and models run against the same cases before release.

  3. 03Measure

    Task success, human correction, latency, and cost are scored together.

  4. 04Promote

    A candidate ships only when it beats the current agent at the agreed thresholds.

  5. 05Expand

    Every reviewed production failure becomes a new test for the next candidate.

Right-size the model around the task

Pay for the outcome, not the largest model.

Frontier capability designs hard cases. Stable production work moves to a specialist open model, narrow context, caching, and deterministic tools.

LayerFrontier-onlyFriday agent
ProductionLargest model every runSpecialist model + routing
ContextBroad prompt every runNarrow evidence + cache
ImprovementAd hoc prompt changesReplayable evaluation
ObjectiveCapability per callCost per successful outcome
Qualified workflows100×

targeted lower inference cost than a frontier-only baseline

Actual savings depend on workflow complexity, volume, model choice, latency requirements, and the quality threshold your team accepts.

Guardrails stay part of the agent

Control is part of the architecture.

Review security and governance

Minimum tool access

Only the connectors and actions required by the selected workflow enter the agent scope.

Human approval gates

Consequential, external, or irreversible actions pause for an accountable reviewer.

Release evidence

Every promoted configuration carries quality, latency, and cost results from the evaluation set.

Sovereign runtime

The model and evaluation harness can run inside infrastructure you control.

Personalized onboarding · 15 minutes

Bring one bottleneck. Leave with a sovereign agent pilot map.

We will map the workflow evidence, your on-premise boundary, the approved open model, the agent's allowed actions, and the test that proves whether the pilot is worth running.

Start personalized onboarding Built around your workflow, not a generic tour
  1. 01Verified bottleneck
  2. 02Sovereign deployment
  3. 03Permissioned agent