{ }Flows as code
Each workflow is a versioned spec with schemas, tools and policies. Changes are diffed, reviewed and rolled back like software.
Agent Flow · by Explore AI Inc. · Irvine, California
Agent Flow builds and runs AI agents that take the repetitive work off your team: invoices, claims, referrals, orders, quotes and lead follow-up. They work inside the systems you already use. Every step is traced, and every risky action waits for a person.
01 · The film
Freight audit, patient referrals, wholesale orders, a sales quote and a new web lead. Narrated in English and Mandarin. Every screen is the product you can try below, running on demo data.
02 · Live demo
Pick a workflow and press run. The agent reads the document, calls tools, checks your rules, and stops for your approval when money or risk is on the line. You make the call.
Demo data. Companies, people and documents in this playground are fictional; timings and costs are representative of a production run.
03 · In the browser
Terminal and carrier portals, payer portals, government sites, webmail. The agent signs in with credentials from your vault, clicks, types, reads and downloads, and stops for you before anything is submitted or sent.
04 · Use cases
We studied the public websites of 126 Southern California companies, from property managers to medical labs, and listed the repetitive work they describe. Pick an industry to see one of those workflows before and after.
Minutes are typical per-item estimates that show the shape of the work, not measured results. Your teardown measures your own.
The same agents work the front of the business too: the replies, quotes and renewals that win deals when they go out quickly.
05 · How we deploy
No platform migration, no rip-and-replace. We start from one workflow your team already runs and the systems it already touches.
30 minutes · free
We map one workflow end to end: volumes, systems, exceptions, and who signs off. You get an automation-potential score and a fixed quote within two business days.
weeks 1–2
Connectors, extraction schemas and your rules written as code. We build a golden test set from your own past cases, so accuracy is measured, not promised.
week 3
The agent runs beside your team on live work and changes nothing. Every case is scored against what your people actually did.
week 4 onward
Low-risk cases go touchless. Exceptions reach a person with the evidence already assembled. Thresholds widen only when the numbers earn it.
Illustrative touchless-rate targets. Your pilot reports the real numbers on your own cases.
06 · Platform
Flows are versioned code. Policies are testable functions. Every run leaves a trace you can replay. Models are routed per step, so you pay frontier prices only where reasoning is needed.
Each workflow is a versioned spec with schemas, tools and policies. Changes are diffed, reviewed and rolled back like software.
Small, fast models extract; frontier models reason over exceptions. Cost per run is metered and capped per flow.
Golden sets built from your history gate each deployment on field accuracy and decision accuracy. No silent regressions.
Dollar thresholds, confidence floors and sensitive fields route work to the right person in Slack, Teams or email, with the evidence attached.
Every prompt, tool call, diff and decision is stored. Open any run, see exactly why it happened, and replay it against a new version.
Supervised browser actions handle carrier, payer and county portals that have no API, inside the same guardrails.
# flows/freight-invoice-audit.yaml
flow: freight-invoice-audit
version: 14
trigger:
inbox: ap@yourco.com
match: { attachment: pdf, sender_in: carriers }
steps:
- extract: { schema: carrier_invoice, model: fast-extract, min_confidence: 0.92 }
- tool: tms.get_load # rate confirmation + stops
- tool: terminal.gate_events # real arrival / departure times
- reason: { model: frontier, task: reconcile_accessorials }
- policy: ap_variance # escalates to a person
- act: erp.create_bill
- notify: carrier.dispute_email
guardrails:
pii: redact
max_cost_per_run: $0.25
approvals: slack:#ap-approvals
evals:
golden_set: evals/freight_2026q3.jsonl # 1,200 past invoices
gate: { field_f1: ">=0.97", decision_acc: ">=0.99" }
from agentflow import policy, Allow, Escalate
@policy("ap_variance")
def ap_variance(run):
"""Pay small, explainable variances; escalate the rest."""
variance = run.billed_total - run.expected_total
limit = max(150, 0.02 * run.expected_total)
if run.min_confidence < 0.92:
return Escalate("low extraction confidence")
if abs(variance) > limit:
return Escalate(
f"variance ${variance:,.2f} over ${limit:,.0f}",
approvers=["ap-lead"],
evidence=run.citations,
)
return Allow()
{
"run": "run_7Hq2x", "flow": "freight-invoice-audit@14",
"duration_ms": 41230, "cost_usd": 0.061,
"spans": [
{ "step": "extract", "model": "fast-extract", "ms": 3810, "fields": 23, "min_conf": 0.96 },
{ "step": "tms.get_load", "ms": 420, "status": 200 },
{ "step": "terminal.gate_events", "ms": 1160, "via": "browser" },
{ "step": "reason", "model": "frontier", "ms": 6240, "tokens": 5812 },
{ "step": "policy.ap_variance", "result": "escalate", "variance": 76.5 },
{ "step": "approval", "by": "ap-lead", "waited_ms": 27900 },
{ "step": "erp.create_bill", "ms": 610, "bill": "BILL-20417" }
]
}
Most of the wall-clock is the human approval, by design: the agent prepares the evidence, a person makes the call.
07 · ROI
Move the sliders to match one workflow on your team. The estimate uses the Production plan below and a conservative ramp.
Estimate only. Uses the starting Production price scaled by complexity, savings that ramp to the share above over three months, 48 working weeks, and the starting pilot fee credited toward Production. Payback includes the pilot month. Your teardown gives you a quote on your real volumes.
08 · Pricing
No two back offices are alike: volumes, systems, exception rules and risk all differ. Every engagement starts with a free teardown and a fixed written quote, so you see the number and the expected payback before you commit.
$0
30-minute working session
from $15,000
fixed scope · one workflow · 4–6 weeks
from $4,500 / agent / mo
quoted per workflow · annual
Custom
several departments or regulated data
what drives your quote
A fixed-price build, then one yearly fee for hosting, monitoring, model updates and changes as your process evolves.
We run, watch and improve the agents for a monthly fee. Scope changes ship within two business days.
One-time delivery into your cloud with source code and runbooks, plus optional yearly support.
09 · Security & control
Our US-hosted cloud, your AWS, Azure or GCP account, or on-premises for regulated data.
Your documents are never used to train shared models. Model providers are used under no-retention terms where offered.
PII and PHI are detected and tokenized before reasoning steps; only the fields a step needs are revealed.
Scoped credentials kept in your secrets vault. Write actions are allow-listed per flow.
Immutable log of every read, write, prompt and approval, exportable to your SIEM.
Pause any flow instantly and fall back to your human queue, with nothing lost in flight.
10 · FAQ
No. Agent Flow works through the systems you already use: APIs where they exist, supervised browser actions where they don't, and your existing inboxes and shared drives.
It stops. Confidence floors, dollar limits and sensitive fields are written as policies; anything outside them goes to a person with the evidence and a recommended action already prepared.
Two ways: a golden test set built from your past cases gates every release, and a shadow week compares the agent with your team on live work before it is allowed to act.
It replaces the copy-paste. Most teams use the hours to absorb growth without new hires, clear backlogs, and move people to exceptions and customer work. How you use the capacity is your call.
We route per step across leading commercial and open models, and can run open models inside your environment when data cannot leave it.
Explore AI Inc., an Irvine, California research and product company building adaptive agents. See explore-ai-inc.com.
11 · Get started
Tell us what your team does over and over. In 30 minutes we'll map it, score it and tell you honestly whether an agent will pay for itself.