Governance is not a feature
It wraps every stage. Redaction happens before storage, not after; permissions are enforced by the runtime, not by a prompt.
The interesting part of an AI system is rarely the model. It is everything around it: how data gets in, what happens when the model is unsure, who is allowed to approve what, and how you prove six months later why a decision was made. That is what we build.
Every deployment is the same shape. What changes is which perception models sit in stage three and which systems stage six is allowed to write to.
It wraps every stage. Redaction happens before storage, not after; permissions are enforced by the runtime, not by a prompt.
Every reviewer fix and every downstream outcome flows back as labelled data. The system gets measurably better at your work, not at a benchmark.
Below threshold, the pipeline does not guess. It routes to a person with the page or clip already open at the right region.
We have deployed into a datacentre with no outbound internet at all, and onto a fanless box bolted to a machine frame. The software is the same; the packaging is not.
| Mode | Where it runs | Egress | Typical hardware | Best for |
|---|---|---|---|---|
| On-premise | Your datacentre, your Kubernetes | none | 2 × L4 or A10 per 1M pages/mo | Regulated document workloads |
| Private cloud / VPC | Your AWS, Azure or GCP account | vpc-internal | g5/g6 or NC-series instances | Elastic volume, existing cloud estate |
| Edge | Beside the camera or the line | telemetry only | Jetson Orin NX / AGX, industrial PC | Real-time inspection, poor connectivity |
| Air-gapped | Isolated network, offline updates | none | Customer-supplied GPU nodes | Defence, critical infrastructure, health |
Perception, structuring and agents scale independently. The bottleneck is never the whole pipeline.
Every stage is a durable queue. Nothing is lost when a node, a network or a camera drops.
Model versions are pinned per workload and rolled forward deliberately, with a shadow period first.
Edge nodes keep working through outages and reconcile when the link returns.
We would rather have the uncomfortable conversation in week one than in month nine. Here is our default posture - all of it configurable tighter, none of it looser without a written decision from you.
# applied at every stage, not bolted on data: residency: "in-country" encryption: at-rest AES-256 · in-transit TLS 1.3 raw_retention: configurable · default 0 days identity: sso: SAML · OIDC rbac: per queue, per schema, per camera retrieval: permission-aware models: weights: run inside your boundary third_party_llm: opt-in, per workload shared_training: never audit: mode: append-only export: SIEM · syslog · S3 covers: [reads, writes, approvals, overrides, model_changes]
An automation that ends in a CSV somebody has to import is not an automation. We connect to the systems of record directly, and where an API does not exist, we build the adapter.
We would rather lose a deal early than deliver a pilot that quietly dies in month four. Roughly one in five scoping calls ends with us saying the problem is not ready for this yet - usually because the data does not exist, or the process it feeds is not agreed.
A working session with the people who do the job today, not only the people who buy software. We want to see the worst documents, the poorest camera angle, the exception nobody has written down. We leave with a defined success metric and the honest failure modes.
A real pipeline against a real sample - typically 500 to 5,000 documents, or a few hours of footage. You get numbers, not a slide: precision, recall, per-field accuracy, latency, and the list of cases it cannot yet handle.
The system runs alongside the current process without touching anything, so you can compare its answers to your team's on live volume. Only when the numbers hold do the writes get switched on - usually one queue at a time.
Accuracy, drift, throughput and queue depth on a dashboard your team owns. Reviewer corrections feed retraining on a schedule. When the world changes - a new vendor layout, a new part, a new camera - you hear it from the monitoring before you hear it from a complaint.
We are pragmatic about models. Open weights where they are good enough and cheaper to run, frontier models where reasoning genuinely earns its cost, and our own training where your data is the advantage.
| Layer | Approach | Runs where |
|---|---|---|
| OCR & layout | Fine-tuned recognition and layout models, trained on your document families | your boundary |
| Detection & segmentation | Real-time detector family, distilled and quantised for the target device | edge / gpu node |
| Tracking & temporal | Multi-object tracking with temporal action models over sliding windows | edge / gpu node |
| Reasoning & agents | Frontier LLMs where they earn it, with strict tool contracts and schema-constrained output | your boundary or opt-in API |
| Retrieval | Hybrid lexical + dense retrieval, permission-aware, citation-enforced | your boundary |
| Orchestration | Durable queues, idempotent stages, replayable pipelines | kubernetes / docker |
| Review UI | Side-by-side evidence and correction interface, embeddable in your apps | browser |
| Monitoring | Accuracy, drift, latency, cost and queue metrics, exported to your stack | prometheus / otel |
The first call is technical, not commercial. Bring the person who will have to run this, and the person who has to sign off on the risk.