Agentic AI observability & assurance

Know what your agents did — and why.

Trace viewers were built for one model call returning one string. VIGIL is built for graphs — and for the applications, infrastructure and network underneath them. Eighteen agents correlate the run, find the root cause and produce the evidence. Deployed inside your perimeter.

Instrumented in hours, not sprints · SaaS, VPC, on-prem or air-gapped

run_9f42c1 · order-remediation · 12 agents · 4m 12s FAULTED
intakeclassifyenrichpolicyplanlookupdrafttool_callverifycommit

Select any node — or use Tab and Enter

Root cause

tool_call

Called crm.lookupAccount with customer_id: null. The API returned an empty result, which the planner treated as a valid negative — so nothing raised, retried or alerted.

1.9sduration
4,180tokens
2retries
  • Running in production at
  • a top-5 Indian private bank
  • a global pharma commercial analytics group
  • a tier-1 telecom operator
  • two EU-resident data platforms

The expensive failures don’t throw errors.

When a multi-agent system goes wrong it rarely crashes. It loops, it drifts, it picks the wrong tool, or it returns something plausible and wrong. None of that appears in a latency chart — and by the time a human notices, the run is three days old.

Failure mode 01

The silent loop

Two agents hand work back and forth until a retry cap ends it. The run “succeeds”. The token bill is thirty times what you modelled and nobody is alerted.

Failure mode 02

The fabricated argument

An agent calls a real tool with an argument it invented. The API answers honestly, the answer is meaningless, and the graph carries on as though it were fact.

Failure mode 03

The quality regression

A prompt change, a model version bump or a drifting data source degrades output slowly. Latency is flat, errors are zero, and your users notice before you do.

Built for prompts. VIGIL is built for graphs.

The unit of analysis changed. A single business transaction is now a graph of agents, tools, retries, branches and shared state that runs for minutes. Instrumenting that with a prompt logger gives you three hundred spans and a search box.

What a trace viewer sees

prompt completion

1 call · 1 latency · 1 cost

What actually ran

12 agents · 38 tool calls · 2 retries · 1 silent fault

The agents, and the stack beneath them.

Agent-only tools stop at the model call. Infrastructure monitoring never sees the reasoning. VIGIL covers both in one plane — which is the only way to answer “was it the agent, the service, or the network?” without three tools and a war room.

Agent & workflow tracing

Every run as a graph: agents, agentic workflows, handoffs, tool calls, retries, loops and the exact state at each step.

LangGraph · LangChain · OpenTelemetry

Application observability

The services your agents call — latency, errors, dependencies, saturation — stitched into the same trace rather than a separate tab.

Traces · metrics · logs

Infrastructure & network

Hosts, GPUs, queues and the network path underneath. When a run stalls, you see whether the agent hesitated or the link did.

One plane, not three tools

Evaluation & output quality

Continuous scoring on live traffic, not just a test set. Regressions surface as a trend, before a user files a ticket.

RAGAS · custom judges · human review

Cost & token intelligence

Spend per run, per agent, per customer, per tenant. Find the thirty-times spike the week it happens, not at quarter end.

Attribution · budgets · anomaly alerts

Safety, drift & evidence

Guardrail monitoring, behavioural drift detection, and audit-ready evidence packs generated from real runtime data.

EU AI Act · ISO 42001 · NIST AI RMF

Eighteen agents watching yours.

Most observability hands you a dashboard and leaves the analysis to you. VIGIL runs its own fabric of eighteen agents that correlate the run, rank the likely cause, score the output and assemble the evidence. You get a verdict, not homework.

  • Trace & correlateStitches agent, application, infrastructure and network signals into one run.
  • Root causeRanks candidate causes and points at the node, the state and the prior runs that match.
  • EvaluateScores output quality continuously against your own criteria.
  • Cost & driftWatches spend and behaviour for the change nobody announced.
  • Safety & evidenceMonitors guardrails and compiles the audit pack on demand.
VIGIL · 18 AGENTS trace · root cause · evaluate · cost & drift · evidence observes YOUR SYSTEM · YOUR PERIMETER agents · apps · infrastructure · network
Agents watching agents — lightAGENTS WATCHING AGENTSThe observability layer is itself agentic.VIGIL · 18 AGENTSTrace & correlateRoot causeEvaluateCost & driftSafety & evidenceobservesYOUR SYSTEM · INSIDE YOUR PERIMETERagents · apps · infrastructure · network

Runs inside your perimeter.

Agent telemetry contains your prompts, your customer data and your business logic. That is why deployment is a first-class product decision here, not a paid add-on — and why VIGIL runs entirely offline when it has to.

Fastest start

SaaS

Hosted by us. Sensible when the data is non-sensitive and you want a result this week.

Telemetry leaves your network

Common choice

Private VPC

Runs in your own cloud account under your IAM, your keys and your logging.

Telemetry stays in your cloud

Regulated default

On-premise

Your data centre, your hardware, your models. The usual answer in BFSI and pharma.

Nothing leaves the building

Sovereign

Air-gapped

No egress at all. On-premise models, offline updates, physically isolated operation.

No egress, by construction

Placeholder figures — replace with approved numbers before launch

What changes once you can see it.

Three engagements, anonymised pending customer sign-off. Every number below is a stand-in until the case study is approved for publication.

4h 20m

Mean time to diagnosis

A twelve-agent servicing workflow at a private bank. The win was not faster dashboards — it was not having to reproduce the failure at all.

31%

Token spend removed

Retry storms and redundant enrichment calls found in the first fortnight of instrumentation at a pharma analytics group.

1 day

Time to first trace

From SDK install to a fully instrumented production graph. Instrumentation effort is the objection we hear most, so we measure it.

The four-week pilot

Pick one production flow. We instrument it.

Not a free trial you have to staff yourself. A fixed-scope engagement with Shyena engineers, on one of your real agent workflows, against a success metric you define on day one. If it does not hit the metric, you owe us nothing further and you keep the instrumentation.

  • Week 1Scope and instrument. One workflow, agreed success metric, SDK in your environment.
  • Week 2Baseline. Full-fidelity traces, evaluation criteria configured, cost attribution live.
  • Week 3Findings. Root causes on real incidents, drift and spend analysis, guardrail gaps.
  • Week 4Decision. Results against your metric, evidence pack sample, and a rollout plan you can take to procurement.

Straight answers.

Including the comparison you are already running in your head.

How is this different from LangSmith, Langfuse or Arize?

Those are good tools and they are ahead of us on ecosystem integrations and community. The difference is scope and shape. They were designed around the model call, and agent support was added on top; VIGIL’s primitives are graphs, nodes, handoffs and state from the ground up. They stop at the model boundary; VIGIL also covers the applications, infrastructure and network your agents depend on, so “was it the agent or the service?” is one question in one tool. And they are SaaS-first — if your telemetry cannot leave your network, that is usually where the evaluation ends.

Do we have to replace Datadog?

No, and we would not recommend it. If you already run a mature APM estate, keep it — VIGIL ingests from it and sits above it, adding the reasoning layer it was never built to see. Teams without an incumbent often consolidate onto VIGIL for the agent-adjacent stack, but that is a choice, not a prerequisite.

How much instrumentation work is this?

For a LangGraph or LangChain application, an SDK install and a few lines of configuration — first traces the same day. For bespoke orchestration, OpenTelemetry semantic conventions plus a mapping we write with you. In the pilot, our engineers do this work rather than handing you a guide.

Where does our data go?

Wherever you decide. In on-premise and air-gapped deployments no prompt, completion or trace payload crosses your network boundary, and VIGIL’s own analysis agents run on models inside your environment. Architecture and data-handling documentation is published rather than sent on request — see the Trust page.

Can VIGIL prove compliance for us?

It produces the evidence; the determination stays yours. VIGIL generates audit-ready packs from real runtime data, mapped to EU AI Act, ISO/IEC 42001 and NIST AI RMF control language. No observability tool can honestly promise more than that, and you should be wary of one that does.

What does it cost?

The four-week pilot is a fixed-price engagement, credited against a first-year licence if you proceed. Platform pricing is published on the pricing page; on-premise and air-gapped deployments are quoted, because the shape of the estate genuinely changes the answer.