- No shell execution
- No cluster-admin
- No secret reads
- No database writes
- No firewall changes
- No autonomous remediation
Observability signals
Metrics · events · log pipelines · traces — unified fleet viewNodes
840
infra
Clusters
31
infra
Agents
30
APM
Error events
4
events
P2
2
events
SLA risks
3
traces
Log drains
4
logs
Approvals
5
gates
Live telemetry
Fleet metrics & monitors
Streaming timeseries widgets and threshold monitors — Datadog-style dashboard layout, vendor-neutral simulated feed (1.5s ticks).
CPU
42.0%
fleet avg
Latency p95
236ms
model gateway
Error rate
0.30%
production
Throughput
862rps
agent invoke
Agents busy
51%
active workers
CPU utilization
Agent hosts · last ~60s
Request latency
Gateway p95 · ms
Live event stream
Rolling ingest from collectors & agents
Retry budget consumed · tool invoke
23:01:08
Latency probe elevated · model gateway
23:01:05
Retry budget consumed · tool invoke
23:01:01
Collector heartbeat · eu-west-1
23:01:06
Error rate
Failed invokes / total
Throughput
Requests per second
Active monitors
Threshold checks on live series — monitors-as-code pattern (no vendor lock-in)
Fleet CPU anomaly
okavg(last_5m):cpu.utilization{scope:agents}
42.0% · thr > 85%
Gateway latency p95
okavg(last_5m):gateway.latency.p95
236ms · thr > 400ms
Error rate spike
oksum(last_5m):errors.rate{env:production}
0.30% · thr > 2.5%
Log pipeline lag
okavg(last_5m):pipeline.lag.p95
1.5s · thr > 3s
Telemetry pipeline
Collectors forward structured container logs from agent sidecars — ops pattern analogous to Fluent Bit DaemonSets tailing /var/log/containers/*.log.
- healthy
Log collectors (DaemonSet)
4/4 nodes
- healthy
Structured JSON parse
cri-o · containerd
- healthy
Export endpoint
EU residency
- healthy
Pipeline lag p95
1.4s · live
Application latency
Model gateway request performance — live p95 overlay on seeded providers
OpenAI
201 ms live · US / EU routing
Anthropic
232 ms live · US / EU routing
Google Gemini
263 ms live · US
Azure OpenAI
294 ms live · EU (Sweden Central)
AWS Bedrock
325 ms live · EU (Frankfurt)
Ollama (on-prem)
356 ms live · On-premise
vLLM Cluster
387 ms live · On-premise (sovereign)
Infrastructure health heatmap
Composite health by customer and environment — see everything in one place
| Customer | production | staging | dev | dr |
|---|---|---|---|---|
| FS Core Banking Platform | 96 | 91 | 86 | 70 |
| Nordic Payments Rail | 99 | 99 | 84 | 79 |
| Card Issuing Services | 99 | 93 | 88 | 83 |
| Grid Telemetry Fabric | 88 | 83 | 78 | 73 |
| SCADA Edge Estate | 85 | 80 | 75 | 70 |
| Clinical Data Platform | 93 | 88 | 83 | 78 |
| Imaging AI Workloads | 99 | 98 | 82 | 77 |
| National Registry Services | 96 | 91 | 86 | 70 |
Error & incident timeline
Historical event volume (seeded) — live series above for last-minute fleet health
Token spend and retry waste
USD per day across all tenants
Active incidents
Ordered by severity and SLA exposure
Why is fs-prod-cs-tool2 NotReady?
P1inc-4821 · FS Core Banking Platform · production · SLA at risk
Payment rail latency above 850ms p95
P2inc-4818 · Nordic Payments Rail · production · SLA at risk
SCADA edge cluster losing telemetry batches
P2inc-4809 · SCADA Edge Estate · production
Clinical DB replica lag exceeding 90s
P1inc-4802 · Clinical Data Platform · production · SLA at risk
Recurring incident patterns
Signature clustering across 30 days
Registry egress reset during image pull
3 tenant(s) · last seen 2026-08-02
Kafka consumer rebalance storm
2 tenant(s) · last seen 2026-07-30
Replica lag during nightly ETL
1 tenant(s) · last seen 2026-08-01
Edge collector telemetry drops
1 tenant(s) · last seen 2026-08-01
Gateway config rollout latency regression
2 tenant(s) · last seen 2026-08-02