Eridian

Eridian status

All Systems Operational

Uptime: 99.95% over the trailing 12 months. Last incident: 14 days ago. Next maintenance: June 14, 2025 at 02:00 UTC.

StatusOperational

Updated every 60 seconds

ComponentStatusUptime (30d)
Inference APIOperational99.98%
Semantic CacheOperational99.97%
Routing EngineOperational99.99%
RAG PipelineOperational99.95%
PII RedactionOperational99.99%
Observability DashboardOperational99.94%
API Key ManagementOperational99.99%
WebhooksOperational99.96%
Eridian DevOperational99.95%
Eridian LegalOperational99.93%
Eridian OpsOperational99.94%
Eridian RiskOperational99.92%

Incident history

Trailing 90 days
Minor/23 minutes

Elevated latency on EU-West inference endpoint

A network degradation in the EU-West region caused elevated p99 latency (2.4s vs. 1.1s baseline) for 23 minutes.

Impact
EU-West inference p99 rose above SLO. Completions continued through the routing engine.
Customers affected
14
Resolution
Traffic was automatically rerouted to US-East. No data loss. No cache invalidation. Latency returned to baseline at 14:55 UTC.
Minor/16 minutes

Webhook retry backlog after consumer restart

A webhook consumer process restarted after a memory limit change, delaying inference.completed and budget.warning deliveries.

Impact
Event delivery lagged by up to 6 minutes. No events were dropped. API responses were unaffected.
Customers affected
9
Resolution
The consumer recovered, the retry queue drained, and duplicate-suppression checks confirmed a single delivery per event.
Minor/8 minutes

Semantic cache hit rate drop

A vector store index rebuild caused a temporary drop in cache hit rate (38% to 22%) for 8 minutes.

Impact
More requests reached the model on a cache miss. Answer quality was unchanged.
Customers affected
6
Resolution
Index rebuild completed at 09:23 UTC. Hit rate recovered to the 36-39% band within 20 minutes.
Minor/19 minutes

Delayed webhook delivery in US-East

US-East webhook workers fell behind after a burst of budget.warning events from a large legal workspace.

Impact
Downstream automation received events late. Inference itself stayed within SLO.
Customers affected
7
Resolution
Worker count was scaled and the backlog cleared. Alerting thresholds for queue depth were tightened.
Moderate/41 minutes

Document ingest backlog in RAG pipeline

A chunking worker stall queued new contract uploads. Retrieval continued against the last complete index.

Impact
Newly uploaded documents were not searchable until ingest caught up. Existing corpora stayed available.
Customers affected
4
Resolution
Ingest workers were scaled, the queue drained, and every queued document completed indexing. No documents were lost.
Moderate/47 minutes

RAG pipeline degradation

A chunking strategy update caused 12% of RAG queries to return empty context. Observability flagged an abnormal rag_injected: false rate.

Impact
Affected queries answered without retrieved context. Structured output parsing still ran.
Customers affected
3
Resolution
Rollback deployed at 04:34 UTC. SLA credits were issued to the three affected accounts.
Minor/14 minutes

PII redaction latency spike

A new NER entity-model revision added roughly 400ms to p95 masking time on legal workflows.

Impact
Legal inference stayed successful but slower. Masking accuracy did not change.
Customers affected
5
Resolution
The entity-model revision was rolled back. p95 masking returned to the 90-120ms band.
Moderate/36 minutes

Routing failover during Gemini region congestion

The primary Gemini route in EU-West exceeded the latency SLO. The routing engine shifted eligible traffic onto the GPT fallback chain.

Impact
Some legal and ops workflows used fallback models. Completions continued. Cost per request rose while failover was active.
Customers affected
11
Resolution
Provider congestion cleared. Primary Gemini traffic resumed after health checks passed for 10 consecutive minutes.
Minor/11 minutes

Observability dashboard metric lag

The metrics aggregator delayed request-volume and cost charts by about 8 minutes. Inference APIs were not affected.

Impact
Operators saw stale charts. Alerts based on the inference API itself continued to fire on time.
Customers affected
12
Resolution
The aggregator was restarted and historical points were backfilled. No metrics were discarded.
Minor/27 minutes

Semantic cache cold start after node recycle

A scheduled replica recycle dropped cache hit rate from 36% to 19% while new nodes warmed.

Impact
Higher model invocation volume during the warm-up window. No incorrect cache returns were observed.
Customers affected
8
Resolution
Replicas reached a steady hit rate. Recycle batches were reduced to avoid simultaneous cold starts.

Upcoming maintenance

Planned maintenance: vector store index migration (HNSW to IVF-PQ). Expected impact: 5% reduction in cache hit rate for 2 hours. No API downtime.