Affected
Major outage from 12:45 PM to 12:55 PM, Degraded performance from 12:55 PM to 1:17 PM
Major outage from 12:45 PM to 12:55 PM, Degraded performance from 12:55 PM to 1:17 PM
Major outage from 12:45 PM to 12:55 PM, Degraded performance from 12:55 PM to 1:17 PM
- PostmortemUTCPostmortemUTC
RCA: Bitmovin Observability query and export outage (27 July 2026)
On 27 July 2026, Bitmovin Observability was unable to serve queries or data exports for approximately 20 minutes, with reduced data completeness for a further 24 minutes. Data ingestion was never interrupted, incoming data was buffered throughout, and no customer data was lost. This post explains what happened, what we have done, and what we are changing.
Impact
Between 12:33 and 12:53 UTC, the Bitmovin Observability query API, dashboards and data export were unavailable. Queries and exports resumed at 12:53. Data collected during the outage was buffered and then backfilled, completing at 13:17 UTC. Until that point, queries covering the affected period could return incomplete results. Data ingestion and collection were not affected at any time, and there was no data loss.
What happened
An internal query executed against our observability database cluster triggered a defect in the database engine. The memory used by this query was not correctly accounted for, which meant the cluster's memory safeguards never engaged. Instead of the query being limited or rejected, memory was consumed until nodes became unresponsive, and the failure propagated across the cluster rather than staying contained to the node serving the query.
Automated alerting notified our on-call engineers within 30 seconds. The cluster was restarted at 12:43 and completed recovery of all shards, resuming full write acceptance, at 12:53. Buffered data was then backfilled and fully available at 13:17.
Timeline (UTC)
Time
Event
12:32
An internal query is executed against the observability database cluster
12:33
First query and export errors observed
12:34
The cluster becomes unavailable, automated alerts fire and engineers begin investigating
12:43
Cluster restart initiated
12:53
All shards recovered, writes fully accepted, queries and exports resume
13:17
Backfill of buffered data complete, all systems operational
What we are changing
The query that triggered the defect was an ad hoc analytical query run by an operator, not a query our application issues in normal operation. We have instructed our engineering team not to run this class of query against the production cluster until the underlying defect is resolved.
We are reporting the defect upstream and validating whether a fix is available in a newer version of the database engine.
Our ingestion buffering worked exactly as designed and protected customer data. We apologise for the disruption and for the incomplete data during the backfill window.
- ResolvedUTCResolvedUTC
As of 15:17 CEST, the backfill has completed and all buffered data is available.
All systems are operational again. No data was lost.We will provide a root cause analysis as soon as possible.
- IdentifiedUTCIdentifiedUTC
The database cluster has recovered. Querying (dashboards and API) and data export have resumed, and incoming data is being written normally again.
We are now triggering the backfill of the buffered data. Until it completes, recent data may be incomplete in queries and exports.
We are monitoring the results and will confirm once the backfill has finished.
- InvestigatingUTCInvestigatingUTC
Starting at 14:33 CEST, our core Bitmovin Observability cluster is experiencing an outage. Ingestion is unaffected and incoming data is being buffered, so no data loss is expected. Querying (dashboards and API) and data export are currently unavailable.
We are investigating and will post an update within 30 minutes.