VOP requests failing

Major incident API
2026-04-01 12:38 CEST · 36 minutes

Updates

Update

What happened

Between 12:38 and 13:13 CEST on 1 April 2026, a proportion of Verification of Payee (VOP) requests on our shared production environment returned 504 errors. Not all requests were affected, and the level of impact varied across customers depending on their VOP request volume during the window. Affected endpoints included incoming payee verification requests, outgoing payee verification requests, and dashboard-facing VOP views. Other Mambu Payments services continued to operate normally. We apologise for the disruption this caused to your operations.

Resolution

We identified the issue and restarted the affected service at 13:13 CEST, restoring normal operation within 35 minutes of the first error. A permanent fix addressing the underlying cause was deployed later the same day at 17:25 CEST.

Root cause

When the VOP service needed to refresh a configuration it holds in memory, a bug caused it to open a competing internal process alongside the one already handling your request. Under concurrent load, this created contention that caused the service to stall and return errors to callers instead of completing the request.

What we are doing about it

Immediate actions taken

  • Restarted the VOP service to restore normal operation.
  • Deployed a targeted fix the same day to eliminate the root cause permanently.

Preventive measures

  • Auditing other service components that follow a similar internal pattern to identify and fix comparable risks before they can surface in production.
  • Expanding automated load testing to catch this class of issue during our release process rather than in production.
  • Improving our monitoring to detect this type of service degradation earlier, reducing the time between a problem occurring and our team being alerted to it.

Communication

Our alerting detected elevated failure rates within two minutes of the first affected request and our team began investigating immediately. We did not notify customers proactively during the incident. That is not acceptable for an incident of this severity and we are putting in place a defined notification process so that customers are informed promptly whenever a significant incident is confirmed, without waiting for resolution.

Lessons learned

  • Proactive customer notification during an incident is not optional. Detection speed means nothing to a customer who found out themselves.
  • Components that manage internal resources under concurrent load require dedicated review and load testing as a standard part of the release process.
April 7, 2026 · 11:14 CEST
Resolved

This incident has been resolved. Some VOP requests sent or received between 10:38 UTC and 11:13 UTC failed due to elevated latency resulting in time out.

April 1, 2026 · 13:14 CEST
Monitoring

A fix has been implemented and we are monitoring the results.

April 1, 2026 · 13:13 CEST
Investigating

The issue has been identified and a fix is being implemented.

April 1, 2026 · 13:10 CEST
Investigating

Some VOP requests are failing. We are currently investigating the issue.

April 1, 2026 · 12:38 CEST

← Back