VOP requests failing
Updates
What happened
Between 12:38 and 13:13 CEST on 1 April 2026, a proportion of Verification of Payee (VOP) requests on our shared production environment returned 504 errors. Not all requests were affected, and the level of impact varied across customers depending on their VOP request volume during the window. Affected endpoints included incoming payee verification requests, outgoing payee verification requests, and dashboard-facing VOP views. Other Mambu Payments services continued to operate normally. We apologise for the disruption this caused to your operations.
Resolution
We identified the issue and restarted the affected service at 13:13 CEST, restoring normal operation within 35 minutes of the first error. A permanent fix addressing the underlying cause was deployed later the same day at 17:25 CEST.
Root cause
When the VOP service needed to refresh a configuration it holds in memory, a bug caused it to open a competing internal process alongside the one already handling your request. Under concurrent load, this created contention that caused the service to stall and return errors to callers instead of completing the request.
What we are doing about it
Immediate actions taken
- Restarted the VOP service to restore normal operation.
- Deployed a targeted fix the same day to eliminate the root cause permanently.
Preventive measures
- Auditing other service components that follow a similar internal pattern to identify and fix comparable risks before they can surface in production.
- Expanding automated load testing to catch this class of issue during our release process rather than in production.
- Improving our monitoring to detect this type of service degradation earlier, reducing the time between a problem occurring and our team being alerted to it.
Communication
Our alerting detected elevated failure rates within two minutes of the first affected request and our team began investigating immediately. We did not notify customers proactively during the incident. That is not acceptable for an incident of this severity and we are putting in place a defined notification process so that customers are informed promptly whenever a significant incident is confirmed, without waiting for resolution.
Lessons learned
- Proactive customer notification during an incident is not optional. Detection speed means nothing to a customer who found out themselves.
- Components that manage internal resources under concurrent load require dedicated review and load testing as a standard part of the release process.
This incident has been resolved. Some VOP requests sent or received between 10:38 UTC and 11:13 UTC failed due to elevated latency resulting in time out.
The issue has been identified and a fix is being implemented.
Some VOP requests are failing. We are currently investigating the issue.
← Back