Database HS following GRA11 maintenance

Hello, we have a database on GRA11 down (host unreachable) for almost 7 hours now. A maintenance was scheduled, with a 30‑minute planned interruption.

This is the second critical DB incident in a few months, that’s a lot. Are other users experiencing the same issue?

Hello @omoise,

We are currently experiencing an incident:
https://public-cloud.status-ovhcloud.com/incidents/bn54j3232tlv

We apologize for the inconvenience caused.

^FabL

We are also affected, still waiting for a resolution.

@FabL can you give us more details and a forecast/estimate? We can’t even fork to another node in a different location and we depend entirely on you to reactivate the service under the current conditions.

We are also impacted on GRA11, with a Managed PostgreSQL instance still unavailable

Beyond the incident, the main problem is the lack of visibility on the service recovery. We have neither an estimate nor the ability to switch or fork the instance to another location.

Could you provide us with a precise status of the situation, an estimated restoration time, and the guarantees regarding data integrity?

Because of this incident, our e‑commerce sites have been completely down since yesterday at 5 p.m. This represents a huge loss of revenue. There's no way to fork to make a temporary switch; we are completely stuck.
OVH gives us no visibility and redirects us to the incident link…
Could we know what's happening? And get an estimate?

What service version do you have?

https://us.ovhcloud.com/legal/sla/managed-databases/

Due to this interruption, our tools and sites have been down since yesterday, which has already caused significant financial losses and a considerable loss of revenue.

Unfortunately, we are unable to implement a fallback solution by forking or duplicating our database to another node or location. This inability prevents us from temporarily shifting our operations and leaves us entirely dependent on the restoration of your infrastructure.

The situation is therefore becoming particularly critical for our business. We would appreciate it if you could promptly provide us with more details about the incident as well as an estimate, even an approximate one, of the time required to restore the service.

We also would like to know what temporary solutions could be considered to allow us to resume our operations as soon as possible.

Yes, until they update their status we won’t know what’s happening....

OVH GRA DC seems to be severely understaffed. Two weeks ago, I waited over 24 hours for a dedicated server failure to be addressed. Today's database issue likely affects hundreds of users, tt appeared at 6:00 PM on Sunday, they released an update at 7:00 AM next morning, and there's been no update or resolution since then? This doesn't look good.

In our case we detected a temporary outage with our monitoring systems yesterday at 18:27, and the service re‑activated a few minutes later, so we attributed it to the scheduled maintenance with an estimated downtime of less than 30 minutes; the problem was when another outage was detected at 21:18 and there has been no news since.

Initially we didn’t give it much importance because it seemed part of the scheduled maintenance and the downtime would be low, but we followed up and around 2 am we opened a specific ticket, for which we still have no response.

In the morning, around 8 am we contacted them by phone and the staff theoretically handling the case mentioned they do not have access to the technical support department’s information and cannot provide more details, and that we should wait until 9 am, which was when the scheduled maintenance was expected to end.

Shortly after the expected closure, OVH opened a new hardware incident: https://public-cloud.status-ovhcloud.com/incidents/bn54j3232tlv

And a few minutes later it appears that they marked the GRA database‑maintenance task as completed, as they changed the maintenance icon to OK:

https://public-cloud.status-ovhcloud.com/

On the other hand, we sent a query this morning via X to the OVH Spain support profile about this incident, and we received a reply a few minutes ago:
"Hello,
I confirm that the Public Cloud database maintenance is still ongoing.
We do not have an exact completion time for the maintenance. Sorry for the inconvenience."

From our side it is not completely indifferent that a small financial compensation be made regarding the cost of the service provided, as the damage caused is much greater; what we request is technical information about the incident and an estimated time for resolution, so we can assess mitigation methods for the damage incurred. We also want to know whether OVH can raise forks of backup nodes while the operation is ongoing, since all options are locked from the interface.

Honestly, I don’t think we’re asking for much: information, an estimated resolution time, and what measures OVH offers to provide uptime even in another geographic zone.

I see that they’ve enabled the forks of the nodes

I just saw it, although the interface sometimes throws an error; testing the fork creation… thanks.

My fork is still "Creation in progress"
Has anyone succeeded?

The one I'm managing is a database with millions of records and it's still in process; I'll let you know if it finishes correctly.

Same for me. It’s always “Creation in progress”

For me, the database is in an update state with no possibility to perform manipulations such as duplication

This incident is impacting our service and we are stuck

No information on the date of service restoration

It seems the fork isn’t working, because in my case it keeps being created; has anyone with a small database had it finish?

Same for me. I think the fork is blocked. Not working.

The fork worked. I managed to get a duplicate of my database.

After more than 24 hours of outage and no reliable communication, our Managed PostgreSQL finally seems to be back.

We really wanted to make the effort to prioritize a European provider. But our companies, our clients, and our business are too important to depend on a service that is unavailable for that long, without clear visibility or a fail‑over solution.

At some point, you also need to keep your business running. Even if that means returning to the famed American cloud that we love to criticize, but which generally knows how to keep this type of service available.

So we will migrate. Because beyond the sovereignty discourse, a cloud that is neither reliable, nor transparent, nor capable of ensuring continuity of a critical service is not a credible alternative. It’s just an additional risk.