Public Cloud and SLA 99.99%

It seems that no team is allowed to intervene on Public Cloud (General Purpose) instances at any time other than Monday–Friday 8 a.m.–6 p.m., even though the SLA is 99.99%.

Is that true?

I would like an explanation regarding ticket CS16180037.

If I had known this beforehand, I would never have chosen Public Cloud; I would have stayed on dedicated because even with a 99.95% SLA, I had an intervention on a Saturday. So, in the end, Public Cloud offers no higher availability.

I've been with OVH almost since its inception, so I'm disappointed and I am indeed looking to move everything to a competitor after this incident.

People will tell me that the SLA is not a response time, but waiting two days to act when it fails on a Saturday morning… that’s not what I imagined.

I would like an official answer so I know whether I should migrate the servers.

Thank you.

Hi,

You intrigue me, give us details. Do you mean you have a Public Cloud instance down since Saturday?
Nothing to report on https://public-cloud.status-ovhcloud.com/

Hi @LibreMaster

Standard Support, which is the default, offers technical support by phone and ticket from Monday to Friday during business hours.

This means that if your server fails on a Saturday morning, the specialized technical team will not be available to intervene directly until Monday. The 99.99% SLA you mention for Public Cloud, which also exists for high‑end dedicated servers, guarantees service availability, but it is not a response time for resolving incidents.

If the hardware fails on a weekend, the OVHcloud team (available 24/7 for monitoring and hardware) can identify the problem and, in many cases, start an automatic intervention or schedule a manual one for the next business day.

Keep in mind, you’ll find this with other providers as well… if you want advice I’d recommend purchasing a higher support level, or if you don’t want to incur that cost, you can use Chat support.

The decision to migrate to a competitor is very personal. However, I encourage you to first explore the support options OVHcloud offers to ensure you’re making a decision with all the information :wink:

Note: If I’m wrong about anything, please have a staff member correct me :folded_hands:

Best regards,
Sergio Turpín

You described very well what they told me.

Note that even an instance with a 99.99% SLA, if it goes down at 6:01 p.m. on Friday evening, it won’t be repaired until Monday morning at 8:00 a.m.

To put it bluntly, the SLA applies to business hours, not 24/7.

However, many people like me completely misunderstand the service provided. When I’m told SLA 99.99%, I picture OVH striving to limit downtime to at most 52.6 minutes per year.

But that’s not what happens.

The reality is: if, “by chance”, the instance is down for more than 52.6 minutes, OVH may refund the monthly cost of the instance (may = you have to claim it; it’s not automatic).

So to magnify the issue, the instance could be down for an entire month and OVH would refund the monthly cost. They aren’t even required to intervene at all.

We are paying for the possibility that it works fine during the contract.

So how can I sell a 4‑hour GTI when the instance could stay down for a month?

Higher‑level support isn’t accessible to me; I would like a 4‑hour GTR option (with an R) that would be mandatory for all instances to spread the cost and make the option available. And if the repair takes longer than 4 hours, it should work like the SLA with a discount, but OVH should actually strive to do something—not just tell me “we’ll look at it on Monday”.

There is a major problem because many people don’t know this and have built their entire service on the assumption that OVH intervenes at night and on weekends even though no timeframe is announced. In people’s minds, I think they assume roughly a 4‑hour window.

I’d like to point out that they still intervened “to do a favor” because I insisted three times on the chat and likewise on the ticket.

My principle is as follows: I don’t implement redundancy and I concentrate quality on the system to avoid or reduce downtime, so I take an instance with a 99.99 % SLA to avoid managing hardware and to have an instance that, statistically, will work.

Even if I take two redundant instances and both go down (which would be unlucky), nothing will change; it will still only be repaired on Monday.

A client had an outage on a dedicated server with a 99.95 % SLA, the intervention was performed on a Saturday, so did they step in out of kindness instead of waiting for Monday?

Where can I get an instance with a GTI or GTR 4 h then without it costing an arm?

My principle is as follows: I don’t implement redundancy and I focus quality on the system to avoid or reduce downtime, so I take an instance with a 99.99 % SLA to avoid managing hardware and have an instance that, statistically, will work.

Well, I have to say I thought like you about the “cloud”. I don’t use the service because it’s just five times more expensive than bare metal.
And by that measure, bare metal doesn’t stay down all weekend, so…
I still can’t believe it…

Yeah, same thing—I thought I was paying to actually avoid this by steering clear of bare metal, but no. I’m planning to call OVH, first to get confirmation of the internal rule and then to request a GTI or GTR 4‑hour option “for everyone in a version 4 of the instances”, otherwise why pay twice the price for one instance… if it’s worse.

Keep us posted, I’m interested (because yeah, I also don’t like it when an SSD dies).

Yes, that’s it — the best way to achieve the availability you’re looking for without relying on 24/7 support is, precisely, to design your architecture so that it is resilient by itself, minimizing the need for a technician to intervene manually.

The solution to avoid this problem isn’t to pay for Business support, which is costly and intended for large critical environments, but to design your system so it doesn’t depend on a single instance.

Precisely for this reason, OVHcloud offers 3‑AZ regions. A 3‑AZ is a region that has three physically separated Availability Zones (AZ) (with its own power, cooling and network), but inter‑connected with very low latency. With this architecture, you can distribute your instances across multiple AZs.

Take a look at this :backhand_index_pointing_down:

https://www.ovhcloud.com/es-es/bare-metal/uc-3-az-resilience/

In addition, OVH already has it designed so that failover is automatic. In fact, this architecture is what allows OVH to offer a 99.99 % SLA for many critical services :wink:

You won’t find an isolated instance with a 4‑hour GTI at an economical price. What you can do, with the same level of availability, is distribute the load across several instances. In the end, the cost of this resilient deployment approaches that of Business support, but it gives you a much higher availability without depending on a technician clicking a button, because recovery is automatic.

Best regards,
Sergio Turpín

For my outage in question, it wasn’t visible to their manager, the status showed “active” but the instance was dead.

There is no need for 3 AZs and, as I said, even with 2 instances the problem remains unchanged. This problem can only be solved one way: have 2 or 3 additional on‑call people in separate time zones and spread the cost across the instances.

The 3 AZs are for staying below the GTI 4h threshold or in case of fire. In the last fire, I restored an instance in 15 minutes with its data in another data centre.

I understand that you're quite angry, but you know you can encounter this with any provider… however, if in the end you decide to migrate everything, let it be for something better, not because you're waiting for a technician over a weekend. Because, let's be honest, doing this is like changing a car because you got a flat tire on a Sunday. Sometimes the problem isn’t the car; it’s that you don’t have a spare tire :slightly_smiling_face:

We’re here to give you a hand with whatever you need. I hope they resolve it for you soon.

Kind regards,
Sergio Turpín

It would be necessary to see exactly what is included in OVH interventions.

An instance down, without the monitoring seeing it, is exceptional.

After that, as several participants here have said, the best solution remains to deploy your own HA infrastructure on a home bare‑metal cluster… But it quickly becomes a convoluted mess, and you’re locked into a single zone… Hence the interest of PCI instances…

And yes, there you go, you’ll anyway have a much higher cost. If the service must absolutely have max uptime because it’s clearly critical, that will of course cost a lot more.
The point here is therefore why an instance can go down without monitoring detection, if I understand correctly @LibreMaster

Performance‑wise, we’re far better on a home‑grown bare‑metal infrastructure, but personally I’ve stopped trying to manage it alone—it’s mentally exhausting.

You need at least two or three people on the team to handle on‑call duties, be able to take “real” vacations, etc., etc…

It’s for freelancers like me (and many others) that PCI‑offered solutions are appealing…
It does cost more, yes, but that’s precisely because we don’t have to manage the entire underlying infrastructure…

That's why I’ve been looking for a partner or a backup for 2 years.
Not easy :frowning:
I now have 51 servers spread across 24 bare metal + a few OVH VPS (I know you have a lot more than me).
Let’s say it’s still manageable on my own, but I’m bothered by hardware issues 2 or 3 times a year.
I also thought, like @LibreMaster, that this constraint would disappear with cloud solutions and that it would “justify” the much higher price for equivalent performance.

Hello @TTY and @Sich

All that’s left is for you to team up together. :clown_face:

Managing 51 servers solo is like having 51 teenage kids :joy:

In the end, the cloud isn’t cheap, but the price you pay is for not having to sleep with your phone on the nightstand in case a hard drive fails. And that, for a freelancer, is priceless. Or does it? It does have a price, but it’s the one they charge on the invoice.

We could always set the servers up again in the garage, but I have no idea what the climate‑control costs will be now :sweat_smile:

Well, I’m on a shifted time zone but I can certainly host an instance in a garage at my parents’ place instead of OVH—they have fiber. In terms of GTI, it could take a month if they’re away cruising around on the roads with the motorhome :wink:

Yes, that’s correct, their monitoring didn’t detect it but I have my own monitoring, I received a flood of alerts :slight_smile: And then I saw that the “console” returned an error, I suspected that the “active” status wasn’t right and that an OpenStack service had crashed.

Then I opened a ticket and, hold on tight....

An AI replied by telling me nonsense. It connected me to an “agent” (perhaps it’s another AI posing as a human by changing its name each time)

But this AI/agent then reassigned the ticket to the wrong Public Cloud project even though it was the correct one from the start.

I asked for it to be corrected.

Whether via chat or ticket, the same “we’ll see on Monday” response with the feeling of getting a gift in case of premature intervention.

I thought I’d request a “hard reset” of the instance, assuming it would lock and be visible to OVH and, indeed, it officially got locked. Then a divine intervention (you have to be realistic, it can only be the hand of the Almighty handling the SLA on weekends) occurred after 1 or 2 hours and brought the instance back online.

In the end the outage lasted a total of 4 h 30 min.

I say that it would merit a brief request for clarification on ML (or not, up to you).