Hello,
I receive a monitoring e‑mail, my server suddenly became unreachable (ping impossible).
I check myself, server unreachable by ping, I trigger a hardware reboot on the disk. Same result: server not pingable.
I trigger a reboot in rescue mode, 10 minutes later, a monitoring e‑mail, rescue‑mode reboot failure, an intervention on your server has been automatically scheduled.
I therefore suspect a hardware problem or an issue with OVH’s internal network.
12 hours later a technician intervenes.
Response from the technician working on the machine:
“I can see the login screen but the server is not pingable, I reboot in rescue mode, the server is reachable again, it is a client‑side configuration error”
How can I be responsible for a rescue‑mode boot failure?
Thank you.
Hi @Sebourg
To fine‑tune my answer a bit:
- What service do you have exactly contracted with OVH?
- Did you make any changes shortly before you got the error? If so, which?
I’ll wait to give you feedback.
Best regards,
Sergio Turpín
Hello @sturpin
It is a bare‑metal/kimsufi server.
Before the first monitoring email (server does not respond to ping)
After the first monitoring email (server does not respond to ping):
- Change of boot type: boot on disk => netboot (boot in rescue mode)
Server still unreachable even after the request to boot in rescue mode
At this point, can you give me the IP address of this server, just so I can run an nmap to see if it really hasn’t rebooted?
That’s when the lack of KVM/IPMI on the Kimsufi range becomes a real pain…
If you end up getting control, the first thing to do is to check the S.M.A.R.T. counters of your hard drive… and if there are bad blocks, do as little activity as possible before you’ve recovered the most important data.
Hello @fritz2cat
Everything has returned to normal since the technician’s intervention. (According to him, he also performed a reboot in rescue mode and that worked, the server is reachable again on the network)
As a result, I had no action to take; the server started functioning again just as it did before it became unreachable. (I checked, there is no disk problem)
What really strikes me is being told that I’m responsible for the configuration of a server that doesn’t boot in rescue mode. How can support make that claim? To me, if a server doesn’t boot in rescue mode it’s either a hardware or a network issue (internal routing at OVH, a network configuration I can’t influence in any way).
After 12 hours of interruption I should be eligible for compensation in view of a 99.9 % SLA on the Kimsufi range.
Best regards.
In your place I’d still analyze the hard drive
smartctl -a /dev/sda
@fritz2cat Done on the 4 disks, no errors reported.
A server with all its disks HS should boot via netboot: rescue mode