1. Home
  2. Knowledge
  3. RAID 5 failure
Knowledge

RAID 5 failure – what to do now

A RAID rarely fails all at once. It fails a second time – and what happens in between decides everything.

+49 160 94791535

Mon–Fri 8 am–6 pm · emergencies and pick-up by arrangement

Why RAID 5 fails the way it does

RAID 5 distributes data and parity information across all drives. One drive may fail without data loss: the content is calculated from the remaining ones. That is the strength of the system – and the source of the misunderstanding that it is a backup.

The problem starts with the rebuild. To restore redundancy, every block of every remaining drive has to be read. In an array that has been running for years, that is the heaviest load the drives have seen in a long time – and exactly the moment a second, long-weakened drive gives up.

Individual drives from a RAID array, numbered and prepared for read-out

The most common sequence

One drive drops out, often unnoticed because the array keeps running. Weeks later a second one follows, the controller starts rebuilding onto a spare, and the process aborts halfway. At that point part of the user data has already been overwritten with reconstructed blocks – and those blocks were calculated from an incomplete set.

This is why the single most useful action is to stop. Do not start another rebuild, do not let the controller initialise the array, and do not swap drives around to “see whether it comes back”.

What to do instead

  1. Shut the system down. A degraded array that keeps running degrades further.
  2. Note the physical order and labelling of the drives before removing anything. The order is part of the reconstruction, and guessing it costs time.
  3. Write down what has already been done: which drive failed first, whether a rebuild ran, whether the array was initialised, whether a file system repair was confirmed.
  4. Leave the drives as they are. Do not attempt a repair with consumer software – it writes to the array.

How the reconstruction works

Every drive is imaged individually first, damaged ones repaired to the point where they can be read at all. The array parameters are then derived from the images: drive order, block size, parity rotation and start offset. Only when all four are correct does a readable file system appear again.

The advantage of working on images is that a wrong assumption costs nothing but time. On the original drives it would cost data.

What this says about backups

A RAID protects against the failure of a drive. It does not protect against deletion, ransomware, a fire, a controller writing nonsense or someone reformatting the wrong volume. Those are exactly the cases a backup is for – on a separate medium that is not permanently connected.

Frequently asked questions

Two drives have failed. Is everything lost?

Not necessarily. Often one of the two is only partly damaged and can be read far enough to complete the set.

The rebuild is still running. Should we let it finish?

No. A rebuild onto a failing member writes over the data you want back. Stop it.

Do you need the controller as well?

Usually not, but it helps in unclear cases. The drives with their slot positions are the important part.

Lost data? Every further start-up attempt can do harm.

Call us or describe the case briefly in the form. The initial assessment is free of charge.

+49 160 94791535