Home›Guides›What to do if…

What to do if…

Last night's backup failed

A backup that fails for one night is not a disaster, it is a delay: the last clean copy is at least 48 hours old if the previous night’s succeeded. Several failures in a row are a protection incident: you look for the cause the same day, rather than simply acknowledging the alert.

Updated October 20263 min read5 sources cited

Key points

  • Read the error message, not just the red light: space, missing source, credentials, locked files, bandwidth, stopped agent.
  • A “successful” job may have copied an empty folder: look at the size copied.
  • Write down the date of the last success and tell the manager of the department concerned.
  • Rerun after fixing the cause, then check the following night.
  • Three failures in a month on the same machine: change the configuration, not just press “retry”.

1. Read the error, not just the red light

The usual causes, in the order in which they appear:

CauseSignFix
No space left on the destination, or quota reachedWrite error, retention getting shorterIncrease the volume, or knowingly shorten the history
Source switched off, off the network, path changedDrive letter or share renamed; job “successful” on an empty folderRestore the source, correct the path, check the size copied
Credentials rejectedService account password expiredRestore the account: otherwise every night will fail
Locked files or database not quiescedPartial copyUse the method intended for open databases; a SQL database in this state cannot be restored cleanly
Link too slow or cutJob interrupted at the end of the windowMore efficient incremental backup, longer window, or less unnecessary data
Agent stopped on the machineNothing sentRestart the service, find out why it stopped

The most misleading case is the green job on an empty folder. The ANSSI, France’s national cybersecurity agency, requires backups to be checked systematically, monitoring in particular for inconsistent data or file volumes, network slowdowns and configuration changes. A copied size that drops sharply from one night to the next deserves as much attention as a failure.

2. Know when the last success was

It is the only date that matters for today’s RPO. If it is more than a few days old, tell the person responsible for the department concerned. They are working without a safety net and need to know. Acknowledging the alert in the console without saying so is what turns an incident into data loss two weeks later.

3. Rerun after fixing the cause

Run a manual job once the cause has been dealt with. Wait for it to finish. If the rerun fails, the cause is still there. The next morning, check the following night’s run: many “obvious” fixes do not survive the second pass.

A successful job proves that a copy was written, not that it can be restored. For a SQL Server database, Microsoft states that the backup verification command does not check the structure of the data it contains. The ANSSI requires backups to be tested regularly, with a written restore procedure; NIST also recommends testing backups to ensure that files can be recovered without errors. After a backup incident, a test restore of a file or a database is the best check. See How do you test that a backup works?.

4. If it keeps happening

Three failures in a month on the same machine: the scope, the bandwidth or the product is unsuitable. Change a setting (exclude a huge, useless folder, split the job, increase storage) rather than just rerunning it by hand every Monday.

Morning checklist

  • Have all of last night’s jobs finished, and not just “not in error”?
  • Is the size copied consistent with that of previous nights?
  • Is the last-success date for each critical machine less than 24 hours old?
  • Have the alerts been read by a named person, and not just received?
  • Does the remaining space on the destination cover the planned retention?

At WeDoBack

24/7 monitoring covers the backups and sends an alert when a backup does not complete. The alert is the start of this page, not the end. With INTEGRAL, two hours of support per month can be used to deal with the cause. With SMART, support is billed per intervention: the failure remains visible to the customer in the console, and it is up to the customer to read it. Support can be reached on +33 9 72 50 78 28, from 9:00 to 13:00 and from 14:00 to 17:30 (Paris time). To size storage, the published rule of thumb is the current volume multiplied by three, then an adjustment after a week of use: a volume that is too tight shows up as failing jobs or a shrinking retention. The encryption key, held by the customer, plays no part in an upload failure: if the job fails, the remote copy has simply not been updated.

Frequently asked questions

Is one failed night serious?

Rarely, if the previous night succeeded and the cause is fixed during the day. The risk lies in accumulation: each failed night increases the amount of work that would be lost in a disaster. Beyond a few days, it is an incident to report to management.

Is a “successful” status enough to prove that a backup is good?

No. The ANSSI, France’s national cybersecurity agency, requires backups to be checked systematically, in particular for inconsistent data volumes, and restore tests to be run regularly. For SQL Server, Microsoft states that verifying a backup does not check the structure of the data it contains: only an actual restore, followed by a consistency check, proves it.

Who should monitor backup alerts?

A named person, with a deputy for holidays. An alert that lands in a shared mailbox that nobody reads is the same as no alert at all. Also decide who informs management when the last success goes beyond a threshold agreed in advance.

Need help now?

Do not restore anything until you have identified a clean copy. We can guide you.

Call +33 9 72 50 78 28or write to us

Dealing with an incident right now?

Our teams help you identify the right copy and restore it, Monday to Friday, 9 am to 1 pm and 2 pm to 5:30 pm (Paris time).