Bacula tape backup software and LTO tape library tips and troubleshooting

Share
Bacula tape backup software and LTO tape library tips and troubleshooting

Environment notes: our setup includes a NeoT48 LTO8 tape library with two tape drives and a PostgreSQL-18 catalog for Bacula 15.0.3.

NeoT48 Tape library tips

# find tape library /dev/sg* path. should show up as "mediumx"
lsscsi -g

# get a status listing of tapes in the tape library
mtx -f /dev/sg3 status

# the following commands may not work if bacula is actively using the tape library

# load a tape into tape drive 0 from slot 16
mtx -f /dev/sg3 load 16 0 

# unload a tape from tape drive 0 into slot 16
mtx -f /dev/sg3 unload 16 0

See error codes associated with tapes at library web interface>service>cartridge memory. Code definitions here:

Table 1. TapeAlert Flags supported by LTO tape drives
Flag Number and Name
Hex code
Description
Action Required
SNMP Trap ID
Call Home
1Read warning01hSet when the tape drive is having a problem reading data. No data is lost, but there is a reduction in the performance of the tape.Isolate the fault between drive and tape by following these steps:
  • Use a known good tape cartridge in the suspect drive. If the drive fails, contact your IBM® Service Representative.
  • Use the suspect tape cartridge in a known good drive. If the test fails, discard the cartridge.
201 (Warning)No
2Write warning02hSet when the tape drive is having a problem while data is being written. No data is lost, but there is a reduction in the capacity of the tape. Isolate the fault between drive and tape by following these steps:
  • Use a known good tape cartridge in the suspect drive. If the drive fails, contact your IBM Service Representative.
  • Use the suspect tape cartridge in a known good drive. If the test fails, discard the cartridge.
202 (Warning)No
3Hard error03hSet for any unrecoverable read, write, or positioning error. (This flag is set with flags 4, 5, or 6.) See the Action Required column for Flag Number 4, 5, or 6 in this table.203 (Warning)No
4Media04hSet for any unrecoverable read, write, or positioning error that is due to a faulty tape cartridge.Replace the tape cartridge.204 (Warning)No
5Read failure05hSet for any unrecoverable read error where isolation is uncertain and failure might be due to a faulty tape cartridge or to faulty drive hardware.If Flag Number 4 is also set, the cartridge is defective. Replace the tape cartridge. If Flag Number 4 is not set, see Error Code 6 in Resolving LTO tape drive errors.205 (Warning)No
6Write failure06hSet for any unrecoverable write or positioning error where isolation is uncertain and failure might be due to a faulty tape cartridge or to faulty drive hardware.If Flag Number 9 is also set, make sure that the write-protect switch is set so that data can be written to the tape. If Flag Number 4 is also set, the cartridge is defective. Replace the tape cartridge. If Flag Number 4 is not set, see Error Code 6 in Resolving LTO tape drive errors.206 (Warning)No
7Media life07hSet when the tape cartridge reached its end of life (EOL).
  1. Copy the data to another tape cartridge.
  2. Discard the old (EOL) tape cartridge.
207 (Warning)No
8Not data grade08hSet when the cartridge is not data-grade. Any data that you write to the tape is at risk.Replace the tape with a data-grade tape.208 (Warning)No
9Write protect09hSet when the tape drive detects that the tape cartridge is write-protected.Make sure that the cartridge's write-protect switch is set so that the tape drive can write data to the tape. See Setting the write-protect switch on an LTO tape cartridge.N/ANo
10No removal0AhSet when the tape drive receives an UNLOAD command after the server prevented the tape cartridge from being removed.Refer to the documentation for your server's operating system.N/ANo
11Cleaning media0BhSet when a cleaning tape is loaded into the drive.None. Status only.N/ANo
12Unsupported format0ChSet when you load an unsupported cartridge type into the drive or when the cartridge format is corrupted.Use a supported tape cartridge.N/ANo
13Recoverable mechanical cartridge failure0DhThe drive has detected a mechanical volume failure and the volume is able to be unloaded. This is normally a broken tape detected during midtape recovery, and is normally a high usage tape that has broken tape at the leader pin.Quarantine the tape cartridge and contact your IBM Service Representative. No
14Unrecoverable mechanical cartridge failure0EhThe drive has detected a mechanical cartridge failure and the cartridge is not able to be unloaded. Normally set when the tape split apart.Do not attempt to extract the old tape cartridge. Call the tape drive supplier's help line.214 (Error)Yes
15Cartridge memory chip failure0FhSet when a cartridge memory (CM) failure is detected on the loaded tape cartridge.Replace the tape cartridge. If this error occurs on multiple cartridges, see Error Code 6 in Resolving LTO tape drive errors.215 (Error)No
16Forced eject10hSet when you manually unload the tape cartridge while the drive was reading or writing.No action is required. Informational message only.N/ANo
17Read-only format11hSet when a write attempt is made on a read-only cartridge. The flag is cleared when the cartridge is ejected.None. Status only.N/ANo
18Tape directory is corrupted in the cartridge memory12hSet when the drive detects that the tape directory in the cartridge memory is corrupted.Re-read all data from the tape to rebuild the tape directory.218 (Warning)No
19Nearing media life13hThe media is nearing its specified usage lifeSchedule a time to migrate the data to another tape cartridge219 (Warning)No
20Clean now14hSet when the tape drive detects that it needs cleaning.Clean the tape drive. See Methods of cleaning drives. N/ANo
21Clean periodic15hSet when the tape drive detects that it needs routine cleaning.Clean the tape drive as soon as possible. The drive can continue to operate, but you must clean it soon.N/ANo
22Expired clean16hSet when the tape drive detects an expired cleaning cartridge.Replace the cleaning cartridge.222 (Warning)No
23Invalid cleaning cartridge17hSet when the drive expects a cleaning cartridge to be loaded and the loaded cartridge is not a cleaning cartridge.Use a valid cleaning cartridge.223 (Warning)No
25Interface19hSet when the tape drive detects a problem with the Fibre Channel interface.See Error Code 8 or 9 in Resolving LTO tape drive errors.225 (Error)Yes
26Cooling fan failure1AhSet when a tape drive fan failed.Contact your IBM Service Representative.226 (Error)Yes
30Hardware A1EhSet when a hardware failure occurs that requires a drive reset to recover. If resetting the drive does not recover the error, note the error code on the single-character display and refer to Resolving LTO tape drive errors.230 (Error)Yes
31Hardware B1FhSet when the tape drive fails its internal power-on self-tests.Note the error code on the single-character display and refer to Resolving LTO tape drive errors.231 (Error)Yes
32Interface20hSet when the tape drive detects a problem with the Fibre Channel interface.See Error Code 8 or 9 in Resolving LTO tape drive errors.232 (Error)Yes
33Eject media21hSet when a failure occurs that requires the tape cartridge to be ejected from the drive and retried.Unload the tape cartridge; and reinsert it and restart the operation.233 (Warning)No
34Download fail22hSet when an FMR image is unsuccessfully downloaded to the tape drive through the SCSI or Fibre Channel interface.Ensure that it is the correct FMR image. Download the FMR image again.N/ANo
35Drive humidity23hSet when the drive humidity sensor indicates the humidity is out of range.  No
36Drive temperature24hSet when the drive's temperature sensor indicates that the drive's temperature is exceeding the recommended temperature of the library.See Error Code 1 in Resolving LTO tape drive errors.236 (Warning)No
37Drive voltage25hSet when the drive detects that the externally supplied voltages are either approaching the specified voltage limits or are outside the voltage limits.See Error Code 2 in Resolving LTO tape drive errors.237 (Error)No
38Predictive failure26hThe drive has predicted a potential failure.Run drive diagnostics 238 (Error)Yes
39Diagnostics required27hA failure occurred that requires drive diagnostics to be run.Run drive diagnostics239 (Error)Yes
49Diminished native capacity31hThe media is either partitioned or shortened and is not configured to use full native capacity.If using partitioning, this is normal. If not, then determine why there is a partitioned cartridge present and unpartition the cartridge.249 (Informational)No
51Tape directory invalid at unload33hSet when the tape directory on the tape cartridge that was previously unloaded is corrupted. The file-search performance is degraded.Use your backup software to rebuild the tape directory by reading all the data.N/ANo
52Tape system area write failure34hSet when the tape cartridge that was previously unloaded could not write its system area successfully.Copy the data to another tape cartridge and discard the old cartridge.252 (Warning)No
53Tape system area read failure35hSet when the tape system area could not be read successfully at load time.Copy the data to another tape cartridge and discard the old cartridge.253 (Warning)No
55Load failure37hA load failure occurred because the media could not be loaded and threaded.Remove the tape and try another. If the problem persists, contact your IBM Service Representative.255 (Warning)No
56Unrecoverable unload failure38hAn unload failure occurred because the media could not be unloaded.Contact your IBM Service Representative.256 (Error)Yes
58Microcode failure3AhA drive panic occurred.Send drive dumps to your IBM Service Representative. No
59WORM Medium - Integrity Check Failed3BhThe drive determined that data on the tape is suspect, from a WORM point of view.
  1. Copy the data to another WORM tape cartridge.
  2. Discard the faulty WORM tape cartridge.
259 (Warning)No
60WORM Medium - Overwrite Attempted3ChThis flag is set when the drive rejects a write operation because the rules for allowing WORM writes are not met. Data can be appended to WORM media only. Overwrites to WORM media are not allowed.Write the data to a WORM tape cartridge or write the data to a non-WORM tape cartridge.N/ANo
61Encryption policy violation3DhThe drive is configured to require encryption, but encryption is not enabled.Contact your IBM Service Representative. No

Bacula tips

Job architecture

Although Bacula proports to be able to resume a "incomplete" job, in practice, most jobs that hit errors just fail outright and must be restarted. You can update the catalog to mark a failed job as incomplete and then resume the job, but this appears to resume from the beginning of the job, rather than where it left off. Even if it does resume where it left off, there's no way to know if some of the files were backed up in a bad state due to the error. So it is best to restart the jobs from the beginning.

As such, it is wise to keep job sizes down to a few terabytes at most. Here's an example fileset configuration which separates out a few subdirs into their own filesets, while still keeping an "other" fileset to allow any new subdirs/files in the parent directory to be backed up:

Fileset {
  Name = "ARCHIVE_DATA"
  Description = "The ARCHIVE_DATA directory under /data"
  Include {
    Options {signature=SHA1; onefs=no; }
    File = "/data/ARCHIVE_DATA"
  }
}

Fileset {
  Name = "data1"
  Description = "The data1 directory under /data"
  Include {
    Options {signature=SHA1; onefs=no; }
    File = "/data/data1"
  }
}

Fileset {
  Name = "other"
  Description = "Other dirs and files located in /data"
  Include {
    Options {signature=SHA1; onefs=no; }
    File = "/data"
  }
  Exclude {
    File = "/data/ARCHIVE_DATA"
    File = "/data/data1"
  }
}

Watching Bacula job progress

Pull out the useful information about the job currently accessing the storage device:

# general information about running jobs
watch 'echo "status storage" | bconsole | grep -e 'Files=' -e 'Writing:' -e 'pool='' -e 'FDReadSeqNo'

# More info about current copied file etc. Swap 6 for the number of the bacula client you want to modify
watch '(echo "st client" && echo "6") | bconsole | grep "Running Jobs" -A 7; echo "st s" | bconsole | grep "spooling"'

# running jobs list:
watch 'echo "st dir" | bconsole | grep "Running Jobs" -A 5'

Get Bacula jobs stored on each tape

Bacula catalog PostgreSQL query to get jobs that exist on each tape

select m.mediaid, volumename, array_agg(distinct(j.name)) as jobs_containing, p.name as pool_name from jobmedia jm join media m on jm.mediaid = m.mediaid join job j on jm.jobid=j.jobid join pool p on m.poolid=p.poolid group by p.name, m.mediaid, volumename order by p.name, m.mediaid;

Purging and reusing volumes

Use above query to double check which jobs you're purging, first.

bconsole
> purge volume
> # Select the volume you wish to erase and reuse
> update volume
> # Select status and set to "Recycle"
> # Rather than recycling, you can also relabel a volume here

Reading a Bacula tape label without Bacula

mtx -f /dev/sg3 status  # Make sure only the tape you want to delete is in a drive to avoid confusion
mt -f /dev/nst1 rewind
dd if=/dev/nst1 bs=64512 count=1 | xxd | less  # print off start of data including Bacula label, if it exists
mt -f /dev/nst1 rewind

Deleting the label on a tape (loses data)

# remove label from bacula
bconsole
> delete volume
> 13 # pick the pool the volume is in
> tape1  # enter the tape's name
> yes

mtx -f /dev/sg3 status  # Make sure only the tape you want to delete is in a drive
mt -f /dev/nst1 rewind  # rewind tape to start
dd if=/dev/nst1 bs=64512 count=1 | xxd | less  # Double check that the data/label doesn't look unexpected
mt -f /dev/nst1 rewind  # rewind to start again
mt -f /dev/nst1 weof  # Delete the first block which includes the tape label
mt -f /dev/nst1 rewind  # rewind to start one last time
# now bacula will not be able to read the tape label

Setting up encryption

# Generate master certs
openssl genrsa -out bacula_encryption_master.key 2048
openssl req -new -key bacula_encryption_master.key -x509 -out bacula_encryption_master.cert

# Generate local fd certs
openssl genrsa -out bacula_encryption_fd.key 2048
openssl req -new -key bacula_encryption_fd.key -x509 -out bacula_encryption_fd.cert

# create pem file
cat bacula_encryption_fd.key bacula_encryption_fd.cert > bacula_encryption_fd.pem

In your bacula-fd.conf file, add encryption info to your filedaemon:

FileDaemon {
  ...
  PKI Signatures = Yes  # Enable Data Signing
  PKI Encryption = Yes  # Enable Data Encryption
  PKI Keypair = "/path/to/bacula_encryption_fd.pem"  # Public and Private Keys
  PKI Master Key = "/path/to/bacula_encryption_master.cert"  # ONLY the Public Key
}

Then restart the filedaemon: systemctl restart bacula-fd

More tips

  • Force bacula to scan every tape for its contents with update slots scan
  • Enter a period (".") to cancel most input prompts in bconsole
  • Restart the storage daemon seems to fix some issues: service bacula-sd restart

Resolving Bacula errors

"Could not connect to client"

Make sure ports are open on the client and that the bacula-fd service is running.

firewall-cmd --add-port=9103/tcp --permanent
firewall-cmd --reload

"bacula-fd: Fatal error: failed to connect to storage daemon"

Make sure ports are open on the device which is running the storage daemon. Make sure the storage daemon is running.

firewall-cmd --add-port=9103/tcp --permanent
firewall-cmd --reload

"Waiting to connect to storage daemon"

See above.

"Unable to open device, no medium found"

Manually load a tape into a drive in the tape library using mtx. Then run update slots in Bacula, then try the job again or let it resume.

Job stuck on running when backing up via NFS mount

More reliable to run a file daemon on the remote device, rather than relying on NFS to serve the files for backup.

"Bad storage"

Double check configured autochangers/drives. Make sure the correct devices are used and that the autochangers have the correct /dev/sg numbers.

"Warning: Director wanted Volume 'X'. Current Volume 'Y' not acceptable because: 1998 Volume 'Y' catalog status is Append, not in Pool."

If it loads a drive with a tape but then reports that the wrong tape is loaded (and the error message mentions a tape that is in the other drive), try swapping the DriveIndex numbers in the bacula-sd.conf and restarting bacula-dir and bacula-sd.

"Bconsole cannot connect to director"

Make sure postgresql is running and then restart bacula-dir.

Despite what the status might suggest, the director may not be running. Try executing bacula-dir -d 100 to see if there's any errors on launch. Also double check bacula-fd -d 100. Should say "bacula-fd is already running" if it's working.

Job not automatically finding volume

Double check that the mediatype listed in bacula-dir.conf, bacula-sd.conf, and reported by the media in bconsole (list media) match.

Volume marked "Full" before actually full

Check the tape library's web interface. The NeoT48 may be marking drives with Tape alert flag (TAF) 49 which means "Diminished native capacity" – "The media is either partitioned or shortened and is not configured to use full native capacity."

NeoT48 marks tapes as "Bad Tape"

Stop Bacula jobs and then stop the bacula-dir service. Clean drive, reboot, move each bad tape one-by-one into the drive slot and then back out. Just doing a move operation isn't enough– seems the drive has to be able to load the tape and read it for it to decide the tape is no longer bad. Restart bacula-dir service. Update volstatus to append for any volumes that Bacula marked error due to this problem. Re-run jobs as needed.