Data Recovery

Whereas Azure's automated backup systems may be relied on for recovery from equipment failure, routine administration of an Omeka site, including accession of new documents, and revision andd enrichment of existing assets carries a risk of catastrophic accidental data loss. In these situations, it is necessary to

  1. Be aware that the data loss has occurred. (Routine monitoring of Omeka integrity will be covered in another document.)
  2. Refrain from any new work that will change the Omeka database or media files collection.
  3. Engage an Omeka System Administrator with Azure credentials and knowledge to initiate recovery procedures.

This document discusses the basics of a point-in-time recovery of a complete Omeka instance. It may be possible to do finer-grained restorations of spefic items, resource templates, etc. But these require a more in-depth understanding of the Omaks-S schema than is covered here.

Topics

Recovery Scenarios and Timeframe

When the Omeka manager realizes that data loss has occurred, all changes to the omeka instance are stopped, because the most typical recovery scenario will be to roll-back all changes to a specific date and time in the past. At this point, the designated Omeka administrator is contacted. This will be a person who understands the fundamental architecture of the CHC's Omeka installation in Azure.

Recovery of an Omeka instance requires the recovery of the entire Omeka schema in the CHCMySQL database, and if there was a loss of document items, it may also be necessary to recover the chcPersist file share which contains all of the media files and thumbnails.

Backup Retention

The backup retention policies for MySQL databases preserve 30 days of database snapshots. The Azure recovery systems allow for recovery of states down to specific times of day during this period. The current backup procedure for files is as follows:

  • Daily snapshots preserved for 35 days.
  • Weekly snapshots preserved for 12 weeks.
  • Monthly snapshots preserved for 12 months.
  • Our monthly journal procedure preserves a MySQL backup that is stored indefinitely. in the chcOffline db backups folder.

Routine Integrity Audits and Prompt Recovery Initiation

The the weekly and monthly audit journal procedures are intended to actively check for data loss and integrity issues. Because data loss and recovery scenarios become more cloudy and complicated as time passes, they shuold be analyzd and executed as soon as practical (definitely within a couple of weeks.) First, is the 30 day window on the easily recoverable database and file-system snapshots, and second, because any changes made to the Omeka database or files hapening after the recovery target date will be lost. Details of the routine audits are provided at Regular Integrity Audits

Restoration Scenarios

There are two types of recovery scenario:

  • File Recovery Required: If media has been deleted (usually because a Document Item has been deleted) then files will need to be restored. Namely the original and thumbnail images that are kept in the CdashPersist/Files folder within the CDASHPersist storage account. In this scenario, a database recovery should also be made which will restore a snapshot of the database that is concurrent with the recoverd file system snapshot.
  • Database Recovery: In cases where item and item set properties and relationships have been altered but no items or media has been deleted, then a restoration can be accomplished by restoring a snapshot of the database only.

Plan, Execute and Verify your Recovery

Every recovery workflow should begin with a clear understanding of the problem, the current situation and what a successful recovery should look like.

  1. Discover the problem
  2. Analyze the cause and the scope of the issue in terms of:
    • Dates
    • Numbers of resources affected
    • As specific as possible, the defining properties of the set of resources affected.
    • The cause of the issue.
  • Plan the restoration:
    • Make a note of the resource counts -- including the counts of thumbnail and original image files -- and outstanding integrity issues from the latest weekly and monthly audits and activity logs.
    • How many resources do you expect to have before and after the restoration?
    • What queries will you make to verify that the recovery has impacted the missing or damaged resources?
  • Execute the restoration on the Stage instance.
  • Verify the restoration according to your plan.
  • Point the production image to the restored dtabase and file system as necessary.

    Database Restoration Technicalities

    The details of MySQL recovery are covered in Point-in-time restore in Azure Database for MySQL with the Azure portal.

    In most recovery scenarios, including routine testing of backup and restoration procedures, a database restoration would be made to a new instance of Azure Database for MySQL and then the Stage instance -- CHCOmekaStage. In such a scenario, the CHCMySQL database may be restored to a new instance of MySQL database for Azure. Linking an instance of CHCOmeka or CHCOmekaStage to the database is accomplished by using Azure Storage Explorer to save copy of thedatabase.ini file located in CHCPersist/Config. All you need to do is alter the to point to the new database instance Host: the database.ini.

    Database Restoration Technicalities

    Details of recovering Azure file share shapshots is covered in Restore Azure Files

    When a recovery requires the restoration of media files, a filesystem restoration of the CHCPresist file share must be carried out. The simplest way to do this would be to do a point-in-time restore directly to the original location. In cases where you wanted to verify that the restoration actually achieves wgat you want, you may choose to restore the file system wi an alternate location. Note that in the case of the production instance of chcPersist, this is a very large number of relatively large files.

    If you need to point an instance of CHC Omeka (Production or Stage) to a file system restored to an alternate location, this is accomplished in the Azure portal, by accessing the Web App's Settings. Under Configuration, add or alter the storage mount for the persist volume, which is linked to the appropriate folders within the web app's running container.

    Verify a Successful Restoration

    Before you walk away from a restoration, you should do a few things to check that the restoration was successful.