Disaster Recovery Methods
AITS managed enterprise systems operate out of two separate data centers in located in different parts of the state for disaster recovery purposes. The primary data center that supports production enterprise systems is a high-availability facility located on the campus of the University of Illinois Chicago. The secondary data center that supports enterprise systems is located on the campus of the University of Illinois Urbana-Champaign.
In the event of a disaster that requires the recovery of AITS managed enterprise systems housed in our on-premise facilities, AITS has developed the following methods of recovery.
Recovery Method 1 – Full Scale Rapid Recovery
This method is the primary method used to restore all production enterprise systems in the event of a failover event being declared. This recovery method is best suited for large scale DR events where all production systems housed in the primary location need to be recovered to the DR environment in the secondary location. Over 99% of all enterprise systems supported by AITS can leverage this recovery method. The procedures used to bring up all systems using this method are similar to those used for routine system and network upgrades. However, this recovery method does not work for restoring individual systems or subsets of applications nor does it have the ability to restore to a particular point in time.
|
Recovery Time Objective (RTO):
|
24-hours for all systems to be restored after a DR failover event has been declared.
|
|
Recovery Point Objective (RPO):
|
In general, data loss of less than 5 minutes is expected. In the event of a sudden loss of services in the primary data center, long-running transaction failures and incomplete transactional processing may result in lost data.
|
|
Systems Using This Method:
|
All on-premise enterprise systems supported by AITS use the system with the exception of those that require specialized equipment or hardware.
|
This method employs storage level DR replication tools built into the storage arrays used by AITS. Through replication, AITS systems, configurations and data on the storage arrays are automatically replicated from the primary location to secondary location in near real-time providing a mirror image of the systems in the secondary data center. This copy is contained on a set of hardware dedicated to disaster recovery which is isolated from all user and network access. The available hardware resources assigned to this secondary environment matches the capacity of the production enterprise systems in the primary data center.
Recovery Method 2 – Recovery From Backups
This method is used to restore enterprise systems from a particular point in time using backups. This method would be used when a smaller set of individual applications must be restored due to a system problem or recovery requires systems to be restored to a specific point in time due to a variety of reasons including: data corruption, security events, or ransomware.
|
Recovery Time Objective (RTO):
|
24 to 72 hours for a service to be restored once a disaster event has been declared. The recovery time may vary depending on the nature of the event, the amount of data being restored, and the effort required to rebuild/restore an application to a point in time. During a major DR event that utilizes Recovery Method 1, systems requiring recovery from backups may wait until resources are available to perform the recovery from backups, thus extending the recovery time.
|
|
Recovery Point Objective (RPO):
|
Up to 24 hours for normal restores. For point in time recovery, the RPO will be dependent on recovery date identified.
|
|
Systems Using This Method:
|
Systems using specialized hardware or any individual enterprise system that requires recovery.
|
Nightly backups of all enterprise development and production systems are created and stored locally at the data center in which the system is normally operating and are copied to the other data center for offsite vaulting. There is also a copy of all backups stored in a tertiary location as an additional level of protection.
This copying procedure ensures that any system can be recovered in the event of an issue at its normal operating location. The restoration of a system will typically occur in the data center in which the impacted system normally resides. However, the determination of where a system will be recovered is based upon the situation.
Recovery Method 3 – Special System Recovery (limited use)
This method is used to recover enterprise systems that require special hardware or are not part of the enterprise storage network. The AITS standard is for all systems to utilize the enterprise storage arrays and virtualized servers which allows them to automatically be part of Recovery Methods 1 and 2. However, some systems require unique components that require an individualized System Recovery Plan (SRP) for each service. The number of systems that fall into this category is extremely small and typically require a duplication of specialized equipment and systems between the primary and secondary locations. These services may also require manual configuration and setup during a failover event. Systems with these requirements may also leverage Recovery Method 2 to restore services from backups.
|
Recovery Time Objective (RTO):
|
24 to 72 hours for all major services to be restored once a disaster event has been declared. The resources required to restore the services may be limited during a full-scale recovery event causing delays in bringing the services up and running.
|
|
Recovery Point Objective (RPO):
|
System dependent
|
|
Systems Using This Method:
|
Systems using specialized hardware that cannot be restored via Recovery Method 1.
|
Recovery Method 4 – Cloud System Recovery
A number of AITS enterprise systems are cloud-based applications supported by an external vendor (or Software as a Service products.) The backup and recovery methods, policies, and recovery objectives are dependent upon the vendor and can vary from system to system. These are typically defined within the agreement between the University and the vendor.