D4IS Operations & Resilience

IS Operations & Resilience: Glossary

Batch Processing
A method of processing transactions where data is collected over a period and processed together at a scheduled time, rather than in real time.
Business Continuity Plan (BCP)
A comprehensive plan that outlines procedures and instructions to maintain critical business functions during and after a disaster or disruption.
Capacity Management
The process of ensuring that IT infrastructure resources are sufficient to meet current and future business requirements in a cost-effective manner.
Change Management (Operations)
The structured process for controlling modifications to IT infrastructure, applications, and services to minimize disruption and maintain service quality.
Cold Site
A disaster recovery facility that provides only basic infrastructure (power, connectivity, space) with no pre-installed hardware or software. It has the longest recovery time.
Configuration Management Database (CMDB)
A repository that stores information about IT assets and their relationships. It supports incident, problem, and change management processes.
Database Replication
The process of copying and maintaining database objects in multiple databases to improve availability and support disaster recovery.
Disaster Recovery Plan (DRP)
A documented process to recover and restore IT infrastructure, systems, and data following a disaster. It is a subset of the broader business continuity plan.
Electronic Vaulting
The electronic transfer of bulk data to an offsite storage facility, typically on a scheduled basis, to support recovery in the event of a disaster.
Help Desk / Service Desk
The single point of contact for users to report incidents and request services. Under ITIL, the service desk manages the incident lifecycle from logging to resolution.
Hot Site
A fully equipped disaster recovery facility with hardware, software, and data that can become operational within hours. It provides the fastest recovery time.
Incident Management
The ITIL process of restoring normal service operation as quickly as possible after an unplanned interruption, while minimizing the impact on business operations.
Job Scheduling
The automated process of managing and executing batch jobs, data transfers, and system tasks according to predefined sequences, dependencies, and timing.
Mean Time Between Failures (MTBF)
The average time between system or component failures. A higher MTBF indicates greater reliability.
Mean Time to Repair (MTTR)
The average time required to restore a failed system or component to normal operation. A lower MTTR indicates faster recovery capability.
Maximum Tolerable Downtime (MTD)
The longest period of time a business function can be unavailable before the organization suffers unacceptable consequences.
Media Library Management
The process of controlling physical and electronic storage media, including tracking, labeling, storing, and disposing of media assets securely.
Mirror Site
A disaster recovery facility that is an exact real-time replica of the primary site, providing near-zero downtime through continuous data synchronization.
Operations Bridge
A centralized location from which IT operations staff monitors and manages IT infrastructure, services, and events across the enterprise.
Problem Management
The ITIL process of identifying and managing the root cause of incidents to prevent recurrence. It differs from incident management, which focuses on restoring service.
Recovery Point Objective (RPO)
The maximum acceptable amount of data loss measured in time. It determines the frequency of data backups required.
Recovery Time Objective (RTO)
The maximum acceptable duration of time within which a business process must be restored after a disruption to avoid unacceptable consequences.
Release Management
The ITIL process responsible for planning, scheduling, and controlling the movement of releases to test and live environments.
Remote Journaling
The process of transmitting transaction journal entries to a remote site in near real time, enabling recovery with minimal data loss.
SLA Monitoring
The continuous process of measuring and reporting on service level agreement compliance, including availability, response times, and resolution times.
System Resilience
The ability of a system to continue operating correctly and recover quickly in the face of faults, failures, or other adverse conditions.
Warm Site
A disaster recovery facility that has hardware and network connectivity pre-installed but requires software installation and data restoration before becoming operational.