Servers
Server management: who is responsible for what?
Most server incidents we see are not technically difficult. They happen because everyone assumed someone else was responsible. Here is a practical way to divide responsibilities between the parties involved in running a business application.
The four parties
- Hosting or cloud provider: physical hardware, data centre, network to the server, and the virtualisation platform.
- Server-management team: the operating system, web server, PHP or runtime, database engine, patching, monitoring, backups and recovery tests.
- Application developers: the application code, its dependencies, migrations and application-level bugs.
- Your business: decisions, priorities, access approvals, vendor contracts and paying for the resources the system needs.
The grey areas to settle in writing
Application dependencies, such as framework and library updates, sit between operations and development. Decide who monitors them and who tests them.
Backups are often split badly: the provider snapshots the disk, the managed team backs up the database, and nobody tests a full restore. Write down which backups exist, where they are stored, how long they are kept, and who restores them.
DNS and certificates often belong to whoever set them up years ago. Move domain registrar and DNS accounts into the business’s name and delegate access.
Security incidents need a named first responder and a named decision-maker, agreed in advance.
A simple responsibility matrix
For each task — patching, monitoring, backup, restore test, certificate renewal, deployment, database tuning, capacity planning, incident response — record who is responsible, who must be consulted and who must be informed. Review the matrix whenever a supplier changes.
Signs that responsibilities are unclear
- Nobody can say when the operating system was last patched.
- Monitoring alerts go to a former employee’s mailbox.
- The only person with root access is an external freelancer.
- Backups exist but have never been restored.