Managed Services & 24/7 Support
Someone on your team is already carrying the pager, usually without having been asked. We take over the monitoring, the out-of-hours cover and the incidents themselves.
The rota is three people and one of them is on holiday. Alerts go to a channel everyone has muted, and the person who knows why the overnight reconciliation job fails is the same person who wrote it years ago. Nobody decided this arrangement; it accumulated. Running systems outside office hours is a job, and it needs people whose actual job it is.
Monitoring that alerts on everything alerts on nothing
Every service has a dashboard and a threshold somebody set once. The channel fills overnight, the team learns to scroll past it, and the one alert that mattered arrives in the same grey as the rest. Meanwhile the runbook is a person, and escalation means ringing them at home.
- Alerts firing on conditions nobody has tuned since the system launched
- A runbook that lives in one engineer's head and leaves when they do
- Escalation that means phoning someone who is asleep and hoping they answer
- The same incident every month, because nobody was given time to fix the cause
Nobody can support what nobody wrote down
Before we take a system on, four things have to exist on paper: what it does, what normal looks like on an ordinary Tuesday, what has broken before and how it was fixed, and who may approve a change at three in the morning. Most of the time at least two of those are missing, and finding them is the first piece of work.
That discovery is chargeable work, not a formality before the real contract starts. We sit with the people who currently get the calls, read back through the alerts and the outages, and write down the failure modes they already carry in their heads. Taking over a system nobody has documented is how a support contract turns into an argument about whose fault the outage was.
What an agreement covers
| Scope | What we do | Stays with you |
|---|---|---|
| Monitoring and alerting | Tune the alerts, watch the dashboards | What counts as business-critical |
| Incident response | Take the call, escalate, write up the cause | Sign-off on customer communication |
| Maintenance and patching | Schedule, test and apply in your accounts | The window we are allowed to use |
| Backup and recovery | Run restores and prove they work | How much data loss is acceptable |
| Change and release | Run the release through your pipeline | Approval, and the release calendar |
| Capacity and cost | Report on spend and headroom monthly | The budget and whether to spend it |
What you get
Cover is easy to promise and hard to evidence. These are the things you can check after a month of it, rather than the things a proposal can say.
Cover that is not one person
The rota is a team, with a documented handover between shifts, so a resignation or a fortnight in Greece becomes a staffing question rather than an outage risk.
Every incident ends in writing
A cause, a timeline and the change that stops it recurring, filed in your repository. Repeat offenders get scheduled as work instead of absorbed by whoever is on call.
A monthly view of what happened
What broke, what changed, what is trending towards trouble and where the spend went. A working document for your review, not a certificate of good service.
An exit that is already written
Runbooks, dashboards and on-call procedures live in your repositories and your accounts throughout. Leaving is a scheduled handover of things you already hold, not an extraction.
Before you sign
Someone is awake and paged, with a second name behind them if the first does not answer. What varies is what wakes them. We agree with you which systems justify a call at night and which can wait until the morning, and that list goes into the agreement rather than being assumed on both sides.
Not sure which one you need?
Describe the problem in a paragraph and we will tell you which service applies.