Stripe webhook retries piling up
inbound
supportmaintenanceautomation
We take over support, maintenance and automation for teams who have a working product and nobody left to look after it.
currently keeping [PLACEHOLDER: 14] systems running for teams in fintech, logistics and health
Support, maintenance and automation are the same job wearing different hats: keeping something that already works from quietly getting worse.
01
Your users email us, not you. We triage in your tracker, reply in your tone, and escalate to someone who has read the code. Nights and weekends are a rota, not hope.
12
open
on callyouus
02
Frameworks go end-of-life whether or not anyone is watching. We keep versions current, patch CVEs before your auditor finds them, and refactor the parts that make every change expensive — in small, reversible steps.
03
Most of what a company calls process is a person retyping data between two tools it already pays for. We find those, and replace them with something that runs at 06:00 and says so when it does not.
Handover is the part most people get wrong. We do it slowly on purpose — one service at a time, with the previous owner still in the room.
01week 1
We read everything — repo, infrastructure, tickets, the runbook nobody updated since the last person left. You get a written map of the system, the risks in it, and what we would do first. It is useful even if you stop here.
output · system map + risk register
02weeks 2–4
Access, credentials, deploy paths, on-call. We shadow whoever holds it now, take the pager one service at a time, and write down the parts that only ever lived in someone’s head.
output · runbooks in your repo
03ongoing
We hold the queue. Every Friday you get a note: what came in, what shipped, what is scheduled, what is getting worse. One call a month. Nothing else for you to manage.
output · weekly note, monthly call
An unqualified metric is marketing. These are the four we report to clients every month, including the months they are not flattering.
uptime across managed services
rolling 90 days · [PLACEHOLDER]
median first response
p50, business hours · [PLACEHOLDER]
systems maintained
across [PLACEHOLDER: 12] clients
longest engagement
still running · [PLACEHOLDER]
Compressed to the shape they actually had: what was wrong, what we changed, what happened next.
note 01
[PLACEHOLDER: logistics SaaS · 60 staff]
rails 5.2 · aws · sidekiq
problem
340 open tickets, a Rails 5.2 monolith, and a deploy process that lived in one person’s terminal history. Nobody left could say which of the four cron boxes mattered.
intervention
Took the inbox in week two. Wrote the deploy down, then automated it. Upgraded to 6.1 across nine weeks in merges small enough to revert, then to 7.1.
outcome
Queue under 20 by the end of the quarter. Deploys went from 40 minutes and a phone call to 6 minutes and a button. [PLACEHOLDER: metric]
note 02
[PLACEHOLDER: healthcare ops · 120 staff]
python · gcp · airflow
problem
The job that sent billing data to their provider died on a schema change and logged nothing. Finance had been reconciling by hand long enough to assume that was normal.
intervention
Rebuilt the job with retries and a dead-letter queue, added alerting that pages a person rather than filling a channel, and backfilled the missing period.
outcome
Eleven weeks of data recovered in four days. Finance stopped reconciling by hand. [PLACEHOLDER: metric]
note 03
[PLACEHOLDER: fintech · 25 staff]
node · stripe · xero
problem
Ops exported invoice lines from Stripe, pasted them into a spreadsheet to fix the mapping, then keyed them into Xero. Month-end close carried whatever they mistyped.
intervention
One small service reads the Stripe API, maps line items, writes to Xero, and puts anything it is unsure about on a single exceptions page. No new vendor, no new subscription.
outcome
Nine hours a week back. Manual errors stopped showing up in close two months later. [PLACEHOLDER: metric]
Most people start on a retainer after the audit. If you still have engineers, on-call is usually the cheaper answer — we will say so.
| Attribute | RetainerYou hand over the whole thing. | On-callYou keep your developers. We cover what they cannot. | Automation projectOne process, fixed and handed back. |
|---|---|---|---|
| best for | No engineers left in-house | A small team that is stretched | A specific manual process |
| scope | Support, maintenance and the upgrade roadmap | Escalation and out-of-hours cover only | Scoped build, then handover |
| response | [PLACEHOLDER: SLA] | [PLACEHOLDER: SLA] | Weekly check-in |
| cover | Business hours plus on-call rota | 24/7 rota | Project hours |
| commitment | 3 months, then 30 days’ notice | Rolling monthly | Fixed scope, fixed end |
| from | [PLACEHOLDER: £X,XXX / mo] | [PLACEHOLDER: £XXX / mo] | [PLACEHOLDER: £X,XXX fixed] |
You hand over the whole thing.
You keep your developers. We cover what they cannot.
One process, fixed and handed back.
not sure which — the audit answers it
30 minutes · no deck · we will say if we are not the right fit