Case notes
Published with client permission, with names and numbers adjusted where required by contract. The failure modes are unchanged — they are the part worth reading.
Tracking portal returned intermittent 520 and 522 errors, roughly 1 in 40 requests, concentrated on one mobile carrier. Six months of retries, three vendors blamed, nothing reproducible from the office.
Reproduced from a carrier SIM, not a laptop. Packet captures showed responses over 1492 bytes vanishing on the IPv6 path. The carrier filtered ICMPv6 Packet-Too-Big, so path-MTU discovery never got the signal — the connection simply stalled until timeout.
Pinned edge to IPv4 as an immediate stop-gap, then clamped MSS on all fourteen tunnels and got type 2 permitted upstream. IPv6 re-enabled in week five with a synthetic check that fails loudly if the filter returns.
Fig. 1 · response path, before and after MSS clamp
Quarterly restore drills were documented as policy. The last recorded execution was 26 months old. Retention claimed 90 days; the lifecycle rule had been silently transitioning objects to a tier the restore tooling could not read.
Ran a cold restore of the largest dataset with a stopwatch and no shortcuts. It took 31 hours against a stated 4-hour RTO, and two of six datasets failed outright on the archive tier. Findings went out the same day, unsoftened.
Rebuilt lifecycle rules against measured read patterns, added a restore-path check to the retention job, and put a quarterly drill into the pipeline so it fails the build rather than the audit.
Fig. 2 · measured restore by stage · drill 1 vs post-rework
CI took 38 minutes on a codebase that compiled locally in four. The team had added more runners twice, which made it more expensive without making it faster.
Instrumented the cache before changing anything. The key included the commit SHA, so every run was a cold run — the cache had been decorative since it was added. A second issue: the base image was rebuilt per job rather than pulled.
Re-keyed on lockfile digest plus toolchain version, pinned and published the base image, and added a hit-rate metric with an alert threshold so the regression cannot happen quietly again.
Fig. 3 · CI wall-clock by stage · median of 20 runs
Register
Clients who asked not to be named appear as sector only. Two engagements are listed as withdrawn — we scoped them, disagreed with the plan, and did not take the work.
| Client | Practice area | Length | Headline outcome | Status |
|---|---|---|---|---|
| Northvane Logistics | Network & edge | 6 wks | Path-MTU blackhole removed; p99 down 15.8 s | Closed |
| Petrichor Bio | Storage | 9 wks | RTO 31 h → 3 h 10 m, verified by drill | Closed |
| Kite & Pallas | Build & release | 4 wks | CI 38 m → 6 m 40 s | Closed |
| Halyard Retail Group | Migration | 14 wks | Datacentre exit, 11-minute cutover window | Closed |
| Sundara Freight | Network & edge | 5 wks | Anycast failover drill; real-IP restored in logs | Closed |
| Cormorant Media | Storage | 7 wks | Replication lag budget defined and enforced | Closed |
| Fintech, Singapore | Second opinion | 3 days | Migration plan rejected; rescoped in-house | Closed |
| Logistics, Thailand | Migration | — | Cutover window unacceptable; declined | Withdrawn |
| Gaming, Vietnam | Build & release | — | Root cause was org, not pipeline; declined | Withdrawn |
| Healthcare, Thailand | Storage | 11 wks | Key custody and break-glass rebuilt | Closed |
Bring the symptom and the last thing you tried. First reply is a real answer, not a calendar link.
Start a scoping call