Angel13 SystemsInfrastructure Practice

Case notes

What the work
actually looks like.

Published with client permission, with names and numbers adjusted where required by contract. The failure modes are unchanged — they are the part worth reading.

Case 01 · Network & edge · 6 weeks

Northvane Logistics — the 520s that only happened on mobile

Problem

Tracking portal returned intermittent 520 and 522 errors, roughly 1 in 40 requests, concentrated on one mobile carrier. Six months of retries, three vendors blamed, nothing reproducible from the office.

Approach

Reproduced from a carrier SIM, not a laptop. Packet captures showed responses over 1492 bytes vanishing on the IPv6 path. The carrier filtered ICMPv6 Packet-Too-Big, so path-MTU discovery never got the signal — the connection simply stalled until timeout.

Outcome

Pinned edge to IPv4 as an immediate stop-gap, then clamped MSS on all fourteen tunnels and got type 2 permitted upstream. IPv6 re-enabled in week five with a synthetic check that fails loudly if the filter returns.

Packet path before and after MSS clamping. Before: 1500-byte responses are dropped by the carrier and the ICMPv6 Packet-Too-Big signal is filtered upstream, so the connection stalls until timeout. After: MSS is clamped to 1452 bytes at the tunnel and the response completes. BEFORE · 1 IN 40 REQUESTS client (v6) carrier B edge origin ✗ 1500 B response dropped ICMPv6 TYPE 2 · FILTERED UPSTREAM stall 15–20 s 520 / 522 AFTER · WEEK 5 client (v6) carrier B edge origin ✓ 1452 B · MSS clamped at tunnel ICMPv6 TYPE 2 · PERMITTED + SYNTHETIC CHECK 200 OK p99 −15.8 s

Fig. 1 · response path, before and after MSS clamp

1 in 40 → 0 errors in 30 days −15.8 s p99 latency 14 tunnels corrected
Case 02 · Storage · 9 weeks

Petrichor Bio — a backup policy that had never been tested

Problem

Quarterly restore drills were documented as policy. The last recorded execution was 26 months old. Retention claimed 90 days; the lifecycle rule had been silently transitioning objects to a tier the restore tooling could not read.

Approach

Ran a cold restore of the largest dataset with a stopwatch and no shortcuts. It took 31 hours against a stated 4-hour RTO, and two of six datasets failed outright on the archive tier. Findings went out the same day, unsoftened.

Outcome

Rebuilt lifecycle rules against measured read patterns, added a restore-path check to the retention job, and put a quarterly drill into the pipeline so it fails the build rather than the audit.

Measured restore time by stage. Drill 1 took 31 hours against a stated 4-hour objective, dominated by a 24-hour archive-tier thaw, with 2 of 6 datasets unrecoverable. After rework: 3 hours 10 minutes, all 6 restorable. LOCATE ARCHIVE THAW TRANSFER VERIFY Drill 1 2026-03-11 31 h 00 m ARCHIVE THAW 24 H TRANSFER ✗ 2 of 6 datasets unrecoverable on archive tier Post-rework 2026-05-02 3 h 10 m ✓ 6 of 6 restorable 0H 6H 12H 18H 24H 30H STATED RTO 4 H

Fig. 2 · measured restore by stage · drill 1 vs post-rework

31 h → 3 h 10 m measured RTO 6/6 datasets restorable −41% storage spend
Case 03 · Build & release · 4 weeks

Kite & Pallas — a cache with a 4% hit rate

Problem

CI took 38 minutes on a codebase that compiled locally in four. The team had added more runners twice, which made it more expensive without making it faster.

Approach

Instrumented the cache before changing anything. The key included the commit SHA, so every run was a cold run — the cache had been decorative since it was added. A second issue: the base image was rebuilt per job rather than pulled.

Outcome

Re-keyed on lockfile digest plus toolchain version, pinned and published the base image, and added a hit-rate metric with an alert threshold so the regression cannot happen quietly again.

CI wall-clock by stage, median of 20 runs. Baseline 38 minutes, dominated by an 18-minute cold dependency step and a 9-minute base-image rebuild. After re-keying the cache: 6 minutes 40 seconds. CHECKOUT BASE IMAGE DEPS COMPILE TEST Baseline CACHE HIT 4.1% 38 m 00 s BASE REBUILD DEPS — COLD CACHE COMPILE ✗ cache keyed on commit SHA — cold on every run Re-keyed CACHE HIT 91% 6 m 40 s ✓ keyed on lockfile digest + toolchain 0M 6M 12M 18M 24M 30M 36M

Fig. 3 · CI wall-clock by stage · median of 20 runs

38 m → 6 m 40 s median CI 4% → 91% cache hits −60% runner cost

Register

Closed engagements, 2024–2026.

Clients who asked not to be named appear as sector only. Two engagements are listed as withdrawn — we scoped them, disagreed with the plan, and did not take the work.

ClientPractice areaLengthHeadline outcomeStatus
Northvane LogisticsNetwork & edge6 wksPath-MTU blackhole removed; p99 down 15.8 sClosed
Petrichor BioStorage9 wksRTO 31 h → 3 h 10 m, verified by drillClosed
Kite & PallasBuild & release4 wksCI 38 m → 6 m 40 sClosed
Halyard Retail GroupMigration14 wksDatacentre exit, 11-minute cutover windowClosed
Sundara FreightNetwork & edge5 wksAnycast failover drill; real-IP restored in logsClosed
Cormorant MediaStorage7 wksReplication lag budget defined and enforcedClosed
Fintech, SingaporeSecond opinion3 daysMigration plan rejected; rescoped in-houseClosed
Logistics, ThailandMigrationCutover window unacceptable; declinedWithdrawn
Gaming, VietnamBuild & releaseRoot cause was org, not pipeline; declinedWithdrawn
Healthcare, ThailandStorage11 wksKey custody and break-glass rebuiltClosed

Recognise any of these?

Bring the symptom and the last thing you tried. First reply is a real answer, not a calendar link.

Start a scoping call