Anatomy of a 200ms API Endpoint
Where does 200ms actually go? Usually not where you think. This is a trace-by-trace teardown of a real endpoint: connection pool, query plans, N+1s hiding in serializers, cache stampedes, and the Octane config that shaved the last 40ms.
Measure before you optimize
We added request tracing with exact per-layer timings. The first finding: 60% of the time was in Eloquent hydration — not the database.
The fixes
Selective columns, eager loading, and a tagged cache with probabilistic early expiration to kill stampedes.