
A Counter Said MariaDB Wrote Its Redo Log Every 5.6 ms. It Wrote It Once a Second.
Somewhere in your monitoring there is a panel built on a status counter you picked because of its name. Maybe the query came from a dashboard written for another database, or for another fork of the same one. Have you ever sampled that counter by hand, next to one you already understand, to check that the two move the way you think? I picked Innodb_os_log_written for a measurement because its name says exactly what I wanted: bytes of redo log written. On MySQL 8.4 it agreed with the crash results. On MariaDB 11.8 it said the redo log left the process every 5.6 ms. A crash test on the same server, earlier that day, had lost 9,123 acknowledged commits, which only happens if the log sits in memory for most of a second. I had measured both numbers myself, and they could not both be true. ...

I Explained a Slow Read and Wrote It Down. The Ratio That Disproved It Was Already in the Table.
You have a benchmark that says something is slow. You also have an explanation for why, and it sounds right. If you have already written that explanation into a design note or a ticket, this post is about the one check I skipped before doing exactly that, and what it cost me. The system is minidb, a small relational database I wrote in Go to learn storage internals. The mistake would have happened with any benchmark that has two rows. ...

I Wrote the Index Escalation Threshold as a Number, Then Measured It and Found It Off by 15x to 100x
Somewhere in a design document you have signed off on, there is a sentence of the form “when it passes X, we will do Y.” Has anyone ever generated X rows to check that the number means what it says? At the end of last month I signed off on a design that runs a lookup against a column with no index, on a table my team does not own, and wrote the condition for fixing it as a sentence with two numbers in it: “when the table passes about 1,000,000 rows or the p95 of the voucher filter passes 300 ms, request a non-unique index from the owning team.” The sentence had one job: get the option “ship without the index” through design review, with “we will add it later” turned into something a reviewer could sign. It did that job. ...

A Merge Has Exactly Two Parents. The Line From a Third Branch Took Down Every API Route.
For 17 hours and 55 minutes, every route under /api/* on staging answered with a fatal error. The merge that caused it landed at 17:18 one evening; the fix that ended it landed at 11:13 the next morning. That duration is from the private repository and the server log, so you cannot verify it; read it as my account. The merge itself was ordinary: a feature branch into staging, one conflict, in the file that registers HTTP middleware. The conflict was real, both sides had added a line at the same spot. What went wrong was not the conflict. It was how the conflict was “resolved”, and then how the result was checked and declared fine. ...

The Formula in the Ticket Subtracted Sales Twice, and Only Two Screenshots Could Prove It
One SKU. Total intake 200 units. Three campaigns hold quotas of 10, 88 and 12 units, and have sold 1, 9 and 0 out of those quotas. How many units are still available to sell? The ticket said: available = total intake − total allocated − total sold. Plug the numbers in and you get 200 − 110 − 10 = 80. The production screen for that SKU, at the same moment, said 90. ...

The Import That Edited the Wrong Row for Years Because Two Tables Shared the Same IDs
An import feature that adjusts campaign quotas from a spreadsheet had been in production for years. Every manual QA pass looked the same: upload a file, open the campaign, see one quota changed, mark the ticket done. Zero automated tests covered the handler. It was editing the wrong row. Not sometimes — structurally, on every run where the two ids happened to line up, and on the staging database they did. The number of production rows it touched over those years is something I have not measured and will not guess at here. What I can show is the mechanism, why three separate safety nets each let it through, and the one test that would have caught it on day one. ...

Ghost Quota: How a Missing Foreign Key Quietly Lowers a Sales Ceiling
A table called campaign_product_variants had no foreign key, no ON DELETE CASCADE, and no deleted_at. Before the shared pool existed, that was a cosmetic problem: a handful of rows pointing at SKUs that no longer existed made a report look a few units smaller than reality. I never saw anyone treat that as a bug, and I did not go back through old tickets to check. Then we shipped a shared stock pool — several campaigns drawing quota from one pool per SKU — and the same orphaned rows stopped being cosmetic. On a dev database I could still poke at, 14 orphan rows were locking 294 units out of sale, permanently, with no admin screen that could select and release them. The missing foreign key had turned pre-existing technical debt into a lowered sales ceiling. ...

Every Increase Failed With 422, and the User Saw Nothing At All
QC’s bug report had two sentences. Clicking Increase in the adjust-quantity modal did nothing. Clicking it again also did nothing. “Nothing” is the worst symptom to be handed. A crash leaves a stack trace, a wrong result leaves a wrong number; an unchanged screen leaves no starting point at all. What was actually broken The screen is an internal admin panel for a promotions backend. An operator opens a modal to adjust the allocated quantity for one SKU in one campaign. The modal has two tabs — Increase and Decrease — and one submit button. Whichever tab is active decides which operation is sent. ...

A Feature Flag as a Cross-Team Release Gate
There’s a kind of bug I never saw in a log, because it never ran: I read it straight out of two lines of code, before writing a single line of my own. The question wasn’t “how do I fix this” — it was whether shipping something you already know will fail every single time, if flipped on too early, counts as done. Context Early this September I owned the backend piece of a second stock mode for promotional campaigns. The mode already in production: a product has to be allocated a fixed quota before it can sell in a campaign. The new mode is the reverse: a product gets no quota of its own — it sells against the warehouse’s available stock, live, at order time. The feature had a fixed ship date, and my part sat inside it. ...

The Lock I Dropped During a Refactor, and the 1184 Green Tests That Didn't Notice
QA filed a one-line gap: no concurrency tests on the stock pool. I opened the file expecting to write four tests and close it in an afternoon. Instead I found that the lock those tests were supposed to cover was not there anymore. What can go wrong here The system is a promotions backend. Several campaigns draw from one shared stock pool per SKU. Before a campaign can reserve units, a service reads the pool balance, subtracts what is already reserved, and writes an allocation. ...