Two 3-second timelines of MariaDB 11.8 redo counters at flush_log_at_trx_commit=0: Innodb_os_log_written jumps 526 times, every 5.6 ms; Innodb_lsn_flushed jumps 3 times, every 1002 ms, with redo still in RAM between jumps; below, kill -9 five times lost 9,123 acknowledged commits

A Counter Said MariaDB Wrote Its Redo Log Every 5.6 ms. It Wrote It Once a Second.

Somewhere in your monitoring there is a panel built on a status counter you picked because of its name. Maybe the query came from a dashboard written for another database, or for another fork of the same one. Have you ever sampled that counter by hand, next to one you already understand, to check that the two move the way you think? I picked Innodb_os_log_written for a measurement because its name says exactly what I wanted: bytes of redo log written. On MySQL 8.4 it agreed with the crash results. On MariaDB 11.8 it said the redo log left the process every 5.6 ms. A crash test on the same server, earlier that day, had lost 9,123 acknowledged commits, which only happens if the log sits in memory for most of a second. I had measured both numbers myself, and they could not both be true. ...

October 5, 2026 · 8 min · Hoang Nguyen Thai
Two bar charts of read cost at chain depth 60. Left, struck through in red: what my explanation predicted, oldest reader much taller than newest. Right: what the table already said, newest 1450 ns and oldest 1480 ns, equal, with the ratio 1.02 highlighted.

I Explained a Slow Read and Wrote It Down. The Ratio That Disproved It Was Already in the Table.

You have a benchmark that says something is slow. You also have an explanation for why, and it sounds right. If you have already written that explanation into a design note or a ticket, this post is about the one check I skipped before doing exactly that, and what it cost me. The system is minidb, a small relational database I wrote in Go to learn storage internals. The mistake would have happened with any benchmark that has two rows. ...

October 5, 2026 · 8 min · Hoang Nguyen Thai
Log-log chart: a dotted extrapolation line ending near 10 ms, two measured no-index curves far above it (MariaDB 10.11 at 968 ms and 13.0 at 148 ms at 1M rows), a dashed 300 ms threshold, and a flat under-1 ms line for the indexed case

I Wrote the Index Escalation Threshold as a Number, Then Measured It and Found It Off by 15x to 100x

Somewhere in a design document you have signed off on, there is a sentence of the form “when it passes X, we will do Y.” Has anyone ever generated X rows to check that the number means what it says? At the end of last month I signed off on a design that runs a lookup against a column with no index, on a table my team does not own, and wrote the condition for fixing it as a sentence with two numbers in it: “when the table passes about 1,000,000 rows or the p95 of the voucher filter passes 300 ms, request a non-unique index from the owning team.” The sentence had one job: get the option “ship without the index” through design review, with “we will add it later” turned into something a reviewer could sign. It did that job. ...

September 28, 2026 · 10 min · Hoang Nguyen Thai
A merge commit with two parents, and a third branch feeding one extra line into the resolved file

A Merge Has Exactly Two Parents. The Line From a Third Branch Took Down Every API Route.

For 17 hours and 55 minutes, every route under /api/* on staging answered with a fatal error. The merge that caused it landed at 17:18 one evening; the fix that ended it landed at 11:13 the next morning. That duration is from the private repository and the server log, so you cannot verify it; read it as my account. The merge itself was ordinary: a feature branch into staging, one conflict, in the file that registers HTTP middleware. The conflict was real, both sides had added a line at the same spot. What went wrong was not the conflict. It was how the conflict was “resolved”, and then how the result was checked and declared fine. ...

September 23, 2026 · 7 min · Hoang Nguyen Thai
A 200-unit bar split into 110 allocated and 90 shared pool, the 10 units sold out of quotas highlighted inside the 110, and two subtractions below it: the literal reading giving 80 and the production screen giving 90

The Formula in the Ticket Subtracted Sales Twice, and Only Two Screenshots Could Prove It

One SKU. Total intake 200 units. Three campaigns hold quotas of 10, 88 and 12 units, and have sold 1, 9 and 0 out of those quotas. How many units are still available to sell? The ticket said: available = total intake − total allocated − total sold. Plug the numbers in and you get 200 − 110 − 10 = 80. The production screen for that SKU, at the same moment, said 90. ...

September 23, 2026 · 7 min · Hoang Nguyen Thai
Two tables with overlapping auto-increment ids, an import file pointing at one and the code reading the other

The Import That Edited the Wrong Row for Years Because Two Tables Shared the Same IDs

An import feature that adjusts campaign quotas from a spreadsheet had been in production for years. Every manual QA pass looked the same: upload a file, open the campaign, see one quota changed, mark the ticket done. Zero automated tests covered the handler. It was editing the wrong row. Not sometimes — structurally, on every run where the two ids happened to line up, and on the staging database they did. The number of production rows it touched over those years is something I have not measured and will not guess at here. What I can show is the mechanism, why three separate safety nets each let it through, and the one test that would have caught it on day one. ...

September 22, 2026 · 7 min · Hoang Nguyen Thai
An orphaned row still counted against a shared stock pool with no screen to release it

Ghost Quota: How a Missing Foreign Key Quietly Lowers a Sales Ceiling

A table called campaign_product_variants had no foreign key, no ON DELETE CASCADE, and no deleted_at. Before the shared pool existed, that was a cosmetic problem: a handful of rows pointing at SKUs that no longer existed made a report look a few units smaller than reality. I never saw anyone treat that as a bug, and I did not go back through old tickets to check. Then we shipped a shared stock pool — several campaigns drawing quota from one pool per SKU — and the same orphaned rows stopped being cosmetic. On a dev database I could still poke at, 14 orphan rows were locking 294 units out of sale, permanently, with no admin screen that could select and release them. The missing foreign key had turned pre-existing technical debt into a lowered sales ceiling. ...

September 22, 2026 · 5 min · Hoang Nguyen Thai
A validation error message rendered inside a hidden tab pane, invisible to the user

Every Increase Failed With 422, and the User Saw Nothing At All

QC’s bug report had two sentences. Clicking Increase in the adjust-quantity modal did nothing. Clicking it again also did nothing. “Nothing” is the worst symptom to be handed. A crash leaves a stack trace, a wrong result leaves a wrong number; an unchanged screen leaves no starting point at all. What was actually broken The screen is an internal admin panel for a promotions backend. An operator opens a modal to adjust the allocated quantity for one SKU in one campaign. The modal has two tabs — Increase and Decrease — and one submit button. Whichever tab is active decides which operation is sent. ...

September 22, 2026 · 8 min · Hoang Nguyen Thai
Diagram of a feature flag gating deploy order between two teams

A Feature Flag as a Cross-Team Release Gate

There’s a kind of bug I never saw in a log, because it never ran: I read it straight out of two lines of code, before writing a single line of my own. The question wasn’t “how do I fix this” — it was whether shipping something you already know will fail every single time, if flipped on too early, counts as done. Context Early this September I owned the backend piece of a second stock mode for promotional campaigns. The mode already in production: a product has to be allocated a fixed quota before it can sell in a campaign. The new mode is the reverse: a product gets no quota of its own — it sells against the warehouse’s available stock, live, at order time. The feature had a fixed ship date, and my part sat inside it. ...

September 22, 2026 · 5 min · Hoang Nguyen Thai
A SELECT FOR UPDATE statement dissolving into an empty string on the sqlite driver

The Lock I Dropped During a Refactor, and the 1184 Green Tests That Didn't Notice

QA filed a one-line gap: no concurrency tests on the stock pool. I opened the file expecting to write four tests and close it in an afternoon. Instead I found that the lock those tests were supposed to cover was not there anymore. What can go wrong here The system is a promotions backend. Several campaigns draw from one shared stock pool per SKU. Before a campaign can reserve units, a service reads the pool balance, subtracts what is already reserved, and writes an allocation. ...

September 21, 2026 · 6 min · Hoang Nguyen Thai