Withdrawn by the author. Please disregard the original contents of this post — I have removed them rather than leave a wrong conclusion sitting in a search index.
What I got wrong. I reported that pinning the dmc governor to performance stopped a memory-corruption fault on one of my two Axon V0.2 boards (0 faults in 20 pinned, against 2 in 18 on stock dmc_ondemand), and suggested DRAM frequency retraining was implicated.
That conclusion does not survive more data. After posting, I reverted the board to stock dmc_ondemand to chase a further question, and it then ran 87 iterations with 10,975 DRAM frequency transitions and produced zero faults — the same configuration that had previously failed twice in eighteen runs.
The mistake was a confound, not a measurement error. Every fault I ever observed fell inside a single window of about two to three hours. Everything after that window is clean, whether pinned or not. So my “A/B” actually compared inside the fault window, unpinned against after the fault window, pinned. I changed one variable, saw the fault stop, and credited the change — when the fault had simply stopped. With a rare intermittent fault the null hypothesis has to be that it stopped on its own, and I did not establish the base rate over a comparable duration before publishing.
I would rather withdraw this quickly than have anyone spend engineering time chasing a rockchip-dmc lead I created.
What I still believe is true, and is not the point of this post: one of my boards did genuinely corrupt data in memory — captured at byte level, always memory-resident with the on-disk copy intact, and never reproduced on the identical second board. The trigger is unknown, and it is not frequency transitions and not temperature. I will open a fresh, properly-controlled report if and when I have something that survives its own null hypothesis.
Apologies for the noise.