Blog
Notes from the bench
Write-ups from the device bench — what CI reported, and what real phones did.
RSSGrade Your AI-SDLC Governance: 10 Dimensions, L0-L4
July 24, 2026
In May, 84 malicious npm package versions shipped with valid SLSA provenance. Provenance proves where code was built, not whether it was authorized. This is the 10-dimension rubric I use to govern AI coding agents across ~25 repos, my own scorecard included.
The tests were never flaky. At 33% red, nobody could tell.
July 16, 2026
I ran my full mobile CI diagnostic against nextcloud/android using only public data. 14,449 runs, 252 root-cause clusters, and two-thirds of every failure traced to one line of test code.
The lab on your wrist
July 15, 2026
What published research says phone and watch sensors can actually do, and the three places sensor projects fail on the way from lab demo to product.
My own tool scored 0/5 on the one task it exists for. I published that first.
July 13, 2026
I built a pre-registered benchmark for mobile-agent tooling, ran my own stack against three competitors, and lost the task I built the thing for. Here is the loss, the fix, the re-run, and the reason you should still be suspicious of the result.
The bench said the phone was idle. The wall charger made that a lie.
July 12, 2026
First full run of my device bench across three real Android devices: what passed, what the emulator had been hiding, and why my idle test was never idle.