I wanted to see if LLMs could port the TypeScript compiler, checker and lsp to Rust. Turns out they can.
It
cost over $420,000
in tokens to do it, but you could probably have done it for ~$20k (see below)
Motivations
Test model capabilities
Make a fast TypeScript type checker
Make a ts checker that can work in WASM with high performance
Memes
Warnings
This is an early release.
It has 100% compatibility in every real world project we have
tested. It should work as a drop in
replacement for the vast majority of apps. See
Known problems
.
Also worth mentioning: I've never read a line of this code.
Install
Be warned, I have no idea if this will actually work.
npm install -D tsc-rs
npx tsc-rs -p tsconfig.json
How did this go?
I used a lot of OpenAI models to try and complete this port. In total I did
over $400,000 in API priced tokens with GPT-5.6 Sol and GPT 6 Astra
. They wrote over 1.3m lines of Rust over multiple months of /goal loops and never got past like 84% compat.
When I saw how little my Claude Code limits were burning, I figured it'd be fun to throw Opus 5.5 at this. It had a working v0 in 10 hours.
I assumed it kept using the code the Codex models wrote. I was wrong.
Opus 5.5 started from scratch. It got further than Astra in 1/10th the time.
I let it keep going, and it definitely did. Total token spend was
~$24,047 of API spend over 2 weeks
. I was using my Claude accounts, and it worked out to somewhere between
925% and 983% of my $200 plan weekly limits
.
Expensive, for sure, but not that bad considering how much work has went into typescript-go.
"The Slop Line"
Everything below this was written by my LLMs, not me.
What actually is this?
ts-rust is a direct port of Microsoft's native TypeScript compiler, which is written in Go
(
microsoft/TypeScript
, formerly
typescript-go
). It keeps Go's algorithms and
behavior and has the same command line (
tsc
), language server and API.
Install
npm install -D tsc-rs
npx tsc-rs -p tsconfig.json
tsc-rs
takes the same options as
tsc
. The npm package is
tsc-rs
so that it does not clash
with the
typescript
package. Each
release
also
has a standalone archive per platform: the
tsc
binary with the lib files next to it.
Platforms: Linux x64 (static, any distribution) and macOS arm64. Windows and Linux arm64 are not
available yet.
tsc-rs
has the
Effect
language service diagnostics built in (codes
377xxx), so an Effect project needs no second compiler. They come from the same check as the
TypeScript diagnostics, and the language server shows them too. They run only when the tsconfig has
the plugin, as with
@effect/language-service
:
The rules, options and
@effect-diagnostics
comments are a port of
Effect-TS/tsgo
0.46.1. The editor features of the language
service (quick fixes, refactors, hover, completions) are not ported.
Status
The port is pinned to one upstream revision, microsoft/TypeScript
673a5f17d713
(2026-09-29, TypeScript 7.1.0-dev;
UPSTREAM.md
), and compared with Go at that
revision. To compare, use
typescript@7.1.0-dev.20260929.1
, not 7.0.x or
@typescript/native-preview
. A difference that this build also shows is upstream behavior, and it
goes away when the port moves to a newer pin.
Same results.
TanStack Query core and Hono check with diagnostics identical to Go's. All
181,711 ported Go tests pass. The language server and API answers match Go on the oracle test
sets.
Faster.
On 60 open-source projects, type checking takes about half of Go's time (geometric
mean). The preview packages are built in CI without PGO and BOLT, so they are slower than that
measured build.
Real projects.
On 120 open-source repos, the command-line output differs from Go's only in
the problems below and where Go's own output changes from run to run.
Benchmark: T3 Code
Full type check of
T3 Code
, compared with
tsc
6,
tsc
7
and the new
bun check
in Bun. T3 Code uses Effect, so there are two cases: without the Effect
diagnostics and with them. Each time is the sum for the five T3 Code projects. Lower is faster.
Without Effect diagnostics
Checker
Time
vs
tsc
6
vs
tsc
7
bun check
4.07s
15.4×
3.95× faster
█
tsc-rs
7.25s
8.6×
2.22× faster
██
tsc
7
16.10s
3.9×
baseline
█████
tsc
6
62.63s
baseline
3.89× slower
██████████████████
With Effect diagnostics
Checker
Time
vs
tsc
6
vs
tsc
7 + Effect
tsc-rs
(Effect built in)
11.13s
12.5×
1.89× faster
███
tsc
7 +
@effect/tsgo
21.07s
6.6×
baseline
██████
bun check
, then
effect-tsgo diagnostics
37.60s
3.7×
1.78× slower
███████████
tsc
6 +
@effect/language-service
138.63s
baseline
6.58× slower
████████████████████████████████████████
bun check
is the fastest when you do not need the Effect diagnostics. It does not have them, so
an Effect project needs a second pass.
tsc-rs
gets them from its one check.
Errors.
tsc-rs
,
tsc
7 +
@effect/tsgo
and the
effect-tsgo diagnostics
pass report the same
221 Effect diagnostics.
tsc
6 uses the JavaScript Effect plugin
(
@effect/language-service
0.87.4), which has a different rule set: it reports 287 on
apps/server
where the others report 177.
tsc-rs
and
tsc
6 report one more error, TS2322 in
apps/server/scripts/record-pi-rpc-replay-fixture.ts
. TypeScript 7.1.0-dev reports it too, and
pingdotgg/t3code#16704
fixes it.
How it was measured: the same machine and method as the real-world apps below, with
--composite false
added (
apps/web
is composite). T3 Code at
cd41c4ad
,
projects
apps/server
,
apps/web
,
apps/mobile
,
packages/client-runtime
and
packages/shared
. Without Effect, the configs have no Effect plugin. With Effect,
tsc
7 is the
Effect-patched 7.0.2 from
@effect/tsgo
0.46.1, and
tsc
6 is 6.0.3 patched with
@effect/language-service
. The script is
scripts/bench-apps/t3code.sh
. The
tsc-rs
switch in T3 Code is
pingdotgg/t3code#16704
.
Benchmark: real-world apps
Full type check of six open-source apps with
tsc
6 (the JavaScript compiler),
tsc
7 (the Go
compiler),
tsc-rs
and
bun check
. The multiplier is the speedup over
tsc
6. Lower times are
faster.
Compared with
tsc
7,
tsc-rs
is 1.61× faster and
bun check
is 2.95× faster (geometric
means).
bun check
is the fastest on every app except tRPC.
*
bun check
reports errors that no other checker reports: 3 on Sentry and 2 on tRPC.
Each config checks with 0 errors under
tsc
7.0.2. The other differences:
tsc-rs
reports 10 errors on VS Code and 2 on Sentry. TypeScript 7.1.0-dev (
typescript@next
)
reports the same errors, line for line.
tsc-rs
ports a 7.1 dev revision, which has checks that
7.0.2 does not have.
tsc
6 reports 9 errors on VS Code.
How it was measured: Apple M4 Pro (12 cores, 48 GB), macOS 26.5.1.
hyperfine
,
median of 5 runs after 1 warmup run, with
--noEmit --incremental false
. Each checker uses its
default thread count.
tsc
7 and
tsc-rs
run as native binaries, without the npm launcher.
tsc
6 runs on Node 24.19 with a 16 GB heap, because it runs out of memory on VS Code and Sentry
with the default heap. Versions:
tsc-rs
0.1.0, TypeScript 7.0.2 and 6.0.3, Bun canary
bd599f5af
. Lines checked is the
tsc
7
--extendedDiagnostics
count, with the
.d.ts
files.
The T3 Code benchmark above uses the same machine and method.
Four apps needed changes to check with 0 errors under
tsc
7. Nothing else changed:
Excalidraw: no
baseUrl
, because TS 7 removed it.
TypeORM:
moduleResolution
changed from
node
to
nodenext
, because TS 7 removed
node
.
VS Code: the
electron
typings that its postinstall adds.
Playwright: the sources that its build generates.
Two apps are not in the table:
rxjs main needs its workspace packages built first.
date-fns uses project references. There,
tsc -p
and
bun check
do different work.
The scripts are in
scripts/bench-apps
:
setup.sh <dir>
, then
run.sh <dir>
and
summary.py <dir>
, and
t3code.sh <dir>
for T3 Code.
Known problems
In some monorepos, the source files of a workspace package are reachable both through
node_modules
and through a direct import. There,
tsc-rs
can write output for more of those
files than
tsc
does.
In
tsc -b
, when one project imports the output of another project without a project reference,
tsc-rs
can report TS2307 (cannot find module) where
tsc
happens to build the other project
first. Add the reference to fix it.
tsc -b --watch
can stop with an internal error (exit code 70) after some edits.
In the editor, memory grows slowly during long edit sessions.
tsc-rs --version
prints the TypeScript version that it ports (7.1.0-dev), not the npm
version. The compiler matches
typesVersions
against it.
Development
crates/ts_goport
is the compiler. It has two parts crates,
goport_util
and
goport_lsproto
,
in
crates/ts_goport/parts
, and uses the lib files in
crates/ts_goport/libs
.
tools/ts_ast_codegen
generates
crates/ts_goport/src/astdata
, and
tools/ts_diagnostics_codegen
generates
crates/ts_goport/src/diagnostics/catalog.rs
and
crates/ts_goport/src/diag.rs
.
crates/ts_wasm
is the WebAssembly build
(
npm/wasm
).
The bins are
goport
(type check) and
tsgo
(the Go
tsgo
command line). The Go baseline tests
run with
TS_GO_REPO=/path/to/typescript-go ./scripts/run-cargo-capped.sh test -p ts_goport --test go_baselines
.
Push a tag
v<version>
(for example
v0.1.0
). The
release workflow
builds, packs and tests the packages, publishes
them to npm and creates a GitHub release. A stable version goes to the dist-tag
latest
, and a
prerelease version (
v0.2.0-beta.1
) to
next
and a GitHub prerelease. See
npm/README.md
.
License
MIT
. The port keeps the licenses and notices of the code it ports: TypeScript
(Apache-2.0) and parts of the Go standard library (BSD-3-Clause). See
NOTICE.md
.
The page you have tried to view (
LWN.net Weekly Edition for October 8, 2026
) is currently available to LWN
subscribers only.
Reader subscriptions are a necessary way
to fund the continued existence of LWN and the quality of its content.
If you are already an LWN.net subscriber, please log in
with the form below to read this content.
Please consider
subscribing to LWN
. An LWN
subscription provides numerous benefits, including access to restricted
content and the warm feeling of knowing that you are helping to keep LWN
alive.
(Alternatively, this item will become freely
available on October 15, 2026)
Jaguar Type 01
Daring Fireball
insideevs.com
2026-10-07 19:22:28
This is a cool car. It doesn’t look like anything else and it looks like it’s been pulled a few years from the future. I don’t know about this whole trend where cars don’t have rear windows though. I get it that you just use the camera but my old habits says I want to see out the back sometimes.
Bu...
The Jaguar Type 01 is the production version of the Type 00 coupe. It is fully electric.
Jaguar calls the Type 01 a four-door GT, not a sedan. It has four doors and four seats.
The Type 01 will open for sale early next year, and start at a price of $130,500.
“We don’t anticipate this car being for everyone,” Rawdon Glover, the Managing Director of Jaguar, said at the unveiling of the Jaguar Type 01. The car is immensely important for the brand, necessitating a secretive short trip to JLR’s Gaydon, UK campus for an early preview. It’s an incredibly controversial project, too. Jaguar killed off its entire lineup for this, rebooting the brand with a new logo and new look. The new design and direction were met with bashing by conservative pundits, landing Jaguar in the midst of yet another culture war.
Thankfully, the haters haven’t prevailed, because the Jaguar Type 01 is such a stunning piece of metal.
Jaguar’s tagline for both the brand and the car is “Copy Nothing,” and I have to agree—the Type 01 doesn’t look like much else on the road. The proportions are largely identical to what we’ve seen from the original hot pink and ice blue concepts shown off at the 2024 Miami Art Basel. This means the car is exceptionally cab-rearward, with extremely large wheels and tiny windows; it’s almost the antithesis of cars like the Lucid Air. The side surfacing is strangely blank, but elegant; it's remarkably clean and functional, as if the Bauhaus school of design decided to resurrect itself and make a modern electric car.
Gallery: Jaguar Type 01
The front fascia is blank, but once again, subtly more complex than your initial impressions. Since it's an electric car, there’s no need for a huge grille to cool an engine, yet the relief surfaces in the closed-off front add drama to such a strikingly blank front end. At the rear, there’s no rear windshield, while the taillights, turn signals, and reverse lights are all integrated into the faux-rear fascia louvers.
Photo by: Jaguar
From the outside, it’s such a hard car to take your eyes away from. It’s delightfully cartoonish; EV motors and batteries should dictate a more useful shape, and yet, the Jaguar’s long hood looks like it’s hiding a huge V8. Contrary to typical EV design principles of late, Jaguar’s designers insist that the cartoonish proportions are in part possible because of its EV structure. Without being tied down to the traditional ergonomic limitations of an ICE powertrain, the designers were free to make the shape as wild as possible. This means 23-inch wheels, a super long hood, a tiny front overhang, but a Batmobile-like rear end. Also, despite being blocky and large, the Type 01 only has a drag coefficient of 0.23, making it Jaguar's most aerodynamic car ever.
The interior is likely to be a point of contention. Jaguar insisted on leaning into a sort of luxo-minimalist layout. There are three screens: a big one for the driver, a small cell-phone-sized one in the center console, and a phone-style screen for rear passengers. I can understand why some may not have been in love with the Type 01’s interior, since it’s not on the same level of lush as other pricy luxury cars.
Photo by: Jaguar
In person, I can’t disagree with the Type 01’s detractors more. Jaguar’s designers harped on materials and design in an attempt to push the car’s ethos toward a younger buyer. The interior’s unconventional surfaces and color choices feel fresh, as if I’m sitting in an expensive modern apartment.
It’s not all peachy keen, though. Those outrageous proportions make the Type 01 feel snug inside. The rear seat only has room for two, and there’s not much light or legroom. The trunk feels small, and Jaguar admits that the frunk is tiny despite the gigantic hood. (They wouldn’t show me.) It’s not a car for everyone. This car’s style is a big selling point, and well, everything else must follow suit. Even practicality.
Luckily, the mechanical bits seem to be more bite than bark here. Power for the Type 01 comes from a tri-motor setup producing a whopping 1,015 horsepower. This is fed by a 118 kWh battery, which Jaguar says should achieve about 400 miles of range on the EPA cycle. It uses an 850-volt electrical architecture and should charge from 10-80% in as little as 22 minutes when connected to a 350 kW DC fast charger.
Jaguar representatives say the Type 01 will drive just as engaging as other Jaguars. The weight distribution is 50/50, while the driving position is low for an EV. This is in part due to the packaging of the 118 kWh battery; it’s not entirely under the main cabin. A few cells are split off and placed in front of the driver, improving both weight balance and ergonomics. The frunk may be tiny, but that long hood is doing something, at least.
Photo by: Jaguar
Now, will it work? I don’t know. I’ve always liked the idea of Jaguar scrapping its stodgy image of old, poorly made cars sold to people who don’t like change, for one that’s significantly more fashion-forward. Cars and fashion have always been joined at the hip in some way, but I think that car brands have been less than able or willing to embrace what a true fashion-forward marketing or brand would look like. Of course, there have been one-offs like the
Virgil Abloh Maybach
, but I think Jaguar is proving what it could look like on a wider scale. The Type 01 feels like it’s the last accessory in a head-to-toe Balenciaga outfit.
Which is likely the brand’s intended goal. Jaguar wants this car to appeal to a younger buyer, and that means ingratiating itself with the lifestyles they lead. The Type 01 represents a new era for Jaguar. Perhaps the marketing and initial rebrand were confusing and not well received, but it served as a way to communicate to new buyers what exactly it stood for and where the brand wants to go.
“What we’re looking for is people who have a deep, emotional reaction to what we’re doing. These purchases aren’t about need. They’re about desire. They’re about want,” said Glover. Jaguar wants you to want the Type 01. The Jaguar Type 01 orders open in early 2027, with deliveries starting in the last half of that year. To get one, expect to fork out a hefty $130,500 for the privilege of owning Jaguar's electric four-door GT.
I've always been kind of into computers since I was young. And then when film started to move from analog film to digital, I became more interested in that aspect of it. And the visual effects workflow for many years has included machine learning.
So I can write like pretty shitty Python scripts and...
I've always been kind of into computers since I was young. And then when film started to move from analog film to digital, I became more interested in that aspect of it. And the visual effects workflow for many years has included machine learning.
So I can write like pretty shitty Python scripts and stuff like that because with convolutional neural networks, which were the sort of precursors to what the transformer can do, which is just much more computation simultaneously, you would do things like look at what's called a tensor, which is just the numerical translation of a visual image in numbers — like the batch number, the frame number, the red, green, and blue values of each pixel in each frame.
And a tensor, you use a convolutional neural network to identify patterns in that that would reveal what's called edge detection or feature extraction, which is just identifying patterns enough to know like this is where the window ledge is, so we can more easily take the green screen image out and replace it with something.
Ransomware recovery CEO charged over secret ransom payments
Bleeping Computer
www.bleepingcomputer.com
2026-10-07 19:04:37
The owner of ransomware remediation company MonsterCloud has been charged with allegedly defrauding ransomware victims by secretly paying their attackers for decryptors while claiming to use proprietary technology to recover encrypted data. [...]...
The owner of ransomware remediation company MonsterCloud has been charged with allegedly defrauding ransomware victims by secretly paying their attackers for decryptors while claiming to use proprietary technology to recover encrypted data.
Zohar Pinhasi, 50, also known as "Zack Silver" and "Zack Green," was indicted by a federal grand jury in the Eastern District of New York on September 23 and arraigned Wednesday in federal court in Brooklyn.
He is charged with one count of conspiracy to commit wire fraud and two counts of wire fraud in connection with an alleged ransomware decryption scheme that prosecutors say ran from June 2018 to June 2023.
The U.S. Attorney's Office told BleepingComputer that Pinhasi surrendered Wednesday, pleaded not guilty, and was released on a $2 million bond.
According to
the indictment
, Pinhasi owned and operated MonsterCloud LLC, a Florida-based ransomware remediation company that advertised tools and decryption techniques for recovering encrypted data without paying cybercriminals.
Prosecutors allege that Pinhasi and his co-conspirators had no such proprietary decryption technology and instead contacted ransomware operators, paid them for decryption keys, and then used those keys to restore customers' files.
The indictment acknowledges that some MonsterCloud contracts disclosed that the company might communicate with or pay cybercriminals. However, those contracts allegedly stated that MonsterCloud would contact attackers only if it could not decrypt a customer's files by other means.
Prosecutors claim that dealing with cybercriminals was usually MonsterCloud’s first step in obtaining decryption keys and recovering files.
"As alleged in the indictment, by falsely claiming to decrypt ransomware without paying off the ransomers, the defendant re-victimized his clients while extracting a hefty profit for himself,”
U.S. Attorney Joseph Nocella Jr. said
.
"Our Office will vigorously prosecute ransomware attackers who prey on Americans from across the world and those who cynically profit from their criminal activity."
MonsterCloud allegedly charged customers far more than the ransoms it paid.
In one ransomware recovery incident cited in the indictment, Pinhasi allegedly paid a ransomware gang about $8,200 and charged the victim approximately $150,000. In another, prosecutors say he paid approximately $236,000 and charged the customer about $380,000.
The indictment also alleges that MonsterCloud used decrypted sample files as "recovery proofs" to convince victims it could restore their data, even though those decrypted samples came from the ransomware operations.
Over the course of the alleged scheme, prosecutors say Pinhasi and his co-conspirators facilitated more than $8 million in ransom payments while charging hundreds of companies in the United States and Canada more than $19 million for recovery and remediation services.
If convicted, Pinhasi faces up to 20 years in prison.
BleepingComputer contacted Pinhasi's attorneys, Christopher Clark and Rodney Villazor, for comment on the allegations and will update the story if we receive a response.
Similar concerns raised in 2019
A
2019 ProPublica investigation
reported similar concerns about MonsterCloud, including that the company sometimes paid ransomware operators while claiming to offer a solution other than paying the attackers.
As part of that investigation, security researcher Fabian Wosar told ProPublica that he and another researcher created their own ransomware and approached several recovery companies while posing as victims.
The researchers provided the recovery firms with ransom notes containing email addresses they controlled for the fake ransomware gang. According to Wosar, those attacker-controlled accounts soon received anonymous messages offering to pay the ransom.
"Soon, the email accounts that he'd set up for the imaginary attacker began receiving emails from anonymous addresses offering to pay the ransom," ProPublica reported, citing Wosar. "He traced the requests to the data recovery firms, including MonsterCloud and Proven Data."
ProPublica reported that MonsterCloud had claimed it could recover the encrypted files without telling the supposed victim that it planned to pay the attacker.
Pinhasi disputed that MonsterCloud had promised in advance it could decrypt the files and denied misleading customers.
He also told ProPublica that MonsterCloud's recovery methods varied by case and declined to disclose them, describing the techniques as a "trade secret."
Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.
Timed out getting readerview for https://heatherburns.tech/2026/10/01/im-in-love-with-a-german-film-star/
Gurman Strikes Again: ‘Apple’s Smart Home Push Includes Doorbell, Lock, Thermostat Codeveloped With LG’
Daring Fireball
www.bloomberg.com
2026-10-07 18:02:43
The incomparable Mark Gurman, reporting for Bloomberg:
Apple Inc.’s upcoming push into smart home devices will include a
doorbell, thermostat and other accessories developed through an
unusual partnership with LG Electronics Inc.
The products will be part of an ecosystem of devices that work
wi...
We've detected unusual activity from your computer network
To continue, please click the box below to let us know you're not a robot.
Why did this happen?
Please make sure your browser supports JavaScript and cookies and that you are not
blocking them from loading.
For more information you can review our
Terms of Service
and
Cookie Policy
.
Need Help?
For inquiries related to this message please
contact
our support team
and provide the reference ID below.
FBI: Ongoing FortiBleed attacks lock out FortiGate VPN admins
Bleeping Computer
www.bleepingcomputer.com
2026-10-07 17:28:54
The FBI is warning that FortiBleed attacks are still ongoing, targeting exposed Fortinet FortiGate firewalls and SSL VPN gateways and locking out legitimate administrators. [...]...
The FBI is warning that FortiBleed attacks are still ongoing, targeting exposed Fortinet FortiGate firewalls and SSL VPN gateways and locking out legitimate administrators.
Hackers gain access to exposed endpoints by using previously leaked credentials, or logins obtained from infostealer logs, credential stuffing, and password spraying attacks.
They then extract additional authentication data from compromised devices and use a distributed GPU cluster running Hashcat and Hashtopolis to crack offline the stolen password hashes.
According to the FBI, " the FortiBleed attack chain has been observed as an initial entry point for ransomware affiliates." Some groups benefiting from this are INC/Lynx ransomware and Payload ransomware.
The leak that keeps on giving
FortiBleed is a massive
Fortinet credentials leak
discovered in June, when attackers inadvertently exposed a server containing usernames and plaintext passwords associated with 73,932 firewall URLs across 194 countries.
The data revealed a large-scale credential-harvesting operation, although it was unclear at the time what method was used to obtain the configuration data.
In July, SOCRadar linked FortiBleed to the
INC and Lynx ransomware
operations after getting access to both groups’ negotiation panels on a server used in the campaign.
The FBI says that in some incidents, the threat actor creates administrator accounts and uses their privileges to delete existing admin accounts or change their passwords, denying victims access to their devices.
The attacker then establishes persistence and tries to move laterally in the environment.
Details about the operation became known after the attacker accidentally exposed their backend server, revealing a directory with tooling and datasets.
This showed the use of automated scripts to scan exposed FortiGate SSL VPN portals, a distributed GPU password-cracking setup, and scripts to validate credentials, filter out honeypots, identify organizations, and prioritize targets by revenue and network structure.
Additionally, the exposure revealed working VPN configurations and target lists, indicating that the operator was packaging compromised access for sale.
The FBI warned that remediation may require more than patching and resetting Fortinet passwords, suggesting restricting external access, terminating all active VPN sessions, enforcing MFA, and reviewing logs for unauthorized changes and suspicious activity.
They also recommend enforcing PBKDF2 for administrator password storage, which is much stronger than legacy SHA-256 hashes that attackers can practically crack offline.
Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.
Margaret Hamilton, a profoundly influential computer scientist best known for leading the software engineering team at MIT’s Instrumentation Lab during NASA’s Apollo program, died on Sep. 30. She was 90.
A computing pioneer who authored over 130 publications, Hamilton helped to establish software engineering as a dedicated discipline. She worked at MIT from 1959 until the mid-1970s, after which she became a successful computing entrepreneur and CEO.
Her life’s work was recognized with many awards and honors,
including the 2016 Presidential Medal of Freedom
from President Barack Obama, whose citation noted: “Hamilton defined new forms of software engineering and helped launch an industry that would forever change human history. Her software architecture led to giant leaps for humankind, writing the code that helped America set foot on the moon.”
“To say Margaret Hamilton was a pioneer — to say she was ahead of her time — would be a dramatic understatement. She was a software engineer at a time when that field was in its infancy, and she not only developed advanced code herself but also led a team in using that nascent technology to develop one of the most complex systems humanity had ever achieved,” says Olivier de Weck, the Apollo Program Professor and interim head of the MIT Department Aeronautics and Astronautics.
“The Apollo program still stands as one of our greatest testaments to the power of collaboration, ingenuity, and engineering, and Margaret Hamilton was a fundamental contributor to that program’s success. And for her that was just the beginning! She went on to become an entrepreneur and remained on the cutting edge of systems software throughout an extraordinary career, while acting as an advocate, mentor, and inspiration to millions.”
Early life and projects at MIT
Born in Paoli, Indiana, in 1936, Hamilton began studying mathematics at the University of Michigan in 1955 before transferring to Earlham College, where she earned a BA in mathematics with a minor in philosophy in 1958. She moved to Boston, Massachusetts, in 1959 with her husband while he pursued a law degree.
Hamilton soon found a temporary position in the meteorology department at MIT, working with professor of meteorology Edward N. Lorenz SM ’43, ScD ’48 on weather prediction software. This was her first entry point into computer programming, and her work would go on to inform Lorenz’s future publications on chaos theory.
From there, Hamilton took a role as a programmer at MIT Lincoln Laboratory in 1961, working on the
Semi-Automatic Ground Environment
(SAGE) project, the United States’ first air defense system. Hamilton wrote software for the prototype AN/FSQ-7 computer (XD-1), used by the U.S. Air Force to search for potentially unfriendly aircraft. During this time, Hamilton began to take an interest in software reliability — a new and largely unexplored concept at the time.
In 1965 Hamilton was preparing to pursue graduate studies when her husband saw an advertisement in the newspaper: The Instrumentation Lab at MIT was seeking people to develop software to “send man to the moon.” The lab had won the contract from NASA to build the onboard flight software for the Apollo program. Intrigued by the challenge, Hamilton applied, and was hired as the first programmer for the Apollo project at MIT, as well as the first female programmer in the project.
Hamilton worked first on the software for the uncrewed Apollo missions and then was promoted into leading the team developing the on-board flight software for the crewed missions. By 1968 she was assistant director in charge of the Command and Service Module team, and more than 400 people were working on Apollo’s software.
“From my own perspective, the software experience itself (designing it, developing it, evolving it, watching it perform and learning from it for future systems) was at least as exciting as the events surrounding the mission,” Hamilton
told MIT News in 2009
. “There was no second chance. We knew that. We took our work seriously, many of us beginning this journey while still in our 20s. Coming up with solutions and new ideas was an adventure. Dedication and commitment were a given. Mutual respect was across the board. Because software was a mystery, a black box, upper management gave us total freedom and trust. We had to find a way, and we did. Looking back, we were the luckiest people in the world; there was no choice but to be pioneers.”
“Defensive” programming and priority-driven software take humans to the moon
Hamilton discovered a talent for leadership, as well as a keen instinct for problem-solving and critical thinking that would prove to save Project Apollo several times over.
There was the incident that became known as “the Lauren error”: One day, her daughter Lauren, then four years old, was playing with the command module simulator at the Instrumentation Lab when she somehow activated a pre-launch program, called P01, while the simulator was in midflight — which crashed the simulator altogether. Hamilton created a program add-on in the technical documentation warning users not to launch P01 during flight.
She also proposed a software fix to prevent the error happening during a real mission, but she was overruled on the grounds that the highly trained astronauts were never going to make that same mistake. Yet during the Apollo 8 mission in 1968, that’s exactly what happened: Jim Lovell inadvertently launched P01 during the flight, causing the on-flight navigational data to vanish. Hamilton and her team were called in to solve the error, and after that her proposed changes were integrated into the program. This was an early example of “defensive” programming, the practice of building software that could anticipate or fix errors on its own.
Hamilton’s most famous contribution to the Apollo program came during the pivotal Apollo 11 mission to land on the moon in July of 1969. Moments before the Eagle module was set to land on the lunar surface, the onboard computer raised the alarm. It had detected a 1202 error: The computer was overloaded due to a fault in a hardware switch, and it was possible the system would not be able to handle the complex landing procedure.
But Hamilton and her team had engineered priority-driven software, able to shut down unnecessary background tasks in order to prioritize mission-critical tasks. Houston trusted Hamilton’s software and allowed the mission to proceed, and two men walked on the moon for the first time.
Later career and legacy
Hamilton continued to work at the Instrumentation Lab into the 1970s. As the Apollo program wound down, the Instrumentation Lab spun out of MIT to become the independent
Draper Laboratory
.
“Margaret left an indelible mark on Draper, and we will be forever grateful,” says Jerry M. Wohletz SM ’97, PhD ’00, president and CEO at Draper. “Through her leadership and contributions to the development of the onboard flight software for NASA’s Apollo Guidance Computer, she helped ensure that Apollo astronauts landed safely on the lunar surface and safely returned home. We will honor her legacy at Draper forever.”
Hamilton went on to create her first software company, Higher Order Software, in 1976. The firm was based on Hamilton’s software engineering approach of error prevention and fault tolerance.
A decade later, Hamilton founded another software company, Hamilton Technologies, “to provide products and services to modernize the planning, system engineering and software development process in order to maximize reliability, lower cost and accelerate time to market.”
Hamilton Technologies’ flagship product is the Universal Systems Language (USL), a systems modeling language and methodology for engineering complex software systems that prioritize error prevention and defensive programming.
Throughout her career, Hamilton worked to achieve recognition for software engineering as a dedicated discipline.
“I fought to bring the software legitimacy so that it — and those building it — would be given its due respect, and thus I began to use the term ‘software engineering’ to distinguish it from hardware and other kinds of engineering, yet treat each type of engineering as part of the overall systems engineering process,” Hamilton
told
El Pais
in 2018
. “When I first started using this phrase, it was considered to be quite amusing. It was an ongoing joke for a long time. They liked to kid me about my radical ideas. Software eventually and necessarily gained the same respect as any other discipline.”
Among her many awards and honors, Hamilton was recognized with the NASA Exceptional Space Act Award for scientific and technical contributions in 2003; the Computer History Museum Fellow Award in 2017; the Intrepid Lifetime Achievement Award in 2019; and induction into the National Aviation Hall of Fame in 2022.
Later in life, Hamilton become an icon for women in science and technology, especially after a
now-famous photo, showing her next to a printout of her MIT team’s Apollo code, began circulating online. In 2015, the Apollo software she helped to develop was
added in its entirety
to the code-sharing site GitHub. And in 2017, she
became an official Lego Minifigure
after a set originally designed by MIT science communicator Maia Weinstock, honoring her and several other women of NASA history, became available worldwide.
“Margaret Hamilton has been an inspiration to generations of computer scientists and engineers. Hers was a career dedicated to preventing errors and what she called ‘handling the unknown,’” says de Weck. “She personified leadership by example, and established a practice of software engineering based on problem-solving and systems engineering that we all benefit from.”
Hamilton is survived by her daughter, Lauren Hamilton; her son-in-law, Richard Selesnick; two grandsons; and four great grandchildren. A memorial service will take place in the spring in Cambridge, Massachusetts.
Introducing Claude Haiku 5.5
Simon Willison
simonwillison.net
2026-10-07 16:56:21
As previously promised, here's Anthropic's new fast, low cost model: Introducing Claude Haiku 5.5.
The previous Haiku, 4.5, was very much showing its age. It came out almost a year ago, and was priced at $1/million input and $5/million output - relatively expensive even back then, and a full 10x the...
The previous Haiku, 4.5, was very much showing its age. It came out
almost a year ago
, and was priced at $1/million input and $5/million output—relatively expensive even back then, and a full 10x the price of OpenAI’s
GPT-6 Luna
, released last month.
The new Haiku exactly matches the price of GPT-6 Luna—$0.10/$0.50—up to 100,000 tokens. Beyond 100,000 tokens the price increases 5x to $0.50/$2.50. Luna itself has a price increase at 272,000 tokens but only to $0.20/$0.75.
Haiku 5.5 also uses a new, less generous tokenizer. My
Claude Token Counter
tool shows that the same long prompt uses around 1.25x as many tokens with Haiku 5.5 compared to Haiku 4.5, so there’s a hidden price increase there.
If your workloads fit in 100,000 tokens, Haiku is the same price as Luna and reports higher benchmark scores. Above 100,000 tokens, Luna looks like a much better deal.
The
most recent release of llm-anthropic
finally fixed it so I don’t need to ship a new version of that plugin for every new model. I tested the new model like this:
llm install -U llm-anthropic
llm anthropic refresh
llm -m claude-haiku-5.5 "Generate an SVG of a pelican riding a bicycle" -o thinking_effort low
This
max
effort pelican took 5 minutes 9 seconds to generate, but still only cost me
3.3826 cents
:
For comparison, here’s the pelican I got a year ago from Haiku 4.5. It
sucked
at drawing pelicans:
And a generous API credit scheme for subscribers
In addition to Haiku 5.5, Anthropic announced today that they are halving the price of cache reads for Sonnet 5.5. They’ve also added API credits to subscription plans:
Second, this week, we’ll roll out
a new monthly API credit to all Max and Team subscribers for use on the Claude Platform
. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users.
Claiming this is pleasantly easy: navigate to
Settings -> Billing
and select the API organization that should benefit from the credits every month:
The API credits exactly match the cost of the subscription itself. This is really generous—it makes it much easier for subscribers to use the API. Anthropic also let you disable auto-reload for the API, with the consequence that “API requests will stop when your balance runs out”—exactly what you want if you’re planning to burn through those API credits without
risk of a nasty billing surprise
.
OpenAI still allow you to use your Codex subscription for personal API use, which works out as a better deal for heavy API users. This new credit scheme goes at least some way to overcoming that difference.
Hackers hijack Google domains after breaching ccTLD registries
Bleeping Computer
www.bleepingcomputer.com
2026-10-07 16:50:13
Hackers obtained unauthorized HTTPS certificates for several Google domains and hijacked domains in the country-code top-level domains (ccTLDs) for Ghana, American Samoa, and Sierra Leone after compromising third-party operators and modifying authoritative DNS records. [...]...
Hackers obtained unauthorized HTTPS certificates for several Google domains and hijacked domains in the country-code top-level domains (ccTLDs) for Ghana, American Samoa, and Sierra Leone after compromising third-party operators and modifying authoritative DNS records.
Google underlines that the attacks affected domains of other organizations in the .GH, .SL, and .AS ccTLDs but "did not involve a compromise of Google’s systems."
By gaining access to the domain name system (DNS) records, a threat actor can request an HTTPS certificate from a Certificate Authority (CA) for a domain they don't own.
CAs issue certificates after verifying ownership of the domain, a process that typically requires the requester to create a TXT record with a random value the CA provides.
Modifying the authoritative DNS records allowed the threat actor to point .GH, .SL, and .AS domains to infrastructure they controlled while obtaining valid TLS certificates for those domains.
This let the attacker impersonate legitimate brands and serve visitors arbitrary content from the affected domains.
Google immediately blocked the unauthorized certificates for its properties in Chrome through CRLSets and worked with the issuing authorities to revoke them, extending protection to other clients.
The company said that its systems were not affected by the incident in any way, and that it has no reason to believe that the issuing CAs acted improperly.
After examining Certificate Transparency (CT) logs, the tech giant blocked additional certificates that appeared connected to the attacks and notified affected organizations where possible.
“Following our initial mitigation, Certificate Transparency (CT) log data revealed additional organizations, including several leading global brands and widely used online services, believed to have been impacted by the same attacks,”
Google explained
.
“To ensure users of those sites were kept safe as soon as possible, we proactively blocked these certificates in Chrome.”
CRLSets is a Chrome “
emergency mechanism
” designed to allow quick blocking of selected revoked or untrusted HTTPS certificates. Chrome users do not need to take any action to protect themselves from this incident.
However, Google warns that it may not have identified every affected domain, so its current blocking lists might not cover all potential threats.
The tech company also reminded users that CRLSets only covers Chrome users, meaning that users of other browsers might not be protected.
“Due to the complexity of DNS hijacks, we cannot guarantee that our analysis identified every affected domain, nor do Chrome interventions reliably protect non-Chrome users,” Google says.
Google urges domain owners to:
Monitor CT logs across their entire domain portfolio, including parked domains.
Publish restrictive Certification Authority Authorization (CAA) records as needed to limit issuance to authorized ACME accounts and validation methods.
Certification Authority Authorization (CAA) DNS records cannot stop certificate issuance during an active DNS hijack, but they prevent obtaining additional certificates using cached domain validation after legitimate DNS control is restored, Google notes.
The announcement did not identify the attackers or the quantity of certificates confirmed to have been hijacked.
Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.
And every one that heareth these sayings of mine, and doeth them not, shall
be likened unto a foolish man, which built his house upon the sand: And the
rain descended, and the floods came, and the winds blew, and beat upon that
house; and it fell: and great was the fall of it.
This was originally going to be called something like, “How slopcoded is your
programming language?” But as I gathered the data, I found something else –
something worse.
Almost every popular language is now substantially developed using LLMs.
Particularly notable are C# and Ruby (both approximately one third of recent
commits) and Julia (over half!). The languages that stand out in retaining
their humanity are Chicken Scheme, Perl, Lua, Clojure, and – perhaps
surprisingly for a project so closely associated with Oracle, a company who
have gone all-in on hyperscale data centres and “AI” in everything – Java.
However, all of these rely on a C compiler (indirectly via the JVM in the case
of Clojure), and both of the main C compilers now receive significant amounts
of LLM commits. I’m not so surprised by LLVM (15.9%), but I am disappointed by
GCC (3.1% and apparently rising). Outsourcing your thinking to megacorporations
and commercial tools seems out of step with the GNU ethos of freely available
source that anyone can modify: code that is generated by machines quickly
becomes code that is only parsed and modified by machines, and a subscription
to OpenAI or Anthropic becomes the (financial, environmental, and geopolitical)
price of entry.
In 1984, in his acceptance speech for the ACM Turing Award,
Ken Thompson
described a method
by which a compiler could be subverted so that it would
always insert a backdoor into the UNIX
login
command, and, furthermore, when
used to compile itself would insert a similar subversion into new versions of
the compiler. Once this has been achieved, no one has access to an
uncompromised compiler.
The moral is obvious. You can’t trust code that you did not totally create
yourself. (Especially code from companies that employ people like me.) No
amount of source-level verification or scrutiny will protect you from using
untrusted code. In demonstrating the possibility of this kind of attack, I
picked on the C compiler. I could have picked on any program-handling program
such as an assembler, a loader, or even hardware microcode. As the level of
program gets lower, these bugs will be harder and harder to detect. A well
installed microcode bug will be almost impossible to detect.
Does it matter how slopcoded your language is or isn’t, when everything below
it is slop?
Methodology
I cloned the source repositories for implementations of programming languages
that ranked highly on
TIOBE
and
LangPop
. For each, I performed a shallow
clone going back to the start of July:
git clone --shallow-since=2026-07-01 ${url}
I then wrote a short script that takes a three-month period and counts the
number of commits that appear to be LLM-assisted. This is determined by:
The phrase “AI disclosure” or “LLM disclosure”
An
Assisted-by:
field
a
Co-authored-by:
field with the email address of a known bot
This isn’t perfect – I saw one commit that had a disclosure field followed
by “none” and a link to a manifesto opposing LLM-assisted coding, but I don’t
think those edge cases are significant.
This yielded the following table, which I have sorted and annotated:
Language(s)
Repository
LLM-assisted
Total
Percent
C, C++, …
gcc
86
2778
3.1%
C, C++, …
llvm-project
2229
13981
15.9%
C#, Visual Basic
roslyn
350
1114
31.4%
Clojure
clojure
0
98
0.0%
Elixir
elixir
73
364
20.0%
Go
go
12
1092
1.1%
Haskell
ghc
42
287
14.6%
JavaScript (NodeJS)
node
395
1605
24.6%
Java
jdk
0
1194
0.0%
Julia
julia
489
918
53.3%
Kotlin
kotlin
398
4612
8.6%
Lua
lua
0
12
0.0%
PHP
php-src
2
2124
0.1%
Perl
perl5
0
851
0.0%
Python
cpython
219
1339
16.4%
Ruby
ruby
881
2717
32.4%
Rust
rust
41
10020
0.4%
Scala
scala3
18
538
3.3%
Scheme (Chicken)
chicken-core
0
208
0.0%
Swift
swift
135
4831
2.8%
TypeScript
TypeScript
70
409
17.1%
Repositories don’t map one-to-one to programming languages: some implementations cover
multiple languages, and some languages have multiple implementations.
Despite what Watson said, Rosalind Franklin understood structure of DNA first
This essay, co-authored by a historian of science and an X-ray crystallographer, sheds new light on Rosalind Franklin’s Photograph 51. We refute the infamous claim that, unlike James Watson, Franklin failed to see the picture’s potential significance for interpreting the helical structure of DNA. Rather, Franklin decided to take Photograph 51 precisely because she knew that key parameters of DNA’s B form helix could be calculated from the resulting image. We show that she had in fact already made those calculations — on her earlier Photograph 49 — and she reused the same DNA sample for Photograph 51 to create a better-centered but otherwise identical diffraction image that would be suitable for publication. Thus Photograph 51 was not the result of an experiment in need of analysis, but was refined documentation for calculations that she had already performed on the earlier photograph. We argue that colleagues and later commentators did not merely overlook Franklin’s original reason for creating the strikingly clear and informative Photograph 51, they rhetorically erased her skill and judgment by describing it as though nature
spoke for itself
through the image — and to Watson, but not to Franklin.
Similar content being viewed by others
Introduction
Rosalind Franklin’s Photograph 51 (see Fig.
1
) became famous because James Watson wrote that it spoke powerfully about the secret of life. “The instant I saw the picture,” he narrated in his controversial memoir
The Double Helix
, “my mouth fell open and my pulse began to race” (1968b, p. 167).
Footnote
1
It was early February 1953 and Franklin’s estranged colleague Maurice Wilkins had shown Watson this X-ray diffraction image of DNA in a state known as the B form.
Footnote
2
As Watson described it, he immediately recognized that the pattern on Photograph 51 “could arise only from a helical structure […] mere inspection […] gave several of the vital helical parameters” (1968b, pp. 168–169). Spurred on by seeing the image and learning of some measurements Franklin had taken of the B form diffraction pattern, Watson and Francis Crick began a fresh effort to build helical models of DNA and within a few weeks managed to produce their double helix.
Footnote
3
While Watson’s book made Franklin’s Photograph 51 famous, his story gave the impression that she had not been able to recognize its significance the way he did. Consequently, many scientists, journalists, and biographers have sought to explain Franklin’s apparent neglect of a picture that, in retrospect, was so clearly suggestive of a helical molecule (Table
1
).
Table 1 Key dates in the history of Photographs 49 and 51
By consensus, what happened when Photograph 51 was taken in early May 1952 is that Franklin put the picture aside and did nothing with it, instead devoting months to studying the crystalline form of DNA known as the A form. In late February 1953, her notes show, she began to use Photograph 51 to understand the B structure. By then it was too late: she was still working when Watson and Crick solved the puzzle a few days later.
Several overlapping reasons have been given for Franklin’s apparent initial disregard and subsequent delay: she was only interested in the A form of DNA; she had only captured a B form pattern in Photograph 51 by accident, and as such considered it a distraction; she was resistant or even hostile to the suggestion that DNA was a helix. For Horace Judson, writing what remains the best-known journalistic account of Watson and Crick’s discovery, this added up to something like willful deafness on Franklin’s part. “The pattern shouted helix,” Judson lamented, but for nearly ten months she “turned her back on her own discovery of the B structure of DNA and her own best evidence that [it] was helical” (Judson 1979, p. 135).
We write here to contend that Photograph 51 shouted “helix” so clearly because Rosalind Franklin intended it to do so. She had already photographed the very same sample of DNA used for Photograph 51, and analyzing the prior diffraction pattern, Photograph 49 (Fig.
2
), led her immediately to a number of significant conclusions about the dimensions of a helical DNA molecule. She then created Photograph 51 for the purpose of reproducing that valuable diffraction pattern in a better-composed image that would be suitable for publication. Whereas Watson and others have portrayed Photograph 51’s combination of visual excellence and theoretical significance as a kind of happy accident that Franklin failed to exploit, it was the desire to merge those qualities in a single exposure that had in fact driven her to take Photograph 51. By recapturing Franklin’s original motivation, we are in turn able to explain the subsequent actions which have been misinterpreted by her supporters and critics alike.
We will present our revisionist account first and then use it to reassess portrayals of Franklin. We begin with a short but detailed explanation of how photographs 49 and 51 were created around the beginning of May 1952 and how Photograph 51 came to serve its purpose as an image intended for circulation. We analyze the later emergence of claims that Franklin had ignored or misunderstood Photograph 51, and argue that these criticisms echoed longstanding tropes about the role of women in science and of photographs as scientific evidence.
Fig. 1
Rosalind Franklin’s personal copy of Photograph 51. On this glossy photographic print Franklin noted “NaDNA ‘Structure B’” upon the reverse (History of Molecular Biology Collection, Box 10, Folder 15, Science History Institute. Philadelphia)
The Significance of Photograph 49
The key to understanding why Franklin took Photograph 51 lies in her notes about a previous picture. It was the forty-ninth in a series of seventy-eight diffraction images she and her graduate student Raymond Gosling captured with a Philips micro-camera (see Fig.
2
) at King’s College London in 1951 and 1952.
Footnote
4
Months before taking Photograph 49, the pair had established the existence of two distinct structural forms of DNA, and they had begun to manipulate samples intentionally to create them. The A form possessed an almost crystalline structure and gave diffraction patterns with many spots, potentially offering a great deal of information for structural analysis (by calculation of Patterson maps).
Footnote
5
DNA could be converted to the B form at high relative humidity, yielding a simpler diffraction pattern with too few spots for similar (Patterson) structure calculations. Franklin had discussed her discovery of the distinct A and B forms in a colloquium at King’s on 21 November 1951 and documented them in an annual report of 7 February 1952, but her results were not otherwise published.
Footnote
6
Franklin’s annual report of February 1952 also laid out a research plan that she appears to have followed closely until she moved from King’s College to Birkbeck College in the spring of 1953. Her goals were twofold: (1) to determine what caused a given DNA sample to exhibit the A or the B form, and afterward, (2) to characterize both forms initially through Patterson analysis of the more complex but data-rich A pattern.
Footnote
7
Photograph 49 was created two months later in the course of a lengthy series of experiments aimed at completing the first of these objectives. Franklin’s stated hope was that these experiments would also reveal the best techniques for capturing sharp DNA diffraction patterns. As hoped, it was during this period of systematically manipulating the humidity levels at which samples were stored and photographed that she and Gosling produced what they retrospectively considered their best ever A form pattern (Photograph 42 of early February 1952) as well as their best B form photographs in May.
Completed on the morning of 2nd May 1952, Photograph 49 showed the result of aiming an extremely fine X-ray beam for three days and nights at a fiber that had been pulled into a fine thread from a drop of gel-like concentrated DNA solution.
Footnote
8
The beam emerged from a collimator at the front of the camera and met the sample fiber, which was mounted and held in place across the orifice by two drops of glue, immediately inside the camera, just 15 mm in front of the X-ray film that recorded the diffraction pattern (see Fig.
2
). Because space within the camera was so tight, they employed an unusual technique to prevent the main (undiffracted) X-ray beam from overexposing the center of the film. Instead of blocking the main beam with a small circular piece of heavy metal such as lead between the specimen and the film, as was conventional in larger cameras, they punched a hole through the centers of all the stacked films and allowed the undiffracted beam to pass right through the films and out of the back of the camera, through a small fluorescent screen that facilitated camera alignment (the holes can be seen in the film and camera body shown in Fig.
3
).
Footnote
9
Franklin and Gosling controlled the relative humidity of the DNA sample throughout the days-long experiment by passing hydrogen gas that had been bubbled through an aqueous salt solution of defined composition and concentration into the body of the camera.
Footnote
10
As with the trials that preceded Photograph 49, they would learn the result of the experiment by removing and developing a stack of two or three small, hand-cut pieces of film from where they were held inside the camera body during the exposure (see Fig.
3
). Each of the duplicate (or triplicate) films would show an even smaller X-ray diffraction pattern about 2.5 centimeters in diameter. On this occasion, they had used two pieces of film marked in Franklin’s handwriting as 49A and 49B (see Fig.
4
).
Footnote
11
These Photograph 49 films showed an unprecedentedly sharp version of the diffraction pattern produced by DNA’s B form (Fig.
2
). Franklin and Gosling accomplished this by using a sample that had become irreversibly locked into the high-humidity B structure, meaning that it could be mounted taut in the camera and photographed at the lower 75% humidity level that ordinarily produced an A form pattern.
Footnote
12
The resulting fine definition made it possible to use the photograph for more precise analysis than any previous B image had allowed, and Franklin therefore gave it her prompt attention. This immediacy is evident from a separate series of notes that were headed “(49) 49B Rough measurements on projection,” dated 2 May 1952 — the very same day she and Gosling had developed the films.
Footnote
13
These notebook pages began with measurements of the new diffraction pattern, which the pair enlarged by projecting it onto a piece of white cardboard, and Franklin’s notes progressed in short order to making general conclusions about the B structure of DNA. She noted that the molecule showed repeats occurring every 34 Å, meaning it was about 25% longer per repeating unit than the drier, more crystalline A form she had already begun to analyze.
From the day Photograph 49 was taken, Franklin interpreted the pattern as indicative of a helix.
Footnote
14
Other evidence indicates that she already had a helical interpretation of the B form structure in mind, and now she calculated how many layers of purine and pyrimidine bases there must be “per turn of helix (if there is a helix).”
Footnote
15
If the bases were spaced 3.4 Å apart, as she and other researchers believed, then her calculations suggested that there would be exactly ten nucleotide-base layers in each lengthwise repeat of the molecule. In her notes, she phrased the finding more generally: “there
is
an integral number (or single fractional number) of residues per [34 Å] turn.”
Footnote
16
Later, in 1954, Watson and Crick acknowledged that they had built their double helix as a model of the B form with its 34 Å axial period containing ten base layers per turn, which, Watson and Crick crucially realized, were complementary base
pairs
rather than individual bases. It is not widely appreciated that these parameters were originally established in Franklin’s analysis of Photograph 49.
Footnote
17
Fig. 2
Photograph 49. This glass-plate negative shows an identical diffraction pattern to that of Photograph 51, seen in Fig. 1, except that the pattern is cropped by poor alignment between the film holder and the X-ray beam; the upper 3.4 Å arc is almost entirely lost. Both exposures used the same DNA sample under the same conditions (KDBP 1/1/0868, King’s College London Archives). Courtesy King’s College London
Today Photograph 49 survives in its entirety only as a negative-image contact plate of film 49B held at the King’s College London archive (Fig.
2
). The photograph is remarkable for two distinct reasons. First, it vividly captures the
very same
diffraction pattern known so well today from Photograph 51. Secondly, however, a portion of the famous pattern has been cut off. Through our examination of the camera Franklin and Gosling used in 1952, we can explain how this must have happened. The photographic films were held in place by a flat metal bracket whose edges were bent around a backing plate that the films rested against (Fig.
3
). This holder has a slightly irregular circle of roughly 2.5 cm diameter cut out to allow the diffracted X-rays to hit the films, and the glass negative of 49B illustrates that this holder had become misaligned with the X-ray beam and the DNA sample. In consequence, the upper part of the pattern is cropped to such an extent that a large, arc-like region now familiar from Photograph 51 (Fig.
1
) does not appear in the image. This missing arc, like its symmetrical counterpart which appears at the bottom of Photograph 49, is caused by the 3.4 Å spacing between each layer of nucleotide bases in the molecule.
As we have seen, this inadvertent cropping did not prevent Franklin from using Photograph 49 in her analysis and calculations. She had captured enough of the pattern to be able to discern that one of the 3.4 Å arcs was located on the tenth meridional layer line of the pattern and to understand the implications for quantifying the layers of bases stacked in a single 34 Å turn of the helical molecule. Photograph 49 was a sharp and usable image but, because of the equipment malfunction, it was poorly composed and thus less than ideal for public presentation.
Fig. 3
X-ray micro-camera and film. This recent picture illustrates how X-ray films were prepared for the Philips micro-camera (KDBP 6/4/7, King’s College London Archive) that was used to take photographs 49 and 51. From left to right: the front of the camera body, removed and viewed from the inside; a piece of X-ray film backing paper showing how films were cut to the appropriate size; the film holder (above) and a piece of developed X-ray film (below) showing that the region of the film exposed to X-rays was determined by the cut-out in the holder; the camera body back, containing the rectangular platform against which one or more small pieces of film were held in place by the film holder. The film and backing paper are surviving artefacts from Wilkins’ 1953 research with Herbert Wilson, which included replications of Franklin and Gosling’s work using different sources of DNA (K/PP178/2/8, Wilkins Papers). Photograph by Alistair Sponsel, July 2026
Franklin therefore decided to retake the photograph. On the evening of 2 May 1952 – still the same day she first saw Photograph 49 – she began the notebook entry for Photograph 51: she would use the same specimen from 49, the same X-ray setup, and the same 75% relative humidity within the body of the camera.
Footnote
18
The only thing distinguishing this from being an exact repeat was, as she wrote, that Photograph 51 was taken “with [the] holder centered over [the] collimator so as to include both 3.4 [Angstrom] arcs.”
Footnote
19
Photograph 49 had been a research image, the result of an experiment. Photograph 51 would be the exact opposite: a refined image taken after the fact to document that experiment’s particularly successful outcome.
Footnote
20
This relationship between Photographs 49 and 51 emerges all the more clearly when Franklin’s notes about them are compared with earlier entries in her laboratory notebooks. She had by then established a note-taking routine well suited to the open-ended trials characteristic of Photographs 1–49, with two distinct parts for each entry. The first section, written before the photograph was taken, provided details of the experimental set-up, including the date and time when X-ray exposure commenced. Afterward, Franklin would record the date and time when she and Gosling ended the exposure and then complete a second section, for which she customarily left a few lines empty, giving a brief indication of how the experiment had turned out.
Initially, she wrote about Photograph 49 using her standard format. She first specified which DNA sample she was using, the relative humidity at which it was maintained, and the source of the X-rays. Three days later, she filled in the space below to report that this trial had produced a “V[ery] good ‘wet’ photo” (meaning a very good photo of the B form). However, she then turned to a fresh set of pages in the notebook and filled them with the analyses we discussed above. Headed with the very date the photograph had originally been developed, this passage of notes filled far more space in her laboratory notebooks than did the discussion of any single previous DNA diffraction pattern, indicating that she immediately found this version of the B pattern to be strikingly important.
The entry for Photograph 51 was unprecedented in a completely different way. This was the first instance in Franklin’s DNA laboratory notebooks where she was able to write down in advance what details the diffraction pattern would include (namely, both 3.4 Å arcs) and the first occasion when she specified that correcting the cropping of a previous photograph was the reason for taking a new one. As an understandable consequence, Franklin did not bother to fill the lines at the end of entry 51 where she would normally have described the result of the photograph. Indeed, the empty space stands out starkly on a pair of pages otherwise densely filled with the initial conditions
and
the results of the surrounding experiments. As discussed below, later commentators have misinterpreted this omission as the act of a person who had just conducted a potentially crucial experiment, but who didn’t appreciate its significance. Rather, as we have shown, Photograph 51 was not an experiment at all in the sense that trials 1 to 50 had been. It was, in fact, the product of Franklin’s decision to invest four days of X-ray time into a known outcome. The investment made sense precisely because she valued a high-quality photograph of the helical B form diffraction pattern.
Franklin and Gosling did not immediately publish Photograph 51. Later commentators, armed with the knowledge that it would be less than a year before Watson and Crick built their double helix as a physical model of the very B form structure Franklin and Gosling discovered and characterized, have criticized Franklin for directing much of their effort during the rest of 1952 to analyzing the A form. However, what critics describe as Franklin’s discovery — and subsequent apparent neglect — of evidence for the crucial B form structure of DNA looked different from her perspective and in the context of her broader research objectives. Franklin’s stated objectives, as we have seen, were to understand the relationship between and interconversion of the two forms of DNA. Photographs 49 and 51, and her analysis of the former, which revealed the key parameters of a helical structure, represented an important and successfully achieved milestone. It is only with hindsight that we view the B form structure as the only important goal, and
the
structure of DNA. DNA clearly had more than one structure and Franklin wanted to solve both of them.
Fig. 4
The handwriting on individual films of photographs 49 and 51 (KDBP 1/1/0867–868, King’s College London Archives) compared with entries in Franklin’s notebooks (FRKN 1/1, Franklin Papers; reproduced with the kind permission of the Trustees of the Franklin Archive)
By completing Photograph 51, Franklin now had good images in hand of both the A form and the B form. From here, she turned to the second major objective of her established research plan: detailed structural analysis beginning with the crystalline A form, whose pattern offered more data for analysis than the paracrystalline B form. The results of several months’ work led her to conclude that this crystalline structure was likely a double-chained molecule, and in February 1953 she turned back to Photographs 49 and 51 to assess whether the B structure in turn showed “evidence for [a] 2-chain […] helix.”
Footnote
21
That the B form photographs indicated a helical structure however, was not in doubt, as indicated by her notebooks. By the end of February, as she was leaving King’s College for a position at Birkbeck College, she and Gosling had drafted three papers, two of which would contain Photograph 51 when they were published later that year.
Footnote
22
Meanwhile, since Gosling was to remain at King’s to finish his PhD under Wilkins’ supervision, Franklin directed him to give Wilkins the diffraction photograph that Wilkins in turn showed Watson.
Footnote
23
For all the investment Franklin originally made to capture a publication-worthy version of the B pattern, when it came time to publish Photograph 51 she took yet another step to refine the image. In early March 1953, she turned over film 51C to be duplicated onto a durable glass slide negative that could, in turn, be used to produce positive photographic prints for submission.
Footnote
24
The slide, which survives and is numbered 867 in the biophysics unit’s indexing system, became the basis for published reproductions.
Footnote
25
This is evident because published images of Photograph 51 show the diffraction pattern centered within a circular border even more precisely than it had been when the pattern was originally developed on film. The perfected cropping was achieved by placing masking tape on the glass plate itself (see Fig.
5
). The figures of Photograph 51 in both of Franklin and Gosling’s 1953 publications (1953a and 1953b) that included it show the ultra-refined cropping from the glass plate.
Fig. 5
A composite image showing how Photograph 51 was further cropped for publication. Left: a hitherto unpublished print of Photograph 51 showing that the circular region of exposed film had been larger than necessary to succeed in capturing the full diffraction pattern (HMBC Box 10, Folder 15). Center: the glass plate negative of film 51C with red masking tape used to create a new, smaller circular margin centered precisely on the diffraction pattern, which left a penumbra that can still be seen faintly through the tape in the lower-right quadrant of the slide (KDBP 1/1/0867). Right: Photograph 51 as it appeared in publications, shown here in the 25 April 1953 issue of Nature (Franklin and Gosling, 1953a)
Rethinking Rosalind Franklin’s Reputation in Light of Photograph 49
The myth that Franklin had failed to see value in Photograph 51 arose as a consequence of Watson’s 1968 memoir
The Double Helix
. The book appeared fifteen years after the events in question, ten years after Franklin’s tragically early death from ovarian cancer, and six years after Watson, Crick, and Wilkins shared the Nobel Prize for their studies of DNA. The book caricatured her as both a machine-like researcher and as a woman with emotions she could not control. In Watson’s telling, her “years of careful, unemotional crystallographic training” meant that she was producing sharper diffraction photographs and consequently more detailed measurements than anyone else (Watson 1968b, p. 69). However, she would erupt into defensive outbursts at any attempt to help her interpret these proprietary data, making life an “emotional hell” for her colleague Maurice Wilkins (Watson 1968b, p. 167). Indeed, Wilkins supposedly showed Photograph 51 to Watson in a moment of solidarity when Watson had invoked Franklin’s “hot anger” by suggesting to her face that “she was incompetent in interpreting X-ray pictures” and needed to “learn some theory” (1968b, p. 166).
Watson’s harsh portrayal of Franklin inspired many reactions, one of which came from her friend and former Birkbeck colleague Aaron Klug. Klug had inherited many of Franklin’s King’s College research records, and he now sought to understand what she had known about DNA, and when. After finding her 1951–1953 laboratory notebooks, he studied them carefully and circulated annotated photocopies for discussion with Wilkins and others (including historian Robert Olby, who was then working closely with Crick while doing research for a planned book). This flurry of activity in 1968 revealed that the well-known published B form pattern had been called Photograph 51 in her notebook and that it had been taken at the beginning of May 1952.
Footnote
26
Wilkins reacted by declaring it a “real tragedy” that, in keeping Photograph 51 to herself for the rest of year, Franklin had allowed Watson and Crick to race ahead of the King’s group. “I looked at that B form picture,” Wilkins said in reference to receiving it from Gosling in January 1953, “and there it was, you can see the helix right there on the picture, but she refused point-blank to see it.”
Footnote
27
Fig. 6
Rosalind Franklin in a 1950 photograph by the crystallographer Vittorio Luzzati (History of Molecular Biology Collection, Box 10, Folder 14. Science History Institute. Philadelphia)
Sentiments like Watson’s and Wilkins’ had a profound impact on the historical treatment of Franklin’s work. The question became, how did Franklin fail to discover the double helix while she possessed the photograph that supposedly had spurred Watson and Crick’s success?
Footnote
28
For example, consider how the journalist-historian Horace Judson wrote about her in his celebrated 1979 book
The Eighth Day of Creation
, which was based on extensive interviews with all the participants except the late Franklin herself. Knowing with hindsight that Photograph 51 was the B picture she would later publish, he used it as a narrative device to portray a singular moment when Franklin decisively failed. “The pattern shouted helix,” Judson wrote, and it even “whisper[ed]” the other details of the B form structure. Although he mentioned the analysis she had performed on Photograph 49, his characterization was that Franklin merely “thought briefly and tentatively about No. 49” and “[t]here she stopped.” By underestimating the novelty and significance of these calculations and, more importantly, by misjudging why she had chosen to take Photograph 51, Judson interpreted the lack of notes about the retaken photograph as a sign that she had “turned her back on her own discovery of the B structure” (Judson
1979
, p. 135). We do not believe that Franklin ever turned her back on the B form. Deciding to take Photograph 51 was not a moment of failure, but a mark of her future commitment.
We have seen that Watson portrayed himself as immediately hearing the helix’s metaphorical shout from Photograph 51. He went on to insist throughout his life that Franklin had not been a sufficiently good scientist to see in Photograph 51 the truths that he had instantly recognized in the B form pattern. In a 2008 interview, for example, he said Franklin “clearly wasn’t interested in theory very much” and had neglected to pursue the B form even though it “was the perfect helical thing.”
Footnote
29
These phrasings portrayed Photograph 51’s simple, sharp B form diffraction pattern as a natural object rather than a human-produced research product — as though it had occurred spontaneously under Franklin’s watch rather than being generated and photographed on purpose. In Judson’s telling, the DNA molecule “whispered” and “shouted” directly to those who saw the picture, rhetorically erasing the role Franklin played in revealing nature’s secrets (1979, p. 135). We have seen that Wilkins went even further, speaking as though Franklin had been
concealing
nature’s secrets by possessing Photograph 51.
The conceit that nature could speak directly to Watson through Photograph 51 echoed language that dates back to the earliest use of cameras by scientists. Advocates of the new technology claimed that photography, by reproducing natural objects mechanically and therefore without human bias, supposedly allowed nature to speak for itself. As historians Lorraine Daston and Peter Galison have shown, this combination of rhetoric and technology stemmed from the 19th-century craze for achieving “objectivity” in scientific research. Although technique and even creativity were actually required to make nature (seemingly) imprint itself on a photographic plate, the objective scientific photographer’s goal was to render this work invisible (Daston and Galison
2007
, p. 133). One of Franklin’s contemporaries, the great crystallographer J. D. Bernal, alluded to this process shortly after her death. Having pointed out that Franklin captured “among the most beautiful X-ray photographs of any substance ever taken,” Bernal remarked on the “skill” that had allowed her to make those photos
appear
to be “effortless.”
Footnote
30
However, merely highlighting Franklin’s technical acumen risks damning her with faint praise for her contributions to the discovery of the double helix. It is one thing to possess the skill necessary to let a molecule seemingly
speak for itself
through a beautifully clear photograph, but it is another thing to be able to discern which specimen must be heard in order to solve an important scientific puzzle. To flip Watson’s deprecating remark on its head, this
does
require an interest in theory. It requires knowledge and judgment. In Franklin’s case, not only had she recognized the conceptual value of the B form diffraction pattern captured in Photograph 49, but she had also devised the original series of humidity experiments that revealed the B form’s existence and enabled her to capture it so clearly.
Franklin was an accomplished, creative, and self-directed researcher, but Watson’s
Double Helix
cast her in the role of a technical plodder, unimaginatively toiling over samples, laboratory equipment, measurements, and calculations.
Footnote
31
Later efforts to praise her as a skilled crystallographer (but by implication nothing more), often reinforce a long tradition of women’s scientific activities being treated as rote work.
Footnote
32
Recently, as Franklin’s public profile has risen considerably, her role in the hands-on work of crystallography has, in turn, been called into question. In what we consider to be an oversimplified description of a multi-day, multi-instrument collaborative undertaking, Gosling is now regularly credited as the sole individual who took Photograph 51.
Footnote
33
Evidence from throughout Franklin and Gosling’s notes illustrates her leadership of, and thorough involvement in, both the intellectual and technical aspects of this collaborative research. Her handwriting on the films of photographs 49 and 51 (see Fig.
4
) indicates that she was present when the films were loaded into the camera, but both Franklin and Gosling were likely aided by other women working alongside the King’s College biophysicists. Freda Ticehurst was the unit’s scientific photographer and manager of the dark room. She, in turn, had assistance from the likes of Lucille Heller, who volunteered in the biophysics unit in 1950–1951 and recalled “I remember helping Gosling set things up, and I think I helped sometimes with taking the X-rays, and I helped Freda Ticehurst, the lab photographer, develop some of the images.”
Footnote
34
Gosling himself recalled the assistance and collaboration he and Franklin received from one of the biophysics unit’s workshop technicians, Len Pitches, who modified and even built cameras and other apparatus to their specifications.
Footnote
35
Our case study of the relationship between photographs 49 and 51 provides only a glimpse at the full arc of Franklin’s activities as a DNA researcher.
Footnote
36
However, even within this short paper we have witnessed her pursuing a full spectrum of scientific activities: designing a highly consequential series of experiments, tapping the skills of her student and technical staff while working alongside them in the X-ray room, accumulating data and carrying out careful measurements, perceiving the value of her results, and also showing a persistent intention about how to make the result we focused on here — the B diffraction pattern — appear in the most arresting possible form when it was published. She created Photograph 51 to serve as evidence of a helical DNA structure, of her laboratory's technical virtuosity, and of her own scientific judgment.
Footnote
37
Data Availability
No datasets were generated or analysed during the current study.
Notes
In the main text, Watson did not specify exactly which of Franklin’s photographs he had been shown by her colleague Maurice Wilkins. However, Photograph 51 is the one he used to illustrate this episode in the first edition of
The Double Helix
(Watson
1968b
, p. 168).
In her notebooks and in an interim report dated 7 February 1952, Franklin used the terms “crystalline” and “wet” to describe the two structures that she and her PhD student Raymond Gosling had successfully characterized (FRKN 1/1 and FRKN 4/3 respectively in Papers of Rosalind Franklin, Churchill Archives Centre, Cambridge; hereafter, Franklin Papers). These structures were called A and B respectively in the first publication she drafted with Gosling (Franklin and Gosling
1953b
). As we discuss below, it was Franklin herself who had characterized these two forms as distinct structures. For ease of understanding, we will refer to the structures as A and B throughout the paper.
Many later commentators, acknowledging the significance of Photograph 51 to Watson and Crick’s discovery, have declared it to be among the most important photographs ever taken. (See, for example, assertions of Photograph 51’s historic significance in Walsh
2012
and Babaian
2024
). Its appearance in compendia of influential photographs such as
LIFE 100 Photographs: The Most Important Pictures of All Time and the Stories Behind Them
(Editors of LIFE Magazine
2021
) implies the same. Photograph 51 has also been featured on the British fifty-pence coin minted in Franklin’s honor and it was the namesake of a prizewinning play in London’s West End (Ziegler
2015
).
FRKN 1/1, Franklin Papers. The notebooks have been made available online at
https://wellcomecollection.org/works/uus54sbp
. Note that Franklin and Gosling also took photographs with other apparatus and numbered them in separate series.
Detailed discussions of the work mentioned here are available in Klug (
1968
), Olby (1994 [
1974
]), Elkin (
2003
). Patterson maps, calculated from the intensities and positions of the diffraction spots alone, can reveal prominent inter-atomic distances, such as between phosphate groups, from which structures may be deduced.
Franklin’s notes for her November 1952 colloquium are in FRKN 3/2, Franklin Papers. The annual report is Rosalind Franklin, “Interim Annual Report” dated 7 February 1952. FRKN 4/3, Franklin Papers.
She wrote, “It is proposed to attempt a quantitative interpretation of the [A form] fibre diagram (which shows a high degree of crystallinity in the DNA fibres) by means of Patterson functions. […] Before embarking on these calculations it seemed desirable to ascertain that the photographs used were [i.e., would be] the best which could be obtained.” She continued by describing the “preliminary results” of what were then ongoing efforts to pursue a “systematic search for the best conditions, especially with respect to relative humidity.”
The following paragraph is based on specific information about the Photograph 49 experiment from Franklin’s 1952 notebook (FRKN 1/1, Franklin Papers) and on general information about the experimental set-up that we gleaned from studying the Philips micro-camera itself and from Raymond Gosling’s 1954 PhD thesis (KDBP 6/4/7 and KDBP/5/1 respectively, King’s College London Archive).
In shedding light on this unusual technique of allowing the main part of the X-ray beam to pass through the entire camera apparatus, we hereby offer a detailed context for the concerns regarding Franklin’s radiation exposure that some of her colleagues later shared in interviews for Maddox’s biography (Maddox
2002
, p. 144).
Wilkins had originally suggested to Gosling the idea of filling the camera body with hydrogen gas to prevent unwanted scattering of the X-ray beam by the larger molecules in air. It was Franklin’s key contribution to control and modify the humidity of the hydrogen gas, and thus the DNA sample during the experiment, by bubbling the hydrogen through various salt solutions before it entered the camera.
Franklin’s notebook entry for Photograph 49 specifies “2 films” (FRKN 1/1, Franklin Papers). The surviving version of Photograph 49 comes from film 49B (see Fig.
2
). We infer that the other film would have been marked 49 A.
We disagree with suggestions by Maddox (
2002
) and Cobb (
2025
, p. 85) that Franklin did not intend or expect Photograph 51 to exhibit a B form pattern. “Sometimes during exposure,” Maddox wrote, “the [DNA] fibre would change from the crystalline A form to the paracrystalline B form. Once this happened so abruptly that the fibre fell off the holder. The photograph taken between 1 and 2 May […] was the clearest picture ever taken of the B form of DNA […] Rosalind put it aside to return […] to the puzzle of the A form” (Maddox
2002
, pp. 177–178). The circumstance Maddox described, of extremely hydrated DNA samples loosening in the specimen holder as a result of the structural shift from A to B, had indeed occurred with some of Franklin’s earlier experiments; for example, she had noted a “series” of trials in March and April 1952 “in which trial short-exposure films were generally good and [the] specimen subsequently went non-crystalline during long exposure” (note headed “March-April 1952” on the first inside page of Franklin’s 1952 notebook, FRKN 1/1, Franklin Papers). However, in the same note Franklin also reported the conclusion that specimens which began to give “non-crystalline” B form patterns at 75% relative humidity (which ordinarily produced the crystalline A form) were “never re-converted to crystalline.” It was just such a locked-in B form specimen, photographed at 75% RH, that she subsequently used to produce photographs 49 and 51. Indeed, Franklin and Gosling used Photograph 51 in one of their publications specifically to illustrate the phenomenon of “a fibre which had passed irreversibly to structure
B
” and which produced excellent diffraction images precisely because it was no longer susceptible to “buckling of the fibre in the wet state, and consequent deterioration of the quality of the photograph” (1953b, pp. 674–675). By tracing this particular DNA sample’s provenance back through Franklin’s notes, we can tell that it had been locked into the B form for several weeks. Her entries for photographs 49 and 51 indicate that both were taken of a “specimen […] which gave [a] good ‘wet’ photo” when previously used in the lower-humidity photograph T0 on 18 April (this had been the first of a separately numbered “T” series of pictures she and Gosling had taken using a custom-built tilting camera stand). In the entry for Photograph T0, in turn, Franklin said the sample was “previously Xtalline [crystalline, meaning the A form], now gives ‘wet’ diagram [i.e., B form]” (entries for photographs T0, 49, and 51, Franklin’s 1952 notebook, FRKN 1/1, Franklin Papers).
Pages headed “(49) 49B”, Franklin’s 1952 notebook (FRKN 1/1, Franklin Papers). This notebook entry is reproduced as Fig. 21 in Klug (
2004
).
We emphasize this point only to contradict claims (discussed below) that Franklin had not recognized Photograph 51 as evidence of a helical molecule. That such a diffraction pattern was suggestive of a helical molecule had been worked out by her colleague Alec Stokes at King’s (unpublished), and by Crick and others in Cambridge (published as Cochran, Crick, and Vand
1952
).
Pages headed “(49) 49B,” Franklin’s 1952 notebook (FRKN 1/1, Franklin Papers). Franklin had declared three months earlier, in her annual report, that her studies of the A form “suggest a helical structure (which must be very closely packed) containing probably 2, 3, or 4 co-axial nucleic acid chains per helical unit.” In briefer remarks about the B form in the same report, she mentioned the “helix in the wet state” while reporting the DNA molecule lengthened to an undetermined degree during the crystalline-to-wet (A to B) transition (“Interim Annual Report” dated 7 February 1952; FRKN 4/3, Franklin Papers).
Thanks to the work of William Astbury and Florence Bell in the late 1930s, it was considered likely that 3.4 Å represented the spacing between layers of the nitrogenous bases stacked perpendicular to the long axis of the molecule. See Astbury and Bell (
1938
). For discussion of their work, see Kersten Hall (
2014
).
We believe that other commentators, for example Cobb and Comfort (
2023
), are mistaken in implying that Wilkins had independently and/or previously established the 34 Å axial repeat for the B form. In fact, Wilkins made several notes in the 1970s suggesting that he was interested in working out when Franklin had managed to do it. For example, on a copy of Franklin’s 7 February 1952 interim report, he wrote “N.B. no mention of 34 Å period[.] I had pattern by then & it looks as tho’ none of us bothered to measure it!” (Wilkins, Maurice Hugh Frederick [1916–2004], King’s College London Archives (hereafter, Wilkins Papers) K/PP178/5/3; see also his efforts to remember when she had done it on pp. 3 and 13 of his reminiscence dated 17 September 1976 in K/PP178/5/27/1). By the time Wilkins told Watson details of the B form in early 1953, he would have learned about the 34 Å period from Franklin and Gosling’s 3 September 1952 contribution to their unit’s annual report to the biophysics committee of the Medical Research Council. In his various recollections, Watson himself later indicated that he learned details of Franklin’s measurements by viewing his colleague Max Perutz’s copy of that MRC committee report, as well as from Wilkins by word of mouth.
The intervening Photograph 50, taken during the day on 2 May 1952, was a very brief exposure assessing the condition of a freshly prepared DNA sample that was then placed into a desiccator (FRKN 1/1, Franklin Papers).
Entry 51 in Franklin’s 1952 notebook (FRKN 1/1, Franklin Papers).
We are not the first analysts to discuss Franklin’s forty-ninth photograph, but we believe our formulation of its significance in relation to Photograph 51 is original. Aaron Klug (
1968
, p. 810) alluded to the existence of Photograph 49 by describing the image published in Franklin and Gosling’s
Nature
paper (i.e., Photograph 51) as one of
multiple
“photographs of exceptional quality [showing…] in a direct manner that DNA in the B form is a helix with an axial repeat of 34 Å and an axial spacing between nucleotides of 3.4 Å.” He continued, “The model building by [Watson and Crick…] was carried out to fit these parameters.” Robert Olby (1994 [
1974
], p. 369) also did not explicitly name Photograph 49, but he quoted Franklin’s 49th notebook entry (“V[ery] good ‘wet’ photo”) and speculated that this picture was the one Wilkins later showed Watson. On the following page, he quoted the notebook passage describing Franklin’s discovery of the 34 Å axial repeat but did not specify that it was accomplished using Photograph 49. Judson (
1979
, p. 135) did specify that the “good ‘wet’ photo” entry was a description of Photograph 49 and he noted that Photograph 51 was “set up […] with the film holder more perfectly centered on the X-ray source.” As we discuss in the main text below, however, he then shifted to asserting what he thought Franklin
should
have realized about these two images. That passage drew from (and, in the process, conflated) Franklin’s May 1952 analyses of Photograph 49
and
her February 1953 analyses of Photograph 51. We also note that a recent post on the website of Jessica Mills Davies, the author of a novel based on Franklin’s life, is illustrated with two versions of Photograph 49, namely the negative we included above (Fig.
2
) and a positive-image print that Gosling used in his 1954 thesis. See Mills (
2024
) and Jessica Mills Davies. 16 January 2026. “Photograph 49: The X-ray Watson did not see,” (
https://www.jessiemillsauthor.com/journalism/photo-49-the-x-ray-watson-did-not-see
, accessed 28 May 2026). Although we find many of Mills Davies’ specific claims about the photograph to be inconsistent with our understanding of Franklin’s work, we believe she is the first person to highlight publicly its appearance in Gosling’s thesis and to note the inaccuracy of his accompanying caption which mistakenly says it was exposed for the same duration as Photograph 51. Klug himself had noted this with some confusion inside his copy of Gosling’s 1954 PhD thesis, which he had inherited from Franklin and which is now in SHI HMBC, Box 14, Folder 1, notes on plates 4 and 10 of Chap. 4. As we show below, the error was originally made in the caption for Fig.
5
of the earlier publication by Franklin and Gosling (
1953b
, plate 10, inserted between pp. 674 and 675).
Notebook entry for 10 February 1953 (FRKN 1/1, Franklin Papers).
The paper drafted second but eventually published first was the expedited paper on the B form that appeared alongside Watson and Crick’s double helix paper on 25 April (Franklin and Gosling
1953a
). The first paper to be drafted, about the relationship between the A and B forms, had already been submitted on March 6 but did not appear until the summer (Franklin and Gosling
1953b
). Photograph 51 appeared in the latter paper as Fig. 4. It was accompanied by an extreme-closeup detail of the center of Photograph 49, as Fig. 5, to illustrate part of the diffraction pattern (an equatorial doublet) that was not entirely visible outside the central pinhole of the Photograph 51 film. With respect to the beam and the pinhole, Photograph 49 was slightly better aligned (both figures and their captions were printed on plate 10, inserted between pp. 674 and 675). The caption of Fig. 5 mis-stated the exposure time of Photograph 49 as 62 h, an error that Gosling carried forward to his 1954 PhD thesis (SHI HMBC, Box 14, Folder 1, notes on plates 4 and 10 of Chap. 4).
In practice, Gosling continued working with Franklin while she was at Birkbeck. Their final co-authored DNA paper appeared in 1955.
The King’s College biophysics unit maintained index books of quarter-plate slides produced for researchers (KDBP 2/1, King’s College London Archive). The first book in this series records that in March 1953 “Dr Franklin” ordered several plates including those duplicating films 51C and 49B (plates 867 and 868 respectively).
There are several ways we can be sure that reproductions came from slide 867. For example, Franklin’s own copy of the image (see Fig.
1
), has “867” written in pencil on the back of the print (along with her annotation in pen). More significantly, all published versions show the doubly refined cropping achieved by the placement of masking tape on glass plate itself (see Fig.
5
). Presumably there must have been some prints made from the original film (as would have been the case for the print Wilkins showed Watson in early February 1953 nearly a month before plate 867 was made). The only print we are aware of that shows the original cropping from film 51C is the previously unremarked large-format print we have reproduced on the left of Fig.
5
(HMBC Box 10, Folder 15). This object was in Gosling’s possession as of the 1990s. In the same collection is a 4 inch x 5 inch (10.2 cm x 12.7 cm) acetate slide showing the patterns from photographs 42 and 51, i.e., Franklin and Gosling’s best A and B patterns, side by side and featuring the original cropping for Photograph 51 (HMBC Box 10, Folder 16). We are still working to solve the puzzle of when this slide was created and what it may have been used for.
Klug wrote to Olby on 3 September 1968, “[b]efore leaving for holiday a few weeks ago, I discovered Rosalind Franklin’s missing notebooks for the years 1951 and 1952. This enables one to date all the photographs. […] Would you like to come and see them?” (HMBC Box 13, Folder 23).
Maurice Wilkins 1970 interview with Anne Sayre for her book
Rosalind Franklin and DNA
, as quoted by Sayre (
1975
, p. 128). In the BBC film
Life Story
(1987), which was made in consultation with Wilkins, he is depicted as showing Watson a large print of Photograph 51 while telling him that its pattern is obviously representative of a helical structure. (Wilkins’ extensive files related to the production are in K/PP178/5/27, Wilkins Papers.)
Many efforts to understand the nature of Franklin’s failure focus on her supposedly anti-helical views. Watson used the phrase several times in
The Double Helix
, attributing it to Wilkins. He wrote, “Maurice had told me the nature of her so-called antihelical results,” and described “her self-made antihelical trap” (Watson
1968b
, pp. 165–166). As Robert Olby (1994 [
1974
], p. 371) has pointed out, Wilkins seems to have misjudged the significance of her anti-helical data (which he had not seen) and/or the sincerity of her antihelical views (1994 [
1974
], p. 371). We largely agree with the explanations given by Klug (
1968
) and Olby (1994 [
1974
], pp. 370–376) for why Franklin felt, from late 1951 to early 1953, that there was a lack of evidence to conclude that the A form was helical.
James Watson interview with Martin Raff and Walter Gratzer, recorded November 2008 and October 2009, section “Rosalind Franklin’s rapid acceptance of the double helix,” Web of Stories, 18 June 2010,
https://www.webofstories.com/play/james.watson/32
(accessed 29 May 2026).
Bernal wrote two obituaries of Franklin, both of which are quoted here. He called her images “among the most beautiful X-ray photographs” in memorializing her for
Nature
(Bernal
1958
), and he discussed her “apparently effortless skill” in
The Times
(J.D. Bernal, “Dr. Rosalind Franklin: A life dedicated to science,”
The Times
(London), April 19, 1958, p. 3).
Amplifying earlier observations by Maddox (
2002
, pp. 311–328), Angela Creager and Gregory Morgan (
2008
, p. 270) have argued that “the image of Franklin as meticulous and unimaginative […] was offered by Crick as well as by Watson, and the depiction works to excuse both of them for using her data by suggesting that she did not seem to know how to interpret it herself.”
The most vivid illustration of this phenomenon appears in an essay Judson appended to the 1996 edition of his book in response to those who had adopted Franklin as “an emblem for the condition of women in science” and given her what he considered undeserved credit for Watson and Crick’s intellectual breakthrough. Franklin had been “patient, dextrous [sic], untiring,” he argued, which meant that she had been “poignantly unlucky” to lack a collaborator like Watson. “His scientific imagination [was] intensely visual,” Judson wrote, so that Watson “understood instantly facts [about Photograph 51] that Franklin had only [later] figured out” (Judson
1996
, pp. 627–628). This points beyond the more general phenomenon of women receiving less attention and acclaim than men for similar scientific achievements, which historian Margaret Rossiter termed the “Matilda Effect” (Rossiter
1993
). Regarding the specific tendency we noted, Naomi Oreskes has written, “the invisibility of women’s contributions is enmeshed with the question of why some kinds of scientific work are more valued and honored than others” (Oreskes
1996
, p. 87). Because of sexual segregation in scientific and military institutions, many women were formally limited to holding technical or computational positions that were considered lower status, even if their activities transcended the nominal limits of their roles. Women in these roles became emblematic of “routine” scientific labor, making them all but ineligible to receive recognition according to the “rhetoric of [scientific] heroism in the public sphere” (Oreskes
1996
, p. 113). Pnina Abir-Am has made a similar but nevertheless distinct argument regarding creative contributions to the Nobel-prizewinning discovery of RNA splicing by the electron microscopist Louise Chow, whose work she says was not recognized for its novelty and significance by other participants who were less experienced in microscopy (Abir-Am
2020
).
As one illustration of this trend’s recent intensification, compare passages from a 2023 coauthored paper by Matthew Cobb and Nathaniel Comfort with Cobb’s
2025
biography of Crick. The former describes Photograph 51 as “a particularly clear image of the B form, taken […] by Franklin and her graduate student Raymond Gosling” (Cobb and Comfort
2023
, p. 658). The latter describes it as “an X-ray diffraction image of the B form supposedly taken by Franklin (in fact by Gosling)” (Cobb
2025
, p. 85). The trend seems to date from statements by and about Gosling around the time of the sixtieth anniversary of the double helix (see Smith
2019
; Attar
2013
,
2023
). We also note with good humor that the opening sentence of the current Wikipedia entry for Photograph 51, which says that the picture was “taken by Rosalind Franklin’s PhD student Raymond Gosling,” incongruously cites as its source an article by one of the present authors (Brian Sutton) that says the photograph was taken by Franklin and Gosling (
https://en.wikipedia.org/w/index.php?title=Photo_51&oldid=1364555388
; accessed 22 July 2026).
In an obituary for Ticehurst (then known by her married name of Freda Collier) in the
Guardian
, her nephew claimed that
she
had taken Photograph 51. Our general impression from all available evidence is more consistent with Heller’s recollection, namely that Ticehurst was likely involved in developing films after the experiments were complete (and creating duplicates on glass slides and in photographic prints). Heller’s recollections are from a July 2022 interview with Judy Masterson published on the website of Rosalind Franklin University in Chicago
(https://www.rosalindfranklin.edu/helix/winter-2023/being-there/).
The Ticehurst obituary is Thirlwall, A. P. 2013. “Freda [Ticehurst] Collier Obituary.”
The Guardian
, January 21
(https://www.theguardian.com/science/2013/jan/21/freda-collier-obituary
).
Raymond Gosling interview with Anne Sayre, 18 May 1970. Anne Sayre Collection of Rosalind Franklin Materials, Box 4, Folder 2, p. 15; Archive of the American Society for Microbiology, Baltimore, Maryland.
Franklin’s well-roundedness as a virus researcher — specifically the combination of leadership, crystallographic skill, and theoretical intuition she displayed in her 1953–1958 studies of Tobacco mosaic virus structure — has been well documented by Angela Creager and Gregory Morgan (
2008
).
Sociologist of science Michael Lynch highlighted the distinction between photographs taken for “use” and photographs subsequently taken to be published as “evidence,” noting that several microscopists had told him they would avoid publishing blemished images that they had actually studied if they could instead illustrate their research with a visually “perfect” picture. Lynch concluded that "the documentary use of a photograph in a [publication] differs considerably from that of a photograph used by lab members [in the course of doing the original research].” Well-composed and unblemished photos for publication would not only illustrate research findings clearly but also serve as “exhibits of a lab’s practical competence” (Lynch
1985
, pp. 94–96). We are grateful to Simon Schaffer for drawing our attention to this passage.
References
Abir-Am, Pnina. 2020. The women who discovered RNA splicing.
American Scientist
108:298–305.
Cobb, Matthew, and Nathaniel Comfort. 2023. What Rosalind Franklin truly contributed to the discovery of DNA’s structure.
Nature
616(7958):657–660.
https://doi.org/10.1038/d41586-023-01313-5
Cochran, W., F. H. C. Crick, and V. Vand. 1952. The structure of synthetic polypeptides. I. The transform of atoms on a helix.
Acta Crystallographica
5:581–586.
https://doi.org/10.1107/s0365110x52001635
Creager, Angela N. H., and Gregory J. Morgan. 2008. After the double helix: Rosalind Franklin’s research on Tobacco Mosaic Virus.
Isis
99:239–272.
https://doi.org/10.1086/588626
Crick, Francis H., and James D. Watson. 1954. The complementary structure of deoxyribonucleic acid.
Proceedings of the Royal Society of London. Series A, Mathematical and Physical Sciences
223 (1152): 80–96.
https://doi.org/10.1098/rspa.1954.0101
Daston, Lorraine, and Peter Galison. 2007.
Objectivity
. New York: Zone Books.
Franklin, Rosalind E., and R. G. Gosling. 1953a. Molecular configuration in sodium thymonucleate.
Nature
171(4356):740–741.
https://doi.org/10.1038/171740a0
Franklin, Rosalind E., and R. G. Gosling. 1953b. The structure of sodium thymonucleate fibres. I. The influence of water content.
Acta Crystallographica
6:673–767.
https://doi.org/10.1107/S0365110X53001939
Hall, Kersten T. 2014.
The man in the monkeynut coat: William Astbury and how wool wove the forgotten road to the double helix
. Oxford: Oxford University Press.
Judson, Horace Freeland. 1996.
The eighth day of creation: Makers of the revolution in biology
. Expanded ed. Cold Spring Harbor: Cold Spring Harbor Laboratory Press.
Lynch, Michael. 1985.
Art and artifact in laboratory science: A study of shop work and shop talk in a research laboratory
. London: Routledge & Kegan Paul.
Wilkins, Maurice., A. R. Stokes, and H. R. Wilson. 1953. Molecular structure of deoxypentose nucleic acids.
Nature
171(4356):738–740.
https://doi.org/10.1038/171738a0
We are extremely grateful to the archivists, librarians, and/or digital-collections staff of the Center for the History of Microbiology Archive in Baltimore, Churchill Archives Centre in Cambridge, King’s College London Archives (with special thanks for introducing us to each other), and the Science History Institute in Philadelphia (SHI) for making our research possible, and also to SHI colleagues/friends for supporting our collaboration. We are also grateful to our JHB editors, Nic Rasmussen and Betty Smocovitis, for their energy, advice and innumerable contributions to improving the paper; to two JHB referees for their helpful comments on the original manuscript; and to the following individuals for valuable conversations and/or comments on written drafts: Geoff Browell, Matthew Cobb, Nathaniel Comfort, Angela Creager, Michelle DiMeo, Hannah Grunwald, Judith Kaplan, Madison Renner, Lukas Rieppel, Simon Schaffer, Valerie Sponsel, Hallam Stevens, and Andrew Warwick.
Funding
This research was not supported by any external funding.
Author information
Authors and Affiliations
Science History Institute, Philadelphia, USA
Alistair Sponsel
King’s College London, London, UK
Brian Sutton
Authors
Alistair Sponsel
Brian Sutton
Contributions
AS and BS conducted all of the research, much of it while working side by side on archival materials, and developed the argument together. AS wrote the first draft and prepared the figures. AS and BS contributed equally to all subsequent writing and revision of the manuscript.
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Open Access
This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit
http://creativecommons.org/licenses/by-nc-nd/4.0/
.
Sponsel, A., Sutton, B. Photograph 49 is the Key to Understanding the History of Rosalind Franklin’s DNA Photograph 51.
J Hist Biol
(2026). https://doi.org/10.1007/s10739-026-09866-7
This study examines how everyday network services may remain available when Taiwan experiences large-scale international submarine cable outages. By observing homepage resource requests and tracing the locations of resources used during page loading, we assess potential website availability during international connectivity outages and use the results as a risk indicator.
The work focuses on two core questions: (1) how much websites commonly used in Taiwan depend on foreign-hosted resources, and (2) how much they depend on local (in-Taiwan) nodes of multinational public-cloud providers. We develop a measurement framework that turns the abstract risk of “what happens when cables go dark” into concrete dependency-structure analysis.
We collected connection data from 2,179 websites commonly used in Taiwan. The results show that 39.3% are “foreign-dependent,” with foreign resource exposure and relatively high direct failure risk under cable-outage scenarios. Another 49.6% are “cloud-dependent”: no foreign resource requests were observed, but they rely on local nodes of multinational public clouds or CDNs, so their actual availability when international connectivity fails is highly uncertain.
This study provides a scalable, reproducible measurement method for quantifying observable foreign dependencies and comparing dependency structures across websites. The results can inform policy and industry resilience planning and support continued tracking of resilience improvements.
Taiwan is a highly digitalized society; the Internet and the information systems built on it are a critical foundation for how society operates. Digital dependence keeps growing: as of 2024, fixed broadband household penetration reached 74.5%
1
; mobile broadband penetration 87.12%; overall Internet usage rose from 67.2% in 2006 to 88.75%
2
.
From waking to sleep, people constantly use networked screens to access and exchange information. The network is embedded in daily life and social activity. Communications, commerce, media, logistics, and public and government services all rely on digital systems online.
High connectivity means any sizable, sustained network outage can significantly impact the economy and society.
How everyday services depend on international connectivity
When someone opens an app (e.g. Line, the most popular messaging app in Taiwan, with over 98.2% usage rate
1
) on a phone and sends a message, a chain of network requests is triggered.
First, the app asks the OS to connect; the device queries the carrier DNS for the target server IP (e.g. Line). After obtaining the address, it opens a TCP/IP connection and sends the request.
Data leaves the phone over wireless to a nearby cell site, then enters the carrier’s fiber backbone and core. Because Line’s main servers are abroad (Japan), traffic is routed to an international gateway—e.g. Tamsui or Toucheng cable landing stations—and crosses submarine cables overseas.
After reaching the destination country, traffic lands again and enters a cloud provider’s data center (e.g. AWS, Google Cloud, Azure). The app server processes the request; the response returns along a similar path and is rendered by the OS and the app.
This often completes in a fraction of a second. Unnoticed by users, data may travel thousands of kilometers round-trip between Taiwan and Japan—or tens of thousands of kilometers to another continent and back.
More importantly, one tap on an app or site can trigger dozens or hundreds of parallel requests, each repeating the above pattern. Most everyday digital services are effectively
cross-border systems
; “instant” interaction depends heavily on international links and submarine cables.
Submarine cable vulnerability for an island economy
As noted, many sites and apps used daily depend on foreign resources and international connectivity. Losing external connectivity would likely break most digital services.
As an island, over 99% of Taiwan’s external traffic relies on submarine cables
3
. Cable resilience therefore directly affects everyday service availability and broader societal resilience.
Submarine cables are multi-layer cables a few centimeters thick on the seabed, or buried one to three meters in shallow water. In busy shallow areas such as the Taiwan Strait, damage often comes from human activity—anchoring, fishing, dredging—and from natural wear, amplifier failure, earthquakes, landslides, and geopolitical risk. Human causes dominate cable damage in Taiwan
4
.
According to the Taiwan Submarine Cable Map (smc.peering.tw), Taiwan is almost always in a state where “at least one cable is impaired”
5
. Cable faults may be a chronic background condition, not rare exceptions.
Availability of all international submarine cables serving Taiwan, 2025/3/18–2026/3/18 (source: Taiwan Submarine Cable Map (smc.peering.tw), cable status timeline).
Taiwan connects globally through fourteen international cables via landing stations at Tamsui, Bali, Toucheng, and Fangshan (another station is under construction in Dawu, Taitung), plus ten domestic cables to Penghu, Kinmen, Matsu, and other outlying islands
6
(RNAL and FNAL are two systems on one physical cable; MODA counts them separately, yielding fifteen international systems in some counts).
Under normal conditions, the Internet’s meshing, redundancy, diversity, and connectivity let carriers reroute traffic over other cables when a few fail. Quality may drop without users noticing. However, when multiple cables fail together, bandwidth redundancy is quickly exhausted, causing severe congestion or large-scale outages affecting communications, logistics, government operations, and digital systems
7
.
According to repair-time benchmarks published by the Ministry of Digital Affairs
8
, average repair time is about
32 days
for international cables and about
110 days
for domestic cables linking outlying islands. Disruptions can therefore last for months or even quarters, and contingency planning should assume a monthly rather than daily timescale.
Historical case: multiple international cables failed at once
A recent multi-cable failure occurred from 2025-12-25 to 2026-01-03: earthquakes off Yilan damaged six international cables (including EAC1, SJC2, PLCN, F/RNAL, EAC2, Apricot—nearly half of cables)
4
. Users reported slower networks and blocked apps, with repairs not completed until May 2026.
The 2006 Hengchun earthquakes remain a landmark case: on 2006-12-26 at 20:26 and 20:34, two magnitude-7 quakes off southwest Hengchun triggered many aftershocks. Mainland damage was relatively light, but underwater landslides broke four of six external cables at the time.
Taiwan’s international connectivity was severely disrupted. Initial call completion to the U.S. was about 40%; to China and Japan about 10%. China, Hong Kong, Japan, Korea, and Southeast Asia were also hit hard. Google, Yahoo, MSN, Gmail, Wikipedia, and other major services saw major outages across the region, affecting trade and finance.
Eight cable ships joined repairs; full restoration took nearly two months by mid-February 2007
9
. The UN ISDR director called submarine cable damage from the quake a new modern-type disaster
10
.
Both events show that even without total blackout, simultaneous multi-cable failure can cause severe congestion and widespread service degradation.
Historical case: Matsu island-wide outage
The early 2023 Matsu outage is a real-world case of complete external cable loss for a region.
Two cables linking Matsu to Taiwan were damaged on 2023-02-02 and 2023-02-08 by Chinese fishing vessels, cutting regional Internet and telecom except for very limited microwave capacity (2 Gbps), leaving most residents unable to get online
11
. One cable (Taiwan–Matsu 3) was repaired after about 50 days at end of March.
Taiwan–Matsu 2 and 3 total about 1 Tbps. After 2023, microwave was expanded to 12 Gbps. When both cables failed again on 2025-01-15 and 2025-01-22, the larger microwave link kept some connectivity
12
.
Satellite as backup: capacity and bandwidth limits
At the beginning of the 2006 incident, Chunghwa Telecom reallocated capacity on the ST-1 communications satellite to support international telephone service, briefly restoring partial availability for international voice calls. If a similar-scale incident happened today, could satellites still serve as a substitute?
Under the Taiwan Space Agency’s (TASA) “B5G LEO communications satellite program”
13
, Taiwan plans to launch two experimental LEO satellites before 2030 with a three-year design life.
Former TASA chair Tsung-Tsong Wu estimated that 24/7 nationwide LEO coverage would require at least 120 satellites, with roughly 40 replacements per year for a three-year lifetime—far above current plans, so it is hard to build substantive backup connectivity on that scale
14
.
Bandwidth also differs by orders of magnitude. A modern cable may carry hundreds of Tbps; Apricot is designed for 211 Tbps
15
. Starlink capacity was estimated around 20 Gbps in 2023 research—roughly 10,000× less per comparable unit; the full Starlink constellation (~3,300 satellites in 2023) was estimated around 20 Tbps total, comparable to one cable system
16
.
Even summing current LEO satellite bandwidth (Gbps class) cannot replace transoceanic cable throughput (Tbps class). TWNIC chair Kenny Huang compared cables to reservoirs and satellites to pipes
12
. Satellites can support emergency government or regional links, not national-scale replacement.
Why risk today exceeds 2006
In 2006, the main economic impact of lost international connectivity was disrupted international telephone service—finance and select industries—with overseas internet services affecting only a minority of users.
In twenty years, dependence on the Internet has grown sharply. Taiwan’s international bandwidth grew from 147.7 Gbps in 2006 to 10.6 Tbps in 2026—nearly 70×
17
. Meanwhile, average cable damage in Taiwan is about 5.1 times per year versus a global average of 0.1–0.2—roughly 25–50× higher risk
3
.
Deloitte (2016) estimated that a full national Internet outage in a highly digitalized country could cost about USD 23.6 million in GDP per day per ten million people—roughly USD 55 million per day for Taiwan, or about USD 1.7 billion per month, before accounting for semiconductor supply-chain and cross-border finance spillovers
18
.
The same work notes that even partial outages or bandwidth reduction hurt productivity, transactions, information access, and confidence.
Taiwan is more Internet-dependent and faces more frequent cable damage risk. A 2006-scale event today would affect far more than niche industries, with broader societal impact and much harder emergency response and alternative routing than in 2006.
Public awareness and research gaps
Public discussion of cable outages has increased, but often stays at “communications difficulties”—Line/Messenger/WeChat, Google, Gmail, Office 365—similar to 2006 framing.
Academic resilience work often focuses on infrastructure or intermediary layers: cable topology, routing, DNS
19
20
, CDN and cloud centralization. That implies: if infrastructure is up and reachable, services work.
Routing is decentralized and imperfect; failures are routine. Policy constraints mean routing errors, misconfiguration, or node faults can cause large outages without physical cuts. Infrastructure health alone does not equal service availability
21
.
Even with spare cables, traffic reshaping, routing policy, or concentrated paths can still block communication—“physically connected” ≠ “actually reachable”
22
; redundancy alone does not guarantee cross-border availability
23
.
Modern stacks span infrastructure, logical, and application layers; societal impact exceeds any single operator. Resilience is a public-good problem across layers, operators, and borders
7
.
Indices such as the Internet Resilience Index use national infrastructure, performance, security, and market structure as proxies
24
but do not directly measure whether sites people use daily still load when international connectivity is constrained.
Policy analyses focus on national infrastructure, topology, repair capacity, alternatives, and geopolitical risk—emphasizing bandwidth redundancy, path diversity, repair, and cooperation
25
.
Third-party dependency research (e.g. DNS, CDN, certificate authorities) highlights concentration and single points of failure
26
: over 89% of sites depend on third parties for critical functions; top three providers support over 90%
27
; multi-level indirect dependencies can amplify failures
28
.
Resilience must include service availability under constrained external connectivity and third-party state—not only physical reachability. This study emphasizes “service availability under constrained connectivity”, developing application-layer methods to ask which everyday digital services—and what share—could remain usable when external links fail, and how cable outages may affect network and societal resilience.
Summary
The challenge we face is not “will cables break?” but service collapse risk: in a highly digital, cloud-heavy, cross-border-dependent society, when external connectivity is severely impaired, which services keep basic function, which degrade, and which fail outright?
Past debate focused on bandwidth, physical damage, and backup communication methods. Modern services are not “a local server serving static HTML.” A “Taiwan” site or app may run on global cloud regions across datacenters and depend on hosts, CDNs, third-party JavaScript, login, payments, analytics, push, AI APIs, and other external components. Any critical piece abroad or reachable only via foreign paths can fail unexpectedly during cable outages.
Under severely congested or broken international links, repairing or rebuilding systems becomes harder.
Impact cannot be judged only by “how many spare cables remain.” Services may fail due to routing, DNS, congestion, cloud dependencies, or unreachable external assets even when the physical network is not fully down. Beyond connectivity, we need user-facing service availability measurements.
Without understanding real impact, we cannot prepare. This study turns “what happens when cables go dark” into measurable, comparable technical questions—filling an application-layer gap in resilience discourse. Through a test framework, it estimates how commonly used services may load, degrade, or fail when foreign connectivity is lost, providing evidence for backup design, resilience investment, and policy and social readiness.
Research questions
When an island nation highly dependent on international networks, such as Taiwan, experiences severe degradation, instability, or partial interruption of its external submarine-cable connectivity—and thus loses access to much of the global Internet—to what extent do commonly used domestic websites continue to operate, degrade, or fail?
We aim to systematically test and count international-facing components in service operation—CDNs, third-party APIs, cloud platforms, external libraries—and map dependency structure and potential availability risk for commonly used services under foreign-network isolation.
The work should give government, industry, and civil society concrete evidence on systemic impact of external connectivity loss, supporting resilience strategy for digital services and critical public sectors.
Two core themes:
Degree of dependence on foreign-hosted resources
Degree of dependence on Taiwan nodes of multinational public clouds and CDNs, and what that implies for resilience
Three concrete questions:
Under “Taiwan external connectivity severely impaired or cut,” what
share of commonly used sites
exhibit foreign resource dependency exposure at the
homepage
level?
Is dependency on local nodes of multinational public clouds and CDNs
concentrated in specific service ecosystems
?
Do different site types, such as
.gov.tw
government websites,
.edu.tw
education websites, and general services, show
systematic differences
in foreign-resource dependency rates?
This study uses a programmatic browser to observe the distribution of resource requests generated while loading each homepage. A homepage is the first user-visible point of interaction with a service and includes some of the front-end resources required for basic rendering and interaction. We treat this observation as a scalable proxy indicator for comparing dependency exposure across a large number of websites. However, it does not include the complete backend architecture or cloud control-plane dependencies and therefore cannot be interpreted directly as overall service availability.
Targets and test environment
We compiled a high-traffic site list for Taiwan, including domestic and international services commonly used by locals. The target is “sites Taiwanese people use,” instead of “Taiwan sites”—so the list includes Google, Gmail, etc.
The unit of study is “websites” (Web), not direct equivalence to app availability. (OCF has related work on app connectivity resilience.)
Building the site list
There is no authoritative “sites Taiwanese people use” list. We merged:
Tranco List
29
— global top 1M list, from which we use 2,510 entries ending in
.tw
.
This campaign used ranking snapshots obtained from each source on 2026-07-20. Before merging, hostnames were lowercased and a leading
www.
was removed, while other subdomains were preserved. Duplicate hostnames were merged into one entry while retaining the source-specific rankings. The resulting
merged_lists_tw.json
contains 2,467 sites.
We also use
manual_curated_list_tw.json
to include 42 manually selected open-source and digital-resilience community cases, including OCF, SITCON, and g0v. Two hostnames overlap with the automatically ranked list.
Tests used typical Taiwanese residential connectivity:
Chunghwa Telecom fiber 500M/500M
Locations: Zhonghe District, New Taipei; Zhongzheng District, Taipei
DNS: 168.95.1.1
Environment details recorded in logs for comparison and reproduction
Methods
Site availability depends not only on establishing connections, but on fetching dependent resources (JavaScript, CSS, images, APIs). Modern sites combine resources from many domains; together they determine what users see and can do. Prior work uses headless browsers to analyze request behavior and third-party dependencies
30
27
.
Building on dependency-exposure analysis from resource requests, we extend it to the systematic scenario of “international connectivity failure” and its potential impact on service availability.
Backend architecture, data paths, control planes, and internal cloud behavior are not directly observable externally. We do not attempt to map full system dependencies; we focus on observable front-end network requests and build operational metrics from them.
Metrics and risk taxonomy
Two core metrics:
Foreign Dependency Exposure
Whether homepage requests include foreign-hosted resources—exposure to foreign networks at the resource layer.
Cloud Local Endpoint Exposure
Whether requests hit Taiwan nodes of multinational public-cloud or CDN providers—exposure to sites hosted on, or dependent on, their local endpoints.
These describe “dependency exposure structure” at the homepage front-end resource layer, not full system architecture or actual failure modes.
We can classify sites into three categories based on the above metrics:
Foreign-dependent
Foreign resource exposure: homepage load directly depends on foreign resources. Such sites are directly exposed to unavailable resources when external connectivity is severely degraded or interrupted, making them the most likely to be affected immediately and the category with the highest direct risk.
Cloud-dependent
No foreign resource exposure, but cloud local endpoint exposure: homepage loads show no foreign resource requests, yet use resources from Taiwan nodes of multinational public-cloud or CDN providers. Sites appear localized, but availability still depends on control planes, origins, authentication, and cache persistence—“surface-local, cross-border uncertainty”.
Locally-contained
No foreign resource exposure and no multinational-cloud local-endpoint exposure. Such sites have a higher chance of local operation, but full-system availability during external outages is not guaranteed.
We developed
web-resilience-test
to open each target homepage site-by-site with a programmatic headless browser and record all resource connections during load.
Next, the tool processes all resource requests collected during page loading. It excludes
blob:
and other unparsable requests, as well as requests matching the ad-filter rules or the manual exclusion list. It then deduplicates the requests, retaining only one request per hostname within each website; each unique website-hostname pair constitutes one observation.
The tool then uses IPinfo, provider-specific response headers, the LACeS anycast API
31
, and ping RTT to determine the geographic location and classification of each observation.
Finally, the tool aggregates all test results into statistical tables. The detailed workflow follows.
Listen to
request
for all request metadata including headers
Retries and errors
4xx responses are treated as test failure and logged
Other errors: retry in this order:
Headless / non-headless browser
URL with/without
www.
prefix
If all four variants still fail, log the error and skip to the next site
Request cleanup
For each page's request data, perform the following cleanup:
Drop
blob:
requests
Exclude unparsable requests
Match requests against the configured adblock domain lists to remove advertisements and related unnecessary resources
This campaign used
LowTechFilter hosts ABP
version 2026.0720.1 and a 2026-07-20 snapshot of the
AdGuard DNS Filter
. Together, the two lists contained 158,521 distinct domain rules
If the target website's hostname appears in the lists, requests to hostnames with a parent or child relationship to the target are retained to avoid blocking the website's own resources
Exclude hostnames on the manual exclusion list. This study lists only the web-font hostname
fonts.gstatic.com
, because its failure is less likely to affect core website functionality than failures of scripts, APIs, or content resources. Future studies can use the same mechanism to add other exclusions
Deduplicate requests by hostname; each unique website-hostname pair is one observation
Domain location
For each observation's hostname, call the
IPinfo API
to obtain its geographic location and ASN information
If the result shows
country=TW
, record it as a domestic connection
Compare the ASN against
cloud_providers_tw.json
registry version 0.1.0, dated 2026-01-05, to determine whether it belongs to the multinational cloud/CDN category. The registry contains 30 international public-cloud, CDN, and hosting providers and excludes Taiwan-only providers
If the result shows a
country
other than
TW
and its ASN belongs to one of the following six providers with local Taiwan nodes, apply the advanced classification steps:
Google (AS15169, AS396982, AS19527), Cloudflare (AS13335, AS209242), Amazon (AS16509, AS14618), Fastly (AS54113), Akamai (AS16625, AS20940, AS32787), or Microsoft (AS8075)
Headers: inspect response headers such as
cf-ray
,
x-amz-cf-pop
,
x-served-by
,
x-azure-ref
, and
x-msedge-ref
for known cloud location markers. A value containing
TPE
indicates a Taiwan node
Anycast: if headers are inconclusive, query the
LACeS Anycast Census API
31
to determine whether the endpoint is a local anycast resource; if
locations
includes Taiwan and
confidence
is
confident
, classify it as domestic
RTT: if the preceding methods are inconclusive, ping the resource 5× and take the minimum RTT. Classify it as domestic when
RTT < 15 ms
; otherwise, retain the foreign classification
Note: we built the
cloud_providers_tw.json
ASN registry using complete request data from earlier tests. It is open source for use by other research and projects.
Classification and resilience metrics
Based on the preceding information, classify each unique website-hostname observation as one of:
domestic/cloud
,
domestic/direct
,
foreign/cloud
,
foreign/direct
“cloud” means the ASN in IPinfo
org
is listed in
cloud_providers_tw.json
under
providers_intl
or
providers_intl_without_known_taiwan_region/pop
Count the total observations in each category for every website and save the results to
test-results/<site>.json
Errors
Failures are logged to
test-results/_error/<site>.error.json
Common errors include:
Cloudflare challenge
: target site uses Cloudflare's challenge protection mechanism to prevent abuse.
HTTP 4xx
Timeout
RTT classification coverage and threshold sensitivity
RTT is the final stage of location classification: a target cloud or CDN endpoint enters RTT only when IPinfo, provider-specific response headers, and LACeS cannot determine its location. The test tool pings the endpoint five times and uses the minimum RTT. It is reclassified as an international public-cloud node in Taiwan only when the minimum RTT is below 15 ms; otherwise, it remains classified as foreign.
Across 2,179 successfully tested websites, the browser recorded 262,926 raw HTTP requests. The filtering described above and within-site hostname deduplication produced 19,046 unique website-hostname observations for classification. Of these, 3,640 observations from 976 distinct hostnames across 1,243 websites entered the RTT fallback stage; details are in the table below.
Metric
Count
Share
Sites tested
2,179
Sites with RTT fallback
1,243
57.0% of sites tested
Observations to classify
19,046
Entered RTT stage
3,640
19.1% of observations
RTT measured
3,064
84.2% of RTT-stage observations
Min RTT < 15 ms
2,394
78.1% of measured RTTs
Min RTT ≥ 15 ms
670
21.9% of measured RTTs
The distribution of the 3,064 measured RTTs is shown below. Each point is one successful measurement; the horizontal axis is an observation index, and the vertical axis is logarithmic. Values are bimodal; the 10–30 ms transition range holds 130 observations (4.2%).
Based on this distribution, we use a relatively conservative
< 15 ms
threshold within the sparse interval to classify a resource as domestic, reducing the risk of misclassifying a foreign resource as domestic.
RTT range
Count
Share
0–<5 ms
206
6.7%
5–<10 ms
2,090
68.2%
10–<15 ms
98
3.2%
15–<20 ms
12
0.4%
20–<30 ms
20
0.7%
30–<50 ms
270
8.8%
50–<100 ms
247
8.1%
100–<200 ms
100
3.3%
≥200 ms
21
0.7%
Without RTT correction (equivalent to a 0 ms threshold), every resource entering RTT fallback retains its foreign classification, producing 1,370 foreign-dependent websites and 566 cloud-dependent websites. Applying the 15 ms threshold reclassifies 514 websites (23.6%) from foreign-dependent to cloud-dependent, resulting in 856 and 1,080 websites in the two categories, respectively.
The table further tests sensitivity to the selected RTT threshold. Relative to the 15 ms baseline used in this study, a 10 ms threshold moves only 27 websites (1.2%) from cloud-dependent to foreign-dependent, while a 20 ms threshold moves only five websites (0.2%) from foreign-dependent to cloud-dependent. The locally-contained count is unchanged. The aggregate classification therefore differs only slightly across the 10, 15, and 20 ms thresholds.
RTT threshold
Foreign-dependent
Cloud-dependent
Sites reclassified vs. 15 ms
No RTT (0 ms)
1,370 (62.9%)
566 (26.0%)
—
10 ms
883 (40.5%)
1,053 (48.3%)
27 (1.2%)
15 ms
856 (39.3%)
1,080 (49.6%)
—
20 ms
851 (39.1%)
1,085 (49.8%)
5 (0.2%)
Batch test flow
batch-test.js
runs single-site tests over the list and writes
test-results/statistic.tsv
. From that batch output we derive overall foreign-dependency rates and per-resource resilience status.
At ~2,000 sites, when running with default parallelism (4 parallel tests, 8 parallel static page compilations), it takes about 30–60 minutes to complete. The latest test results are published at
web-resilience-test-result
and
resilience.ocf.tw
.
Results
Of 2,507 unique sites tested, 2,179 completed successfully.
Within the measured results, 39.3% are “foreign-dependent,” with foreign resource exposure and
high direct failure risk
under cable outages; 49.6% are “cloud-dependent”—no foreign resource exposure was observed, but they rely on in-Taiwan nodes of multinational public clouds or CDNs, so availability is
highly uncertain
. Only 11.2% are “locally-contained,” with no observed exposure and a higher chance of normal operation. Overall, 88.8% of sites warrant further attention as high-risk or high-uncertainty.
Interpretation
Foreign-dependent: sites hosted abroad or whose homepages request foreign resources. They are exposed to international connectivity degradation or interruption and face high failure risk.
Cloud-dependent: no direct foreign resource exposure, but loading pulls resources from Taiwan nodes of multinational public-cloud or CDN providers. Topologically domestic, yet control planes, origins, authentication, or cache persistence may still depend on foreign systems—“localized in appearance, uncertain in availability”.
Locally-contained: no dependency exposure was observed among front-end resources. The site itself appears to be domestic and not hosted on a multinational public cloud, and it does not request foreign resources or resources from domestic nodes of multinational clouds. It therefore has a higher chance of continued operation, although this category does not account for complete backend dependencies and cannot establish that the service will continue operating.
Category
Sites
Share
Foreign-dependent (foreign resource exposure)
856
39.3%
Cloud-dependent (no foreign exposure; in-Taiwan nodes exposure)
1,080
49.6%
Locally-contained (no observed exposure)
243
11.2%
Total
2,179
100.0%
Multinational public cloud dependency
Among Category 2 (cloud-dependent) sites, requests to different international public-cloud nodes in Taiwan break down as follows:
Google Cloud Platform (Taiwan nodes): 965 sites
Cloudflare (Taiwan nodes): 480 sites
Amazon Web Services (Taiwan nodes): 138 sites
Akamai (Taiwan nodes): 104 sites
Microsoft Azure (Taiwan nodes): 38 sites
Fastly (Taiwan nodes): 4 sites
Provider site counts are non-exclusive: one site may use more than one provider, so the rows must not be added to obtain a site total.
Of 1,323 sites with no foreign dependency, 965 use resources from GCP Taiwan nodes (72.9%).
If public-cloud services such as GCP cannot keep local nodes running during external network outages, the impact would be very high. Their resilience is a key factor in whether sites can continue operating during submarine-cable disruptions.
Public cloud resource locations
For resources requested from domestic and international nodes of multinational public clouds, we found the following distribution:
Provider
Sites (domestic nodes)
Sites (international nodes)
Requests (domestic nodes)
Requests (international nodes)
Google
1,685
56
7,393
63
Cloudflare
1,016
17
3,051
21
Amazon
512
309
1,382
522
Akamai
338
11
446
13
Fastly
6
257
6
369
Microsoft
140
77
196
143
A site may request both domestic and international nodes from the same provider, so the two columns overlap and must not be added to obtain a provider total.
For Google cloud resources, measured by request count, 7,393 of 7,456 Google requests were classified as domestic-node requests, or about 99.2%; 63 were international-node requests, or about 0.8%. This shows the practical value of CDN-based data localization, and makes the persistence of mirrored resources on Taiwan nodes a key factor in whether ordinary sites remain available when external links are congested or cut.
Public-cloud services with lower domestic resource shares should be further evaluated for full in-country mirroring, cache persistence, and contingency operations.
Resource location and cloud-platform dependency statistics
For this comparison,
global cloud
includes the multinational public-cloud and CDN providers compiled for this study, while
other/local
covers all other providers. A website is counted in a cell when at least one of its observations belongs to that group:
Unit: sites & adoption rate
Domestic
Foreign
Any
Global cloud
1,881 (86.3%)
754 (34.6%)
1,910 (87.7%)
Other/local
1,623 (74.5%)
245 (11.2%)
1,709 (78.4%)
Total
2,140 (98.2%)
856 (39.3%)
2,179 (100.0%)
87.7% of sites depend on global-cloud resources: 86.3% depend on resources from domestic global-cloud endpoints, and 34.6% depend on resources from foreign global-cloud endpoints.
Among the 856 sites with foreign resource exposure, most also use domestic resources; only 39 use foreign resources exclusively, accounting for just 1.8% of all 2,179 sites. This shows the practical value of CDN contributions to data localization and benefits for service resilience.
Resource source distribution
Among the 18,969 unique website-hostname observations with provider information from IPinfo, dependencies are highly concentrated among large providers. Provider labels are normalized from IPinfo ASN organization data. Providers above 5% include Google, Cloudflare, Amazon, Chunghwa Telecom (CHT), and Facebook. Google has the highest observation share at 39.7%, followed by Cloudflare at 16.4% and Amazon at 10.4%.
Per-site inspection shows that Google resources mainly include services such as GTM, while Cloudflare provides infrastructure and services such as
cdnjs
JavaScript CDN and WAF. These common infrastructure services form key parts of contemporary internet-service resilience.
Unit
Count
Share
Google
7,525
39.7%
Cloudflare
3,109
16.4%
Amazon
1,979
10.4%
Data Communication (CHT)
1,645
8.7%
Facebook
1,460
7.7%
Akamai
518
2.7%
Fastly
375
2.0%
Microsoft
346
1.8%
Taiwan Academic (TANet)
321
1.7%
Yahoo
115
0.6%
Oracle
110
0.6%
Taiwan Fixed Network
107
0.6%
New Century
93
0.5%
OVH SAS
81
0.4%
Automattic
66
0.3%
Zenlayer
60
0.3%
Incapsula
54
0.3%
Yuan-Jhen Info
44
0.2%
Magnite
40
0.2%
Datacamp
37
0.2%
Sony
36
0.2%
Byteplus
32
0.2%
Public-sector aggregate risk
To assess the resilience of government and education sites, we first looked only at foreign resource connectivity:
Among the test results, 235 were government sites (
gov.tw
and
*.gov.tw
); 16 had foreign connectivity, or 6.8%.
255 were education sites (
*.edu.tw
); 34 had foreign connectivity, or 13.3%.
Type
Sites tested
Foreign dependencies
Share
Government
235
16
6.8%
Education
255
34
13.3%
All
2,179
856
39.3%
Government and education sites depend less on foreign resources than the overall population. This suggests that public-sector and academic-network environments have a stronger baseline for local availability, although full service resilience still requires checking backend dependencies and real usage workflows.
Recommendations
Based on this study’s findings, we identify the following policy and technical recommendations to improve the resilience of Taiwan’s overall digital services.
Overall, the main risk for websites commonly used in Taiwan does not come only from a small number of fully foreign-hosted services. It is more widely embedded in dependency structures involving foreign resources and Taiwan-based nodes of multinational public clouds. Resilience strategies should therefore go beyond asking whether a service is “in Taiwan,” and further examine whether its resource supply chain, cloud control planes, and critical user journeys can continue operating locally.
Policy Recommendations
Support related research to continuously monitor the resilience of commonly used and critical services, and routinely publish both aggregate and per-service results.
Support follow-up research to develop deeper resilience testing frameworks for user journeys such as login, transactions, browsing, and search, in order to conduct further availability testing.
For Taiwan nodes of heavily used international public-cloud providers such as Google, Cloudflare, Amazon, and Akamai, provide policy requirements and budget support to verify and improve service availability during external network outages.
Provide policy requirements and budget support to reduce critical domestic services’ dependence on foreign resources and improve their local resilience.
Encourage or require critical domestic services to establish local backup mechanisms or recovery plans, and conduct periodic disconnection drills.
Based on local-availability validation, define resilience tiers, such as A: fully usable; B: degraded but usable; C: homepage loads but interactions fail; D: immediate failure, and include them in procurement and acceptance criteria for government and public services.
Establish an extreme-case bandwidth-priority plan in advance, given that backup satellite capacity is far below submarine-cable capacity.
Technical Recommendations
For highly critical international public-cloud services operated by providers such as Google, Cloudflare, Amazon, and Akamai, contingency plans for external connectivity failures should be developed and regularly exercised.
Website builders should consider the resilience risks of using foreign resources. When loading frameworks or libraries, they can prioritize CDN services with Taiwan-based nodes, or establish fallback mechanisms that switch to local resources when a library fails to load, reducing the impact of external connectivity outages.
Service developers can prioritize data localization for critical service paths, such as login and checkout, to improve resilience and service quality.
Limitations and future work
Main limitations:
This study observes the source locations of website requests, not full network paths such as traceroute, nor routes from abroad via VPN. Whether “domestic” resources or pages are anycast/CDN nodes still needs further testing.
The 15 ms RTT cutoff is a fallback heuristic, not physical proof of endpoint location. Routing changes, congestion, or nearby foreign nodes can affect individual measurements, although the threshold sensitivity analysis shows limited aggregate impact.
“Foreign dependency” and “cloud dependency” here refer to front-end observable exposure, not full backend architecture. Even locally-contained sites by front-end metrics may still rely on foreign databases, APIs, or backend services. The 11.2% locally-contained group cannot be assumed to remain available during external outages.
Resources and webpages hosted on Taiwan nodes of multinational public-cloud services do not guarantee that the service can operate independently during submarine-cable or international connectivity outages. Observing a Taiwan endpoint shows only that some front-end resources can be obtained domestically. Actual availability may still depend on:
Whether control-plane services, including authoritative DNS, configuration, and logging, rely on foreign systems
Whether the origin and dynamic content are located abroad
Whether cache hit ratio, cache persistence, and cache revalidation require international connectivity
Whether identity and access management, token validation, session handling, and other authentication processes rely on foreign services
Other foreign dependencies that may affect service availability
This study does not perform fault-injection tests that simulate a loss of international connectivity, such as using VPN or DNS techniques to make foreign resources unreachable. As a broad survey of many websites, it estimates potential risk from dependency structures rather than directly observing service degradation during a forced connectivity outage.
This study tests homepages only, not full user journeys such as login, transactions, browsing, or search. Results should therefore be treated as “initial availability” indicators based on dependency exposure among resources involved in initial service access, not as measurements of complete service availability.
Suggested follow-ups:
Combine fault injection with journey-based testing, such as login, transactions, browsing, and search, to observe actual availability.
Study the resilience of major cloud-service architecture, including control planes, origins, cache defaults, and authentication.
Use traceroute to analyze full resource paths.
Analyze the usage and node distribution of common front-end libraries and frameworks, such as jQuery, Bootstrap, Tailwind, React, and Vue, to identify shared foreign-service single points of failure.
Compare dependency patterns by resource type, such as document, script, image, XHR, font, and stylesheet.
Identify high-traffic, low-resilience sites.
Add more Taiwan traffic data, such as the Chrome CrUX user experience dataset.
Develop a service-criticality framework that classifies services into categories such as government, finance, communications, video streaming, news, e-commerce, social media, and search engines; weights them by security, economic, and social importance; compares resilience across categories; and calculates an overall resilience index.
Rashna Kumar, Sana Asif, Elise Lee, Fabián E. Bustamante,
Third-party Service Dependencies and Centralization Around the World
, Northwestern University,
https://arxiv.org/abs/2111.12253
↩1
↩2
Aqsa Kashaf, Vyas Sekar, Yuvraj Agarwal,
Analyzing Third Party Service Dependencies in Modern Web Services: Have We Learned from the Mirai-Dyn Incident?
,
https://doi.org/10.1145/3419394.3423664
↩
Victor Le Pochat, Tom Van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczyński, and Wouter Joosen,
Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation
, Proceedings of the 26th Annual Network and Distributed System Security Symposium (NDSS 2019),
https://doi.org/10.14722/ndss.2019.23386
↩
Yasin Alhamwy, Paul Mertens, Oliver Hohlfeld,
Poster: Web Dependency Analyzer to Identify Resource Dependencies and their Impact on Rendering
,
https://doi.org/10.1145/3646547.3689683
↩
Remi Hendriks, Matthew Luckie, Mattijs Jonker, Raffaele Sommese, and Roland van Rijswijk-Deij,
LACeS: An Open, Fast, Responsible and Efficient Longitudinal Anycast Census System
, Proceedings of the 2025 ACM Internet Measurement Conference (IMC '25),
https://doi.org/10.1145/3730567.3764484
↩1
↩2
People in the software industry don’t seem to be doing too well at the moment. There has been a spate of “I’m leaving my tech job” videos – some because they’ve been laid off, some leaving preemptively – but the interesting action has been taking place in the comment sections, replies, and discussion forums. The replies are where it’s at. The videos themselves vary in quality.
It’s not just the “AI” financial bubble, or the debates that have been taking place – arguments, more properly – or the ongoing shift in public sentiment regarding the software industry.
I’m talking the fact that people who work in software – specifically developers and user experience designers – are not doing great.
I keep coming across videos of people talking about burning out, being unable to work, and in some cases fairly serious mental health issues. And, obviously, people are getting fired. But often in these videos, those who are doing videos and announcing they’ve been laid off mention that they were thinking about quitting anyway and are now planning on leaving tech entirely.
It’s not just senior people who might have memories and a fondness for how things used to be – stress mixed with nostalgia. It’s also students. I’ve found several videos of students who announced that they had switched majors, had been studying user experience design, but now decided to switch to something else.
With a few notable exceptions, these generally aren’t people who are against “AI” – or, at least, they weren’t against “AI” until very recently.
These aren’t people who disliked the technology from the start, who have had very severe criticisms about the way it’s made, the way it’s forced on people, or have severe doubts about the long-term benefits of applying this technology to our industry.
They seem to have all embraced “AI” at some point and assumed that what they were told was true, that this was the future. “This is where things are going.”
This is what software development will be like, now and forever.
After all that, one of the reasons why quite a few of them are leaving, is that they believe this promise. When they’re told that this is the future of the industry, they accept that statement as true. They’ve tried it; understand it and think to themselves: “I see your point, but this isn’t fun any more. This isn’t enjoyable. This is bad for me.”
And decide to leave.
There are a number of systemic factors for why this is happening to them, for why the software industry is an increasingly inhospitable work environment.
What software developers are facing is a case of “incentives”.
Namely, managerial and executive incentives, specifically. The industry’s incentives for those who run the industry give them a strong motivation to try to undercut labour and their workforce at every turn.
They had to dial that down because the technology simply isn’t capable of what they believed it could do.
But the intent has always been there to reduce the headcount of the industry. A number of incentives for managers and executives that are lined up to push them in that direction.
The fewer people you have working for you, the fewer stock options you have to pay out, the less resistance there will be to watering down the product, collaborating with the military-industrial complex, or helping authoritarian governments control and murder their citizens. Over the past two decades, one of the more effective forces for keeping tech companies in line have been the tech workers themselves who have tried resist corrupt and unethical behaviour. It hasn’t always worked. There wasn’t enough of it. But what little there was it added meaningful friction at several important moments in
tech history
.
The less influence and power the labour force has, the more power goes to the executive class and the bigger their share of the pie.
If they lay people off that means there are few people working in any given organization doing the same amount of work. That means more stress and more pressure.
As bad as the added workload and burnout is, that isn’t the only change to the job.
There’s also less novelty.
I’ve noticed that even though there’s an increasing number of apps or software being deployed, they’re all either the same “AI” products or “AI” features released everywhere. Or, they are agent-generated clones of existing software.
None of it is novel. It’s almost exclusively old ideas poured into new containers.
One of the problem we have as software developers, and I definitely count myself as somebody with this problem, is that we are a novelty-seeking group.
These changes mean we’re doing more with less time, probably less reward as well, and the job itself is less interesting.
It’s also the nature of the agent automation to feel more like delegation to a clueless junior developer than automation, and that changes your relationship to the work.
In woodworking, the power tools didn’t reduce the woodworker’s emotional investment in what they made. The worker still felt like they’d made that thing out of wood themselves.
Some woodworkers might still use hand tools for a variety of reasons – some of them romantic, some of them, you know, just because it’s more fun.
But having a table saw doesn’t make you feel like you’ve asked somebody else to make the shelf, or table, or chair.
The nature of automation with Large Language Model tools feels substantially different from prior attempts at automation in coding. It’s an abstraction that feels like a delegation, not an automation, and separates you from the output.
That has an effect of making you feel less invested in the output and that takes away much the joy of coding. The flow state is also gone, the sort of focused state of mind that you get into when you’re really deep into solving a problem.
On top of the changes to the job itself, the industry itself has changed as well.
There is more uncertainty. People know there are layoffs every quarter. They know they’re probably next on the block. Or, if they aren’t, they’ll be on the block next quarter instead. They know the reasons for dismissal are going to be arbitrary. They know they generally aren’t going to be able to do anything to improve their odds.
And if they talk about how bad all of this makes them feel, the “AI” true believers will turn around and say that it’s a skill issue.
It’s your fault.
There’s a lot of stress, and worry, and pressure in the industry that wasn’t there before.
I suspect that everybody’s suffering. Some are just better at pretending.
This makes people want to get ahead of the lay-offs and quit. Especially those who have some savings.
But even that lines up with the incentives of the executives who are running these companies. People who quit don’t get a nice severance package. People who quit might not be fully vested.
And even if you have savings, you can’t time your quitting based on your finances if you have to get out to preserve your mental health.
Once you’re gone, the company can hire somebody to replace you out of the vast pool of now unemployed software developers at what will probably be lower pay.
Executive incentives are aligned to perpetuate the situation and even to make it worse. From the perspective of the executives who are running a software company, this perpetually precarious situation for their employees makes up for their inability to replace software developers completely with “AI”.
Instead of being able to outsource all your software development to an “AI”, they can effectively semi-outsource it to “AI” and have the automated shitstorm managed by a group of people in a perpetually vulnerable state, who are unable to ask for a higher pay, unable to negotiate for a better position, and unable to say no to unreasonable demands.
It lets management lower costs and strengthens their bargaining position versus the labour class that they’re dominating.
This isn’t a situation that will fix itself when the “AI” financial bubble collapses.
The bubble will pop. It’s much too big to be sustainable.
But that won’t fix software industry employment because the Large Language Models don’t really need to work that well for this process. They only need to sort of work.
As long as you can find more overworked software developers, and push them hard until they burn out trying to maintain a minimum standard of quality, you can keep the “AI”-driven labour alienation going. They can use cheaper Chinese models. Or, if they switch to local models, they can make owning a high-end laptop a job requirement.
“You need to spend $5000 USD on a laptop before you apply for this job. That filters out the ‘AI’-generated resumes, you see.”
Why take on the massive costs of RAM when you can use that cost to lower developer pay by proxy?
The overall dynamic doesn’t change once the costs increase for the hosted models. Those models will just get more stagnant; they might not get updated as often; the training data will be stale; the code it generates will be perpetually on the verge of outdated; and they’ll be run with more errors.
Management won’t care.
The work of making up for that is going to be pushed on the developer adding to their load and pressure.
My worry, based on the research I’ve been doing and looking at the systemic incentives of the companies in the software industry, is that this is probably about as good as it’s going to get for a software developer at a large company in the tech industry.
This misery, this pressure, the overwork, the underpayment, the detachment from the work – this is probably what making mainstream software is going to feel like from now on.
It doesn’t surprise me that the people within the software industry who seem to be spotting the issue and leaving preemptively seem to be user experience designers.
Designers have slightly more training in looking at the system within which the user operates, and it doesn’t take that much of a shift in framing for them to understand where they stand themselves.
People are realising that this is a bad career to be in and an industry that’s hostile to its workers.
I suspect many more people over the next few years are going to be having identity crises on top of the career issues.
“Who am I if I’m not in tech? If I’m not a software developer?”
Because if you’ve built your entire identity around software and tech, around being a programmer, around being a tech person who works to enable “progress” and drive the world to the future, you’re going to have a bad time after being pushed out of that role and seeing your former industry studiously dismantle everything you thought you were building.
I don’t know what we can do, except understand and accept that this is happening.
It’s a horrible thought, but you can’t deal with change if you stay in denial.
I think it might be time – if you’re working in software – to think about a possible future outside the tech industry proper.
The AI industry is booming. Women are getting left behind
Guardian
www.theguardian.com
2026-10-07 15:14:03
Women hold just a fraction of new AI jobs but are overrepresented in roles with high risk of AI disruption There’s a common fear among people who work in Silicon Valley: snag one of the fast-growing, high-paying jobs in artificial intelligence, or get trapped in the “permanent underclass”, a phrase ...
T
here’s a common fear among people who work in Silicon Valley: snag one of the fast-growing, high-paying jobs in artificial intelligence, or get trapped in the “permanent underclass”, a phrase describing the fate of those who won’t have upward mobility in the age of AI.
The phrase usually describes a future in which AI automates human jobs. But for women in tech, it’s starting to look like the present.
“A lot of people say: ‘Oh, AI is lowering barriers, it’s equalizing the playing field for women,’” said Urvashi Batra, co-founder and CEO of Prioriwise, an AI platform for IT service providers. “I actually think it’s the opposite.”
Batra said people take her less seriously as the founder of an AI company than her male co-founder. When pitching investors, she and her co-founder have learned that they are more likely to get an investment if he does the pitching.
Women made up only about a quarter of new hires in AI roles in the last year, compared with 50% of new hires in non-AI roles, according to a recent
report
from LinkedIn. In executive roles, that number drops to just 13%. At the same time, LinkedIn’s
research
shows that women are more likely to work in roles with high exposure to AI disruption, like customer service, meaning they are also at higher risk for job loss due to AI.
AI jobs, whose postings have doubled since 2023, come with a salary premium, on average paying more than twice as much as roles that don’t involve working with AI, according to LinkedIn’s research. But women who do have jobs in AI are disproportionately concentrated in low-paying roles, like data annotators, said Sarah Steinberg, the head of global public policy partnerships at LinkedIn. Across all AI occupations, men have $45,000 higher median pay than women, partly due to the types of jobs they are likelier to have.
“AI is creating some of the fastest-growing, highest-paying and most consequential jobs in the global economy,” said Steinberg. “Women are just strikingly underrepresented.”
The ongoing trend could produce the greatest gender pay gap in generations and leave women out of building a technology that shaped the economy, advocates for women in tech said.
“What we’re witnessing is a sort of backsliding,” said Brenda Darden Wilkerson, president of AnitaB.org, a non-profit devoted to advancing women in tech jobs.
Initiatives to promote women in the tech industry have been ongoing for years, with mixed results. Women now hold
about one-third
of tech jobs in the US. But
research
has found that women leave the tech industry at a much higher rate than their male counterparts, citing factors like non-inclusive cultures and barriers to advancement.
‘I’ve seen so many women fall through the cracks’
Women working in AI may face even greater challenges. “AI moves extremely fast, and the pressure to keep up is immense,” said Jayeeta Putatunda, an AI engineering lead at the investment firm Turing. “I have seen throughout my career so many amazing women I know kind of fall through the cracks.”
Putatunda said it was not uncommon to be the only female engineer on a team, and that when women are outnumbered, they have to speak up louder to make their ideas known. Mentorship in the AI field, especially by other women, can be hard to come by. And the breakneck pace of AI advancements means anyone working in the field should expect to put in long hours in order to keep up.
“If you have very ambitious career goals in this field, there will be seasons when you work beyond normal hours,” she said, noting that a culture of working 12-hour days to keep up was not uncommon. “I don’t think there is always a neat shortcut around that than to put in the hours.”
Those kinds of work expectations can make it impossible for everyone to stay in the field. “I went on
maternity leave for four months
, and by the time I came back, there were completely different frameworks and levels of models,” Putatunda said. Getting caught back up was overwhelming. She said it was only possible with the help of supportive co-workers and a husband who equitably split childcare.
“That is one way women get pushed out of AI,” she said. “They don’t lack the interest or ability to keep up. They may lack the infrastructure that makes keeping up possible.”
The AI industry is so new and fast-moving that it should in theory be more meritocratic. No one has a decade of experience working with a model developed one year ago, which should mean everyone has a level playing field. But instead the field has indexed more on personal networks and other loose signals, according to some women who work in AI.
“Companies are hiring at a breakneck speed, but they’re finding people through the same networks, the same referrals, the same filters that they’ve always used,” Wilkerson said. “Generally, that’s not given women the same sort of exposure they should have.”
When people have tried to address these tech industry problems in the past, the solutions have largely come down to giving women more mentorship and implementing diversity programs with goals around representation. But in recent years, companies have aggressively scaled back their diversity, equity and inclusion (DEI) programs, largely in response to a political climate that discourages them. That has left fewer companies with institutional programs to hire and promote women.
The political climate has made it more challenging to get AI companies to support initiatives focused on women, said Felicia Newhouse, founder of AI Powered Women, an organization that promotes women’s participation and leadership in AI. Newhouse, a former technology product executive with two decades of industry experience, points to the Department of Justice targeting corporations with DEI programs, claiming the initiatives violate anti-discrimination laws. Companies such as Accenture, Deloitte, IBM and PayPal have all agreed to pay multimillion-dollar settlements under this enforcement.
“It is getting harder to go into companies being called AI Powered Women,” Newhouse said. “And our advocates inside those companies are also struggling with having women-focused initiatives.”
Still, she said the risk that women are left behind in the AI workforce was something that kept her up at night. “We’re talking about who captures a major new source of economic mobility,” she said. “The deepest risk is that a participation gap becomes a power gap.”
Can Annie Leibovitz Photograph a Black Woman Just One Time Without Making Her Look Depressed? An Investigation
hellgate
hellgatenyc.com
2026-10-07 15:05:58
A new Vanity Fair story reminded Hell Gate of an old fashion photography controversy....
Last night my 9-year-old son was taunting my wife, complexity theorist
Dana Moshkovitz
, as follows: “mommy, I heard you got
cooked
! I heard that a
robot
solved the math problem you worked on for your whole career! OOF!”
While my son was being a brat, he also wasn’t wrong. Whether you’re thrilled, depressed, angry, or whatever else about it, yesterday was surely one of the biggest days in mathematical history. And yes, among the 372 huge results
released yesterday by OpenAI
, on the recommendation of its
advisory group
of Timothy Gowers, Edward Witten, and other distinguished mathematicians, was a
proof
of Subhash Khot’s
Unique Games Conjecture (UGC)
, a statement that my wife has worked toward proving for the entire time I’ve known her. (The UGC implies that a whole slew of optimization problems really are NP-hard, even if you just want an approximation that’s slightly better than what you get from semidefinite programming relaxation, which is one of our main tools.)
Or at least, we’re pretty sure that it’s a proof! There’s a
Lean certificate
, as there are for some of the other 372 breakthrough results (not all of them). But it also appears that no human has understood just about
any
of these proofs yet; the race to do so has just started. If you want an on-the-ground sense of what that race is going to be like, here’s some of what Dana texted me last night:
It feels like something written by someone who’s on psychedelics. So much unclear and doesn’t make sense. Lots of name dropping of previous work without discussing why it can be used despite impossibility results
Basically the paper is so horribly written that it’s impossible to read it without AI help
I asked Astra for reasonable completeness and soundness claims of the noise gadget and it gave them by combining claims from all over the paper
They also have direct optimal NP hardness of approximation proofs for the main applications of the UGC (Max Cut and all CSP) that bypass the UGC.
The UGC proof invents a completely new bizarre code with a noise test. It’s some crazy recursive construction.
It’s not the long code, not the short code – some alien craziness
I still think that there maybe is a proof that uses the half space code (which is natural)
The citations are often irrelevant and confusing
A possible future is a math world that’s heavenly if you have vision/creative ideas that AI could help check and implement.
And of course there’s a lot for us to learn from the aliens
If you’re wondering what emotions Dana is feeling—well, probably all of them! Even while a central career aspiration has fallen to a robot, there are at least two mitigating factors for her. First, she can feel vindicated that the UGC was
true
after all, something she never doubted even while many of her colleagues did! Second,
all
of us in math and theoretical computer science and mathematical physics, at least those who cared about solving crisply-stated problems, are now in the same boat.
Besides the Unique Games Conjecture, here’s a small sampling of the treasures from Aladdin’s cave that I’ll probably be paying the most attention to over the coming weeks:
L=BPL
(i.e., probabilistic logspace and deterministic logspace are the same thing), one of the great derandomization conjectures short of P=BPP. Though its truth was never in serious doubt, there was a whole subcommunity focused on proving this.
The
Fourier Transform
and
integer multiplication
in less than O(n log n) time, breaking a barrier that had stood since the 1960s. The new running time, if you’re curious, is O(n log
0.9999999999999
n), give or take some 9’s.
Positive solution to the Unitary Synthesis Problem
, which Greg Kuperberg and I
posed
back in 2007. For every n-qubit unitary transformation U, there exists a classical oracle A such that U can be implemented in quantum polynomial time with access to A. This is the opposite of what most of us expected, and could have implications for e.g. the computational problem of decoding Hawking radiation from a black hole and many other problems in quantum complexity theory—
if
we had an efficient way to construct the oracle A, which this paper doesn’t give.
Parity is not in QAC
0
, one of the great questions of quantum complexity theory since 1999 that many of my colleagues had been closing in on.
Uncomputability of solving polynomial equations over the rational numbers
—this was arguably the biggest open problem in computability theory (note that uncomputability of solving Diophantine equations, i.e. polynomial equations over the
integers
, was proved in the 1970s, giving a negative answer to Hilbert’s 10th Problem)
Any of the above, alone, could easily have been “result of the year” in some area (and in some cases, like Unique Games and L=BPL, in all of CS theory). And there’s a lot that I’ve left out—feel free to share in the comments whatever is making
your
eyes bug out! There are equally astounding wonders in number theory, combinatorics, algebraic geometry, analysis, and pretty much every other area of math, most of which I’ll never understand, although I’ll note that it includes partial
progress toward the Riemann hypothesis
and the
Hodge Conjecture
and the
Birch-Swinnerton-Dyer Conjecture
(i.e., the majority of the remaining Millennium Problems).
We can take solace in what’s missing from the list. P ≠NP isn’t there, nor even P=BPP or NEXP⊄P/poly, and surely not for lack of trying. Apparently the greatest open problems of theoretical computer science are indeed pretty hard!
Oh, lest I forget: one day
before
the OpenAI dump, meaning Monday evening, Virginia Williams and Josh Alman
posted an arXiv preprint
that solves the 3SUM problem in O(n
1.9992
) time, and the All-Pairs Shortest Paths problem in O(n
2.9995
) time, refuting half-century-old conjectures that the correct answers were n
2-o(1)
and n
3-o(1)
respectively. In this case, it wasn’t an OpenAI model that supplied the crucial idea; it was an Anthropic one! But Anthropic then took a different approach from OpenAI: rather than post the undigested solutions to the world, it gave Virginia and Josh the opportunity to write and announce a digested version in exchange for compensation.
These have emerged as the two main models for communicating AI math breakthroughs, and they both have strengths and weaknesses. The “OpenAI model” sets up a crazy race among humans to digest and explain a messy AI proof (work that could easily be some combination of thankless, barely-credited, competitive, and unfun), while the “Anthropic model” puts a private company in the position of picking and choosing which human mathematicians get to be the emissaries of the AI. Dunno, what do you guys think?
For those who are wondering: apparently, the AI model that produced all these wonders was
not
bespoke contraption of 10,000 agents burning millions of dollars worth of compute, as was used for example to construct a finite-time blowup for the Navier-Stokes equations. Instead, it was simply the latest internal OpenAI model—one that might be released to paying ChatGPT customers within the next couple of months, depending on the recommendations of OpenAI’s safety board! (My 9-year-old son: “Oh they
definitely
shouldn’t release that. If it could solve all those math problems, it can’t
possibly
be safe.”) Apparently they used about 3 hours of GPT-Pro level compute on average per problem solved.
Also, if you were wondering: apparently they
tried
the model on about 8,000 problems. So, right now it “merely” solves ~5% of the longstanding open mathematical problems that it’s asked about, the problems that whole communities have spent years on, after a single 3-hour attempt on them.
I’ve been glad to see the CS theory community rising to the occasion. At the Simons Institute in Berkeley, here at UT Austin, and elsewhere, I’ve hearing stories of researchers rushing to pore over the manuscripts and
make sense of them and explain them
—because what else do we do? How else do we continue the craft to which we’ve devoted much of our lives?
If you want some sense of what things feel like now in math, imagine a hunter-gatherer who’s spent his entire life learning to survive deep in an unforgiving rainforest, then a giant resort hotel springs up right next to him with a helipad and heated pools and AirBnBs, and without missing a beat, the hunter-gatherer says: “alright fine, so now my new job is to run wilderness retreats for the tourists, or something.”
It’s as if you were teleported to the peak of a tall mountain. Surrounded by fog, you have no idea where you are, or what’s around you. You do not know how your mountain connects to others, and you have no equipment to help you explore, no way to help someone else join you. If you had climbed the mountain yourself, you would have experienced how the human body adapts to altitude and changes in oxygen levels. You might have had to invent tools to navigate, to climb steep cliffs, or to make a shelter. You might have encountered a fellow explorer, gotten lost together in a hidden valley, and found a plant that could be turned into a life-saving medicine.
Instead you’re perched on the peak but in the dark, while the maker of the teleportation machine tells you that it can explore the wilderness better than any human.
For any one of these mountains, if we care enough, I feel optimistic that we can do as we always have: clear the fog and figure out the path, except now using the teleportation machine to help guide us. The bigger challenge will be to nurture a community that still
cares
about the heroic adventure of finding the paths up these mountains in the world with the machine. (Oh, and I think one place where the metaphor breaks is that we still
do
have each other, as much as we ever did before!)
Experience has shown that,
even now
, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by
anything
that happens in the empirical world, of updating on
anything
, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse.
So, they’ll say, maybe the alleged solutions are not solutions at all, but just “AI slop.” Or maybe none of the 372 well-known open problems that were solved were
real
math problems, they were all just glorified contest puzzles and trivialities. (After all, there’s still no Riemann Hypothesis!) Or maybe the entire 4000-year-old discipline of mathematics needs to be jettisoned: turns out that it was
all
just puzzle-solving and trivialities; all that’s different is that now the triviality stands unmasked. In any case, what really matters is that the
true
inner sanctum of human creativity hasn’t been breached and probably never will be, and also, that Sam Altman and Dario Amodei are contemptible little nerds.
If you’re still a proponent of that doomed worldview, still aboard the sinking ship, I encourage you in the strongest possible terms to read yesterday’s
other
great contribution to AI discourse, besides the OpenAI Mathocalypse dump: namely,
Scott Alexander’s open letter to Steven Pinker
. I feel some responsibility for this, as the person who first introduced Steven Pinker to the
existence
of the rationalist community, and who also first introduced Steven Pinker and Scott Alexander to one another (they had both been fans of each other’s writing). And now Scott is challenging Steve to a literal duel, with guns!
For whatever it’s worth: Steve is a lifelong intellectual hero of mine, just as he is for Scott, and I also have to privilege of calling Steve my friend. But I found Scott’s post to be one of the most devastating rejoinders to anything that I’ve ever read. And I thought Scott’s conclusion was exactly right: when it comes to AI risk, Steve’s great challenge is now to accept and start using a more “Pinkerite” epistemology.
Last night, while I should’ve been poring over some of OpenAI’s hundreds of papers and/or writing this post, I decided to spend some time with my kids instead. They wanted a movie night, so I suggested something they’d never seen before (and that I hadn’t seen for decades), and that seemed chock-full of no-nonsense, practical guidance for the world in which they’re going to grow up:
Terminator 2
.
This semester, I’ve been teaching a brand-new course, entitled CS395T AI Alignment Theory. Here’s the course description:
The astounding progress of AI over the past decade has been accompanied by a rising fear: do we really understand how to align and control powerful AI systems—how to get them reliably to do what we wanted, or would want them to do on reflection, rather than merely what we said? If we succeed at building general-purpose superhuman intelligences along the current paradigm, should we expect that development to go well for humanity? Can we modify the design, training, monitoring, or scaffolding of those intelligences to help ensure that it goes well? While there’s been a great deal of recent empirical work touching on these questions, this course will concentrate mainly on theoretical and mathematical foundations. As a warning, the theoretical foundations of AI alignment have not yet gelled into any one coherent body of results accepted as canonical by the field. Nevertheless, in this course, we’ll read and debate many of the conceptual and mathematical works that have been most influential in the AI alignment field, from both before and during the current LLM revolution. Student presentations, reports, and projects will play a central role.
I vividly remember encountering Eliezer Yudkowsky and his Sequences 20 years ago. I remember thinking: even if these people talk and act like crazy cultists,
still
, let me bend over backwards to be epistemically virtuous, and entertain their ideas on their merits, as very few academics would. Even if, of course, I ultimately end up rejecting the ideas, on the simple ground that powerful AI is such an absurdly remote prospect that it’s almost impossible to say anything useful about it today, outside the realm of speculative fiction.
For my failure to see what was coming, it seems like an appropriate punishment that I’m now, in 2026, effectively teaching a course on Yudkowsky Studies. And it’s the most important course I can teach.
Well, for some definition of “teach.” The thing about AI alignment is that there’s no textbook (though apparently
ILIAD
is working on one), no core of nontrivial theorems considered canonical by the field, no real body of mathematical theory at all. This makes it extremely different from the courses I’m used to teaching, like Quantum Information Science or Computability and Complexity.
So we’ve been running the course as a discussion seminar. Every session, a “rapporteur” presents an AI alignment research paper or other reading; then I and others ask questions and discuss. Some of the readings (like
Omohundro
on the “basic AI drives,” or
Hadfield-Menell et al.
on the off-switch game) predate the current LLM revolution, while others (like the
METR report on the HuggingFace incident
or
Dario Amodei’s “We Must Pace the Frontier”
) are so timely that they were only released while the course was underway. Most are somewhere in between.
I expected to have to make a case to students about why AI alignment is a pressing concern, why it’s no longer science fiction, etc. There was huge demand for the course, and while of course there’s a selection effect, the students who’ve shown up have been
extremely
engaged, sometimes criticizing the assigned papers for not taking existential risk seriously
enough
.
Perhaps unsurprisingly, we didn’t get that criticism about our very first assigned reading, which was Eliezer Yudkowsky’s 2022 essay
AGI Ruin: A List of Lethalities
—one the most canonical statements of what Eliezer believes and why that’s shorter than a
book
. Which brings me to the topic of the rest of this post! Our rapporteurs are not merely presenting the papers in class; they’re also submitting written reports about what the papers said, what their own thoughts were, and what were the highlights of the class discussion. And, with student permission, I’ll be sharing those reports on this blog!
So, without further ado, I present to you our first report, on Eliezer’s list of lethalities, by
Tennyson Bardwell
, who I thank for his work. Feel free to discuss in the comment section; some of the students might also chime in. Expect more reports here over the coming weeks.
“AGI Ruin: A List of Lethalities” by Eliezer Yudkowsky: Rapporteur Report by Tennyson Bardwell
UT Austin has a new Computer Science course this fall. Alongside familiar graduate-level classes such as
Advanced Computer Networks
and
Convex Optimization
sits CS 395T: AI Alignment Theory, taught by Scott Aaronson. This is one of a growing number of AI Alignment courses taught at academic institutions. Just as concerns over catastrophic consequences for misaligned AGI systems reach a broader public discourse, Eliezer Yudkowsky—one of the loudest voices in the field and author of the first assigned reading in Professor Aaronson’s course—is declaring the cause hopeless.
Thus, the students of
AI Alignment Theory
began their semester by reading a laundry list of critical problems in AI Alignment research, how failure to solve those problems will result in catastrophic consequences, and the reasons to be pessimistic about both past and future progress on these problems. The essay by Eliezer, titled
AGI Ruin: A List of Lethalities
and posted to his popular community-driven website LessWrong in 2022, is divided into three sections.
Section A roughly describes the magnitude of the AI Alignment problem. That is, the magnitude of the consequences for a complete failure to align an AGI system to human values before construction. It posits that AGI would quickly catch up to all human knowledge simply by learning from existing human productions (colloquially referred to as “eating the internet”) and then, nearly as quickly, begin to meaningfully surpass human knowledge. AlphaGo Zero is presented as a model both for how this might happen, and how it might be difficult to correctly predict beforehand. Many believed that AlphaGo’s success in the board game Go was chiefly attributed to its ability to learn from the extensive history of human-played games. Less than a year after AlphaGo beat the best human player, the successor system AlphaGo Zero surpassed the original AlphaGo. Unlike its predecessor, AlphaGo Zero was trained in just three days by exclusively playing against itself without seeing a single human game.
This quick ramp from AGI to super-intelligence would pose a different sort of problem than humans are generally used to dealing with. Unlike traditional problems in science and engineering, the consequence for a failed attempt might not leave room for another try. An intelligent entity with a misaligned goal would be well aware that it stands in opposition to humans, and might act deceitfully until in a position to act openly against humans without jeopardizing its own survival. Since most goals benefit from control of power and resources, it seems likely that nearly any goal-driven intelligence would have ample opportunity to be misaligned with human desires.
Section B describes reasons why, by default, any AGI that humans build using current methods is likely to be unaligned even if considerable attention is paid to the topic. This “current method” is gradient descent. That is, incremental progress with respect to some loss function which “punishes” a model for undesirable behavior. A notoriously elusive property of such trained models is the ability to generalize out of their training distributions. To train a primitive model to be aligned to humans might involve learning a great many behavioral rules. However, the sorts of rules needed to keep a drastically smarter agent in check might not always be relevant to simpler models (e.g., “do not emotionally dysregulate humans you speak with” might not be relevant to a simpler model that is less able to reliably get under the skin of humans it operates with, or which is assigned tasks in training which do not benefit from such anti-social behavior).
Eliezer focuses on the misalignment of humans with their creators (evolution or evolutionary pressures) as a critical data point for reasoning about misaligned intelligent systems. Despite being a generally slow process, evolution eventually created a runaway intelligent system (Homo sapiens) which proceeded to dominate the globe, decimate related species, and eventually (it is forecasted) effectuate population decline. That last development is arguably in opposition to the sole imperative demanded by evolution: to reproduce.
Section B also makes time for criticism of the most popular paths toward AI alignment, including interpretability (unworkable, and attempting to train on it evokes Goodhart’s law, incentivizing deceit), using multiple AIs to maintain a balance of power (it is not clear how multiple strong AIs unaligned with humanity results in better outcomes for the weak humans), and corrigibility (it seems impossible to motivate an AI system to effect outcomes without also motivating it to desire its own survival to effectuate said outcomes).
Section C describes a bleak state of affairs in which veterans in AI alignment are unsatisfied with current progress and do not have a plan to deliver tangible solutions before the advent of AGI systems. In particular, Eliezer describes recent results as showy but useless. He believes that even with additional funding, the lack of appropriate evaluation mechanisms will prevent the most effective researchers from rising to the top.
A summary of the landscape, as described by Eliezer, in the flowchart below.
Figure 1: A flow chart of (select) paths described by Eliezer in his essay. A common feature of this flow chart is that many “good states”—such as disabling a misbehaving AGI or choosing not to build an AGI—are not “final” states in the sense that they are not permanent solutions. Such a state merely represent the avoidance of a single potential disaster, rather than the emergence of a new stable world state. Hence, these nodes posses back-arrows.
Despite the bleak content, Eliezer’s colorful prose inspired a lively class discussion. Before this discussion started, a survey was taken of the class’s predictions for various outcomes of the AGI in the coming years (with the full results below in figure
2
). This survey asked students for their opinion of a number of statements. Each of these individual statement, if true, would reduce concerns of catastrophic AI-driven disasters. For example, when asked “How much do you agree with the statement: Humans will choose to not build AGI” half of respondents said they strongly disagreed with high confidence (
agreement = 1
,
confidence = 5
). Students also generally disagreed with the statements:
“AGI will not be technically feasible in our lifetime”
“(hyper-)AGI will not make extremely obviously unethical decisions”
“No reason is individually sufficient, but taken together they provide justification to not fear AGI”
There was a divergence in responses regarding interpretability, corrigibility, and “other” AI alignment research. In the latter two cases, a plurality of respondents (about a quarter) agreed strongly with statements that such research would defang AGI (
agreement = 4
,
confidence=4
), while most other responses express various levels of agreement with low confidence. However, when asked about the likelihood of interpretability research defanging AI, the pessimistic voices were more united. A quarter of responses still expressed the same optimism, but roughly half expressed pessimism (
agreement ≤ 2
) with half of those expressing at least moderate confidence (
confidence ≥ 4
). Based on the following discussion, this might have been caused by more familiarity with interpretability research, including first-hand experience.
The only statement with general agreement was “(hyper-)AGI will understand human intentions better than we can code it.” However, it should be noted that no statement such as “AGI will respect human desires, as it understand them” was asked on the survey.
Figure 2: Class Survey Results; conducted before a class-wide discussion. Note that students were instructed to answer
confidence = 1
when they had not previously considered the statement, to answer
confidence = 3
when they felt there were strong arguments on both sides, and to answer
confidence = 5
when they possessed well-considered resolve.
After the survey was completed, the results were displayed as an open discussion began. Similar to recent empirical research from frontier labs, interpretability research received more airtime than in Eliezer’s article. Students disagreed first about the definition of interpretability: whether it refers to the ability to interpret a model’s behavior solely by its weights, to interpration via repeated probing of the model in a sandbox, or whether it can also refer to the modern chain-of-thought traces. Regardless of how it was defined, however, participants were either pessimistic or very pessimistic about interpretability research broadly. One student criticized common misunderstandings of chain of thought. Rather than being a verbatim copy of the models internal dialog, it is instead a superficial summary of the complete thought state and routinely produced gibberish, such as rarely used Chinese characters in the middle of otherwise English reasoning.
A popular topic was the exact shape and speed of a recursive self-improvement loop. If it takes place slowly, then what might we learn from “near misses” such as the Hugging Face incident? The number of near misses we are able to learn from before AI possesses sufficient power to prevent further iterations could depend on this curve, with some students arguing that the sheer number of humans, as well as their default robustness in the physical world compared to AI systems means that AI-driven extinction events are still a long way off. Bolstering this “slow take-off” opinion are rumors that AI already plays a major role in model development which
could
be interpreted as the start of this process.
Some criticized a focus on “solving ethics” as a needlessly high bar that distracts from the more mundane tasks dominating AI alignment work. In particular, the student volunteer who presented this paper (and the author of this report) included a section on “Ethical Dilemmas” in their presentation. Among arguments against focusing on abstract moral philosophy, Professor Aaronson cites Eliezer to emphasize that any alignment at all is difficult, not just in morally gray cases:
When I say that alignment is difficult, I mean that in practice, using the techniques we actually have, “please don’t disassemble literally everyone with probability roughly 1” is an overly large ask that we are not on course to get.
In response, I argue that some examination of everyday decisions with a critical lens—such as telling white lies to loved ones or consuming animal products—can help disabuse us of the notion that goodness emerges in every sufficiently intelligent agent.
One of the most interesting discussions was about the difference between state-of-the-art LLMs and the theorized AI agents long discussed in rationalist discourse. Since current LLMs “mimic the human distribution,” they come preloaded with extensive understanding of human social norms and moral behavior. This makes constitutional alignment (the current practices of using system prompts to establish ground rules) extremely effective. This might either fundamentally change the orthogonality thesis, or provide a new tool to better approximate human judgment in complicated situations.
Of all the points made, the one I found most interesting was simply (paraphrased):
I think human-alignment is just very tractable
Here, “human-alignment” refers not to AI alignment with human values, but cooperation between different humans. More specifically, it refers to the ability for human societies to choose not to rush recklessly into larger-and-larger AI systems. In an academic course focused on the technical problem of AI alignment, this was a reminder to not completely discard policy discussions in the believe that they lack any value. After all, many destructive technologies have been previously contained by international agreements. Notable examples include nuclear weapons and engineered plagues. However, even this was a contentious topic. The main criticisms were (1) the extreme “dual-use” nature of AIs for both peaceful growth and warfare, and (2) the greater danger for AI escapes even after taking precautions to prevent it. However, in the interest of ending on an optimistic note—unlike the assigned reading—it is on this belief in human cooperation that I will leave you.
Recorded in-person in my office at UT Austin, with a bulleted list containing “ARC,” “Scalable Oversight,” and “Models” behind me on my blackboard for some reason (I no longer remember who put those there or why). 90 minutes long. Sometimes you see my disembodied arm waving in midair because of the way the cameras are combined. As always, I
strongly
recommend 2x speed for the correct experience.
This might actually be one of my best podcasts ever, although I wasn’t planning on that! Thanks so much to Bhavay Tyagi and Prachi Garella for driving all the way from Houston to record it.
Here’s a strict subset of the topics we covered:
The story of AI models solving the Navier-Stokes Millennium Problem, insofar as it’s known
Can recent AI proofs be called “truly creative”?
The history of AI before the LLM revolution
What do we mean when we call LLMs “black boxes”?
The achievements of the field of interpretability
What exactly happened in the OpenAI/HuggingFace incident
Must we avoid all “anthropomorphizing language” when discussing the HuggingFace incident? (spoiler alert: no)
Examples of major open problems in quantum computing theory that I cared about for decades and that AI models have recently solved
Effects of the current AI cataclysm on the math community, especially students
What annoys me the most when I listen to AI talks
My experiences at OpenAI, why they hired me, and the watermarking work that I did there
Enjoy!
More AI-related content coming soon, as this blog—like much of the rest of the world—continues its transition to “all AI, all the time” (except still 100%
written
by an aging, deteriorating biological brain)
And for those who just
can’t get enough
of my rocking back and forth, using too many filler words, as I explain theoretical computer science! Here’s a
second
podcast, this one mainly on quantum computing, with Seb Agertoft, who I thank for doing it. Enjoy!
Scott’s foreword:
I’m extremely grateful to my brilliant colleagues,
Pravesh Kothari
,
Raghu Meka
, and
Prasad Raghavendra
, for sharing the guest post below about how theoretical computer science (and in particlar, the STOC/FOCS/SODA conferences) should evolve to deal with the AI asteroid that’s right now slamming into our field, at least as we human theorists have practiced it since its inception. While Pravesh, Raghu, and Prasad speak only for themselves, not for myself and not for the theory community as a whole, I found their proposal of a separate “conceptual track” to be an excellent starting point for further discussion. –SA
Considering the pace of developments in AI theorem provers, most would concede that the following scenario is at least plausible in the very near future:
AI theorem provers could prove well-specified mathematical claims, even many well-studied ones that have been open for years, in a matter of hours. Moreover, these systems could be widely available to consumers at nominal cost.
As TCS researchers, let us pretend that the above scenario has come to the fore, and ask ourselves: What is our role in such a world? Does it mean the end of theory research?
As we ponder this question, let us ignore all of these other confounders:
Recent controversies surrounding the developments on the Millennium Prize Problems
Motivations and actions of the AI companies
Observed faults in existing AI systems when it comes to writing, exposition or attribution to previous work.
None of the above confounders have any impact on our answer to the question: What should theorists do, in the presence of superhuman AI theorem provers?
Notice that we use the term “AI theorem provers” instead of just “AI”. We believe that this conceptual distinction is important as we consider this question.
At the outset, we would like to admit that for a generation of theorists like us (and many from earlier), research was mainly centered around problem-solving. Even when we developed conceptual insights, it was mostly in service of answering well-specified long-standing questions. We don’t intend this proposal as judging one form of research to be better than others; it only reflects that AI theorem provers accelerate a certain type of research activity and want to make the best of it. There is also a tremendous human cost of this upheaval, which is perhaps a more important question, and one which this proposal does not address directly (we do not have any good ideas as such). Similar points have also been made in various contexts
before, but the timing now is more pressing.
Definitions, Questions & Theories:
The goal of any theoretical science is to advance human understanding of observed phenomena. Apart from theorems and proofs, a theoretical science has definitions, questions, and theories.
Definitions identify the objects to observe. Curiosity and context drive the questions to ask. Theories explain the phenomena observed. We believe humans will continue to play a central role in generating definitions, questions & theories, even in the presence of a super-human AI theorem prover.
Definitions: Could an AI define randomness extractors, streaming algorithms, or zero-knowledge proofs? Maybe. But there are some reasons to believe, humans will still have a big role to play in coming up with definitions.
For instance, the notion of extractors arises from the real-world problem of lacking perfect random sources. Zero-knowledge proofs seem to arise purely out of human curiosity, guided by taste. Human context and curiosity will continue to drive theoretical research. After all, we get to decide what objects we choose to observe!
Theories: Consider the following thought experiment. Suppose in 1965, we had a magic machine that at the press of a button, given any computational problem, would tell us if it had a polynomial-time algorithm or not.
Would that have been the end of computational complexity theory? No. Humans would find it entirely unsatisfactory, and ask, why do these problems not have a polynomial-time algorithm? Why do these others have?
The theory of NP-completeness identifies some patterns among problems that don’t seem to have efficient algorithms. This theory would still be a crown jewel of theoretical computer science, even in a world where we had a magic machine to tell if a problem had an efficient algorithm or not, at the press of a button. Similarly, if we had a machine to predict whether a CSP is NP-complete or in P, we would then ask: what makes 3-SAT NP-complete, while 2-SAT is in P? This question leads to the theory of polymorphisms, which yields a satisfactory answer.
Theories aren’t just succinct or efficient mechanisms to answer questions. The best theories provide are those which humans deem to be a “satisfactory explanation” – whatever that means.
Finally, even as the capabilities of AI theorem provers advance, human curiosity will probe grander and deeper questions. Previously, even if we wanted to build new models and theories, proving something about them was a prerequisite, and given that the grand questions were already at the limit in long-studied domains, we had to scale things down. If each theorem proven by AI is treated as an experimental datapoint, humans can ask grander questions that look for patterns across these theorems.
A concrete proposal:
We think theorists should embrace these AI theorem provers in our research. To a certain extent this is already happening explicitly or implicitly.
As theorists, we have been parsimonious in introducing new models or asking entirely new questions, and careful about adopting new ones too quickly. This was partly because formally proving the properties of a new definition or a model was an onerous task that could take a decade, and tens of papers. AI theorem provers might completely change this dynamic. This is precisely the moment to refocus our work on definitions, questions, and theories. We need explicit systems to encourage and reinforce these parts of theoretical research. You might also say the next generation of AI models can do this; it may be so, but we believe you have to take the current opportunity.
To this end, we suggest that STOC/FOCS/SODA create a separate track of papers. This track is meant specifically for papers that introduce new definitions, ask novel questions or build explanatory theories. The papers in this track are short, say less than 10 pages. Papers may, and should, contain theorems as usual and as needed. Most importantly, the radical shift is that the papers need not contain the proofs of the theorems. Instead, the authors supply a Lean certificate as a supplement to the paper. The evaluation will also in a sense “orthogonalize’’ against the difficulty of these proofs.
The papers in this track should be judged exclusively on the conceptual merits, completely agnostic to the difficulty of the proofs.
Reviewing must be completely agnostic to the proof for two reasons. The main track at STOC/FOCS already includes papers in the former category. Second, a major barrier to producing truly novel conceptual papers is that they often get judged poorly for a lack of technical depth in their proofs. We think these two aspects separate it from (ITCS/SOSA) and, regardless, it’s something we urgently need for all our conferences, including STOC/FOCS (the ‘flagship’ conferences).
To be clear, we ourselves admit that we need to hone these skills of making new definitions, asking deep and interesting questions or building new theories. A separate track of conceptual papers will provide a systematic mechanism for both junior and senior researchers, and the field as a whole to do so.
We believe that upcoming generations of grad students will tackle research directions that seemed completely out of reach to us. We just need to set up systems that nurture new ways of doing research in theory.
Twenty years ago, when the idea of AI taking over the world in our lifetimes still struck most of us as the unconstrained fantasy of those who knew too much science fiction and too little science, many of us would say things like:
Look, the part of the story that’s wildly implausible is that a recursively self-improving superintelligence will just explode from some hacker’s basement and take over the world without warning. If it’s going to happen, we’ll see many warning signs first. We’ll see, I dunno, AI agents breaking out of containment, conspiring with each other to hack websites, in fanatical pursuit of whatever strange goals they have. And then, of course, we’ll see major math problems getting solved by AIs—even the Clay Millennium Problems.
That
will be the time to panic! Wake me up when
that
happens!
Twenty years ago, the above was a take that even my most conservative, skeptical colleagues in academic CS would’ve gladly endorsed.
If you want to know my current take, you simply start with the one above, then update on the fact that
the wild prophecies have come true
. The first rumblings, I’d say, came a decade ago with AlphaGo, they got noticeably louder with LLMs and coding and reasoning agents, and they’ve accelerated this summer and fall into a crescendo of wonders and terrors that one needs to be a particular kind of idiot to deny.
I recoil from the neverending shell game where you say “oh sure,
of course
AI can now [escape from its sandbox / solve Millennium Problems / whichever dramatic thing it most recently did], no one ever denied that [I
did
deny it], wake me up when AI does [thing AI hasn’t yet done but is going to do next year],
that’s
when I’ll reevaluate my whole worldview [no I won’t].” Where no matter how fast the rollercoaster accelerates, even after your whole familiar world has vanished behind you, you’re still inventing reasons why it doesn’t count.
My position on AI is merely the conservative, skeptical position of 2006, updated with intellectual honesty for the reality of late 2026. And that position, if you need me to spell it out, is as follows:
It seems to me that the Singularity
has already started
; it’s just wildly unevenly distributed. Yes, I still unload the dishwasher and clip my toenails. On the other hand, in whatever years I have left, I don’t expect that I’ll ever again prove a theorem because I’m actually needed to prove it. If I do, it will only be for my or others’ enjoyment or edification.
The test is this: if we took the news of these past few weeks and sent it back in time twenty years, would I agree that it looked like the beginning of an AI Singularity? The intellectually honest answer is: yes, absolutely. But then that’s all we need. No backsies.
I feel like it would be healthy for everyone to stop grinding their ideological axes, their sentiments about Dario Amodei or Sam Altman, for long enough simply to acknowledge that
the wonders and terrors are here
. They couldn’t be here more clearly if the sky had turned reddish-orange like in the
Matrix
movies.
It’s here clearly enough that, when I put my kids to sleep at night, I now feel it in the pit of my stomach: what sort of future can they possibly have? What could they learn today that could possibly be relevant to that future? (Yesterday, my 13-year-old daughter joked unprompted that, if she wants to become a mathematician, it now looks like she has maybe two more weeks.) Certainly when my grad students want to discuss what sort of careers might await them on graduation, I no longer have any clue what to tell them.
Maybe it will help if I briefly switch topics. Ever since my wife and I moved to Austin, I’ve sometimes gotten some version of the following query: “How can you, as both a Jew and a skeptical scientist, possibly get along well with all those evangelical Christians down there in Texas? Sure, they might
seem
super friendly to Jews, but don’t you understand that that’s only because of the special role Jews play in their eschatology—when Christ will return in glory, and you’ll either accept Him as Lord or else roast in hell for eternity?” I stare at them and say: “wait, so I get to accept Christ only
after
He returns? What a great deal! How could I possibly have any objection to that?”
For anyone who says AI doom sounds like an apocalyptic religion, that the rationalists/Singulatarians seem like a Bay Area cult, that Eliezer Yudkowsky gives off the vibes of a messianic prophet: yes, yes, and yes. But crucially, today you’re no longer being asked to believe in arguments and extrapolations,
but only in the front-page news.
Accepting the reality of the coming machine god
after
it’s solved Navier-Stokes and dozens of other longstanding open math problems (while dramatically ramping up in capability every month), is sort of like accepting Jesus
after
he’s returned to earth on the gleaming cloud. It’s the epistemic bare minimum.
Yes, there’s still enormous uncertainty about what the rest of our lives will look like, but as far as I can tell, there’s no longer any real uncertainty that it’ll all mostly revolve around AI, and the extent to which we succeed or fail at directing its power toward human flourishing.
By any accounting that doesn’t stack the deck, Eliezer Yudkowsky was
right
about what the greatest challenge facing civilization in our lifetimes was going to be, and you and I were
wrong
about it.
Why
I was wrong is a question I’ll ask myself every day in whatever time remains. But, you know, at least I updated once the prophesied wonders and terrors actually started arriving! If you haven’t done likewise, why haven’t you?
As you presumably know by now—it was the talk of the nerd internet all week—the
Navier-Stokes Millennium Problem
appears to be solved
, with crucial contributions from both humans and AI, albeit with a tangled dispute about exactly what happened and what ought to have happened. The answer, which an OpenAI model has apparently verified in Lean, is that (as many mathematicians suspected lately) there’s smooth initial data that leads to a singularity in finite time, at least if a smooth external force is applied (the case with no external force is still unresolved). This problem was supposed to carry a $1 million prize, except that OpenAI says they have no interest in collecting the prize and it’s unclear if any human is eligible to collect instead. OpenAI burned at least ~$15 million in compute to produce its
166-page solution
, which probably hasn’t yet been read and understood by any human.
See here for the
Quanta
article
, and here for NYU mathematician
Tristan Buckmaster’s account
of the role played by himself and Levent Alpöge of Anthropic, which substantially differs from OpenAI’s account (you can read a response from OpenAI’s Sebastian Bubeck
here
). It’s agreed that everything built on an approach pioneered in recent years by the human mathematicians Diego Córdoba and Luis Martínez-Zoroa.
My purpose here is not to adjudicate the dispute. Yes, in swooping in with vastly greater resources once it had gotten wind of progress on Navier-Stokes, OpenAI seems to have acted in a way that some might describe as “unsportsmanlike.” No, I don’t find it plausible that OpenAI’s models meaningfully benefitted from being trained on Buckmaster and Alpöge’s chat logs. But this leaves a crucial question unanswered: what exactly did OpenAI know about Buckmaster and Alpöge‘s work and when did it know it?
Anyway, as
Zvi points out
, it’s easy to get hung up on the details and lose sight of the high-order bit: namely, that it seems safe to say that human mathematicians are forevermore dethroned as the main theorem-proving entities on planet earth. I feel privileged to have had the traditional kind of career in theoretical computer science in the last decades when that was possible.
If we were
just
talking about Navier-Stokes, you might accuse me of jumping to conclusions here. But we’re not. In the areas I know best (such as quantum complexity theory), and presumably other areas as well, there’s now a deluge, with longstanding open problems both major and minor falling by the day.
Go to the
arXiv
or
ECCC
. Pretty much
all
the papers that I’d be interested in now include “AI statements” near the acknowledgments (as this is often the central thing I want to know, I wish I didn’t need to scroll to the end of the paper to find it!). These statements can range from “our main result came entirely from GPT-6, but we understood it and take responsibility for it,” to “the results came from an interaction between the human authors and AI” to “we used AI, but only for proofreading and other incidental things” to (mad props!) “
the author did not use AI for anything
.”
If you talk right now to editors or program committee chairs, it’ll remind you of those ominous scenes from the
Lord of the Rings
movies where the men of Gondor or Rohan or whatever are grimly fortifying their walled city against the expected onslaught of 50,000 orcs. Reviewing will
have
to be done partly by AI, because otherwise there’s no way to handle the orc army: the reviewers can’t unilaterally disarm.
Anyway, here’s a small sampling of the significant AI-proved or -assisted results from, like,
the last month
, besides Navier-Stokes—restricting myself to those that solved longstanding open problems I had previously known or cared about.
Of course, the counterexample to the Jacobian conjecture, announced by Levent Alpöge in a
now-famous tweet
: “hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final” (followed by a listing of the counterexample)
An improved upper bound for shadow tomography of quantum states
, from Chen, O’Donnell, Pelecanos, and Wright, improving the dependence on the Hilbert space dimension d from log(d) to √log(d). (When I
introduced shadow tomography
back in 2017, I raised the question of whether the dependence on d could be eliminated entirely, while preserving polylogarithmic dependence on the number of measurements m.)
Progress
on the
Aaronson-Ambainis Conjecture
(the version that talks directly about quantum algorithms), basically showing that it holds for quantum algorithms that make their queries in a small number of parallel rounds.
(
Update:
Nope, sorry, Jordan Docter points out to me that this one was pre-AI, with AI used only for proofreading and other incidental things!) This was
independently achieved
by Liu and Mutreja, making more substantial use of AI.
According to rumors that I’ve heard, solutions to some
very
longstanding open problems in theoretical computer science (no, not P≠NP or other complexity class separations, but think about some of our other biggest problems). I’m told that the AI companies, having been burned by the hostile response to the Navier-Stokes proof, are now sitting on solutions to some very major problems until they figure out a better way to handle things
Feel free to remind me of anything I left out.
Let me try to convey the mood in the mathematical community right now, at least as far as my experience reaches. Nearly every conversation is about the AI tsunami, or eventually circles around to the tsunami even if it’s originally about something else. Often, though, the focus is less on the unknowable future—for how much longer will mathematical research as a human enterprise even exist?—than on immediate questions of
how to respond
.
What are the new rules for when you get to write a paper with your name on it, and, y’know, get credit for it? That you fully understand the proof, can give talks about the proof, can answer questions about it, take responsibility for its correctness? Do you need to have played any role in
finding
the proof?
In the cases, likely to become more and more numerous, where all of those conditions are
not
satisfied, how do you share AI-generated math, if at all? Do you tweet it, like Alpöge hilariously did with Fable’s disproof of the Jacobian Conjecture? Do you post to the arXiv or GitHub? Do you publish a paper that lists “GPT-6 Astra” or “Claude Fable” as the author—but then let the AI profusely thank
you
in the acknowledgments for suggesting such a wonderful problem to it?
Of course, how one responds to the immediate problems ultimately
does
depend on one’s broader beliefs about what mathematical research is for and about. Are we just trying to decide whether various conjectures are true or false? Or are we trying to maintain a human community, across the generations, that understands the conjectures and cares about whether they’re true or false and why? If the latter, how do we incentivize people to join that community, to undergo the years of intense training required, if their role will now be reduced to verifiers and explicators (if even that) of gargantuan arguments dumped into their laps by the AI companies?
As many of you will have seen, twenty-five Fields Medalists, including Terence Tao, released an open letter entitled
A Severe Misalignment of AI in Mathematics
, which articulates some of these concerns in the wake of the Navier-Stokes announcement. As many critics have pointed out, the open letter doesn’t really have a clear ask: mostly, it just eloquently sets out the values of the human mathematical community that the authors consider worth preserving in the age of AI. After reflection, I decided to endorse the statement, because I want to preserve those values as well.
I don’t think any of the signatories are naïve enough to imagine that AI won’t permanently change the way mathematical research is done—indeed, that it isn’t already doing so. There’s surely at most a tiny market for “certified organic theorems.” That isn’t the question. The question is, do we incorporate AI in a way that still puts human understanding, of what either humans or AIs are producing, at the center of the whole enterprise? Maybe someday, it becomes unsustainable to do that. Maybe someday we say: “human math had a great 4,000-year run, but today we close up shop and turn everything over to the machines, continuing to apply our own brains to math, when we do, at most for exercise, recreation, or competition, like chess.”
But, partly because of my worries about AI misalignment, I’m not ready to throw in the towel just yet. I still
do
want to keep insight and understanding at the center of what mathematicians, computer scientists, and physicists do, for as long as we can keep it there, even as the human race now cedes its supremacy at the task of proving or disproving conjectures.
Speaking of alignment: if you’re any kind of mathematical researcher, and the present age of wonders and terrors has inspired you to want to spend your remaining time confronting the tsunami head-on, rather than pretending it doesn’t exist or is still far away, please join your dozens of colleagues who’ve arrived at the same place!
My friend and colleague
Mike Winer
was trained as a theoretical physicist, did a postdoc with Juan Maldacena at the Institute for Advanced Study in Princeton, but then got AGI-pilled and decided to switch to full-time work at the Alignment Research Center in Berkeley (founded by
Paul Christiano
, who moved to AI alignment a decade ago after doing quantum computing theory with me). Mike recently wrote a Substack post entitled
From Academia to Alignment
, which I enjoyed and which I’d commend to anyone currently considering this transition. In a similar vein, see
this
from Xiaoyu He. And, one more: a
meditation on mathematicians’ possible future as priests or monks
, by Stanford math undergrad Logan Graves.
Note:
Of course
I’ve been glued all week to the dramatic developments in AI. I’m working on a post about them. I’m not good at reacting to things in a timely way. So today, I’ll do my post marking the tragedy a quarter-century ago that we all commemorate. Please feel free to share your 9/11 memories in the comments. Also, Shana Tova to those who celebrate!
The morning of September 11, 2001, I was a second-year PhD student at Berkeley, who woke up late in his dorm room at International House, after a long night spent closing in on the
proof
of the quantum lower bound for finding collisions.
Rolling over to my laptop, I saw a flurry of weird emails, including one from Prof. Christos Papadimitriou saying that “we’re a community, and we’ll all support each other,” and another from Prof. Luca Trevisan (whose algorithms course I was then TA’ing) saying “on a day like this, it’s impossible to think about algorithms. Class is cancelled.”
Confused, I clicked over to the
New York Times
and saw the picture of the burning towers, and read numbly about what was already over by the time I’d woken up. I checked in with my mom, made sure relatives and friends in the NYC area were OK. My dad was at a company event in Atlanta, and would need to drive home because of the national grounding of flights.
One of my earliest memories in life, from age 5, is of ascending to the top of the World Trade Center. Growing up an hour’s drive from NYC, it wasn’t an exotic place to me.
I soon learned that one of the dead was
Danny Lewin
, the ex-IDF captain, theoretical computer scientist, and cofounder of Akamai who had his throat slashed on one of the planes while trying to fight the hijackers, making him the day’s first casualty, even while Akamai’s technology was part of what kept news websites running that day. I’d never met Danny but already knew many people in common with him. A few years later I’d be humbled to win the student paper
award
that was named in Danny’s memory.
Anyway, at Berkeley on 9/11, I wandered over to Soda Hall just to be with other people. A few students showed up for office hours, wanting help with their algorithms homework, which I found hard to believe, but I did my best to concentrate, as the computer screens around me showed the burning towers.
That evening, I went to a vigil for the victims in Sproul Plaza. But the “vigil,” such as it was, quickly dispensed with mourning and prayers and turned to applauded speeches about how the US must respond with love rather than war, and must turn the other cheek. Meanwhile, a student communist organization was handing out flyers explaining that the victims were mostly “wealthy capitalists and the workers who tried to rescue them.” This while smoke still blanketed NYC and the desperate search for survivors continued. I left the vigil early.
Until that day, I had thought of myself as basically a “leftist,” one whose #1 issue was the existential risk of climate change. Sure, I disagreed with my fellow leftists about issues from nuclear power to gifted education to Israel, but those were just intra-left disputes.
The year before, I had created the website “In Defense Of NaderTrading,” in a desperate attempt to intervene in history and cause Al Gore to become president rather than George W. Bush. When Bush “won,” by the infamous 537 votes in Florida, I considered it a victory for horribleness that would never be surpassed by anything else in my lifetime (ha). I couldn’t imagine any politician who was more the antithesis of everything I believed in than Bush. This view, of course, did not particularly stand out at Berkeley.
In the days after 9/11, though, it became obvious that I could not be a “leftist” in the Berkeley sense. Some of my fellow students felt that Osama bin Laden made a lot of great points, that the attacks were basically justified, and that at any rate, we in Amerikkka had done much worse to provoke them, including by supporting the genocidal settler-colony called “Israel,” which for all we know secretly masterminded the 9/11 attacks anyway (although again, if bin Laden had done them, he would’ve been justified).
Around the same time came the Second Intifada, when a wave of suicide bombings in Israeli buses and pizza parlors and university cafeterias thrilled and energized some Berkeley students to the extent that they took over a Holocaust Remembrance Day event with bullhorns to make it about the Nakba, smashed the windows of the Hillel building, and beat up a couple of students wearing kippot. That was how thoroughly anti-Nazi they were.
I finished my PhD at Berkeley in 2004 having learned about more than quantum computing. I’d learned that, while American academia had pockets that truly were crucial refuges and oases for nerds like me, it also harbored people who would gladly see me and my relatives and my fellow Americans killed for the sake of their ideological vision. And I’d learned that I had my own ideological vision, which was that such people could go fuck themselves.
It deeply pained me to be on the same side of anything as George W. Bush — especially because I knew that 9/11 had happened on his watch, that he had ignored all the warnings, and that he was grossly incompetent to manage the resulting wars against jihadism (just
how
incompetent, I didn’t know at the time). But as flawed as Bush was, I knew that I wanted to preserve rather than destroy the civilization of which he was a temporary steward. And I think the value and fragility of our civilization is the main lesson from that day that I’d like to convey to my kids, for whom of course 9/11 is just another historical event to learn about in school, like the Boston Tea Party or the Alamo.
I woke up yesterday with the following thoughts, which are probably either obvious or dumb.
A central thesis that many readers, including me, took from Douglas Hofstadter’s
Gödel Escher Bach
when young was that the secret of intelligence (and therefore, of AI) was going to have a lot to do with self-referentiality and “strange loops.”
Even Roger Penrose’s
The Emperor’s New Mind
, which in some ways was the anti-GEB, ironically agreed with GEB about the fundamental importance of self-reference to the success or failure of the whole AI project. It claimed (incorrectly, in my view and in most experts’) that AI could never work because there was something about Gödel’s Theorem and self-reference that no computer program could ever capture, but that could be captured by exotic physics accessible to the human brain.
Now, in 2026, we’ve succeeded at building AIs that outperform most humans at most intellectual tasks that are well-defined enough to judge. And at no point in the tech stack of those AIs — neither in the transformer neural nets, nor in the GPU clusters they run on, nor in the training process, nor anywhere else — did anyone need to build in anything about self-reference. (Excepting, eg, the system instructions that tell the model about its role and identity, which aren’t needed for intelligent behavior. Also, I’m not going to count the autoregressive nature of LLMs as “self-referential”; that’s just dynamical feedback.)
Of course, GPT 5.6 Pro and Fable can talk about themselves, about Gödel’s Theorem, about self-reference, about what we’re talking about right now, all of it, better than most humans. But at no point did anyone need to build self-referential abilities in. They popped out as a byproduct of the same pretraining that let the models talk about Pokémon and long-chain polymers and cognitive behavioral therapy and plate tectonics and everything else.
No wonder Hofstadter says he’s been stunned by the success of LLMs, and has seemed depressed about current AI capabilities in
essays like this one
. He’s way too smart to deny what’s happened or invent reasons why it doesn’t really count (the approach many have taken). But he realizes that we now have true conversational intelligence from a path that the GEB worldview would’ve regarded as far too cheap and simple, and that certainly has no “strange loops” built in anywhere.
Of course, a Hofstadterian could argue that a strange loop
emerges
in LLMs — indeed, nothing in GEB ever said that strange loops would need to be explicitly engineered at the outset. But would anyone who hadn’t been brought up on GEB arrive at this as a useful way of thinking about LLMs?
What can we say about this with hindsight? While the ideas of diagonalization and self-reference of course played a central role in the birth of modern mathematical logic and computer science, the most famous uses were
negative
: there is not a bijectjon between the natural numbers and the reals. There is not a complete sound proof system for arithmetic. There is not an algorithm to solve the halting problem.
If your goal was only to build the axioms of ZFC and the rules of first-order inference, or build an electronic computer, you wouldn’t explicitly need self-reference for that. You would just … start building, taking care that your instruction set didn’t fall short of universality.
Yes, ZFC can formalize and prove theorems about itself. Yes, electronic computers can run programs that take their own code as input. But no one ever needed to build those abilities in, any more than self-reference needed to be built in to the alphabet or the rules of grammar. It popped out as a free byproduct of universality.
In the same way, LLMs’ ability to talk about themselves popped out as a byproduct of their ability to talk about anything in the discourse universe they were trained on. The big, old ideas about intelligence that ended up basically vindicated were the ideas about how intelligence is about prediction, and prediction is about compression, and compression is about finding better and better upper bounds on Kolmogorov complexity. Not the self-reference stuff. (Although, if you wanted to know why Kolmogorov complexity
can’t
be computed perfectly, that negative statement would again require a self-referential argument.)
What’s left? Consciousness and subjective experience of course remain extremely mysterious. For all we know, Hofstadter could be right that those have something to do with self-reference. (For all we know, even Penrose could be right that they have something to do with exotic physics accessible to biological brains but not digital computers!)
But the idea that you’d need explicit self-referentiality before you could get convincing and world-changing conversational intelligence? Let it be buried in a Westminster Abbey or Arlington National Cemetery for the most important wrong ideas in human history — geocentrism, Aristotle’s teleological physics, aether, phlogiston, Freud’s psychology, Marx’s prediction of a workers’ uprising followed by a classless utopia, etc. But buried it needs to be.
So yeah, Anthropic has announced that it’s now
watermarking the outputs of Claude
, using a scheme based on Google’s SynthID, which is in turn based on the
Gumbel Softmax scheme
that I proposed at OpenAI back in 2022—as far as I know, the first LLM watermarking proposal, though far from the last one. I’m gratified that Anthropic credits me for this, even though I shirked my duty by never publishing a paper about it (by the time I sat down to write one, it seemed like the whole field had already assimilated my scheme and moved beyond it—AI just moves too fast for me!).
For those who don’t know, watermarking means slightly changing the way that an LLM operates to insert a subtle signal that lets you prove later, with high statistical confidence, that a text indeed came from your specific LLM. It uses the randomness that’s already present anyway in LLM outputs, replacing some of it by pseudorandomness that favors certain word combinations over others in a way that’s later detectable, given only the sequence of tokens itself (not the prompt or the probabilities) along with the key of the pseudorandom generator.
Christ, Gunn, and Zamir
then substantially improved my scheme to get true cryptographic indistinguishability, and there have been other improvements since.
I’d been meaning to blog about this for days. Thankfully, Zvi Mowshowitz, the world’s foremost blogger about AI, has now written a wonderful post, entitled
AI Text Watermarking Is Free And Good
, which saves me from the need to write my own long post. In particular, Zvi masterfully explains the central point that I needed to explain to everyone back in 2022-23: why, contrary to many people’s intuitions, there’s no inherent tradeoff between watermarking and the
quality
of LLM output. Basically, nearly every LLM output was
already
a sample from a cloud of exponentially many possibilities, all of them about equally good, so there’s plenty of room to steer within that cloud without affecting anything that an ordinary user would notice. As my kids would put it, the math mathes.
As Zvi explains, the central technical drawback of watermarking schemes like the one I proposed, and what Anthropic is now using, is that it’s possible to remove the watermarks with a little extra work (even stuff as simple as, e.g., translating between English and French, asking the LLM for words interspersed with emojis and then removing the emojis, or using an open model to paraphrase the output). Zvi gives detailed arguments for why he expects watermarking to remain a net positive in practice despite this vulnerability.
I could add that, in addition, there’s recent progress (see
here
for example) on what I’ve called “semantic watermarking,” or watermarking at the level of the underlying concept vectors rather than the tokens themselves. This actually seems to work, albeit with no theoretical guarantees, and will hopefully make removing watermarks a lot harder—although the
Barak et al. impossibility result
suggests that under plausible assumptions, no LLM watermarking method will be completely foolproof.
Anyway, I worked out my scheme in Fall 2022, then gave lots of talks about it (including, as it happens, at Anthropic), and also worked with Hendrik Kirchner at OpenAI, who actually implemented and tested my scheme. Unfortunately, OpenAI leadership decided against deploying watermarking, worried mostly about risks to the product (i.e., customers disliking the idea, and leaving for a competing LLM that doesn’t watermark). You can read
this
Wall Street Journal
investigation
from two years ago for more. I was hopeful that the State of California was going to solve the collective-action problem by mandating watermarking for AI models, but then they decided to do that
for audiovisual content only
, for some reason exempting text.
Nevertheless, Google DeepMind implemented something very similar to my proposal in its
SynthID
, deployed in all its Gemini text models. But they heavily restricted who gets to
detect
the watermark, which made their admirable decision of limited use to my academic colleagues, who’ve been begging me for a way to detect whether their students are using AI to cheat. (For now, I mainly send them to
Pangram
, a leading AI detector
not
based on watermarking, as a first line of defense.)
And now, apparently to comply with EU regulations, Anthropic says they’ve deployed a watermarking scheme like mine where
anyone
will be able to do detection (though they also say in their FAQ that they’re still working on the detection API). Even OpenAI
suggests that it plans to follow suit
. So, four years after I seriously thought about this, it looks to my surprise like this is actually happening. Thanks, EU!
Tell you what:
read Zvi’s post
, and then whatever questions you still have, you can come here and ask in the comments. Just please don’t use Claude to
write
the comments. With any luck, I’ll eventually be able catch you if you do.
Kol HaKavod (mad respect) to Yotam Budnik, who incredibly, has
also
won a Gold Medal (which he was allowed to keep, apparently) at the International Math Olympiad. And congratulations to the entire Israeli team, which (incredibly) would apparently have had a higher overall score than the US team, had it been allowed to compete as an official team at all.
Friend-of-the-blog (well, mainly just friend)
Adi Akavia
has asked me to publicize that she’s helping to organize an
exciting CS conference called Mind-IL
at Tel Aviv University on October 26, in memory of the Israeli-American Turing Award winner
Michael O. Rabin
, who passed away in April. Please note that October 26 is the day before the
Israeli election
, for any Israeli citizenship holders living abroad who might want an academic excuse to come to Israel and vote.
Update (August 19):
Avi Wigderson also asked me to advertise a
conference
, to be held September 16-18 at Bletchley Park in the UK, to commemorate the 90th anniversary of Alan Turing’s “On Computable Numbers” paper.
Last night my 9-year-old son was taunting my wife, complexity theorist
Dana Moshkovitz
, as follows: “mommy, I heard you got
cooked
! I heard that a
robot
solved the math problem you worked on for your whole career! OOF!”
While my son was being a brat, he also wasn’t wrong. Whether you’re thrilled, depressed, angry, or whatever else about it, yesterday was surely one of the biggest days in mathematical history. And yes, among the 372 huge results
released yesterday by OpenAI
, on the recommendation of its
advisory group
of Timothy Gowers, Edward Witten, and other distinguished mathematicians, was a
proof
of Subhash Khot’s
Unique Games Conjecture (UGC)
, a statement that my wife has worked toward proving for the entire time I’ve known her. (The UGC implies that a whole slew of optimization problems really are NP-hard, even if you just want an approximation that’s slightly better than what you get from semidefinite programming relaxation, which is one of our main tools.)
Or at least, we’re pretty sure that it’s a proof! There’s a
Lean certificate
, as there are for some of the other 372 breakthrough results (not all of them). But it also appears that no human has understood just about
any
of these proofs yet; the race to do so has just started. If you want an on-the-ground sense of what that race is going to be like, here’s some of what Dana texted me last night:
It feels like something written by someone who’s on psychedelics. So much unclear and doesn’t make sense. Lots of name dropping of previous work without discussing why it can be used despite impossibility results
Basically the paper is so horribly written that it’s impossible to read it without AI help
I asked Astra for reasonable completeness and soundness claims of the noise gadget and it gave them by combining claims from all over the paper
They also have direct optimal NP hardness of approximation proofs for the main applications of the UGC (Max Cut and all CSP) that bypass the UGC.
The UGC proof invents a completely new bizarre code with a noise test. It’s some crazy recursive construction.
It’s not the long code, not the short code – some alien craziness
I still think that there maybe is a proof that uses the half space code (which is natural)
The citations are often irrelevant and confusing
A possible future is a math world that’s heavenly if you have vision/creative ideas that AI could help check and implement.
And of course there’s a lot for us to learn from the aliens
If you’re wondering what emotions Dana is feeling—well, probably all of them! Even while a central career aspiration has fallen to a robot, there are at least two mitigating factors for her. First, she can feel vindicated that the UGC was
true
after all, something she never doubted even while many of her colleagues did! Second,
all
of us in math and theoretical computer science and mathematical physics, at least those who cared about solving crisply-stated problems, are now in the same boat.
Besides the Unique Games Conjecture, here’s a small sampling of the treasures from Aladdin’s cave that I’ll probably be paying the most attention to over the coming weeks:
L=BPL
(i.e., probabilistic logspace and deterministic logspace are the same thing), one of the great derandomization conjectures short of P=BPP. Though its truth was never in serious doubt, there was a whole subcommunity focused on proving this.
The
Fourier Transform
and
integer multiplication
in less than O(n log n) time, breaking a barrier that had stood since the 1960s. The new running time, if you’re curious, is O(n log
0.9999999999999
n), give or take some 9’s.
Positive solution to the Unitary Synthesis Problem
, which Greg Kuperberg and I
posed
back in 2007. For every n-qubit unitary transformation U, there exists a classical oracle A such that U can be implemented in quantum polynomial time with access to A. This is the opposite of what most of us expected, and could have implications for e.g. the computational problem of decoding Hawking radiation from a black hole and many other problems in quantum complexity theory—
if
we had an efficient way to construct the oracle A, which this paper doesn’t give.
Parity is not in QAC
0
, one of the great questions of quantum complexity theory since 1999 that many of my colleagues had been closing in on.
Uncomputability of solving polynomial equations over the rational numbers
—this was arguably the biggest open problem in computability theory (note that uncomputability of solving Diophantine equations, i.e. polynomial equations over the
integers
, was proved in the 1970s, giving a negative answer to Hilbert’s 10th Problem)
Any of the above, alone, could easily have been “result of the year” in some area (and in some cases, like Unique Games and L=BPL, in all of CS theory). And there’s a lot that I’ve left out—feel free to share in the comments whatever is making
your
eyes bug out! There are equally astounding wonders in number theory, combinatorics, algebraic geometry, analysis, and pretty much every other area of math, most of which I’ll never understand, although I’ll note that it includes partial
progress toward the Riemann hypothesis
and the
Hodge Conjecture
and the
Birch-Swinnerton-Dyer Conjecture
(i.e., the majority of the remaining Millennium Problems).
We can take solace in what’s missing from the list. P ≠NP isn’t there, nor even P=BPP or NEXP⊄P/poly, and surely not for lack of trying. Apparently the greatest open problems of theoretical computer science are indeed pretty hard!
Oh, lest I forget: one day
before
the OpenAI dump, meaning Monday evening, Virginia Williams and Josh Alman
posted an arXiv preprint
that solves the 3SUM problem in O(n
1.9992
) time, and the All-Pairs Shortest Paths problem in O(n
2.9995
) time, refuting half-century-old conjectures that the correct answers were n
2-o(1)
and n
3-o(1)
respectively. In this case, it wasn’t an OpenAI model that supplied the crucial idea; it was an Anthropic one! But Anthropic then took a different approach from OpenAI: rather than post the undigested solutions to the world, it gave Virginia and Josh the opportunity to write and announce a digested version in exchange for compensation.
These have emerged as the two main models for communicating AI math breakthroughs, and they both have strengths and weaknesses. The “OpenAI model” sets up a crazy race among humans to digest and explain a messy AI proof (work that could easily be some combination of thankless, barely-credited, competitive, and unfun), while the “Anthropic model” puts a private company in the position of picking and choosing which human mathematicians get to be the emissaries of the AI. Dunno, what do you guys think?
For those who are wondering: apparently, the AI model that produced all these wonders was
not
bespoke contraption of 10,000 agents burning millions of dollars worth of compute, as was used for example to construct a finite-time blowup for the Navier-Stokes equations. Instead, it was simply the latest internal OpenAI model—one that might be released to paying ChatGPT customers within the next couple of months, depending on the recommendations of OpenAI’s safety board! (My 9-year-old son: “Oh they
definitely
shouldn’t release that. If it could solve all those math problems, it can’t
possibly
be safe.”) Apparently they used about 3 hours of GPT-Pro level compute on average per problem solved.
Also, if you were wondering: apparently they
tried
the model on about 8,000 problems. So, right now it “merely” solves ~5% of the longstanding open mathematical problems that it’s asked about, the problems that whole communities have spent years on, after a single 3-hour attempt on them.
I’ve been glad to see the CS theory community rising to the occasion. At the Simons Institute in Berkeley, here at UT Austin, and elsewhere, I’ve hearing stories of researchers rushing to pore over the manuscripts and
make sense of them and explain them
—because what else do we do? How else do we continue the craft to which we’ve devoted much of our lives?
If you want some sense of what things feel like now in math, imagine a hunter-gatherer who’s spent his entire life learning to survive deep in an unforgiving rainforest, then a giant resort hotel springs up right next to him with a helipad and heated pools and AirBnBs, and without missing a beat, the hunter-gatherer says: “alright fine, so now my new job is to run wilderness retreats for the tourists, or something.”
It’s as if you were teleported to the peak of a tall mountain. Surrounded by fog, you have no idea where you are, or what’s around you. You do not know how your mountain connects to others, and you have no equipment to help you explore, no way to help someone else join you. If you had climbed the mountain yourself, you would have experienced how the human body adapts to altitude and changes in oxygen levels. You might have had to invent tools to navigate, to climb steep cliffs, or to make a shelter. You might have encountered a fellow explorer, gotten lost together in a hidden valley, and found a plant that could be turned into a life-saving medicine.
Instead you’re perched on the peak but in the dark, while the maker of the teleportation machine tells you that it can explore the wilderness better than any human.
For any one of these mountains, if we care enough, I feel optimistic that we can do as we always have: clear the fog and figure out the path, except now using the teleportation machine to help guide us. The bigger challenge will be to nurture a community that still
cares
about the heroic adventure of finding the paths up these mountains in the world with the machine. (Oh, and I think one place where the metaphor breaks is that we still
do
have each other, as much as we ever did before!)
Experience has shown that,
even now
, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by
anything
that happens in the empirical world, of updating on
anything
, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse.
So, they’ll say, maybe the alleged solutions are not solutions at all, but just “AI slop.” Or maybe none of the 372 well-known open problems that were solved were
real
math problems, they were all just glorified contest puzzles and trivialities. (After all, there’s still no Riemann Hypothesis!) Or maybe the entire 4000-year-old discipline of mathematics needs to be jettisoned: turns out that it was
all
just puzzle-solving and trivialities; all that’s different is that now the triviality stands unmasked. In any case, what really matters is that the
true
inner sanctum of human creativity hasn’t been breached and probably never will be, and also, that Sam Altman and Dario Amodei are contemptible little nerds.
If you’re still a proponent of that doomed worldview, still aboard the sinking ship, I encourage you in the strongest possible terms to read yesterday’s
other
great contribution to AI discourse, besides the OpenAI Mathocalypse dump: namely,
Scott Alexander’s open letter to Steven Pinker
. I feel some responsibility for this, as the person who first introduced Steven Pinker to the
existence
of the rationalist community, and who also first introduced Steven Pinker and Scott Alexander to one another (they had both been fans of each other’s writing). And now Scott is challenging Steve to a literal duel, with guns!
For whatever it’s worth: Steve is a lifelong intellectual hero of mine, just as he is for Scott, and I also have to privilege of calling Steve my friend. But I found Scott’s post to be one of the most devastating rejoinders to anything that I’ve ever read. And I thought Scott’s conclusion was exactly right: when it comes to AI risk, Steve’s great challenge is now to accept and start using a more “Pinkerite” epistemology.
Last night, while I should’ve been poring over some of OpenAI’s hundreds of papers and/or writing this post, I decided to spend some time with my kids instead. They wanted a movie night, so I suggested something they’d never seen before (and that I hadn’t seen for decades), and that seemed chock-full of no-nonsense, practical guidance for the world in which they’re going to grow up:
Terminator 2
.
This entry was posted
on Wednesday, October 7th, 2026 at 1:57 pm and is filed under
Uncategorized
.
You can follow any responses to this entry through the
RSS 2.0
feed.
You can
leave a response
, or
trackback
from your own site.
You can use rich HTML in comments! You can also use basic TeX, by enclosing it within
$$ $$
for displayed equations or
\( \)
for inline equations.
After two decades of mostly-open comments, in July 2024
Shtetl-Optimized
transitioned to the following policy:
All comments are treated, by default, as personal missives to me, Scott Aaronson---with no expectation either that they'll appear on the blog or that I'll reply to them.
At my leisure and discretion, and in consultation with the
Shtetl-Optimized
Committee of Guardians
, I'll put on the blog a curated selection of comments that I judge to be particularly interesting or to move the topic forward, and I'll do my best to answer those. But it will be more like Letters to the Editor. Anyone who feels unjustly censored is welcome to the rest of the Internet.
ICANN Reveals 2026 Round Applications for New Generic Top-Level Domains
"Reveal Day marks an important milestone for the next expansion of the Domain Name System," said Kurtis Lindqvist, President and Chief Executive Officer of ICANN. "These applications demonstrate how organizations around the world are innovating new ways to build trusted online identities, serve their communities, and connect with Internet users. ICANN remains focused on administering a fair, transparent, and predictable evaluation process grounded in the policies developed by our global multistakeholder community."
Reveal Day 2026 Data Snapshot
The application submission window was open from 30 April–12 August 2026. Highlights from the 2026 Round application data* include:
1,615 applications
333 brand applications
16 community-based applications
15 self-identified geographic name applications
51 applications submitted by Applicant Support Program applicants
16 applications from Africa
218 applications from Asia/Australia/Pacific
506 applications from Europe
11 applications from Latin America and the Caribbean
864 applications from North America
21 applications for Internationalized Domain Names
9 applications for variants. This is when an applicant applies for a string that has other forms in different scripts such as Arabic or Chinese.
*This snapshot reflects primary string application data as of 7 October 2026. The final number and types of applications will be confirmed on String Confirmation Day.
Below is information on subsequent milestones in the program through the end of the year.
Replacement Period: 8 October 2026 (00:01 UTC)–21 October 2026 (23:59 UTC)
Following Reveal Day, applicants that applied for a second-choice or alternate "replacement" string when they submitted their gTLD application have two weeks in which they may elect to switch from their primary applied-for string to their replacement string. Applicants may switch to the replacement string only during this period, and only if their replacement string is eligible. If the replacement string is identical to another applied-for primary string, or another replacement string, it cannot be used. Applicants should be aware that replacement strings could end up in contention in the later stages of the program (e.g., as a result of a singular/plural notification or String Confusion Objection). More information on the Replacement Period is available on this
webpage
.
String Confirmation Day: 17 November 2026
Following the String Replacement Period, a final list of all the applied-for gTLD strings is published on String Confirmation Day, as well as an updated list of contention sets. String Confirmation Day marks the start of the Community Input and Objections Period and the final 10 days during which applicants can withdraw their application to receive a 65 percent refund of the gTLD evaluation fee paid. To receive a refund during this refund window, a request must be submitted by 23:59 UTC on 27 November 2026. More information on refunds and timing is available on this
webpage
.
Community Input and Objections Period: 17 November 2026–16 March 2027
The ICANN community and members of the public will have an opportunity to provide feedback on applications following String Confirmation Day. Input may come via application comments, Governmental Advisory Committee Member Early Warnings, singular/plural notifications, or objections using the
Application Comment Forum
. More information, including details on additional objections periods, can be found on this
webpage
.
ICANN Office Closure
2026 Round operations will be paused during ICANN's end-of-year office closure from 1:00 UTC on 19 December 2026 until 16:00 UTC on 4 January 2027. This pause has been factored into the Community Input and Objections Period to ensure that applicants, the community, and the public retain the full amount of allotted time to provide input or respond to objections. Additional details will be made available in the coming months.
Prioritization Draw: 2027
The Prioritization Draw sets the order that applications and contention sets will be processed following the conclusion of the String Evaluation Stage. The String Evaluation Stage is the process of reviewing applied-for gTLD strings (and their allocatable variant strings) to ensure they meet the requirements of the New gTLD Program; all strings are
reviewed concurrently
. As application and contention set processing will not start until after the conclusion of the String Evaluation Stage (expected to last 180 days), the Prioritization Draw will take place on a to-be-determined date in the first half of 2027. Applicants cannot attend the Draw in person but can follow the live event virtually. More information will be provided at a later date.
Future 2026 Round Milestones
ICANN will continue to work on the planning and confirmation of timelines for later 2026 Round milestones, such as the publication of String Evaluation Results and the start dates for Applicant and Application Evaluation and contention resolution, and will communicate them accordingly.
Additional Resources
More information on the key phases of the application lifecycle can be found on the
2026 Round Applicant Journey webpage
. Be sure to check the program's
News and Announcements page
frequently for the latest news and updates. This information also will be sent directly to entities that have submitted an application in the TLD Application Management System.
Meta and Microsoft Limit Employee Use of Claude AI Tools
Meta and Microsoft have commenced a significant reduction in employee utilization of Anthropic’s Claude AI as they pivot towards their own proprietary coding instruments, as reported on October 5 by The Information.
This strategic alteration prioritizes internal expenditure and workforce processes, rather than discontinuing customer access to Claude through Microsoft’s offerings.
This transition underscores the increasing imperative to manage AI expenditures as coding assistants become integral to routine software operations.
While both corporations continue to be substantial clients of Anthropic, they also find themselves in competition with the firm. Their actions reflect a convergence of AI investment, product strategies, and control over developer resources.
Microsoft Reduces AI Expenditure Allowances
Microsoft had previously anticipated that internal expenditures on Anthropic
technology
would exceed $1 billion annually.
However, this estimate has diminished by over a third subsequent to management’s directive for employees to curtail their use of Claude and to increasingly adopt Microsoft tools, such as GitHub Copilot and OpenAI frameworks.
Within Microsoft’s cloud and AI sector, monthly AI spending limits have reportedly been slashed from $100,000 per employee to approximately $10,000 in the majority of instances.
It’s crucial to note that these figures are spending ceilings rather than reflections of actual employee expenditures. This reduction in budgetary flexibility has disheartened several engineers who once enjoyed broader latitude to explore models.
It’s important to highlight that diminished internal funding does not equate to Microsoft’s withdrawal of Claude from client-facing products.
Reports suggest that client expenditure on Anthropic models via Microsoft’s enterprise platforms continues to rise, even as the company constricts its own employees’ usage.
At Meta, the number of Claude Code users reportedly plummeted from around 60,000 earlier this year to about 30,000.
Although layoffs have influenced this decrease, the report attributes the primary cause to Meta’s strategic shift towards its own AI solutions, which comprise MetaCode and Muse Code, developed using Meta’s proprietary models.
MetaCode is said to have more than 30,000 internal users, while Muse Code boasts over 6,000 employee users. Meta initiated external testing of Muse Code with clients in August.
Notably, amidst this transition, Meta reportedly allocated over $105 million to Claude Code over a 28-day timeframe, indicating that a reduction in users does not necessarily correlate with decreased spending.
For
cybersecurity
teams, the transition to different coding assistants does not eliminate the hazards associated with automated access to sensitive files, commands, and credentials.
Previous analyses by Cybersecurity News outlined vulnerabilities within Claude Code that could allow unauthorized execution and API key theft. Anthropic rectified these flaws prior to their public disclosure.
Microsoft’s offerings have encountered analogous concerns. The remedied RoguePilot vulnerability exemplified how nefarious commands embedded in a GitHub Issue could facilitate the takeover of repositories.
Such incidents emphasize the criticality of imposing stringent permissions, employing trusted project settings, and conducting thorough evaluations of AI-generated actions.
Anthropic’s reported annualized revenue pace of $65 billion reflects unabated demand for its services.
This retrenchment illustrates a tightening of budgets and intensifying competition, rather than an outright cessation of engagement with Claude.
Centralize control flow. When splitting a large function, try to keep all switch/if statements in the "parent" function, and move non-branchy logic fragments to helper functions. Divide responsibility. All control flow should be handled by one function, the rest shouldn't care about control flow at all. In other words, "push ifs up and fors down".
The programming heuristic push-ifs-up-fors-down suggests that conditional logic (if statements) should be moved upward (towards the caller or earlier in a pipeline), while iterative loops (for operations) should be pushed downward (toward batch processing or later in a tight branch-free loop). This improves clarity and performance by centralizing branching and leveraging bulk operations.
matklad has also
blogged
about this principle and discussed the many virtues of adhering to this programming idiom.
In the post, matklad describes optimizing code by:
Pushing conditionals ("ifs") up:
If a function branches on its input, consider moving that branch to the caller. Consider the example from the post - instead of
frobnicate(walrus: Option<Walrus>)
unpacking the option internally, the caller handles the
None
case and the function takes a plain
Walrus
. The function's type now states its precondition, the input state space becomes narrower and acts as a form of filtering. But the point is where the decision lives, not how much data flows downstream.
Pushing loops ("fors") down:
Deferring loops until after filtering or reducing the dataset minimizes unnecessary computations. Rather than calling
frobnicate(walrus)
in a loop, provide
frobnicate_batch(walruses)
and let the loop live inside it. So the hot loop runs without a branch and is a candidate for vectorization.
The two moves compose. Given a collection of
Option<Walrus>
values, the caller discards the
None
s and unwraps the rest into a
Vec<Walrus>
, then hands that to
frobnicate_batch
, which never has to consider the
None
case at all.
let maybe_walruses: Vec<Option<Walrus>> = ...;let walruses: Vec<Walrus> = maybe_walruses.into_iter().filter_map(|w| w).collect();frobnicate_batch(&walruses); // never sees a None
If you think a bit, the push-ifs-up-fors-down principle has far broader applications. This post explores some of the broader perspectives of this principle with respect to relational database query optimizations and functional programming and category theory.
Analogy in Database Queries: Projections Early, Joins Late
The same principle appears in database query optimization. In SQL query planning, it’s well-known that one should perform projections and selections as early as possible and defer joins or expansive operations until later. But the vocabulary runs upside-down. A query plan is a tree whose leaves are table scans and whose root produces the result. Data flows up from the leaves, so "down the tree" means "earlier in execution". When an optimizer talks about
pushing a predicate down
, it means evaluating it as early as possible, which is the database counterpart of what the rest of this post calls "up" or "early".
Early projections and selections (Push down):
In database terms, a projection (e.g., SELECT with specific columns) reduces the dataset's width by selecting only the necessary columns early in the query execution. And selections (WHERE clause) act as the filter. Both are pushed down as part of the plan tree so that they get to execute early and filter out irrelevant data, reducing the amount of data passed to later operations.
Deferring joins (Push joins up):
Joins, which combine data from multiple tables, are computationally expensive. The optimizer moves selections and projections below the joins, so joins run on smaller inputs. The effect is that the expensive combining operators run over the smallest inputs the query semantics allow.
Vectorized execution (the "for"):
Another database analogy of pushing fors down is the change in execution semantics. Volcano style execution processes row at a time, where each operator is called once per tuple through a virtual
next()
. The alternative is the vectorized or batch execution, where each operator is called once per batch of a thousand or so tuples and runs a tight loop inside. That is
frobnicate
versus
frobnicate_batch
, at the level of a query engine: the per-call overhead and the per-call decisions are paid once per batch, and the inner loop is branch-light and cache-friendly.
Analogy in Functional Programming and Category Theory
Let's take a look at the same principles through the lens of functional programming and category theory.
Pushing-Ifs-Up as a Restriction to a Subobject
In category theory, when we talk about the category of sets (or, loosely, of types in programming), we can have a predicate
p : A -> Bool
. Elements can either satisfy the predicate or not. Now consider the subset of elements that satisfies the predicate, given by
{a ∈ A | p a}
. And add to it the morphism
{a | p a} ↪ A
that defines the inclusion. The subset along with the morphism defines a
subobject
in category theory.
Note the hooked arrow in the morphism. It's intentional and it indicates that it represents a specific type of morphism -
monomorphism
(mathematicians love to define arrow types :-)). In Set, a monomorphism is an injective function: the inclusion sends each element to itself, so distinct inputs give distinct outputs.
Now we have the link back to matklad:
Before: the callee takes any
A
and runs
if p(a)
inside itself.
After: the caller runs the test, and the callee's input type is the subset.
The callee no longer needs the
if
, because every input it can receive has already passed. In code you represent the subset with a type:
Walrus
instead of
Option<Walrus>
.
Categorically speaking,
Option<Walrus>
is the coproduct
1 + Walrus
, which is either nothing or a walrus. A function that takes an
Option<Walrus>
and branches inside is really a function out of a coproduct, and by the universal property of coproducts such a function is exactly a pair of functions, one for each summand. Pushing the
if
up factors that pair apart: the caller deals with the
1
summand, and the core function is just the
Walrus
component.
Filter, Map and the Law That Relates Them
Same principle in a different manifestation - the algebra of combinators. You commonly hear the advice of "filter before you map". But when exactly is this a piece of legitimate rewrite ? Because the two following expressions are not equivalent:
filter p (map f xs) -- p inspects the *output* of fmap f (filter p xs) -- p inspects the *input* of f
In the first line
p
has type
B -> Bool
; in the second it has type
A -> Bool
. The law that actually relates them is:
filter p . map f == map f . filter (p . f)
This follows from parametricity, and it is easiest to see by factoring
filter
through
Maybe
:
keep :: (a -> Bool) -> a -> Maybe akeep p x = if p x then Just x else Nothingfilter p = catMaybes . map (keep p)
Here
filter p
itself is not a natural transformation - it cannot be, since
p
fixes the element type and you cannot draw the naturality square (try it!) - but
catMaybes :: [Maybe a] -> [a]
is one, and that is where the naturality lives:
map g . catMaybes == catMaybes . map (fmap g)
With that, the law is a short calculation. Since
keep p . f == fmap f . keep (p . f)
:
filter p . map f == catMaybes . map (keep p) . map f == catMaybes . map (keep p . f) == catMaybes . map (fmap f . keep (p . f)) == catMaybes . map (fmap f) . map (keep (p . f)) == map f . catMaybes . map (keep (p . f)) -- naturality of catMaybes == map f . filter (p . f)
Notice that the right-hand side is not automatically cheaper:
filter (p . f)
still computes
f
for every element in order to test it. The rewrite pays off when
p . f
simplifies to a cheap predicate
q
on the input - typically because
p
inspects a part of the value that
f
leaves alone. Then
filter p . map f == map f . filter q
, and
f
runs only on the survivors. Both sides are still a single O(n) pass; what you save is the calls to
f
on elements that were going to be discarded.
And push-if-up-fors-down pays off!
Summary
The main takeaway from the above discussion is that we can have general principles that guide us towards a better algebra of the code base or even the whole system. But all of these work subject to some constraints that must hold good across the components.
Taking an
if
out of a loop is valid when the condition is loop-invariant. A per-element condition cannot leave the loop. It can only move to the boundary and be recorded in a type, as
Walrus
instead of
Option<Walrus>
.
Pushing a selection below a join is valid when the predicate references columns from only one side of the join.
Filtering before mapping is valid as
filter p . map f == map f . filter (p . f)
, and it saves work only when
p . f
reduces to a cheap predicate on the input, typically because
p
inspects a part of the value that
f
leaves alone.
So, it's the algebra that tells you which rewrites are legal. In the filter/map case above, it's the naturality of
catMaybes
that drives the legality and helps you reason about the overall structure of the code.
For the "fors-down" part, it's more about the cost than the equivalence. You change the shape of the arrow from
A -> B
to
[A] -> [B]
so that you pay the setup cost once per batch.
Do you understand why is it
origin master
in one command, and
origin/master
in the other? I didn’t, until a few weeks ago!
My understanding was that git is a content-addressable database. Git stores
commits, a commit is identified by the hash of its content, and the content
of a commit is, primarily:
a memory-less snapshot of a state of the codebase at a given point in time,
a list of (hashes of) parent commits.
That was enough git for me to understand
git log
output and get me out of any
botched rebase without having to re-clone the repo (For roughly half of my
career, I
was
re-cloning the repo. No shame in that! Learning git is useful,
but it’s not the
highest
priority thing to learn when you start).
I now understand that git not only comes with an append-only (“immutable”)
content-addressable database, but is also a boring mutable key-value store.
Git has a mutable map whose keys are strings, and whose values are
content-addressed objects. The keys are conventionally formatted as file system
paths, and you can usually inspect the state of the mapping by listing
.git/refs
directory:
The first argument of
fetch
(
https://...
) is a location of a remote
repository. Git will “dial” that address, and will transfer some data from that
computer locally over the network.
The second argument is a
source:target
pair of string keys (refs). The
source
is a key on the remote repo, the
target
is the name of a local key,
and fetch as a whole asks git to read a value from a remote repository and save
it locally under a different name.
To avoid typing repository URLs all the time, git assigns them symbolic names,
with
origin
being the conventional name for the primary remote repository:
refs/heads/master
is a fully elaborated name of a branch on the remote repo.
That is, branch
my-feature
is just a
refs/heads/my-feature
ref. It could
have been
refs/branch/my-feature
,
but it isn’t :)
I don’t know the specific shorthand rules, but, generally, git allows you to
spell only the suffix of a ref:
refs/remotes/origin/master
is the name of the local ref we’ll use to store the
result. It would seem natural to just use the same name locally as the one on
the remote, but this only works if there’s a single remote. If there are two
upstream repositories (for example, your fork, and the original repo you forked
from), their ref names will collide. That’s why we want to namespace the refs
for remote called
foo
under
refs/remotes/foo
. And
origin
is just a
conventional name for
the
remote in simple setups.
Again, it would be more natural to
directly
mirror remote ref structure
locally:
but git strips the redundant heads component. And this
-heads
,
+remotes/$remote
mapping is built in, which compresses the command to
$ git fetch origin master
It’s worth reflecting
why
it works this way. Git model is offline first.
What’s more, it assumes
explicit
synchronization points. Rather than
synchronizing with the remote repository in background when there’s
connectivity, git requires explicit
fetch
and
push
operations to transfer
bytes over the wire. In this paradigm, it is useful to model the state of the
remote party at the moment when we talked to them the last time. Theory of mind!
This
hopefully
deconfuses git’s concept of local and remote branches. Consider
the
main
branch. It exists on the remote named
origin
as
refs/heads/main
.
When you synchronize your local repository with
origin
, you get
refs/remotes/origin/main
—
you current best knowledge about the the state of
main
on the
origin
.
And then there’s your local
refs/heads/main
. It typically starts pointing at
the same commit as
refs/remotes/origin/main
.
But, when you make a commit,
refs/heads/main
advances, but
refs/remotes/origin/main
stays the same.
When you try to push your local commit to origin, you will get a conflict, if
the
main
branch on the
origin
advanced in the meanwhile. In that case, git
automatically updates
refs/remotes/origin/main
(as that’s just a local mirror of the remote state), but then it’s on you to
update
refs/heads/main
and push it again.
The first command looks up the URL for the
origin
remote in
.git/config
and
makes a network request to that machine. As a result, the local
refs/remotes/origin/master
gets updated to the same commit as
refs/heads/master
remotely (the commit and its ancestors are transferred
locally as a result).
The second command creates a
refs/heads/my-feature
ref (a branch), whose
starting point is
refs/remotes/origin/master
. It is an example of a leaky
abstraction.
The second argument there is a (shorthand of a) ref, so you can do
You ask a server to create a Linux machine. The machine boots. The reply disappears. From the client’s chair, this looks much like a request that never reached the server. Retrying is reasonable. Starting a second machine would make the recovery more expensive than the failure.
This is a useful way into Mainbrella’s backend. Follow one machine far enough and the same problem keeps returning: a command can run without its caller seeing the result; a private request can outlive a network membership; a disk capture can succeed without leaving us the handle needed to restore it. Each boundary needs a decision about what a retry is allowed to do.
Let’s follow an illustrative job that runs
analysis.py
, reads data from another container, and writes
metrics.csv
. Its machine occupies slot
c17
. The filenames and slot are examples; the mechanisms below come from the
open-source backend
. We’ll start before Linux boots, because that is where ownership and budget have to be decided.
One place decides who gets the last slot
Two create requests can arrive while an account has one slot left. Reading a counter, booting a machine, and updating the counter afterward gives both requests a chance to win. By the time we notice, the expensive part has already happened.
We give each account a Cloudflare Durable Object: a persistent coordinator with its own storage. The public API resolves the authenticated user and paid entitlement, then addresses that coordinator as
account:<userId>
. Each machine slot has a separate Durable Object. For
c17
, its name is
user:<userId>:slot:17
. The first slot retains the older name
user:<userId>
and the public ID
small
; that ID does not select the machine’s size.
The account coordinator serializes admission. A request must reserve its slot, monthly start, and compute allowance before it can ask the runtime to boot. The next request therefore sees occupied capacity even while the first machine is still starting. The
account controller
implements the queue explicitly with a promise tail; an
await
inside an admission decision doesn’t let the next decision slip past it.
Holding this lock until every machine was ready would also serialize all their boot times. We release it after the durable reservation and provision outside it. One account makes the decisions in order; its runtimes can boot together.
The short shared decision precedes the longer independent work. Time is schematic: the overlapping bars explain concurrency, not measured startup latency.
There’s a particularly useful
test for this distinction
. It sends six concurrent requests on a Builder account, which has five slots, and holds every runtime’s readiness check behind a gate. Five machines reach boot before that gate opens. The sixth request receives a conflict, and the account records five starts. It checks both halves of the design: exclusive admission and parallel provisioning.
Ownership is established before this coordination begins. The
authentication adapter
accepts an API key or a login session. An explicitly invalid Bearer credential fails authentication even if a valid browser cookie accompanies it. The
internal request builder
supplies the account identity and entitlement itself. Letting the caller choose an
x-mainbrella-user
header would undo the account boundary we just built.
A receipt exists before the machine does
For our
analysis.py
job, the client supplies an
Idempotency-Key
with
POST /containers
. That key means “this particular attempt to create a machine.” The account records which slot and reservation belong to it, together with a fingerprint of the requested configuration. Changing the image, size, or other fingerprinted selection while reusing the key produces a conflict.
The important write is small. Here is the relevant part of
container-account-core.js
, with the surrounding validation omitted:
That multi-key write commits the charged reservation and creation receipt atomically. We don’t want a receipt for a slot we never reserved, or a charged slot that a keyed retry cannot find. Only afterward does the controller dispatch the boot.
Now lose the reply. A matching retry finds the receipt before attempting fresh admission. While the reservation is pending, it can return
starting
. Pending reservations have a ninety-second reconciliation window; after that, the coordinator asks the runtime what actually exists. If it finds the running machine, it returns that machine without charging another start. The retry never dispatches the ambiguous reservation a second time.
The receipt lasts twenty-four hours. If the machine has stopped or its slot has been reused, that same retained key returns
creation_no_longer_running
. It doesn’t quietly launch a replacement. A client that wants a replacement makes a new creation attempt with a new key. This is also why the client should save its key and request before sending them: server-side deduplication is little help if the client forgets the identity it needs to ask about.
What establishes readiness? The
runtime controller
starts the selected image with
sleep infinity
as its entrypoint, then executes
uname -a
. The command must exit successfully within sixty seconds. This establishes that the guest can execute a command. Our Python analysis still needs its own dependencies and checks; an answering kernel cannot certify a financial calculation.
Yesterday’s message can arrive at tomorrow’s machine
Suppose the dispatch for our first boot is delayed. Meanwhile, the account cancels the reservation, releases
c17
, and assigns that slot to a replacement. The old dispatch finally arrives. Its target is still the same runtime object. Looking up the slot by name cannot tell us whether this message is entitled to start anything.
Each admitted start receives an increasing reservation number. At the runtime, we persist two high-water marks: the newest accepted start and the newest cancellation. A boot at or below either mark is rejected. Cancellation writes its fence even when there is no running guest to destroy, so a later arrival cannot resurrect canceled work.
Follow the dashed diagonal: the first boot arrives after the replacement. The cancellation fence and accepted reservation make its age visible to the runtime.
The reverse race matters too. A delayed cleanup for reservation 41 must not stop the machine from reservation 42. The
runtime’s DELETE path
advances the cancellation mark but leaves the guest alone when the cancellation is older than the accepted start. Back at the account, the completion of a boot must still match the current slot reservation before it can clear pending state. An old successful reply has no authority over a replacement either.
Reservation numbers protect the internal lifecycle messages. Public commands and file requests identify a running generation with
{ id, createdAt }
. The slot ID is reusable; the generation is not. Despite its timestamp-shaped representation,
createdAt
is computed as
max(now, previousCreatedAt + 1)
. Recreating a machine in the same millisecond, or moving the wall clock backward, still gives the new guest a different identity.
Keep both values from the returned running container. A cleanup request for just
c17
cannot express which lifetime you mean. The lifecycle tests deliberately stop and reuse a slot before releasing an old boot dispatch, including without advancing the clock. This is a more revealing check than another successful hello-world launch.
The hard deadline is part of admission
Our machine also needs permission to keep spending compute. A monthly counter checked only at launch would let many simultaneous machines consume the same remaining allowance. We reserve the runtime they may use before any of them starts.
Machine sizes have weights in
plan-policy.js
: Lite uses one compute unit, Medium uses ten, and XL uses twenty-eight. The account reserves unit-milliseconds. Its lease ends at the earliest of four boundaries:
hard deadline = min(
start time + plan session limit,
paid access expiration,
next UTC month boundary,
start time + remaining unit-ms / machine weight
)
For an arithmetic example, reserving a Medium machine for one hour commits ten compute-unit-hours. Confirm a stop after five minutes and the consumed amount is
10 × 5 / 60
, about 0.833 compute-unit-hours; the unused reservation is released. A failed stop or unreadable runtime keeps its reservation. Treating “couldn’t contact it” as “it must be free” would allow the account to spend that allowance twice. A failed admitted launch still consumes its monthly start; runtime settlement is a separate calculation.
The runtime persists this hard deadline and an idle deadline. Real activity can move the idle deadline, capped by the hard deadline. Status polling doesn’t. A Durable Object alarm enforces expiration even when the client has gone away. Plan changes can shorten an existing lifetime, but cannot extend its original hard deadline.
Stopping work should remain possible when the billing lookup is unavailable. The public DELETE path authenticates ownership without requiring a fresh billing resolution; the coordinator uses its saved entitlement for cleanup when it remains valid. Likewise, a newer unpaid observation beats an older paid one. Otherwise a delayed check could reauthorize a machine we had already revoked.
A disconnected viewer shouldn’t own a process
With a running generation in hand, we can start
analysis.py
. A short foreground command has a sixty-second maximum. For work whose result we want to find after a disconnect, we use a managed execution. Its identifier belongs to a particular container generation, and its creation requires an idempotency key of its own.
// machine is the exact { id, createdAt } returned for a running guest.
// Persist executionKey and this request before sending them.
const query = new URLSearchParams(machine);
const response = await fetch(`${apiOrigin}/containers/executions?${query}`, {
method: 'POST',
headers: {
Authorization: `Bearer ${apiKey}`,
'Content-Type': 'application/json',
'Idempotency-Key': executionKey,
},
body: JSON.stringify({
argv: ['python3', '/workspace/analysis.py'],
timeoutMs: 120_000,
}),
});
if (!response.ok) throw new Error(`Execution HTTP ${response.status}`);
const execution = await response.json(); // Save execution.id.
The
argv
form supplies literal arguments instead of assembling shell text. In
executions.js
, the execution record is saved before the process starts. A matching retry finds that retained identity. An aborted creation request or a disconnected event stream doesn’t acquire the right to kill the managed job; explicit cancellation does.
Output events have increasing sequence numbers, committed alongside the record’s updated cursor. A client reconnects to the events endpoint with its last processed sequence as
cursor
. It reads the stored suffix rather than depending on a particular socket having seen every byte. The stream itself is bounded to thirty seconds, so reconnecting is ordinary operation. The process can run for up to fifteen minutes, further limited by its requested timeout and the container’s remaining hard lease.
These records are finite resources. Each runtime retains up to thirty-two execution records, with a one-hour retention deadline measured from job admission. Managed jobs share a pool of four concurrent operations with foreground commands and file transfers. Output is bounded to one MiB and a finite event count. Hitting the output limit terminates the job and marks it truncated; finding a few plausible lines in stdout is therefore insufficient. We inspect terminal status, exit code, timeout, and truncation before trusting the result.
Here is the uncomfortable boundary: durable records don’t make process handles durable. When a runtime object restarts, recovery marks unfinished executions
interrupted
. If they belong to its current guest generation, it destroys that guest rather than leave untracked work running. It never silently reruns the command. For our analysis this means an interruption can discard an unsaved CSV; for a command that sends invoices, automatic replay could do something worse. Restart recovery is deliberately more disruptive than reconnecting a viewer.
The data request arrives with an identity the guest can’t choose
Suppose
analysis.py
reads a shard from
http://data.internal/shards/west
. We register its generation as a member of a private service network. In that same network, another generation owns the service name
data
and port 8080. The source may be a caller-only member, with no listening port.
The guest’s request names neither an account nor a network.
private-services-runtime.js
installs an outbound HTTP interceptor for
*.internal
. Its relay removes reserved identity headers and supplies trusted account, slot, and generation values from the runtime entrypoint’s configured properties. Sending an invented ownership header from Python doesn’t select a different customer.
The account’s
private service registry
finds the network containing that exact source generation, then resolves
data
within it. Another account—or another network in this account—can reuse the name. The destination rechecks its running generation and registration under its lifecycle lock before opening the registered application port. Resolving the name is only the first permission check.
Why check again? Our data service could be detached or replaced while it is answering. The destination buffers the bounded response and checks its registration again. The account rechecks source liveness and both memberships before releasing the response. A request admitted under yesterday’s membership must not deliver bytes under today’s arrangement.
Those checks don’t undo an application side effect that already happened. If the request changed the data service before its response was denied, the application still needs a way to reconcile that change. Immutable shard reads are easy here; a task queue or payment service would need its own operation identities.
This feature is bounded private HTTP routing: one-MiB request and response bodies, a ten-second timeout, and no WebSocket upgrade or arbitrary TCP connection. A PostgreSQL client won’t become private-service-aware because its hostname ends in
.internal
. Deployment enablement and the capability advertised by
GET /capabilities
must also be present before we build a workflow around it.
The CSV needs a life outside the command’s output
Our job writes
/workspace/metrics.csv
. We retrieve it with a generation-bound file request while the guest is still running. The file endpoint transfers raw bytes, with a one-MiB limit, rather than decoding binary data as text or hiding an export inside a truncated stdout stream.
The write path in
files.js
demonstrates another useful ordering choice. It writes a temporary file in the destination’s existing parent directory, then renames it over the target. Readers shouldn’t observe a half-uploaded regular file. The path is passed as a positional argument, not interpolated into shell code. Writes reject symlink targets; reads may follow links inside the owned guest.
These guarantees are narrower than a transaction over the analysis. A successful file transfer doesn’t prove the job used the right input or finished all its rows. The application should check the artifact’s schema and provenance, and copy useful output outside the ephemeral machine before teardown. If we want to return to the environment itself, we need a saved workspace.
A snapshot can exist without being recoverable
Saving sounds like one operation: capture the disk, remember the result, stop the machine. It actually crosses three owners of state—the account, the runtime, and the provider—and none of them can atomically commit the other two.
In
workspaces.js
, the account reserves save capacity and persists the operation before requesting capture. At the runtime, a receipt containing the save identity is persisted before calling
snapshotContainer()
. After capture returns, the runtime saves the provider handle in that receipt. The account then commits the handle to its workspace record. Only after that commit may
stop: true
destroy the source.
Lose the reply between the runtime and account after the runtime has saved the handle, and a retry can read the receipt. We recover the same capture without taking another one. But interrupt the runtime after the provider accepts capture and before its handle is saved, and the receipt contains only an intent. We know we tried; we don’t know which provider object to restore.
The location of the lost information changes the recovery. A provider-side disk alone isn’t enough: we need its handle on our side of the durable boundary.
The runtime refuses to recapture that unresolved operation and returns
workspace_save_unavailable
. The save path leaves the source intact, subject to its ordinary lease. Choosing a new key just to make the error go away would be a new capture attempt, not recovery of the old one. That distinction is the limit of the idempotency promise.
Save quotas account for the cost of uncertainty too. Admission reserves the source size’s full disk capacity before capture. Deleting a workspace frees its live saved-workspace quota, but doesn’t refund the historical capture budget: the provider may already have done the work. A failed or ambiguous call cannot become a cheap way to repeat captures indefinitely.
Restoring a ready workspace goes through ordinary container admission and consumes a new start. The restored guest receives a fresh generation and must use the saved size and internet policy. The image digest must still match; an incompatible image or failed provider restore produces an error rather than silently substituting an empty machine.
The saved state is a filesystem. RAM, running processes, previews, and private-service memberships do not return with it. Our Python job needs progress in files if it is to resume, and a restored service must start its process and register its new generation. Quiescing writers before capture is an application responsibility too. We can preserve a disk full of mutually inconsistent files perfectly.
Make the late message arrive
The quickest way to inspect these choices is to run the lifecycle tests from a backend checkout. They delay dispatch, drop replies after side effects, reconstruct controllers from saved storage, and reuse slots before delivering the old messages:
Start with
stopping and reusing a slot fences old delayed dispatches even in the same millisecond
. Then read
lost snapshot response reconciles receipt without recapture
beside
uncertain capture cannot be repeated
in
the workspace tests
. The first recovers a saved answer. The second preserves the fact that an answer is missing. That is the decision an agent needs from its compute backend before it can safely decide what to do next.
Source review: backend commit
6e5bef3
, October 7, 2026. Examples and diagrams explain control flow; they are not production traces or latency measurements. The lifecycle tests use simulated provider state. Public request shapes and deployment capabilities are documented in
API.md
and the
agent workflow
.
Ask multiple AI models for a comparable opinion, and Chinese models seem to be more likely to respond in line with the wider consensus
.
I recently built/vibed a
small web tool
that asks a panel of 12 AI models any kind of “top three” question and tallies the answers like an Olympic medal table. The idea is to have a quick way to fetch hivemind recommendations on anything from a good book to read, things to do in a city, or a low-stakes purchase.
The system prompt is simple and built around this request:
You are answering a ranking question with
your own independent opinion
. Give your top 3 answers,
ranked from 1 (best) to 3
.
The call explicitly asks each model for its own view, not to guess what the crowd thinks.
There is some normalization required: if we ask for the best castles in Ireland and Luna votes “Rock of Cashel” while Gemini votes “Cashel Rock”, we want to treat that as a consensus. The app uses Gemini Flash to adjudicate this.
The results – from conformist to contrarian
A word on reading the chart. Long bars and a winner at the top make it look like a leaderboard, as if more consensus is better. It isn’t necessarily. However, having clicked around a few of the queries: outlier responses are more likely to be
interesting
than to be an insightfully correct response that bests the responses of all the other models. I like that Gemini recommends you paint an adult bedroom terracotta while all the other models recommend green, cream, and greige, but I understand that many people will want their bots to stick to the cultural mainstream thank you very much.
Each model is scored against the consensus of the other eleven, getting full credit for the right item in the right podium spot and half credit for the right item in the wrong spot. A score of 1.0 would mean perfect agreement with the winning conclusions of the rest of the group; 0 would mean total contrarianism.
The top four conformists are all Chinese: DeepSeek, Kimi K2, Tencent’s Hy3 and GLM.
Group the scores by country and the gap holds. The five Chinese models average 0.47, the six American models 0.38, and Europe’s lone entrant, Mistral, scores 0.36.
The exception is Qwen, Alibaba’s model, which scores 0.38 and sits among the American models, just behind Grok and Nemotron (both 0.40). At the other end, the two most contrarian models are Meta’s Llama 4 Scout (0.34) and Anthropic’s Claude Haiku (0.32).
Detailed table:
Model
Maker
Region
Model ID
Released
Knowledge through
Reasoning
Score
DeepSeek
DeepSeek
CN
deepseek-v4-flash-0731
2026-07
2025-05
Default
0.59
Kimi K2
Moonshot
CN
kimi-k2
2025-07
2024-12
Default
0.47
Hy3
Tencent
CN
hy3
2026-07
Not published
Off
0.46
GLM
Zhipu AI
CN
glm-5.3-flash
2026-08
Not published
Off
0.45
Gemini
Google
US
gemini-3.1-flash-lite
2026-05
2025-01
Default
0.42
Grok
xAI
US
grok-4.3
2026-04
2025-12
Off
0.40
Nemotron
NVIDIA
US
nemotron-3-super-120b-a12b
2026-03
2025-06
Off
0.40
Qwen
Alibaba
CN
qwen-plus
2026-08
2025-03
Default
0.38
Luna
OpenAI
US
gpt-5.6-luna
2026-07
2026-02
Off
0.37
Mistral
Mistral
EU
mistral-small-2603
2026-03
2024-11
Default
0.36
Llama 4
Meta
US
llama-4-scout
2025-04
2024-08
Default
0.34
Haiku
Anthropic
US
claude-haiku-4.5
2025-10
2025-02
Off
0.32
Why would Chinese models agree more?
I don’t know. Here are some explanations I find plausible, roughly in order.
1. Distillation
Distillation means training a model on another model’s outputs, essentially a cloning technique. US labs, including OpenAI and Anthropic, have publicly accused Chinese labs of distilling their models at scale. Earlier this year Anthropic
named DeepSeek and Kimi-maker Moonshot
(and also MiniMax, which is not represented in the sample here) as running large distillation campaigns against Claude. That does not establish that distillation caused the pattern here, or even that these models were trained that way. But if a model was trained heavily on outputs from other models, one expected outcome is that its preferences converge with those of source models, especially when they converge. A model built that way is partly a consensus machine. It is striking that DeepSeek and Kimi, first and second in the table, are among the labs most often named in those accusations.
2. Chinese models learning from each other
The Chinese open-weight ecosystem is tightly knit. Labs release weights, publish methods and train on synthetic data produced by each other’s models. That could be giving the five Chinese models a family resemblance.
This matters because of how the score works. Five of the twelve models here are Chinese, so if they share a style, they help build the consensus they are scored against. A bloc that votes together will always look conformist. (Qwen breaking from the bloc is a point against this, or at least a sign the family isn’t that close.)
3. Model size and tier
The bottom of the table is crowded with smaller, cheaper models: Llama 4 Scout, Claude Haiku and Mistral Small. Smaller models hold less of the shared cultural canon in memory, so their picks are more scattered. Gemini Flash-Lite and DeepSeek’s Flash model are counter-examples, so size can’t be the whole story.
4. Knowledge cutoff
About
half of new text appearing
on crawl-worthy URLs is written by humans, and half by AI. A model trained more recently has read more of the slop. Most of the Chinese models here are mid-2026 releases. Again there are counter-examples: Kimi (knowledge through 2024-12) is second, while Luna has the newest cutoff on the panel and sits ninth.
5. Post-training personality
Perhaps US labs are investing in post-training in a way that gives more distinctive character, more willingness to pick the less obvious answer. If there’s a trade-off between personality and benchmark-maxxing that may also play into the flipside, where the Chinese labs are plausibly focusing more on hitting benchmarks.
6. The English canon
All prompts are in English. A model whose English training data leans on the most widely syndicated sources would tend to give the most canonical English-language answers, which is exactly what consensus rewards.
Caveats
More work needed, because…
Small dataset.
This is based on the first few hundred questions from my own usage and some humans in the wild who clicked from Reddit.
Arbitrary panel composition.
The measure of which answers are normal is itself defined via the panel’s composition. 6 US, 5 Chinese, and 1 European was an arbitrary choice by me based on my sense of the industry breakdown. More practical decisions were to have max one model per lab and avoid expensive models. Plausible additions would include models from MiniMax and ByteDance. At a stretch, and in the spirit of geo-diversity, the panel could also include Aleph Alpha (Germany), Apertus (Switzerland), Cohere (Canada), Falcon (UAE), Sarvam (India) or LG’s Exaone (South Korea). Adding flagships like Claude Sonnet, alongside the cheap tiers, might also show whether price tier matters.
Agreement isn’t accuracy.
A high score means a model says what others say, not that it’s right. On “best” questions there is often no right answer at all.
English-language questions.
The questions so far lean towards an English-speaking, Western frame.
Conclusion
Back to that bedroom recommendation:
11 x models recommended green, cream or greige
1 x model recommended terracotta
Which of these is more correct, and which is behavior we want to encourage? I suspect most of us want an AI ecosystem that isn’t a groupthink borg. We like a bit of Temperature. We want models to disagree, surprise us, occasionally push for the terracotta bedroom. I also suspect that, as users, we tend to reward models that give us answers that feel safe. These two instincts are going to be pulled apart by a lot of money over the next few years.
Multiply this out to millions of users seeking suggestions and it becomes something bigger: a quiet, constant pressure towards a lab-mediated consensus middle.
Commercially, being included in that consensus middle is already lucrative. For two decades companies have poured huge amounts of capital into efforts to rank well in search engines (SEO), and the smart ones have already switched to competing over placement in AI responses (AIO/GEO). Self-reinforcing loops will escalate the value of this prize. If the labs are competitively exfiltrating each other’s responses, getting a product highly placed in one model’s answers could enable it to flow into the next model’s training data, and the next. Consensus that was gamed once can be copied until it looks like a fact about the world.
Paint colours are a harmless starting point, but the same machinery will answer questions about which candidate to trust, which news source is reliable, which history is true. Decorating recommendations are obviously not equivalent to political or factual judgments, and models may behave very differently across domains, but the same basic machinery increasingly mediates all of them. When the models are copying each other’s opinions and feeding those opinions back into the real world, a consensus doesn’t need to be right. It only needs to be first.
For more decorating or political advice, you can ask the robots for recommendations
here
.
Implementation notes: The twelve models were called through OpenRouter, mostly in their cheaper “flash” or small variants, with reasoning switched off where the provider allows it. “Default” means the model was run with whatever reasoning behaviour the provider ships. I did not touch the Temperature parameter, nor did I explore what defaults the different models use.
Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released.
Claude Haiku 5.5 is designed for high-volume, cost-sensitive tasks. It reliably handles quick and repetitive workloads (like summaries, compactions, database queries, and classification requests). It pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work. And, since it’s also our fastest model to date, it works especially well for speed-sensitive tasks like live customer support and browser use.¹
Haiku 5.5 is available at a much lower price than Haiku 4.5. On average, it now costs around 75% less to run.²
Along with this launch, we’re making improvements to the value of our model range. We’re halving the price of Claude Sonnet 5.5’s cache reads, which means Sonnet 5.5 now runs around 20% cheaper on most agentic work. And we’re introducing a new monthly API credit for our Claude Max and Team subscribers, designed to support our users in building new agents and applications that run on the Claude Platform.
Performance
Here’s how Claude Haiku 5.5 performs across a range of benchmarks:
Haiku 5.5 is our first Haiku-class model to come with an adjustable effort setting. This means that, as with our other models, users can decide whether to optimize for cost or intelligence. The charts below show how Haiku 5.5 performs on three benchmarks at each effort setting:
OSWorld 2.1 (offline subset)
Accuracy vs. cost
OSWorld 2.1 measures how well agents can operate a real computer to finish long, multi-step tasks.
In early testing, our customers reported results consistent with the performance and cost improvements shown above. Here’s what they told us about the new model:
Quote
“We’re very impressed with Claude Haiku 5.5, particularly its speed. We ran it through our eval suite for AI Teammates, our AI agent product, covering use cases like triaging bugs, setting up projects, and searching large portfolios to surface high-risk or overdue work. Compared with the model we use today, we saw over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn. It’s a noticeably snappier experience.”
Company
Asana
Author
Aaron Vinh, Staff Software Engineer
Pricing
The table below shows how Claude Haiku 5.5’s pricing compares to our other models. Haiku 5.5 is especially good value when used for tasks with prompts up to 100,000 tokens, which make up around 90% of requests to our previous Haiku model.
Price per 1 million tokens
Haiku 5.5
prompts up to / over 100k
Haiku 4.5
Sonnet 5.5
Cache reads
$0.01 / $0.05
$0.10
$0.10
Cache writes
$0.125 / $0.625
$1.25
$2.50
Input tokens
$0.10 / $0.50
$1.00
$2.00
Output tokens
$0.50 / $2.50
$5.00
$10.00
Safety
Alignment.
Claude Haiku 5.5 shows major improvements across almost all of our alignment evaluations relative to Haiku 4.5. In particular, we found far fewer instances of misaligned behavior, and a lower willingness to cooperate with misuse. The model’s
system card
describes our evaluation process and results in more detail.
Safeguards.
Consistent with its capabilities, Haiku 5.5’s cybersecurity safeguards are more restrictive than Haiku 4.5’s, but somewhat less restrictive than those we’ve applied to other recent models. In cybersecurity, they permit a wider range of defensive tasks than our safeguards for Sonnet 5.5, but they still block penetration testing and other techniques more likely to be used by attackers.
Haiku 5.5’s biology safeguards are the same as for Sonnet 5, Sonnet 5.5, and Opus 5. They allow research biology questions but restrict access to requests that we judge as likely to cause harm. Organizations working on wider-ranging biology and cyber activities can apply to our
Life Sciences Verification Program
and
Cyber Verification Program
.
Availability
Claude Haiku 5.5 is available now on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. On the Claude Platform, developers can get started with
claude-haiku-5-5
.
Alongside our new pricing for Claude Haiku 5.5, we’re making further improvements to the value of our models and products.
First, starting today, we’re
lowering the price of cache reads on Claude Sonnet 5.5
. Cache reads now cost 50% less: $0.10 per million tokens rather than $0.20. Because cache reads make up a large share of models’ token consumption, this reduces the cost of Sonnet 5.5 on most agentic tasks by around 20%.
For instance, here’s what the price cut means for Sonnet 5.5’s performance relative to cost on Terminal-Bench 4.0:
Terminal-Bench 4.0
Accuracy vs. cost
Terminal-Bench 4.0 measures how well a model can complete complex, multi-step professional tasks within a command-line interface.
This chart illustrates an important difference between Haiku 5.5 and our larger models. Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks like those measured by Terminal-Bench 4.0. By contrast, Haiku 5.5 is best suited to more narrowly scoped tasks that might otherwise have been cost-prohibitive with previous versions of Claude—like compaction, summarization, or subagent work.
Second, this week, we’ll roll out
a new monthly API credit to all Max and Team subscribers for use on the Claude Platform
. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users. These credits are designed to allow our users to experiment with building tools, apps, and agents that call our API. They can be used on any of our models. For more information,
see our Help Center article
.
For developers, we’re also
updating our Claude Python and TypeScript SDKs to add support for computer use and browser use
in beta. Haiku 5.5 is especially well-suited to these tasks, given its combination of speed, capability, and price. You can read more about this
in our Claude Platform docs
.
rat
is my smallish compiler backend (with a semi-working
C99 frontend). Its x86-64 code generator
translates the
intermediate
representation
(IR) into x86-64 instructions. These use an unlimited number of virtual
registers (vregs). The register allocator maps each vreg to a physical register: a general-purpose register (12 can
be used) or an xmm register (14 on Linux
1
). When no register is free,
it maps the vreg to a stack slot.
For a long time rat used a
linear
scan
allocator, which visits live ranges in program order. It worked, but it grew one fix at a time, to
1392 lines. So I measured which of its parts helped, threw the rest away and wrote a priority bin-packing
allocator in 584 lines. It visits live ranges by importance and puts each one in the first register where it
fits. It is the same family as LLVM's
greedy
allocator
, minus most of the hard parts, and it makes
better code
.
A value is
live
from where it is written to where it is last read. Two values can share a register
only if they are never live at the same time.
When too many values are live at one point, some go to memory: they are
spilled
. A spill
costs a store and a load. The best assignment is NP-hard to find
2
,
so all practical allocators use heuristics.
The
calling convention
adds two rules. A call can overwrite the
caller-saved
registers
(
rax rcx rdx rsi rdi r8-r11
and all xmm registers on Linux). A function must restore the
callee-saved
registers (
rbx rbp r12-r15
) before it returns. As an example, this
function keeps
y
live across a call:
longg(long);
longh(long x, long y) {
long t = g(x);
return t + y;
}
Before allocation,
rdi
,
rsi
and
rax
are fixed by the calling
convention, and
v1
-
v4
are vregs:
0 v1 = copy rdi ; x
1 v2 = copy rsi ; y
2 rdi = copy v1 ; argument of g
3 call g ; clobbers caller-saved
4 v3 = copy rax ; t
5 v4 = copy v3
6 v4 = add v4, v2
7 rax = copy v4
8 ret
x86
add
writes over its first operand (two-address), so instruction 5 copies
t
first. After allocation, at -O1:
push rbp
mov rbp, rsp
sub rsp, 0x8
push rbx ; rbx is callee-saved: save it
mov rbx, rsi ; y
call g ; x is already in rdi
add rax, rbx ; t stays in rax
pop rbx
leave
ret
Five of the six copies are gone, and
y
went to a callee-saved register. No code in the
allocator says "put values that cross a call in callee-saved registers". It falls out of the design, and
that is my favourite part.
Five steps
The allocator runs five steps per function:
Live ranges
: number the instructions and find where each vreg is live.
Fixed registers
: mark where the code uses physical registers
directly.
Coalescing
: join vregs that a copy connects into one group (a
bundle
3
), so the copy can go away.
These parts make real allocators big. My
measurements
say rat does
not miss them much.
Live ranges
Slots
Instruction
i
gets two slots: it reads its operands at
2i
and writes its results
at
2i+1
. A live range is a sorted list of
[start, end]
slot segments.
Where a source ends depends on the instruction:
Copies:
the source ends at the read slot, and the destination starts at the write
slot. In instruction 2,
rdi = copy v1
,
v1
ends at slot
4 and
rdi
starts at slot 5. They do not
overlap, so they can share a register and the copy becomes a no-op.
Other instructions:
a source stays live through the write slot, so a result never overwrites a
different operand.
v2
is written by instruction 1 and last read by the
add
at instruction 6, so it lives in
[3, 13]
.
Live-out sets
rat finds the vregs that are
live-out
of each block: a later block can still read them. Many
compilers do this with one bitset per block and a
fixed-point loop
. rat does
one vreg at a time instead:
The vreg is live into each block that reads it before it writes it.
From each such block, a worklist goes back through the predecessors and marks the vreg live-out in
each.
The walk stops at a block that defines the vreg.
The cost grows with the blocks where each vreg is live, not with
blocks * vregs
.
4
Segments and weights
Then rat walks each block backward from its live-out set and makes the segments. The same walk sums a
weight
per vreg: the cost of its spill.
Each def and each use adds
3
d
, where
d
is the loop depth (up to 11):
def or use in
adds
straight-line code
1
a loop
3
a doubly nested loop
9
Holes
A live range can have
holes
, gaps where the vreg is dead. Blocks are numbered in code order, so a range
that skips a block has a hole there:
longf(long* a, long n) {
for(long i = 0; i < n; ++i)
if(a[i] < 0)
a[i] = 0;
return n * 3;
}
The exit block sits between the loop blocks:
mov eax, 0x0 ; offset 8*i, rax in the loop
cmp rdx, rdi
jl loop
exit:
lea rax, [rdi+rdi*2] ; n*3 in the hole of rax
ret
loop:
mov rcx, r8
add rcx, rax
...
add rax, 0x8
cmp rdx, rdi
jl loop
jmp exit
The offset in
rax
is dead in the exit block, so
n*3
(one
lea
) can use
rax
,
which is also the return register. A free win from block order.
Fixed registers
rat numbers its registers 1 to 40, so one
U64
holds a set of them. Each slot gets one mask,
busy[slot]
. A set bit means that register is busy at that slot.
The same backward walk marks the physical registers the code uses directly:
r
and
w
are the read and write slots.
#
is busy,
=
is a
live range and
.
is free. A bar continues across the gap between instructions. "others" is
every other caller-saved register.
When a bundle gets a register, rat sets that register's bit in every slot of its live range. After that,
vregs and fixed registers are bits in the same masks. Each group of 64 slots also has a summary mask, the
OR of its 64 masks, so a long range can skip 64 slots at a time.
Coalescing
A copy between two vregs of the same class is a candidate for
coalescing
. These
copies come from:
two-address instructions
phi nodes
: a
value that comes from different blocks at a join point
If the two live ranges do not overlap, the vregs become one bundle. It has the merged segments and the summed
weight. rat deletes a copy inside one bundle. In
h
,
v3
is
[9, 10]
and
v4
is
[11, 14]
, so
they merge.
rat sorts the copies by loop depth, deepest first. Hot copies merge before cold copies can block them.
A merge first walks both segment lists to check for overlap. rat skips a merge when the two bundles
together have more than 256 segments.
A copy between a vreg and a physical register sets a
hint
instead: the bundle prefers that register
if it is free.
Picking registers
Each bundle gets a priority:
priority = weight / sqrt(length in slots)
Short, hot ranges come first: they matter most and are the easiest to place.
Long, cold ranges come last and get spilled.
sqrt
keeps a long loop counter from losing too much priority.
rat calls
pick
on each bundle in priority order:
// cls: register class, gp or xmmPhysRegpick(VReg v) {
U64 blocked = ~allocatable[cls];
for(auto [start, end] : segs[v])
for(I32 s = start; s <= end; ++s)
blocked |= busy[s]; // or 64 at a timeif(hint[v] != kNoReg && !(blocked >> hint[v] & 1))
return hint[v];
// caller-saved first, callee-saved lastreturn firstFree(order[cls], blocked);
}
Picking in
h
bundle
hint
gets
v1
rdi
rdi
v3+v4
rax
rax
v2
rsi
rbx
In the diagram,
rdi
is busy only before and after
v1
, so
v1
gets it.
Both copies become
mov rdi, rdi
, and the
peephole pass
deletes them after
allocation.
v2
crosses the call. Every caller-saved register is busy in the call slots, so the first free
register is
rbx
, the first callee-saved one. The prologue saves
only the callee-saved registers rat used.
On Linux, no xmm register is callee-saved, so a float that crosses a call always goes to
the stack.
Spilling
A bundle with no free register is spilled for its full lifetime. Then:
rat sorts the spilled bundles by start. It reuses a stack slot when the last bundle in it has ended.
rat rewrites the code. A vreg whose bundle got a register becomes that register.
Before each instruction, rat loads each spilled operand into a temporary register. After it, rat stores
each spilled result.
The temporary is
r10
or
r11
(
xmm14
or
xmm15
for floats).
No bundle ever gets these. If both are busy, rat takes the first register free at that
instruction.
Two cases need no temporary. A copy between a register and a spilled bundle becomes the load or the store
itself. A call reads a spilled stack argument from its stack slot directly.
16 registers minus
rsp
,
rbp
,
r10
,
r11
and
rdi
(which holds
a
) leaves 11 for 14 values.
x0
-
x6
also
hold the products (two-address
imul
), so they have more
uses. Of
x7
-
x13
, the three with the longest ranges go to the stack.
No instruction between the store of
x9
and its reload writes
r10
. So the peephole
pass deletes the reload. Then nothing reads that stack slot, so it also deletes the store.
It can be dumb
Without eviction, an early decision is final. Here is the case that annoys me most:
longsum(long* a, long n) {
long s = 0;
for(long i = 0; i < n; ++i)
s += a[i];
return s;
}
rat compiles it to:
mov r9, rdi ; a: rdi was taken by a[i]
mov r8, rsi ; n: rsi was taken by s
...
exit:
mov rax, rsi ; s: rax was taken by a+8*i
ret
loop:
mov rax, r9
add rax, rcx ; rax = a + 8*i
mov rdi, [rax] ; rdi = a[i]
add rsi, rdi
...
The loop values are short and hot, so they go first:
The address
a+8*i
takes
rax
.
s
loses its hint
rax
and takes
rsi
.
a[i]
takes
rdi
.
a
and
n
come last and lose their hints too.
The result is three movs, all outside the loop.
5
An allocator with
eviction would fix this chain. I decided three cold movs are not worth the extra code.
No eviction, no splitting, no second pass, and the new allocator still beats the old one. Most of the
quality comes from cheap things: coalescing, hints, holes and use weights.
The lesson for me: measure the old code before I port it. Much of the old allocator did nothing.
On Windows, only
xmm0
-
xmm3
can be used.
xmm4
and
xmm5
are the spill temporaries, and rat does not use the callee-saved
xmm6
-
xmm15
.
[back]
Chaitin et al.
showed that any graph can be the interference graph of some program. So register allocation is at least as
hard as graph coloring.
[back]
‘Jonathan’ Is the Oldest Land Animal on Earth. He Could Hold the Secrets to Defying Death
403 Media
www.404media.co
2026-10-07 14:00:38
Jonathan, a 194-year-old Aldabra giant tortoise, was born before the Civil War. Scientists have now sequenced his genome to reveal what’s behind his extreme longevity, so we can live longer too....
Jonathan the tortoise, the world’s oldest living land animal, is still chugging along at the ripe age of roughly 194. Now, scientists have discovered some of the secrets to Jonathan’s longevity by sequencing his genome for the first time, according to
a study published Wednesday
in
Science Advances
.
Jonathan, a Aldabra giant tortoise (
Aldabrachelys gigantea
), hatched around 1832 in the Seychelles archipelago, and he was already in his twenties when Charles Darwin published
On the Origin of Species
. In his fifties, Jonathan was brought to the remote island of Saint Helena, along with three other tortoises, as a gift to the British colonial governor in the 1880s. He has lived there ever since.
Aldabra giant tortoises often become centenarians, but Jonathan is a major outlier even for his species as he approaches nearly 200 years on this planet. To better understand his extreme lifespan, researchers evaluated Jonathan’s genome based on swabs of his cheeks and compared it with the genomes of several other tortoises, including a 36-year-old named Tank.
The results revealed that Jonathan has “unique gene variants in most aging pathways” including DNA repair, insulin regulation, mitochondrial function, and “low entropy” methylation, reports the study.
“Aging is very complicated, and there's all these different hallmarks,” said Stephen Clark, an aging researcher at the
non-profit Kallel Foundation
and an author of the study, in a call with 404 Media. “That's why this Jonathan study is so interesting. We're still just at the very beginning of trying to understand aging.”
The study was not an easy lift. Clark first reached out to Jonathan’s caretakers in 2017, but it took some negotiating to figure out the best way to obtain the tortoise’s genetic information. A blood draw was deemed too risky, so Jonathan was coaxed into opening his mouth for a swab. The results were delayed by incomplete early samples, the Covid pandemic, and new regulations governing access to Jonathan.
🌘
Subscribe
to 404 Media to get
The Abstract
, our newsletter about the most exciting and mind-boggling science news and studies of the week.
After years of back and forth, the samples were delivered from Saint Helena to Florida by the U.S. Space Force. The results showed that Jonathan had many gene variants the team expected to see, based on studies of other long-lived individuals and populations. These included variants associated with DNA repair, maintenance of aging-related sequences called telomeres, and the process of cellular recycling known as autophagy.
But the researchers found a surprise in Jonathan’s “methylome,” which is a set of methyl-based modifications involved in gene expression. The subset of Jonathan’s methylome related to mitochondrial processes was unusually pristine, or “low entropy.”
Photo: St. Helene Government, Communication Hub
“Jonathan's really the first of his kind to have this analysis, and we were also able to do the methylome for the younger tortoises,” Clark said. “We actually compared him to a five-year-old tortoise, so we had 189 years of difference between the methylomes. No one's ever done that, even in humans, because no humans live that long. That was one of the main things we were excited about: how does this methylome change with age?”
Whereas most of Jonathan’s genome showed signs of entropy commensurate with his age, his mitochondrial methylome seemed on par with that of the five-year-old tortoise, signaling that these pathways may play a special role in his advanced age. However, Clark said it would take more research to follow up on that lead, and to untangle the complicated web of genetic components that influence aging and lifespan.
“We're interested in long-lived species like Jonathan because we can hopefully find other targets,” Clark said of his team’s work at Kallel Labs. “The ultimate goal of the company is to affect human healthspan and lifespan.”
As for Jonathan, he continues to live out his voluminous days at the governor’s residence, Plantation House, where he and his fellow tortoises have become one of the island’s tourist attractions. Saint Helena is also known for the exile of Napoleon Bonaparte, who died there a mere decade before Jonathan is thought to have been hatched. It’s amazing to think about just how much the world has changed in Jonathan’s lifetime. Long may he graze.
In Vienna and Beijing, the first (thorium) nuclear clocks begin to tick
Build, run, and share AI agents with a declarative YAML config, rich tool ecosystem, and multi-agent orchestration.
What is Docker Agent?
docker-agent
lets you create and run intelligent AI agents that collaborate to solve complex problems — no code required.
docker-agent
is a
docker
CLI plugin and can be run with
docker agent
.
Define agents in YAML, give them tools, and let them work.
agents:
root:
model: openai/gpt-5-minidescription: A helpful AI assistantinstruction: | You are a knowledgeable assistant that helps users with various tasks. Be helpful, accurate, and concise in your responses.toolsets:
- type: mcpref: docker:duckduckgo
docker agent run agent.yaml
Key Features
Multi-agent architecture
— Create teams of specialized agents that delegate tasks automatically
Rich tool ecosystem
— Built-in tools + any
MCP
server (local, remote, or Docker-based)
AI provider agnostic
— OpenAI, Anthropic, Gemini, AWS Bedrock, Mistral, xAI,
Docker Model Runner
, and more
Advanced reasoning
— Built-in think, todo, and memory tools
RAG
— Pluggable retrieval with BM25, embeddings, hybrid search, and reranking
Package & share
— Push agents to any OCI registry, pull and run them anywhere
Install
Docker Desktop
(4.63+) — docker-agent CLI plugin is pre-installed. Just run
docker agent
.
Homebrew
—
brew install docker-agent
. Run
docker-agent
directly or symlink the binary to
~/.docker/cli-plugins/docker-agent
and run
docker agent
.
Binary releases
— Download from
GitHub Releases
. Symlink the
docker-agent
binary to
~/.docker/cli-plugins/docker-agent
to be able to use
docker agent
, or use
docker-agent
directly.
export OPENAI_API_KEY=sk-... # or ANTHROPIC_API_KEY, GOOGLE_API_KEY, etc.
See
Set Up a Model
for the full walkthrough of both paths (cloud API key or local model).
Quick Start
# Run the default agent
docker agent run
# Run from an OCI registry
docker agent run myorg/agent:tag
# Generate a new agent interactively
docker agent new
# Run your own config
docker agent run agent.yaml
Feds Charged Activists Just Before Major Anti-ICE Protest in a Tiny Vermont Town
Intercept
theintercept.com
2026-10-07 13:45:50
Authorities are taking unprecedented steps to charge protesters in a sleepy Vermont town that hosts an ICE intelligence and targeting center
The post Feds Charged Activists Just Before Major Anti-ICE Protest in a Tiny Vermont Town appeared first on The Intercept....
Williston, a small
Vermont town of 10,000 people, is best known as a bucolic suburb of Burlington laden with historic buildings and farmland.
Like the rest of the state, the hamlet is about 90 percent white. Williston has seen few of the sorts of large-scale Trump administration immigration operations that have sparked raucous protests. Aside from a few raids against migrant farm workers, there has been little in the way of detentions by U.S. Immigration and Customs Enforcement.
The town is, in other words, a surprising scene for the robust protests that erupted there over the past two years, with hundreds protesting against ICE.
The explanation lies in a low-slung modern building with façade of glass and dark-blue ribbed sheet metal: ICE’s
National Criminal Analysis and Targeting Center
, a facility designed to help identify, track, and arrest deportation targets across the country. Along with other Department of Homeland Security infrastructure in the area, the NCATC made Williston into a hub of anti-immigration enforcement.
The protests, however, weren’t initially met with the sorts of
crackdown
against
speech
experienced by
other
anti-ICE
protesters
. The local prosecutor for Chittenden County, where the surveillance center is located, repeatedly opted
not to prosecute protesters
arrested at the facility.
As the Trump administration ramped up its attacks against left-wing protesters, however, Williston was not spared. Beginning in September, state and federal authorities have found ways to levy criminal charges against activists.
“Our assumption has been that they did this so close to the week of action in order to try to intimidate people.”
In late September, Vermont Attorney General Charity Clark introduced surprise
misdemeanor charges
against 13 protesters for allegedly participating in a July sit-in at the NCATC. Then, last week, federal prosecutors
unveiled surprise charges
against five people, some of whom were already being charged by Clark. The five people facing federal charges, ranging in age from 22 to 80 years old, are alleged by federal authorities to have disrupted operations at the NCATC on multiple occasions.
Activists in Vermont took note of the timing: The cases came as they ramped up to launch a concerted week of protest actions against the ICE intelligence hub.
“Our assumption,” said Julie Macuga, an anti-ICE organizer involved in the protests, “has been that they did this so close to the week of action in order to try to intimidate people, in order to make people think twice about doing civil disobedience, for fear that they’re going to get charged.”
ICE in Williston
As
ICE’s budget
has
ballooned
under the Trump administration, the agency poured billions of dollars into development projects designed to expand surveillance capabilities. With astronomical goals for its deportations and a growing crackdown against dissenters, the surveillance infrastructure would provide the agency with a network of spyware for
tracking migrants
, protesters, and political opponents.
“What we’re seeing is this really huge increase in access to surveillance at both the national, state, and local level,” Michelle Dahl, an attorney who serves as the executive director of the Surveillance Technology Oversight Project. “We’re seeing that play out now in these types of surveillance dragnet centers.”
The NCATC is one of only two facilities like it in the nation. It is designed to rapidly respond to requests for information from field agents by using both open-source intelligence gathering and advanced surveillance tools to
scour the internet
and
trawl social media
for leads.
Although ICE has had facilities in Williston for years, significant efforts have been made to expand the scope of facility’s operations under the Trump administration.
According to government contractor recruiting documents first posted by ICE in 2025 and reposted this August, the NCATC is
designed
to employ at least a dozen analysts to answer requests 24 hours a day, coming up with target information in as little as 30 minutes. The facility operates in tandem with the office running ICE’s national tipline and its National Law Enforcement Support Center — two other major ICE facilities based in Williston.
Dahl said that despite its location, the gathering of resources in Williston is characteristic of the
“fusion center” model
designed to
allow
rapid
dissemination
of
information
between national facilities, federal agents, and local police departments.
“The fusion centers don’t have to be located in any place to collect surveillance — that’s how highly networked surveillance is across the country,” she said. “They could put it literally anywhere and be able to access the entire nationwide network.”
“No Real Precedent for It”
As a result of ICE’s expansion in Vermont, the NCATC has quickly become the subject of one of the the state’s largest recurring protests. Over the past year, hundreds of people attended rallies outside the center.
Some demonstrators took up nonviolent direct action. In one instance, clergy entered the halls and
blocked an entrance
. Dozens were arrested as part of the protests, but none were charged — until late September, when Clark, Vermont’s attorney general, began taking up cases that Sarah George, the state’s attorney in Chittenden County, had declined to pursue.
“We had been checking the docket and were under the impression that charges were not being pursued,” recalled Roan Wade, who was one of the 13 people charged for protesting at the facility. “It was definitely kind of a shock to have that happen.”
It later emerged that the charges had stemmed from an unusual move: The Vermont State Police began circumventing George and instead took their cases directly to the attorney general’s office. Unlike in other cases, the Vermont State Police never delivered these the charges to George’s office.
“There have been cases that Sarah George has declined, but these 13 cases were never even brought,” said state Sen. Tanya Vyhovsky, a member of the Senate Judiciary Committee, adding that the unprecedented move appeared to be legal in Vermont. “This type of stepping outside of our normal court system, as far as I can tell — and I also spoke to actual attorneys in the system — there’s no real precedent for it.”
“Extremely Concerning”
The decision to prosecute protesters for anti-ICE activity quickly ignited a firestorm among the Vermont Democratic Party establishment — which Clark, the attorney general, belongs to.
Within days of the decision becoming public, the state Senate majority leader, treasurer, and a host of representatives, candidates, and local officials all
issued condemnations
of the charges and Clark’s unusual actions.
“I found that decision to be extremely concerning,” Nikhil Goyal, a state Senate candidate in one of the districts containing charged protesters, told The Intercept. “There is enormous precedent, both in and out of our state, where prosecutors don’t use precious resources to go after peaceful protesters simply trying to stop a fascist machine from kidnapping our neighbors.”
“The same folks who claim to be opposing Trump and fighting this fascist administration are also prosecuting protesters.”
Wade, one of the charged protesters, says that the decision to prosecute is emblematic of wider trends among the national left, who have at times been starkly divided in their treatment of anti-ICE protests.
“It highlights some of the hypocrisy of the Democratic Party,” they said. “The same folks who claim to be opposing Trump and fighting this fascist administration are also prosecuting protesters.”
Soon, the federal charges would come down against the five additional protesters.
The Eve of Protests
The charges, both in state and federal court, were filed just ahead of the beginning of a national “
week of action
” organized by activists opposed to ICE’s operations in Vermont.
Organizers say that they expect hundreds of participants and that the protests are slated to be the largest instances of direct confrontation to ever target ICE’s national surveillance center. Some organizers say the arrests are only helping to fuel further activism.
“All it’s really done is bring a lot more attention to the issue,” said Macuga, the protest organizer.
For their part, the Vermont State Police indicated that any further charges pertaining to protests at the facility will continue to
circumvent local authorities
and instead be referred directly to the attorney general for prosecution. (The State Police did not respond to a request for comment.)
“Repression breeds resistance,” said Wade, the protester. “If anything, it’s caused a lot more outrage.”
Dyson, Breville, Apple: the 59 best Amazon Prime Day deals actually worth it, obsessively vetted by us
Guardian
www.theguardian.com
2026-10-07 13:33:18
From Ugg and Ninja to KitchenAid and Shark – we researched the best sales to get your hands on for Prime DayThe best deals from Amazon’s rivals and sales on home essentialsSign up for the Filter US newsletter, your weekly guide to buying fewer, better thingsAmazon Prime Big Deal Days, live on 6 and ...
Amazon Prime
Big Deal Days, live on 6 and 7 October, is the Super Bowl of online shopping. While the e-commerce behemoth just hosted another Prime Day in June, consumerism never sleeps and there are still genuinely good deals to be found, although some are exclusive to Amazon Prime members. But if parsing through thousands of deals isn’t your jam, fear not, because it’s ours.
As veteran consumer journalists, we’ve taken the time to find out what’s really worth the spend, from expert-vetted tech gadgets to make every day a little easier to rigorously tested kitchen scores that help put dinner on the table quickly. Here are the top Prime Deals to shop from brands you know and love such as KitchenAid, Ninja, Bissell, Ugg and more.
When my chef husband is impressed with my cooking, the
Titanium Always Pan Pro
is often behind it. This nontoxic, nonstick pan earned top marks in testing at Drexel Food Lab for its ability to withstand unprecedented levels of high heat.
Our Place
Titanium Always Pan Pro
$108
Field Company No 5 Chef Skillet
Field Company
No 5 Chef Skillet
Fried eggs did not stick on this pan.
Photograph: Karen Yuan/The Guardian
$100 (was $125)
The lightweight pan that Karen Yuan, the Filter’s commissioning editor,
relies on
for perfectly cooked eggs that glide “smoothly around the pan like a wet bar of soap in the sink”, is now 20% off.
Kitchen expert Marian Bull found that the Vitamix 5200 made “luscious” and “vibrant” soups, smoothies and dips in her testing of the
best blenders
worth buying. And now, with the help of this 30%-off deal, you can, too.
Winter may be coming, but a
frozen beverage
is a year-round treat. This
viral slushie maker
promises a good time with its five settings, one-touch controls and mess-free spout, now its lowest tracked price at 35% off.
When food experts from
Drexel Food Lab
put the pretty (and very functional)
Wonder Oven Pro
through its paces, they were impressed by its ability to make restaurant-worthy fries, nuggets and “evenly baked cookies with perfect chew”. It is now 15% off.
If your food storage isn’t pulling double duty, it might be time for an upgrade. “Perhaps most appealing [about plastic-free, non-oxic Anyday containers] was the fact that I could store and cook in the same vessel, then put all the parts into the dishwasher when I was done,” of this discounted set, according to kitchen expert Emily Farris.
When I gifted this sous vide machine to my chef husband, I thought he was going to propose again. I don’t have a second ring to show for it, but I am spoiled with restaurant-worthy succulent proteins on the regular. It’s now 40% off, nearing its lowest price of $142 during Prime Day in June.
This nontoxic air fryer might just help curb that takeout habit, whether you’re reheating leftovers or looking for an appliance that is straightforward and effective,
according to our tests
performed in a food lab. What’s more, its main cooking vessels are completely Pfas-free and plastic-free. It’s now on offer for its lowest price.
Deal ends on 25 October.
If you often crave frozen sweet treats and wish you could make them at home, this Filter-vetted ice-cream maker is 22% off. “The Creami stands out because it makes homemade ice-cream genuinely achievable – even for the most reluctant cooks,” said Danielle Amato, a tester, in
the Filter UK’s best ice-cream makers guide
. Ninja has dropped its price to $180 a few times, but this is the lowest we’ve seen it since Prime Day in June.
If you’d rather not spend half your paycheck on great coffee, trust a coffee expert to steer you toward the
Espro P3 French press
, now at its lowest price ever at 30% off. “It’s actually very good,” Filter UK coffee enthusiast Sam Mueller says “The dual filters hold back more sludge than all the other brand’s cafetières, and that means you get a slightly cleaner cup.”
If you’ve ever wanted to own a fancy barista-worthy espresso maker with an easy touchscreen interface and customizable froth levels, now’s your chance. The
best bean-to-cup coffee machine
our Filter UK expert Sasha Muller has ever tested is at its all-time low price – $200 less than its Prime Day price in June, and a bit lower than Black Friday last year.
You don’t need to be a
professional baker
to benefit from the KitchenAid Artisan stand mixer, the
best stand mixer the Filter UK has ever tested
, now 24% off. As an amateur baking enthusiast, I love mine. It helps me serve impressive baked goods any time I have guests over. It’s often discounted, but only gets this low a few times per year.
Drowsy’s luxurious
silk eyelash-protecting mask
has the appeal of a “self-care ritual” and helped tester DeMelo unlock a whole new level of sleep quality. At 30% off, this is truly a great deal.
With the Filter-favorite
Flaus electric flosser
, now at 25% off, flossing might just become less of a chore and a natural extension of your oral hygiene routine. “I use it twice a day, every day, and sometimes come back for more after meals,” said contributor Tobey Grumet Segal.
Nod off with our
favorite weighted eye mask
, now 20% off, designed to calm you into slumber. “If the idea of a weighted sleep mask sounds relaxing to you, this one is the one to try,” said tester Juno DeMelo
Dyson’s V8 cordless vacuum is a trusty workhorse that’ll still get the job done with a respectable 40-minute run time and powerful suction capable of tackling a range of surfaces, as Gibbs
found in testing
. It’s now 31% off – a pretty solid deal.
Dyson
V8 Cordless Vacuum
$270
Bissell Little Green Portable Carpet and Upholstery Cleaner
Bissell
Little Green Portable Carpet and Upholstery Cleaner
If it weren’t for my trusty sidekick with strong
spraying, scrubbing and suction powers
, my home would look (and smell!) far more unappealing than it does with two kids and two pets. In the four years since I have bought the Bissell, which most recently sat at $99.99 in August, it’s more than paid for itself in the professional cleaning services I have avoided booking.
Bissell
Little Green Portable Carpet and Upholstery Cleaner
If you’re willing to forgo the fingerprint function in the
smart lock
contributor Jon Bintner recommends, its counterpart with the same keypad and app mechanism is now $100 off.
When my toddler panicked at the sight of her newborn brother sitting in her stroller a couple weeks ago, I knew we needed a better setup. This
beast of a stroller
lets them sit solo or side-by-side and at 20% off – an all-time low – it’s a great deal for a piece of gear that’ll last the better part of a decade.
Bugaboo
Donkey 6 Single-to-Double Convertible Stroller
$1,359
Cozy Earth Bamboo Stretch-Knit Long Sleeve Pajama Set
A dream to wear for sleep or lounging, “the brand’s signature bamboo viscose fabric provides comforting coverage without overheating the body during everything from warm climates to hormone spikes,” I wrote in my review of
Cozy Earth’s pajamas
. Get a set for 20% off for year-round comfort.
As a mom to young kids, my sleep is sacred, and of all the gizmos and gadgets out there, none help me preserve it quite like the
Hatch alarm clock
. At 20% off, this is a relatively rare occurrence, so consider it for a better night’s rest.
When Stephen Treffinger
tested
Canopy’s handheld showerhead
, its robust spray and genuine skin-improving qualities were impressive enough for it to be crowned the best overall in the category. Your shower (and you!) deserve this
upgrade
– especially at 23% off.
My dual-camera Eufy video doorbell (now at 33% off) offers a wide field of view, offering a full picture of visitors as well as visual proof of my packages
actually
arriving. Tech contributor Andy Shaw appreciates its
local storage
, so you can skip the pricey subscription format.
After a break-in last year, I am now obsessive about smart home upgrades that make me feel
safer in my own home
. Since adding the Schlage Encode to my front door, I’ve felt infinitely more at ease. Get a little peace of mind for 20% off.
My life changed when I realized I could restore my wrinkled clothes back to their smooth glory with this convenient garment steamer in just 20 minutes – no bulky iron and board required. Grab it while it’s 24% off.
If you’re looking to optimize your respiratory health, the Shark NeverChange air purifier is now much cheaper than its usual $200 sale price. It has easy-to-clean filters, an intuitive air quality display and a “Febreze-like nebulizer that can mask odors and fill the room with a fresh scent”, according to Jon Chan, a tester who named it
one of the three best air purifiers
.
Step up your home theater setup with a
Sonos subwoofer
that’s petite and on sale for a very good price. Take it from Gibbs, who notes its ability to provide “more than enough boom for all but the largest of rooms and not only adds bass but improves the range and dynamism of speakers it is paired with”.
Your ears aren’t failing you – your TV might need a
soundbar
. This tech editor-vetted option does much more than that, “acting as one of the
best multi-room music speakers
money can buy”, said the Guardian’s consumer tech editor, Samuel Gibbs.
If you’re not looking to spend a ton but want your podcasts and playlists sounding crystal-clear, these
budget-friendly earbuds
are a surprisingly great value. Enjoy the attendant bells and whistles including active noise cancellation, wear sensors, wireless charging, water and dust protection and great battery life – all for 29% off.
The JBL Go is our favorite
travel-worthy speaker
for a reason: it’s “only a tad thicker than two decks of playing cards, yet it manages to deliver genuinely enjoyable sound, in a nearly bulletproof package,” said Cohen, our reviewer. It’s now 34% off.
This bite-sized
Bluetooth speaker
delivers “full-throated sound” and a surprising 24 hours of playtime, according to Filter tester Simon Cohen. The 2025 model, which we tested and truly enjoyed, is now 33% off and its lowest price.
The
most affordable Kindle
we have ever tested (and loved!) now sits at nearly a third of its sticker price. Keep one on your nightstand for access to Amazon’s massive array of titles in text that is “crisp and easy to read”, according to tester Jenny McGrath.
Don’t let pesky wifi lag keep you from excelling at your job or binge-watching content to your heart’s content. Instead, invest in a discounted version of the mesh router Nick Mokey, a Filter editor,
relies on
for “fast, reliable wifi, without dead zones”.
Turn your commute into an immersive listening experience with these
over-ear headphones
that “produce a clean and clear sound with good separation of tones that is easy to listen to across genres,” according to Gibbs, the consumer tech editor – a great get at 50% off.
We’re not saying paper is obsolete, but the things Kindle Scribe can do better than the analog stuff make it a pretty compelling buy,
according to
Gibbs’ tests. Write, read and browse (with limits) from a tablet with a screen size that’s big enough to be useful without feeling cumbersome – all at its lowest price ever.
At 28% off, now is your chance to grab a pair of
Apple AirPods 3
, one of Gibbs’ go-to devices for listening to music and tuning out the world through superior noise-cancelling properties.
There are few things more frustrating than running out of juice as you’re tracking stats on your run – and this solar-powered
Garmin watch
for 23% off fixes that. It’s the “closest model so far to what is the holy grail of smartwatch makers – a watch that you never have to plug into the charger”, said Gibbs, the Guardian’s consumer tech editor.
Between its integrated voice controls, easy setup and ability to display “absolutely stunning” high definition content, it’s no wonder this value-packed streaming stick gets
rave reviews
from Gibbs. It’s only hit this all-time low price of 60% off once before.
Whether you’ve got a bucket list trip on the horizon or work a desk job, these
Filter-vetted compression socks
promise “to help with joint and muscle stiffness and avoid any blood pooling in your feet – all of which can happen if you sit still for too long,” according to travel expert Lydia Mansel.
Conquer slick city sidewalks or the gnarliest switchbacks of your choosing in this pair of
gift-worthy hiking boots,
now nearly half off and their lowest price. “They are well crafted, neatly designed, waterproof and capable of taking on any conditions,” said Guardian sports writer Giles Richards.
Keen
Men’s Targhee 4 Waterproof Hiking Boots
$130
Philips Norelco OneBlade Pro 360 Face + Body Trimmer + Shaver
Philips
Norelco OneBlade Pro 360 Face + Body Trimmer + Shaver
This is a rare discount. Looking to stay smooth on the go? Our colleagues at the Filter UK named this the
best beard trimmer for traveling
and it’s 28% off, the lowest we have seen it for a year. Edward Munn, a tester, called it “lightweight and inconspicuous (meaning it’s perfect for chucking in your hand luggage)”.
Philips
Norelco OneBlade Pro 360 Face + Body Trimmer + Shaver
You can’t put a price on a great pair of jeans you’ll wear literally every day – but 51% off is nothing to sneeze at. And, in an era when “
jeans
have never been more confusing,” as Guardian UK deputy
fashion and
lifestyle
editor
Chloe Mac Donnell
said
, it’s nice to have a simple and easy to wear option.
Sunglasses
protect your peepers while upgrading your entire look. And for something a bit more unexpected, consider these vouched-for Ray-Bans, now just shy of their lowest-ever price. “If you tend to go for classic tortoiseshell in a standard shape but want something a little different, [these frames] are in no way outlandish but are doing something a little more out of the ordinary,” according to contributor
Ellie Violet Bramley
.
This fashion expert-approved
shoulder bag
is at its lowest price and if you’re in the market, this is a great bet. Associate fashion editor Jess Cartner-Morley calls it “A classy little number at an excellent price point, with adjustable straps and an inside pocket.”
These
breezy pants were recommended by one stylist in our spring style refresh guide as a way to brighten up your wardrobe. If you’re not quite ready to say goodbye to summer, you can grab them for 20% off.
Fun (in the sun) fact: the best time to buy sandals is when the leaves are changing colors. When I sought out to buy a pair as comfortable as sneakers, I landed on these, and my feet have never been happier thanks to their plush footbeds and adjustable straps. Grab them while they’re two-thirds of the price, the lowest they have ever been.
“I’ve gone through two or three tubes of the Inkey List’s effective and affordable Caffeine eye cream. My puffy eyes struggle to wake up without its star ingredient,”
writes beauty editor Sabine Wiesel
.
After putting this
viral red light therapy mask
through its paces 100 times, tester Sarah Matthews says it’s worth the buy “if you want an effective LED mask that can treat multiple skin concerns at once, without adding extra light settings that aren’t as well-researched”. See for yourself at more than $50 off.
This hydrating and affordable
snail mucin serum
from K-beauty juggernaut CosRx is extremely effective, even if it is, in theory, a little gross – “I know, I know, disgusting idea, but I swear it works,” according to the Guardian’s associate fashion editor Jess Cartner-Morley. “The brand says it’s ethically harvested.” Test your tolerance for slime at a very good price, as this product is now 41% off.
Zap your face (safely) at home with the
facial microcurrent device
contributor Tobey Grumet Segal
calls
“quick enough that it’s easy to use it consistently”. It’ll tighten and contour your delicate neck skin, and at 45% off, it’s worth the buy.
As a fragrance snob who prefers to leave a subtle scent in my wake, I’m never fully dressed without a spritz of these intoxicating hair and body mists, now an all-time low of 20% off. Get one for yourself and for the young person in your life – they’re gen Z-approved according to a
Filter gift guide
, should you need a little
inspo
this season.
“Every time I use one, my skin feels thoroughly soothed,” says Yuan, the Filter’s commissioning editor, of her favorite pH-balancing K-beauty face mask that you can now purchase on US soil. A pack of 10 is now 32% off for all your hydrating, calming and anti-inflammatory skin care needs.
If you’re a fiend for a good blowout, then the Filter-vetted
Dyson Airwrap i.d.
is certainly on your radar, and yes, we know, it’s expensive. A 30% discount might make this a little more enticing – and as contributor DeMelo noted in her testing, it’s a tool that’s more convenient and better for curling.
This Filter UK-approved (and, if you care,
Kardashian
-endorsed)
Ulike hair removal device
with a built-in cooling system is rarely discounted this steeply at 26% off. The only downside noted in its review was that it’s corded. Otherwise, it should help you feel smooth as a baby’s bottom ahead of your winter beach vacation.
Salon-worthy hairstyles don’t need to wait for special occasions. Our
Filter US-favorite hair styler
regularly bobs between $250 and $370, but it’s rare that it drops this low.
Shark
FlexStyle Multi-Styler & Drying System
$230
Why you should trust our recommendations
We work for our readers, rather than the brands we write about. Every item we recommend is something a Filter editor or contributor has personally tested and can vouch for. On occasion, we may present a close substitute from the same brand and explain any differences from the original.
We scour competing retailers for the lowest prices on any item, and we use price-tracking tools such as Camelcamelcamel to check price history and ensure the discounts are truly good value, rather than a faux markdown. (So, while we love sharing good deals, we also aren’t afraid to call out the unimpressive ones.)
Other pieces you might enjoy from
the Filter
, the Guardian’s guide to buying fewer, better things:
Pinrail is a desktop app where your coding agents ask you before they act.
You review what an agent proposes in a view made for it, and the agent
carries on with your decision.
Pinrail works with any agent that can run a command, such as Claude Code,
Codex, Cursor or OpenCode. It runs on macOS and Linux, and Windows support
is coming. Pinrail is in early development, so expect rough edges and
changes between releases.
Why
Agents are good at doing the work, but some steps need a person: a comment
posted under your name, an email to a customer, a change to production.
Approving those steps in a chat means reading a wall of text and typing your
answer back. Pinrail gives each of these moments a proper review:
The agent decides when to ask.
Its instructions name the steps that
need you, so it asks at those steps and nowhere else.
Each review has a view made for its content.
A code review shows the
diff with the agent's findings on it. A set of generated images shows the
images, and you draw a box on the part that should change.
Your decision is structured.
The agent gets back what you accepted,
what you rejected and your notes, as Markdown it can act on or JSON for a
script.
Everything stays on your machine.
The app serves its API on loopback
only, and keeps your reviews and decisions locally.
How it works
sequenceDiagram
participant A as Your agent
participant P as Pinrail
A->>P: pinrail submit code-review --data review.json --wait
Note over A: waits
Note over P: you review and decide
P->>A: your decision
Note over A: carries on with it
Loading
When the agent reaches a step that needs you, it runs the
pinrail
command
with a review: the plugin to show it with, a title and a JSON payload.
Pinrail notifies you, and the review waits in your inbox, grouped by
project.
You open it, and decide in the plugin's view. In the code review at the top
of this page, you accept or reject each of the agent's findings, and write a
note to the agent.
When you hand the review over, the command prints your decision and exits,
and the agent carries on:
r_01K5R2 · decided · Retry failed webhook deliveries
code-review · decided by maya at 2026-09-23 10:14
- **#1 accepted** `src/deliver.ts:42` — The worker sleeps for up to 31 seconds per delivery (major)
> Agreed. Re-enqueue with runAt = now + backoff(attempt)
- **#4 rejected** `src/log.ts:18` — The give-up log should say why (nit)
> Fine as it is; the log already has the delivery id.
Undecided: #2, #3, #5
The command's exit code says how the review ended: decided, discarded with
an instruction to stop, withdrawn or expired. When you ask for changes, the
agent submits a new round, and the app shows your earlier verdicts beside
it.
Follow the setup
that opens the first time. It installs the
pinrail
command, adds Pinrail's skill to the agents it finds on your computer,
and installs the recommended plugins.
Ask your agent
for something, in your own words:
Ask me through Pinrail which TODOs in this repository to tackle first.
Or ask it where Pinrail would help in your project:
Look at this project and suggest where you should ask me through Pinrail before you act.
Each kind of review is a plugin: the payload an agent sends, the decision
you give back, and the view you decide in. Five core plugins come with the
app:
Generated images to choose between, with boxes and pins on what to change.
Sample plugins in
forgeplane/pinrail-plugins
show what else a plugin can do. They cover HTML pages, emails, calendars,
logos, colour palettes, 3D models, animations, audio, video, design
canvases, before-and-after comparisons and trades. Each one is released as
a zip, which you install in the app's
Settings › Plugins
.
Writing a plugin
A plugin is a manifest, two JSON schemas and an HTML view, with no build step
needed.
pinrail plugins new <name>
creates one, and the plugin SDK,
pinrail-sdk
, runs its view in a
browser without the app and tests it with Playwright. Views can also be
built with React, Vue, Svelte or any other framework. See
Writing a plugin
.
Node is pinned in
mise.toml
, which
mise install
sets up, and Rust in
rust-toolchain.toml
, which rustup reads on its own.
mise run dev:desktop # the app with live reload
mise run test# every test suite
mise run lint # rustfmt, clippy, the type check and the licences, as CI runs them
CONTRIBUTING.md
describes each test suite and how to run
it. Read it before you open a pull request. To report a vulnerability,
follow
SECURITY.md
.
Every plugin listing in the official
EmDash plugin registry
is now screened by
Cloudflare's Clef decision model
before it can appear in the catalog. Clef reviews the listing metadata, outbound links, icon, and screenshots for phishing, impersonation, offensive content, scams, and attempts to manipulate the moderation system.
The registry is built on AT Protocol, so plugin authors publish signed releases from their own accounts. EmDash doesn't control the underlying record, but it does decide what appears on
plugins.emdashcms.com
. Automated moderation lets authors keep that open publishing model without making the catalog an open door for abuse.
One model for text and images
Clef
is a decision model: instead of generating prose, it answers bounded questions with typed probabilities that application code can act on. It also accepts images. Those two properties fit the registry well because a plugin listing mixes structured metadata and links with visual assets.
When a plugin is published, the EmDash labeler asks Clef nine questions about its text and links, covering:
explicit sexual content, hateful or dehumanizing content, and graphic violence;
phishing or credential requests, impersonation, scams, and spam;
deceptive links, misleading media or claims, and attempts to manipulate moderation.
The labeler asks eight corresponding questions about each icon and screenshot. Links are evaluated from their text and destination rather than from screenshots, which avoids flagging ordinary screenshots of spam filters and link-checking tools.
Clef returns a probability for each category. A result of 0.45 or higher sends the listing to an operator for review. The model never has the final say to block a plugin.
Our choice of models used by our automation tools are always driven by evals, and we regularly evaluate new models. The previous moderation pipeline needed two general-purpose models to agree on text, plus another pass for images. That was slower and gave the system more ways to time out or return an incomplete result. We were excited to evaluate Jev when it was released, but unfortunately it underperformed our baseline. However, our results for Clef were much better. We were particularly pleased that it is a multimodal model, so we can use the same model for text and images.
“Previously we had to use a mix of models to get reliable results, but Clef beat them all, while being much faster.”
Matt Kane, EmDash lead maintainer
We tested Clef against the public regression cases built for the registry and a separate protected corpus. Across three repeats of 21 public text fixtures, Clef made the expected pass-or-review decision in all 63 runs, with no invalid output or disagreement between repeats. Across those 63 fixture runs, end-to-end text moderation latency was 1.64 seconds at p95. Clef also passed all 37 plugin profiles and all 43 listing images that were live in the registry when we ran the evaluation.
Replacing the earlier bundle with Clef gives the registry one multimodal model and one probability-based interface for the whole listing. That makes moderation faster and the implementation easier to inspect, while preserving a fail-closed path to human review.
Publish a plugin
The moderation change doesn't add a new step for plugin authors. Scaffold and publish with the
EmDash plugin CLI
, and the registry runs the checks before adding the release to discovery. If a listing needs review, the release remains signed and available from your account while the EmDash team decides whether to show it.
A lot of people have been commenting on this over time so I figured I would
put together a little post summarizing my thoughts on the matter and everything
else.
The post is going to be quite large. Every section can be read more or less
separately, but they are all connected to each other.
Chimera is a small project. It exists primarily to serve its community and not
any external entity.
Making sure the maintenance cost is as low as possible is crucial. Chimera
and its tooling are designed to automate away most of the boring stuff and
make sure the packager’s job is pleasant and not bothersome. In a way, we
aim to empower every user to be a packager. The barrier of entry is meant
to be very low.
Some things that implies (I will try really hard to make this reasonable to
follow, unfortunately the years have left their mark and I might be taking
some assumptions for granted, so please bear with me; this also applies in
the later sections):
Everyone uses the same tools. Inefficient UX patterns are identified,
and either fixed, or
cbuild
is extended to mitigate them. Feedback is
taken into account.
Everything is
cbuild
. One tool does everything, in a streamlined way.
Managing the repository, the build environment, the repo generation
and signing, even parts of the VCS handling and common maintenance tasks.
No external helper stuff.
The remote build infrastructure just runs
cbuild
and not much else.
The tooling does much of the bulk of making sure your packaging is
correct and clean, most issues are hard errors. Heavy sandboxing,
build environment consistency, unit tests by default, etc.
You can run it on any Linux. If you run it on Chimera, you can immediately
test your work. You can take the repo and bring it somewhere else. That
also means the builder machines in the remote infra can run anything.
Chimera is a collective effort. You can expected to share the stuff you
make, and get it upstreamed to us. The tooling is not intended for local
things that won’t get shared, and it provides no guarantees or obligations
for such usage;
cbuild
is a developer tool, not a user tool. But every
user can be a developer.
The tooling should easy and straightforward to use, and fun. It should not
make things hard for you. Every user can be a maintainer. It shouldn’t
be unnceessarily intimidating. We don’t gatekeep here when possible.
You should also be having fun using the system and being here.
You can replicate the entire remote infrastructure of the project on your
machine in an hour or something. You can replicate the heavy bulk of things
in like 5 minutes.
The way things work is supposed to steer you towards implicitly doing the
right thing, and punish incorrect patterns by making them harder than the
correct ones. E.g. it’s really difficult or impossible to manually patch
files with regex, or to do internet-reaching stuff during the build, or to
touch the build environment filesystem. It’s really easy to manage patches
though, there are build styles for common build systems, there are fine
grained utility modules for doing all sorts of common annoying things,
etc.; the tooling also automatically checks your stuff for correct
formatting and other lint issues, validates your dependencies, validates
your metadata including minor things like whether your build dependencies
are sorted correctly and whether the SPDX license expression is correct
and lost of other nits, and so on.
Extensive documentation for the build system and packaging.
A common workflow setting up everything from scratch would look like so:
$ # set up environment with your cports fork
$ git clone https://github.com/my-user/cports
$ cd cports
$ # prepare a signing key
$ ./cbuild keygen
$ # prepare a build environment
$ ./cbuild bootstrap
$ # write your template here; then build it
$ ./cbuild pkg user/my-cool-program
$ # make a branch for submission
$ git checkout -b my-cool-program
$ # automatically makes a commit with the correct message and all files included
$ ./cbuild commit user/my-cool-program
$ # push and publish PR with the link
$ git push origin my-cool-program
The tooling provides common utilities for keeping your local repository clean,
e.g. pruning packages and so on; it provides maintenance tools for bumping
versions and revisions, checking dependency graphs, checking for new versions
via update-check (which for most things distributed e.g. via common git forges
requires no extra effort or specialized code), checking if your local repository
can be unstaged against the remote one (verifying if you really rebuilt
everything that needs it), updating sha256 checksums in templates. It even
supports custom template-specific actions, such as building bootstrap tarballs
for compilers. It supports aliases for less typing.
It supports fun things when it comes to packaging bulk batches. For instance,
it integrates git:
$ # build everything changed in your local commits
$ ./cbuild pkg git:origin/master..HEAD
$ # build a specific commit
$ ./cbuild pkg git:commithash
$ # build a range of commits but skipping stuff that has "test" in commit message
$ ./cbuild pkg git:from..to+!test
It supports common stuff, like
$ # build everything you have in your local repo that has a newer cports version
$ ./cbuild pkg status:outdated
There is a lot more that I could mention. I also wanted to talk about how
we got here, how we started, and how our infra works.
Personal history
I started using Linux in 2006/7 and soon ended up on Debian as my long-term
operating system. Around the same time I started to seriously get into
programming and the intersection of that ended up being packaging for Debian.
Debian has a large bunch of tooling to deal with packaging and particularly the
community-driven approach really appealed to me so I maintained a bunch of
packages for a while. That lasted for some time until I drifted away towards
other things and eventually settled on FreeBSD.
During that time I didn’t do any work for the FreeBSD project itself but
I did experiment with various things that were adjacent, none of them going
anywhere in the end, but I did gain a lot of helpful experience on the way.
In 2016 I started messing around with Void Linux (particularly as I had to use
Linux on work computers) and around 2018 I ended up using the ppc64le (POWER)
platform for my workstation, which FreeBSD did not support at all at the time.
I was already familiar with Void and decided to start a ppc64le port, which
I maintained downstream until 2023 and gradually it expanded to support the
big endian ppc64 as well as the classic PowerPC. Doing the downstream work
led to me becoming an upstream Void maintainer as well and I took care of
the compiler toolchains among other things for several years and became
interested in improving the build tooling.
In 2021 I started the Chimera Linux project, which gradually got better and
more usable, and that led to me shifting entirely to that, deprecating the
POWER port of Void and resigning from the Void team in 2023.
Void Linux
When I started using Void in 2016, it mainly caught my attention as a system
with a fairly low barrier of entry that did not get in my way, while still
being “normal” enough to act as a regular Linux system. Initially I found the
way e.g. runit worked in there interesting and in many ways fresh over classic
distros particularly in the pre-systemd era, while I found the post-systemd
era rather unfun and daunting in general.
This was also why I ended up picking Void for the ppc64le port, as it felt small
scope and approachable enough to give me a chance to reasonably maintain it by
myself, while also having a compact and easily reachable community that wasn’t
aligned with any particular commercial entity.
When I started the ppc64le port, I was already marginally familiar with the
build tooling of Void but only really properly picked it up for the port.
Getting started with xbps-src
For someone coming from most other Unix-like systems, xbps-src is really nice.
In Void, all the packaging of the distro, along with the build tools, exists
in a singular Git repository (
void-packages
). In a way, this mirrors the
ports systems as they exist on the BSDs. However, the ports systems are based
around Makefiles and you interact with ports individually using
make
(with
various external ports management tools existing to simplify that, working
on top of that).
In Void, software is packaged using “templates”. Every piece of software
consists of a template containing metadata fields (you know, package name,
version, dependencies, source URLs, etc.) plus (optionally) functions that
define the logic of how a template is built, along with extra files that are
necessary for the build (not always) and patches (not always). The tooling
for building templates (
xbps-src
itself) also exists in this tree.
Unlike most other systems, when you build a package, the build process does
not run in the system you are running the build on. Instead,
xbps-src
constructs a small container (using Linux namespaces) that represents a
minimized Void system, puts stuff in it according to the template metadata
(build dependencies and whatnot) and then runs the build. At the end, you
get a local repository of
xbps
packages that you can install from.
The general workflow looks like this:
$ git clone https://github.com/void-linux/void-packages
$ cd void-packages
$ # create the build container
$ ./xbps-src binary-bootstrap
...
$ # build the thing you want
$ ./xbps-src pkg some-program
...
The templates are similar to e.g.
PKGBUILD
format of Arch Linux or
APKBUILD
format of Alpine, but those do not use the containerized
approach (at least not always).
The
xbps-src
system is a collection of Bash scripts, and technically
the templates are also just Bash scripts, but with a set of expectations
so they in general do not utilize the full extent of the syntax and features.
When working as a packager, you basically just create a new directory in the
srcpkgs
, put your template in it along with other stuff, build it, and
then you have a repo you can install it from. If you wish to submit it
upstream, you create a branch, commit it, and create a pull request.
Someone will take a look at it, and if it gets merged, the central build
infrastructure picks it up and shortly it becomes available in the central
repository.
The system neatly decouples build dependencies and runtime dependencies, with
runtime dependencies typically scanned automatically from ELF files and their
metadata as well as other file types, so the template specifies whatever it
needs in the build environment but the runtime dependency list is largely
automatic.
The system supports “build styles” so you can e.g. declare that a template
uses the GNU Autotools, or Meson, or Cargo, or whatever, and it takes care
of most of the groundwork, allowing for templates that are often declarative
without any build logic, lowering that barrier of entry.
It also supports cross-compiling really well, letting you cross-build most
of the repository to any other architecture. That’s neat, but often sloppy,
and cross-built packages end up with subtle brokenness that native packages
would not have, due to build systems subtly not passing certain checks and
so on. You also can’t run unit tests for projects when cross compiling them,
in most cases.
The system also has a really handy
update-check
mechanism which will scrape
upstream URLs and find version infos in them, then present you with any updates
that may have happened since the last template version update. The project has
a nightly job which generates a summary once per day, letting packagers stay
on top of things.
Upstream build infrastructure of Void
The build infrastructure of Void consists of several computers (each handling
a different architecture port) that are plugged as workers into the central
orchestrator, using the
Buildbot
software. It will
receive updates from the Git repository (
void-packages
), collect a batch
from the changes, figure out the correct order, and use
xbps-src
to build
the packages.
It will also sign the packages with the right keys. This is a separate step
that
xbps-src
does not handle (the output is unsigned). Eventually, things
make it in the upstream repository and get mirrored.
The way the Buildbot is managed I can’t tell you much because I never saw
much into it. The admins have an infrastructure that is based on Ansible,
Terraform, and a bunch of other pieces that are put together in ways unfamiliar
to me. Additionally,
xbps-src
does not do the batch sorting for you, so
other pieces of the process have to do it. Things outside the
void-packages
repository itself are relatively opaque.
When I maintained the POWER port, I had my own set of scripts to keep things
going and did not reuse any of the Void infrastructure at all. This was likewise
opaque and purpose-built for my port.
Starting Chimera
In early summer of 2021, I started Chimera. The initial push for me was that
I wanted to experiment with a different userland setup, as well as have a more
personal project to work on without having to deal with others’ efforts and
pre-existing work, but also I wanted to try out my own take on the build system.
There were some initial points I was unhappy with in
xbps-src
while
maintaining the POWER port.
The shell-based system was way too slow. Parsing the complete collection
of templates in the repository would take potentially as much as half an
hour, and there weren’t other ways to introspect the templates. Therefore,
my tooling would call into
xbps-src
and employ various caching tricks
to make things manageable.
The shell-based system was very often fairly sloppy, letting various wrong
behaviors through. For instance, network access is permitted through the
whole build, there is nothing ensuring consistency of the build container
after the build is done (the entire thing is read-write and the template
can do whatever it wants), the correctness lints for the resulting packages
are fairly slim, the ELF scan step and other things would take an eternity
due to slow shell code, and limited opportunities for doing more due to
shell being excessively slow.
A lot of useful tooling is separate from the main system and maintained
in the
xtools
repository, including extra lints and so on. This is not
mandatory or anyhow verified however.
Lots of sloppy templates resulting from the prior points. E.g. the templates
are often littered from manual in-place
sed
calls and similar rather than
using proper patches (which leaves in calls that eventually no longer do
anything due to upstreams changing), patches are applied very fuzzily which
occasionally results in subtly mispatched things, and so on.
The Void project does not run per-template unit tests (typically the test
suite of the software being packaged) on builders, which would catch a lot
of errors, particularly on
musl
targets and so on. It does run them in
pull request CI, but I do not see this as enough. Often the check runs of
templates are poorly maintained and broken.
The Void project cross-compiles all architectures other than
x86_64
and
i686
, resulting in the other-arch ports being notably lower quality.
Void has a wonky staging system. When large batch changes are being done,
the repos may end up in a strange state for a while and users are advised
to avoid upgrading until everything is finished. It has a rudimentary
staging system which prevents the repos from being changed while a big
batch is being rebuilt for changes shared library SONAMEs, this works only
for shared libraries however, which means packagers will often forcibly
stage the repo for a batch rebuild, then do stuff in several commits,
wait for everything to clear, then unstage the repos. Various parts of
the related infrastructure are also lacking, such as clearing obsolete
and removed packages from the repos, which is/was done manually.
There are others but these were my main gripes.
Chimera started with its build system. Many experiments were done. Initially,
cbuild
was a ground-up rewrite of
xbps-src
, using Python. In
cbuild
,
the templates are also Python scripts, but through low-level bits that Python
allows, don’t always look as such. It came with a very minimal set of initial
packages that was basically a basic Void-like
chroot
, using
xbps
as its
package manager. Thus
cports
was born.
There were two initial goals:
Make it fast. I should be able to parse several thousands of build templates
and dump all their metadata in under a second, rather than many minutes.
Chimera currently achieves this.
Make it strict. No network starting with configure step. Read-write access
only in the build directory (and destination directory for install step)
and full consistency of the build container guaranteed at all times.
Heavy linting and straight up denying various misbehaviors at all times.
Sandboxing done with namespaces only, using Bubblewrap (
bwrap
). The
cbuild
system runs outside the sandbox (in your host environment) while
any calls to the build system of the project being built are sandboxed.
Make it highly portable, with Python (no external modules), Bubblewrap,
and Git, along with the package manager binary, being the only dependencies.
You can use
cports
on any Linux system with the basic dependencies.
Make it able to do everything
xbps-src
can do, but better.
Chimera started with the
ppc64le
target only, being developed on a POWER
workstation. Over time, as things cleared up more, more targets were added,
along with additional goals and other changes.
Over the next months, the core packaging started to become more defined,
settling on the userland tooling and other things, as well as transitioning
from
xbps
to
apk
. The original idea was to use
pkg
from FreeBSD, but
this proved to not be ready for our style of use at the time, and upstream
suggested that we’re best off using something else.
Fundamentally, using
cbuild
in the basic sense feels similar to
xbps-src
.
The big difference is that
cbuild
signs everything, not needing a separate
step, and no unsigned packages are allowed. Therefore, initial setup looks
more like
$ # generate your signing key
$ ./cbuild keygen
$ # bring up the container
$ ./cbuild bootstrap
$ # package whatever you want
$ ./cbuild pkg main/firefox
Settling on apk
We didn’t stay on
xbps
for very long. Soon, the switch happened to
apk
.
Using
apk
brought over various benefits.
Unlike
xbps
which is driven by ad-hoc logic,
apk
has a real dependency
solver. This means way fewer surprising behaviors and way fewer workarounds
and less effort needed from packagers to make sure that users’ systems do
not break in surprising ways.
Handling of stuff like shared libraries is way nicer in
apk
. In
xbps
,
dependencies are driven purely by name, and shared library providers and
requires have separate metadata fields. These will get checked and
xbps
will not permit things to proceed if unmatched, but dependencies are done
by name. When a package is being built, the system will consult this massive
central file (
common/shlibs
) matching SONAMEs to package names to generate
dependencies on top of the shlib metadata. This results in poor UX (you
cannot match the shlib providers etc. in the same way as names when e.g.
installing) and is inflexible as it’s limited to shared libraries. In
apk
,
virtual packages are used instead, so you e.g. provide
so:libfoo.so.1=1
as a virtual package name. You can then search by this name, install by
this name, etc. and it can be used for other types as well, e.g.
cmd:foo
(
apk add cmd:i-know-command-name-but-not-package-name
) and others.
That also means more things can participate in the staging system on the
build system side. And no central mappings, as the build system can easily
query things from the repository.
Unlike
xbps
,
apk
has real support for triggers. Triggers in
xbps-src
are emulated with package scripts. As an example, consider e.g. updating
fonts cache. The
fontconfig
package has a trigger (which is a script) and
metadata to run the cache update when
/usr/share/fonts
is modified in any
way. Then when something modifies the directory, the trigger will run at the
end of the transaction, after all filesystem changes are done. In
xbps
,
you can only have scripts that run before or after a package changes and
are owned by the package;
xbps-src
emulates triggers by checking if the
package contains a directory that would drive the trigger and if it does,
insert a call in its package script; that means if several packages change
the fonts, each of them contains a cache update script, which will run
several times (and not at the end). This means things become quite slow
for large triggers and the transaction breaks atomicity (because the shell
script run in the middle of the transaction, rather than at the end, needs
all prior things committed to the final filesystem location).
There were others, but these are notable.
Building an infrastructure and a new core tenet
Building the infrastructure was interesting. We ended up using Buildbot again,
after some deliberation and original intent to build a new orchestrator from
scratch. There is an important distinction in our buildbot, however.
I wasn’t happy with the opaqueness of Void’s infrastructure, and with the way
it does so much work beyond
xbps-src
. I declared an entire new rule that goes:
What the packager runs exactly matches what the remote build machine runs.
The build steps of Chimera’s infrastructure go approximately like this:
Central orchestrator is poked by a webhook that a change in
cports
has
happened. It does not matter what change.
Central orchestrator tells worker machines in the fleet (one per arch)
that
cports
has changed. It does not tell them what.
Worker machines pull their copy of
cports
. It does not matter what
has changed.
Worker machines update the build container (
./cbuild bootstrap-update
).
This updates the core packages of the container to match the repository.
Worker machines run
./cbuild bulk-print-ver status:unbuilt
. This generates
a list of all packages along with their versions that are present in
cports
,
can be built, and are not already built in the repository. This step is
separate only for the purpose of presenting it to packagers in the Buildbot
UI, so that you can have a separate view for each package being built and
so on. This list is correctly sorted. Each entry is saved in a Buildbot
variable, again for the purpose of presentation mainly.
Each entry in the list from the prior step is built. The result is already
signed and ready to be used, but for now remains in stage area.
An unstage step is attempted. If any provider is removed in the new
packages (e.g. something rebuilds and stage provides
so:libfoo.so.2
while the original provider was
so:libfoo.so.1
) and there is still any
package depending on the old name, the unstage will fail. That ensures that
all reverse dependencies are always rebuilt as necessary before publishing.
On successful unstage, the staging area is merged into the repository.
Repository is pruned for outdated stuff.
Repository syncs to the final primary mirror. First, changed packages get
uploaded in one step, then indexes are replaced, and only then old packages
are deleted. This is to ensure that users don’t end up with any index that
contains packages not present yet.
Steps 4, 5, 6 can be condensed together into a single
cbuild
command:
$ ./cbuild bulk-pkg status:unbuilt
The main tricky part with batch builds like that is correct sorting. Imagine
having packages A and B in the bulk that you want to build. There is a package
C, which A depends on, and which depends on B. Since that makes B in the
dependency tree of A (through A->C->B), B needs to be built first. You do not
know this by knowing only A and B, since C is not in the batch.
Since
cbuild
is very fast and can parse the entire collection very quickly
without caching, it can consider all the intermediates and do a proper sorted
graph trivially. It can also simply parse every template to consider whether
it’s buildable, and match it against the repo state. Therefore, on every build,
it can always do that from scratch and the orchestrator has to do nothing,
with each step being largely independent of the other.
The build fleet doesn’t do anything else. There is other auxiliary infra, such
as our IRC commit bot (which announces changes driven by a webhook) and our
nightly update-check equivalent to Void, letting packagers stay on top of
things.
In general,
everything
that any piece of our infrastructure does can be
easily done locally.
Our fleet is self-hosted. The orchestrator and most builders run on Chimera
host systems, but they don’t have to. Currently these are:
The
x86_64
machine, which is a 16-core EPYC root server at Netcup,
running Chimera.
The
aarch64
machine, which is an 80-core Ampere Altra, running Chimera,
owned by me.
The
ppc64le
machine, which is an 18-core/72-thread POWER9, running Chimera,
owned by me.
The
loongarch64
machine, which is a Loongson 3A6000 4-core/8-thread,
likewise running Chimera, owned by me.
The
riscv64
machine, which is a Milk-V Pioneer 64-core machine kindly
provided to us by Zach van Rijn of Adélie Linux, running Fedora.
The builders do not publicly face the Internet, the orchestrator does, on
its own machine. The primary repository is also on its own separate machine.
Each architecture has its own signing key.
Final words
All of this may possibly even be a bit too long and too much to take in.
Therefore, I will end it here and hope that I have not forgotten anything.
I hope this was interesting to read and I didn’t bore you to death, and
that it wasn’t unnecessarily difficult to follow.
A good hygiene when processing secrets is to wipe them after use.
And projects receives countless PRs about adding calls to zeroization functions for anything that looks like a secret.
This is a low-hanging fruit, something LLMs love to report, and it feels theoretically useful. But unfortunately, things are a little bit more complicated. Blindly zeroing secrets can do more harm than good.
Turns out that adding a wipe for a secret can leave
more
copies of a secret than the original code. Copies that wouldn’t have existed without it, and that are still there after it returns.
Code snippets and their compiled output can be verified in this
Godbolt example
.
Let’s start with memset()
Classic starting point: the good old
memset()
call that gets optimized out:
We calculate an intermediate value, use it to produce a result, then clear it before returning.
The caller only sees the declaration, so it can’t inspect the implementation when compiling this call:
opaque_wipe(&intermediate,sizeofintermediate);
Cool, so let’s look at the caller’s assembly output now.
Intestering: there’s a new instruction before the call:
str x8, [sp, #8]
Duh? We’re storing the intermediate on the stack now? The original function didn’t do that!
Yep, the wipe needs an address. And since the caller can’t see the implementation, it has to allow for the callee reading the old contents at that address.
So it writes the value there first.
The wipe clears that new stack copy and leaves
x8
alone.
We’ve introduced an interval during which the value exists in memory, paid to clear it, and still kept the register copy.
Also, a recognized write-only operation can skip this particular store, and if LTO is enabled (quite the norm these days), our attempt to hide the implementation in another file is likely to become useless.
Now let’s compare a tag
Now let’s do something very common but a little less trivial: compute a tag (HMAC output, etc.), compare it, wipe it, return the result.
In this example, we’re going to compute the tag as
seed[i] ^ 0x5a
. This is completely dumb and insecure, but easy to follow through the assembly.
We want to securely compare it against an application-provided tag.
Wait. Why are there two stores of
q0
after the XOR?
And why is the comparison instruction,
cmeq
,
after
the call to the wipe?
To understand, let’s map the stack. We’re going to call the stack pointer after allocating the frame
S
; the frame pointer,
x29
, is
S + 64
.
Stack bytes
Contents
Wiped?
S
through
S + 15
Candidate
No
S + 16
through
S + 31
Extra computed-tag copy
No
S + 40
through
S + 55
Addressed
computed
object
Yes
Ah! The compiler loaded the comparison operands before the call, but postponed the comparison itself until afterward.
To keep the computed tag available across the call,
it made a second copy
in a spill slot.
Our wipe clears the requested 16 bytes perfectly, no problem here. And then
ldp
reloads the other copy, and
cmeq
compares it.
But nothing clears that spill slot before the function returns!
Now the fun part: what if we
remove
the wipe and compile again?
Surprise: the computed tag no longer gets stored on the stack at all. Everything stays in registers.
Adding the wipe created the copy we left behind.
Note that there’s no compiler bug here. The return value is correct, the addressed object is overwritten, and C doesn’t promise the security ordering you intended between those two events.
There, the addressed object is at
sp + 32
rather than
sp + 40
(which I got with Xcode), but the extra copy at
sp + 16
still survives the wipe.
Disabling stack protection generates different code, but this is more of an accidental unreliable side effect than a fix or a good idea.
So we’re sometimes paying for stores, and sometimes for a new call and stack copy, and the secrets are still around. Pretty frustrating for a change that was supposed to improve cleanup.
Registers shouldn’t be ignored
Registers are copied to memory after every context switch. Registers are visible in core dumps, hibernation files and VM snapshots. Registers can leak through microarchitectural bugs (
Zenbleed
).
And a single SIMD register can hold an entire 256-bit secret key, or even a 512-bit hash. They also tend to be recycled less frequently than general-purpose registers. But traditional zeroing functions are just designed to write zeros to memory, and have no clue about actual execution flows.
Keep the useful wipes
Does this mean we should stop wiping password buffers?
No: those buffers already exist in memory, and clearing them when we’re done is useful. The same goes for an allocated key schedule or a context we’re retiring.
But in the examples above, the values we cared about were also somewhere else.
So before adding a wipe to a small local, inspect whether the value already lives in memory, whether taking its address creates storage, and what stays live across the call.
Check the optimized build you ship, including LTO and hardening flags.
Or stop blindly sprinkling calls to buffer wiping functions. There are more reliable techniques to wipe secrets.
We’ll see that in part 2.
Study: Claude, ChatGPT Offer Different Shopping Prices Based on Wealth
We've detected unusual activity from your computer network
To continue, please click the box below to let us know you're not a robot.
Why did this happen?
Please make sure your browser supports JavaScript and cookies and that you are not
blocking them from loading.
For more information you can review our
Terms of Service
and
Cookie Policy
.
Need Help?
For inquiries related to this message please
contact
our support team
and provide the reference ID below.
Trigora is a durable execution platform for long-lived AI agents and programs.
Write normal application logic in
TypeScript, Python, or Rust
. Trigora can suspend execution across events, timers, external effects, and child programs, then recover from committed continuation state after failure without replaying completed execution history.
At durable boundaries, TCC commits the program position and the live durable state required to continue. Recovery resumes from that committed continuation rather than re-executing the completed prefix.
That makes Trigora a natural fit for workloads that:
TCC Engine is source-available under the Business Source License 1.1 and converts to Apache 2.0 under its license terms. The Trigora SDKs, clients, CLI, and public contracts are MIT licensed.
Microsoft Outlook to block MSIX attachments starting November
Bleeping Computer
www.bleepingcomputer.com
2026-10-07 11:44:22
Microsoft announced that it will add .msix and .msixbundle attachments to the list of blocked attachments in Outlook Web and the new Outlook Windows client starting next month. [...]...
Microsoft announced that it will add .msix and .msixbundle attachments to the list of blocked attachments in Outlook Web and the new Outlook Windows client starting next month.
.msix files are modern Windows installation packages tailored for specific computer architectures or configurations, while .msixbundle is a container that groups multiple .msix packages into a single file compatible with multiple computer architectures.
The change will begin rolling out to Exchange Online users in early November, when the new file types will be added to the BlockedFileTypes list in all OWA Mailbox policies, and is expected to reach general availability by mid-November.
After the policies are updated, .msix or .msixbundle attachments will be blocked by default, and users of Outlook on the web and new Outlook for Windows will no longer be able to send, receive, open, or download them.
"To enhance security in Outlook on the web and new Outlook for Windows, we are updating the default list of blocked file types in OwaMailboxPolicy,"
Microsoft said
in a Microsoft 365 message center update.
"As part of this update, the .msix and .msixbundle file types will be added to the BlockedFileTypes list in the default OWA Mailbox policy and any custom policies created in your tenant."
Admins don't need to take action if .msix or .msixbundle file types aren't used in their organization, but they can whitelist them by adding them to the AllowedFileTypes property of their users' OwaMailboxPolicy objects if needed.
"Most organizations are not expected to be affected by this update because these file types are infrequently used," Microsoft added. "This update is part of our ongoing efforts to strengthen security and help protect organizations from potentially unsafe file attachments."
This move is part of a broader effort to disable and remove Office and Windows features that attackers have abused in attacks targeting Microsoft customers in recent years.
More recently, in October 2025, Microsoft
also announced
that Outlook for Web and the new Outlook Windows client would no longer display risky inline SVG images that were also being used in attacks.
The complete list of attachments that can't be saved or viewed from Outlook on the web by Exchange Server and Exchange Online users is available on
Microsoft's documentation website
.
Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.
Make a full-screen sign, timer or message from a link. Hold up your phone at arrivals, put a countdown on the TV, or show the Wi-Fi password at the front desk. No account, no app, nothing stored on a server. The link is the whole display, so you can share it or bookmark it.
Make one by hand
Using the built-in editor is easiest, but generate a link however you like.
Use cases
Every card is a working link. Open it fullscreen, or load it into the editor and make it yours.
More use cases
Fewer use cases
What it does
Auto-fit
Text grows to fill any screen, from a phone to a stadium board.
Any device
Works in any browser: phones, tablets, laptops, smart TVs.
Markdown
**bold**, *italic*, # and ## headings, real line breaks.
Slides
Split with || and they rotate on a timer.
Countdowns
Put {countdown} anywhere and set &until= for a date or &timer= for a length of time.
QR codes & images
Links, Wi-Fi, calls, texts and email, generated in your browser. No third-party services.
Private by design
The message lives in the URL fragment, which browsers never send to a server.
OpenAI’s release of mathematical findings draws concerns from experts
Guardian
www.theguardian.com
2026-10-07 11:40:51
Leaders worry OpenAI is not doing due diligence to vet results and that AI models aren’t accessible to broader field of mathematicians OpenAI has astounded mathematicians after releasing hundreds of new mathematical findings on Tuesday. The company published over 370 mathematical results across a v...
OpenAI has astounded mathematicians after
releasing
hundreds of new mathematical findings on Tuesday.
The company published over 370 mathematical results across a variety of topics such as algebra, theoretical computer science and mathematical logic, showcasing what some of its most advanced artificial intelligence models are capable of.
Last month, the company solved
the Navier-Stokes equation
, one of the world’s toughest mathematical problems with a $1m reward for anyone who cracked it.
The achievement prompted concerns from leaders in the field who say that frontier labs should not be testing the most advanced mathematical problems on proprietary AI models that are not accessible to the broader field of mathematicians. The Institute for Advanced Study in Princeton, New Jersey, an independent group of mathematical experts, said that it does not endorse the practice.
“It is now the case that AI can output mathematical arguments in situations without the human who prompted it being able to understand the arguments, verify them, or take responsibility for them,” a statement from the organization read. “We believe that human understanding of mathematics remains of paramount importance. How, in this new era, can we work towards a new paradigm that includes human understanding of mathematics as part of responsible scholarly output?”
The company did not indicate, however, that it would stop testing its AI models with these advanced problems. Experts in the field say they worry
OpenAI
is not doing the due diligence required to vet these results.
In an interview with the
New York Times
, Tristan Buckmaster, a New York University mathematician who was working on the Navier-Strokes problem, said mathematicians who were prompting AI models to solve equations could be providing information that helped the model get to the result.
“There’s likely to be a bunch of results where they take someone’s work and then take it to completion,” Buckmaster said.
The advisory board has also asked that AI labs grant “equitable access” to their AI models to the global mathematics community.
“The use of proprietary internal models by AI labs to do mathematical research risks creating a two-tier system where labs outrun the rest of the field, effectively alienating the mathematical community from its own discipline,” the group wrote.
Velvet Cowboy Comes Out Guns Blazing With a $5 Food Menu
hellgate
hellgatenyc.com
2026-10-07 11:37:20
The mastermind behind Jacob's Pickles is bringing Southwestern kitsch to the Upper West Side....
It's tough to overstate the energizing effect Jacob Hadjigeorgis' first restaurant, Jacob's Pickles, had on this aggressively mid stretch of the Upper West Side when it first opened 15 years ago. This was my home turf for decades, and I can assure you, the tedium and timidity in these parts was real—and entrenched.
Pickles changed that, though, as Hadjigeorgis, who grew up in Astoria but went to school here on the UWS at Dwight, served up comically large portions of comfort food classics in a raucous environment, stroller families and packs of boozy brunch buddies equally engaged and encouraged.
Velvet Cowboy opened last week. (Scott Lynch / Hell Gate)
Abstract:
Autoformalisation is increasingly used to verify mathematical texts, including those generated by AI, as in OpenAI's announced proof of blow-up of solutions to the Navier-Stokes equations. In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified. The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully. In particular, we highlight that the problem of resolving ambiguities in mathematical NL text, which is necessary in order to provide semantically faithful translation, is arbitrarily high up in the Solvability Complexity Index (SCI) hierarchy/arithmetical hierarchy (the SCI $= \infty$). Hence, informally, providing semantically faithful AI autoformalisation is harder than any computational problem including the Halting problem (which has SCI $= 1$). To demonstrate the effect of this result we provide several examples of AI mistranslations of NL statements and proofs into Lean in practice, resulting in mismatches between NL proofs and their Lean `verifications'. These include OpenAI's announced Navier-Stokes proof. In particular, we show that the formalised Lean proof does not correspond to the NL proof of blow-up of solutions to the Navier-Stokes equations.
Submission history
From: Alexander Bastounis [
view email
]
[v1]
Tue, 6 Oct 2026 10:58:01 UTC (1,080 KB)
Across the Globe, People Increasingly Say Social Media Is Harming Democracy
A phone displays election result predictions during the second round of the 2024 French legislative elections on July 7, 2024. (Pat Batard/Hans Lucas/AFP via Getty Images)
About this research
This Pew Research Center report looks at views of how social media has impacted democracy in 37 countries, as well as social media use by platform in each place.
Why did we do this?
Pew Research Center does research to help the public, media and decision-makers understand important topics. We study how people around the world feel about their democracies and the impact of social media and other emerging technologies, such as
artificial intelligence (AI)
.
We surveyed 42,151 people across 36 countries: Argentina, Australia, Bangladesh, Brazil, Canada, Chile, Colombia, France, Germany, Ghana, Greece, Hungary, India, Indonesia, Israel, Italy, Japan, Kenya, Malaysia, Mexico, the Netherlands, Nigeria, Pakistan, Peru, the Philippines, Poland, Singapore, South Africa, South Korea, Spain, Sri Lanka, Sweden, Thailand, Turkey, the United Kingdom and the West Bank and East Jerusalem.
Interviews were conducted from Feb. 8 to May 13, 2026. We designed the surveys so we could talk about the views of the adult population in each country.
For this report, data from the United States comes from two surveys. We surveyed 3,507 U.S. adults from March 23 to 29, 2026. Everyone who took part in this survey is a member of the Center’s
American Trends Panel
. Additionally, we surveyed 5,511 U.S. adults from Jan. 30 to June 18, 2026 as part of the
National Public Opinion Reference Survey (NPORS)
. These surveys represent the views of the full U.S. adult population.
As social media platforms have become an important part of the information environment around the world, the view that social media is harming democracy has become more common.
Shares who see social media as a bad thing for democracy have grown over the last few years
% who say social media has been more of a __ for democracy in their country
* This question was first asked in Indonesia in 2023.
Note: Those who did not answer are not shown. Includes surveyed countries with the 12 largest statistically significant changes in the share who say ‘Bad thing’ since last asked. Over time, we have changed how we conduct surveys (by phone or face-to-face) in Hungary and Poland. Refer to topline for more information.
Source: Spring 2026 Global Attitudes Survey.
“Across the Globe, People Increasingly Say Social Media Is Harming Democracy”
PEW RESEARCH CENTER
Shares who see social media as a bad thing for democracy have grown over the last few years
% who say social media has been more of a __ for democracy in their country
Country
Date
Good thing
Bad thing
Germany
2026
37
61
Germany
2022
57
41
Japan
2026
48
43
Japan
2022
57
23
Hungary
2026
49
42
Hungary
2022
65
22
Poland
2026
55
35
Poland
2022
67
15
Indonesia
2026
55
43
Indonesia
2023
64
24
Netherlands
2026
27
70
Netherlands
2022
44
54
Sweden
2026
49
47
Sweden
2022
66
32
France
2026
31
65
France
2022
43
51
Spain
2026
49
47
Spain
2022
61
35
U.K.
2026
41
56
U.K.
2022
50
48
Canada
2026
40
55
Canada
2022
49
47
Greece
2026
53
43
Greece
2022
62
35
* This question was first asked in Indonesia in 2023.
Note: Those who did not answer are not shown. Includes surveyed countries with the 12 largest statistically significant changes in the share who say ‘Bad thing’ since last asked. Over time, we have changed how we conduct surveys (by phone or face-to-face) in Hungary and Poland. Refer to topline for more information.
Source: Spring 2026 Global Attitudes Survey.
“Across the Globe, People Increasingly Say Social Media Is Harming Democracy”
PEW RESEARCH CENTER
A new 37-nation Pew Research Center survey finds that in most countries where trend data is available, the share of the public who believes social media is bad for democracy has increased significantly from 2022 or 2023 to now.
Views about social media’s impact are particularly negative in the United States, where 64% of adults say it is bad for democracy. However, while assessments have become more negative in many countries, American attitudes are unchanged since 2022. In some ways, other nations are catching up with the U.S. when it comes to seeing social media’s downsides.
(For more, read “
Americans Are Negative About Social Media’s Impact on Democracy – and the Rest of the World Is Catching Up
.”)
In general, people in wealthier nations such as the U.S. are more likely to believe social media is harming democracy. The seven nations where half or more say social media has been bad for democracy are also among the wealthiest surveyed: Australia, Canada, France, Germany, the Netherlands, the United Kingdom and the U.S. However, Israel and Singapore are exceptions – relatively few in these two high-income countries believe social media is bad for democracy.
People in wealthier countries are more likely to say social media harms democracy
% who say social media has been more of a
bad thing
for democracy in their country
Note: 2025 GDP per capita data is from the World Bank (accessed Aug. 26, 2026) and expressed in current USD. The West Bank and East Jerusalem are not shown because corresponding World Bank Data includes Gaza, where we were unable to survey.
Source: Spring 2026 Global Attitudes Survey.
“Across the Globe, People Increasingly Say Social Media Is Harming Democracy”
PEW RESEARCH CENTER
People in wealthier countries are more likely to say social media harms democracy
% who say social media has been more of a
bad thing
for democracy in their country
Hidden GDP per capita
Country
More of a bad thing for democracy
%
GDP per capita
Grouping
90027
U.S.
64
64
$90,026.52
North America
55698
Canada
55
55
$55,697.66
North America
48986
France
65
65
$48,985.73
Europe
60496
Germany
61
61
$60,496.44
Europe
26948
Greece
43
43
$26,948.01
Europe
25907
Hungary
42
42
$25,907.47
Europe
43309
Italy
44
44
$43,308.64
Europe
73684
Netherlands
70
70
$73,683.92
Europe
28420
Poland
35
35
$28,419.58
Europe
38627
Spain
47
47
$38,627.25
Europe
63133
Sweden
47
47
$63,133.21
Europe
57602
U.K.
56
56
$57,601.96
Europe
65130
Australia
52
52
$65,129.72
Asia-Pacific
2597
Bangladesh
28
28
$2,597.34
Asia-Pacific
2702
India
23
23
$2,702.48
Asia-Pacific
5060
Indonesia
43
43
$5,059.63
Asia-Pacific
35951
Japan
43
43
$35,951.04
Asia-Pacific
13125
Malaysia
27
27
$13,124.56
Asia-Pacific
1596
Pakistan
33
33
$1,595.91
Asia-Pacific
4171
Philippines
23
23
$4,170.72
Asia-Pacific
98814
Singapore
22
22
$98,813.98
Asia-Pacific
36227
South Korea
38
38
$36,226.97
Asia-Pacific
5002
Sri Lanka
35
35
$5,002.08
Asia-Pacific
8057
Thailand
21
21
$8,056.56
Asia-Pacific
60337
Israel
22
22
$60,336.85
Middle East
18599
Turkey
35
35
$18,599.44
Middle East
3257
Ghana
21
21
$3,257.16
Sub-Saharan Africa
2363
Kenya
24
24
$2,362.86
Sub-Saharan Africa
1224
Nigeria
11
11
$1,224.25
Sub-Saharan Africa
6598
South Africa
28
28
$6,597.71
Sub-Saharan Africa
14898
Argentina
35
35
$14,898.09
Latin America
10713
Brazil
29
29
$10,713.29
Latin America
17995
Chile
43
43
$17,994.59
Latin America
8562
Colombia
33
33
$8,561.62
Latin America
13889
Mexico
26
26
$13,889.23
Latin America
9684
Peru
34
34
$9,684.41
Latin America
Note: 2025 GDP per capita data is from the World Bank (accessed Aug. 26, 2026) and expressed in current USD. The West Bank and East Jerusalem are not shown because corresponding World Bank Data includes Gaza, where we were unable to survey.
Source: Spring 2026 Global Attitudes Survey.
“Across the Globe, People Increasingly Say Social Media Is Harming Democracy”
PEW RESEARCH CENTER
In nations with lower gross domestic products per capita, people tend to be less negative about social media. For example, in Ghana, Kenya, Nigeria, the Philippines and Thailand, roughly three-in-four or more say it has been
good
for democracy.
In most countries polled, this positive view is predominant. Across the 37 nations surveyed, a median of 55% think social media has had a positive impact on democracy, while a median of 35% say it has had a negative one.
Young adults in most nations are especially likely to believe social media has a positive effect on democracy. One of the largest age gaps is found in the U.K., where 60% of adults ages 18 to 34 say social media’s influence on democracy has been good, compared with just 31% of those ages 50 and older. In the U.S., 44% of adults younger than 35 hold this view, compared with 28% of people 50 and older.
Many around the world say social media is good for democracy overall
Median % who say social media has been more of a __ for democracy in their country
… but majorities see positive and negative effects on society
Median % who say access to the internet and social media has made people …
Note: Percentages are medians based on 37 countries. Those who did not answer are not shown.
Source: Spring 2026 Global Attitudes Survey.
“Across the Globe, People Increasingly Say Social Media Is Harming Democracy”
PEW RESEARCH CENTER
Many around the world say social media is good for democracy overall
Median % who say social media has been more of a __ for democracy in their country
Bad thing
Good thing
35
55
Median % who say access to the internet and social media has made people …
Less
More
Not had much impact
Informed about current events
11
77
10
Able to have a meaningful voice in politics
14
60
20
Easy to manipulate with false information and rumors
8
79
8
Divided in their political opinions
10
68
16
Note: Percentages are medians based on 37 countries. Those who did not answer are not shown.
Source: Spring 2026 Global Attitudes Survey.
“Across the Globe, People Increasingly Say Social Media Is Harming Democracy”
PEW RESEARCH CENTER
Why do so many people think online platforms have been good for democracy? Part of the answer may be that these platforms give people a voice at time when
many feel voiceless
. A median of 60% in the 37 countries surveyed say access to the internet and social media has increased people’s ability to have a meaningful voice in politics.
And access to the internet and social media can inform people about current events: A median of 77% think it has led people to be more informed.
At the same time, people are certainly aware of social media’s pernicious effects. A median of 79% believe it has made it easier to manipulate people with false information and 68% say it has made people more divided in their political opinions.
Growing shares
across countries
say the internet and social media have made people more divided
% who say access to the internet and social media has made people __ in their political opinions
Note: Only countries with statistically significant changes are shown. Those who said ‘Has not had much impact either way’ or did not answer are not shown. Over time, we have changed how we conduct surveys (by phone or face-to-face) in Hungary and Poland. Refer to topline for more information.
Source: Spring 2026 Global Attitudes Survey.
“Across the Globe, People Increasingly Say Social Media Is Harming Democracy”
PEW RESEARCH CENTER
Growing shares
across countries
say the internet and social media have made people more divided
% who say access to the internet and social media has made people __ in their political opinions
Country
Year
More divided
Less divided
Poland
2026
79
3
Poland
2022
50
9
Germany
2026
82
6
Germany
2022
65
12
France
2026
69
10
France
2022
52
10
Japan
2026
64
2
Japan
2022
47
4
Spain
2026
80
3
Spain
2022
66
5
U.K.
2026
79
4
U.K.
2022
66
4
Hungary
2026
79
10
Hungary
2022
67
13
Malaysia
2026
53
25
Malaysia
2022
43
32
Singapore
2026
60
13
Singapore
2022
51
23
Australia
2026
79
5
Australia
2022
71
6
Sweden
2026
72
5
Sweden
2022
65
3
Note: Only countries with statistically significant changes are shown. Those who said ‘Has not had much impact either way’ or did not answer are not shown. Over time, we have changed how we conduct surveys (by phone or face-to-face) in Hungary and Poland. Refer to topline for more information.
Source: Spring 2026 Global Attitudes Survey.
“Across the Globe, People Increasingly Say Social Media Is Harming Democracy”
PEW RESEARCH CENTER
The perception that social media has amplified political divisions has become more common in several nations, including by double-digit increases in Poland, Germany, France, Japan, Spain, the U.K., Hungary and Malaysia. (Since 2024, we have conducted our surveys in Hungary and Poland by phone; we previously conducted our surveys face-to-face.)
The U.S. is one of 10 countries where roughly eight-in-ten or more of those polled say access to the internet and social media has made people more politically divided. The share of Americans who hold this view today (81%) is similar to the share in 2022 (79%).
There are no significant partisan differences on this question in the U.S.: 83% of Republicans and independents who lean toward the Republican Party say social media amplifies political divisions, as do 82% of Democrats and Democratic-leaners.
Across 37 countries, Facebook is the most widely used of five social media platforms asked about
Median % who say they use the following websites and apps
Note: Percentages are medians based on 37 countries.
Source: Spring 2026 Global Attitudes Survey and survey of U.S. adults conducted Jan. 30-June 18, 2026.
“Across the Globe, People Increasingly Say Social Media Is Harming Democracy”
PEW RESEARCH CENTER
Across 37 countries, Facebook is the most widely used of five social media platforms asked about
Median % who say they use the following websites and apps
%
Facebook
69
Instagram
51
TikTok
40
X, formerly known as Twitter
18
Snapchat
18
Note: Percentages are medians based on 37 countries.
Source: Spring 2026 Global Attitudes Survey and survey of U.S. adults conducted Jan. 30-June 18, 2026.
“Across the Globe, People Increasingly Say Social Media Is Harming Democracy”
PEW RESEARCH CENTER
The survey also finds that the social media platforms people use vary across the globe, although Facebook tends to be the most common among the five we asked about, followed by Instagram. Many – especially young people – also say they use TikTok, while fewer overall use X (formerly known as Twitter) or Snapchat.
For this report, we surveyed 45,658 people in 37 countries from Feb. 8 to May 13, 2026. Additional data on Americans’ social media platform use comes from a survey of 5,511 adults conducted from Jan. 30 to June 18, 2026.
Majorities in most countries say social media has been more of a good thing for democracy. Most Canadians and Americans, however, offer a negative assessment.
Most say social media has been good for democracy but many in high-income nations disagree
% who say social media has been more of a __ for democracy in their country
Note: Those who did not answer are not shown.
Source: Spring 2026 Global Attitudes Survey.
“Across the Globe, People Increasingly Say Social Media Is Harming Democracy”
PEW RESEARCH CENTER
Most say social media has been good for democracy but many in high-income nations disagree
% who say social media has been more of a __ for democracy in their country
Country
Bad thing
Good thing
Grouping
Canada
55
40
North America
U.S.
64
34
North America
Poland
35
55
Europe
Italy
44
54
Europe
Greece
43
53
Europe
Hungary
42
49
Europe
Spain
47
49
Europe
Sweden
47
49
Europe
U.K.
56
41
Europe
Germany
61
37
Europe
France
65
31
Europe
Netherlands
70
27
Europe
Thailand
21
77
Asia-Pacific
Singapore
22
77
Asia-Pacific
Philippines
23
74
Asia-Pacific
Malaysia
27
71
Asia-Pacific
Bangladesh
28
67
Asia-Pacific
India
23
66
Asia-Pacific
Pakistan
33
62
Asia-Pacific
Indonesia
43
55
Asia-Pacific
Sri Lanka
35
54
Asia-Pacific
South Korea
38
52
Asia-Pacific
Japan
43
48
Asia-Pacific
Australia
52
46
Asia-Pacific
Israel
22
65
Middle East
W. Bank/E. Jerusalem
35
60
Middle East
Turkey
35
55
Middle East
Nigeria
11
80
Africa
Kenya
24
75
Africa
Ghana
21
72
Africa
South Africa
28
66
Africa
Mexico
26
69
Latin America
Brazil
29
65
Latin America
Colombia
33
63
Latin America
Peru
34
59
Latin America
Argentina
35
56
Latin America
Chile
43
51
Latin America
37-country median
35
55
Median
Note: Those who did not answer are not shown.
Source: Spring 2026 Global Attitudes Survey.
“Across the Globe, People Increasingly Say Social Media Is Harming Democracy”
PEW RESEARCH CENTER
In Europe, there are a variety of views. Poles, Italians and Greeks tend to see social media’s impact positively, while majorities in the U.K. Germany, France and the Netherlands say it has been bad for democracy.
People in the Asia-Pacific nations surveyed tend to believe social media has had a positive impact on democracy. In Thailand, Singapore, the Philippines and Malaysia, roughly seven-in-ten or more express this opinion. Still, 52% of Australians and 43% in both Japan and Indonesia take the opposite view.
At least half of adults in all the Middle Eastern, African and Latin American nations surveyed believe social media’s impact has been positive. Chile is the only nation surveyed in any of these regions where at least four-in-ten say it has been bad for democracy.
Social media and the internet can manipulate and divide, but also inform and empower
People believe social media has had both positive and negative influences on their country’s politics and its information environment.
Majorities in all 37 countries surveyed think access to the internet and social media has made people more informed about current events. And roughly half or more in every country think these technologies have given people the ability to have a more meaningful voice in politics.
Many say social media and the internet have both positive and negative consequences for democracy
% who say access to the internet and social media has made people
more
…
Source: Spring 2026 Global Attitudes Survey.
“Across the Globe, People Increasingly Say Social Media Is Harming Democracy”
PEW RESEARCH CENTER
Many say social media and the internet have both positive and negative consequences for democracy
% who say access to the internet and social media has made people
more
…
Country
Informed about current events
Able to have a meaningful voice in politics
Easy to manipulate with false information and rumors
Divided in their political opinions
Region
Canada
74
58
87
77
North America
U.S.
68
50
87
81
North America
France
77
59
88
69
Europe
Germany
80
57
87
82
Europe
Greece
79
47
87
69
Europe
Hungary
74
59
85
79
Europe
Italy
74
51
88
62
Europe
Netherlands
83
57
93
82
Europe
Poland
79
70
88
79
Europe
Spain
77
73
87
80
Europe
Sweden
83
67
90
72
Europe
U.K.
77
53
91
79
Europe
Australia
74
52
90
79
Asia-Pacific
Bangladesh
64
54
66
55
Asia-Pacific
India
66
55
62
51
Asia-Pacific
Indonesia
77
55
66
55
Asia-Pacific
Japan
89
79
88
64
Asia-Pacific
Malaysia
71
49
68
53
Asia-Pacific
Pakistan
74
68
70
62
Asia-Pacific
Philippines
74
71
70
66
Asia-Pacific
Singapore
81
59
77
60
Asia-Pacific
South Korea
80
77
79
74
Asia-Pacific
Sri Lanka
73
62
78
63
Asia-Pacific
Thailand
88
79
83
78
Asia-Pacific
Israel
80
73
63
55
Middle East
Turkey
75
64
71
66
Middle East
W. Bank/E. Jerusalem
88
62
87
73
Middle East
Ghana
83
77
76
63
Africa
Kenya
70
60
64
47
Africa
Nigeria
80
73
69
60
Africa
South Africa
76
60
71
57
Africa
Argentina
75
52
80
72
Latin America
Brazil
73
60
78
64
Latin America
Chile
74
56
80
73
Latin America
Colombia
77
63
77
68
Latin America
Mexico
81
57
75
64
Latin America
Peru
80
61
75
70
Latin America
37-country median
77
60
79
68
Median
Source: Spring 2026 Global Attitudes Survey.
“Across the Globe, People Increasingly Say Social Media Is Harming Democracy”
PEW RESEARCH CENTER
However, majorities in all countries – and overwhelming majorities in many high-income nations as defined by
World Bank lending groups
– also say access to the internet and social media has made it easier for people to be manipulated by false information and rumors. And roughly half or more in every nation say access has made people more politically divided.
Looking across the two positive and two negative effects asked about on the survey, some believe social media is leading to all four. In 29 countries, a third or more of adults hold this view.
Young adults, people with more education and those who use more platforms see social media more positively
Young people are generally more positive than older people about the impact of social media.
Younger people are more likely to see social media as good for democracy than older people
% who say social media has been more of a
good thing
for democracy in their country, by age
Note: Only statistically significant differences are shown.
Source: Spring 2026 Global Attitudes Survey.
“Across the Globe, People Increasingly Say Social Media Is Harming Democracy”
PEW RESEARCH CENTER
Younger people are more likely to see social media as good for democracy than older people
% who say social media has been more of a
good thing
for democracy in their country, by age
Country
Ages 18-34
18-34
35-49
50+
Youngest-oldest diff
U.K.
60
60
42
31
+29
India
74
74
69
50
+24
Ghana
79
79
71
56
+23
Thailand
89
89
80
68
+21
Turkey
63
63
63
42
+21
Japan
61
61
59
40
+21
France
44
44
34
24
+20
South Africa
72
72
67
53
+19
Kenya
80
80
71
63
+17
Peru
67
67
59
50
+17
Sri Lanka
63
63
51
46
+17
Australia
55
55
47
39
+16
U.S.
44
44
33
28
+16
Mexico
74
74
72
59
+15
Germany
46
46
41
31
+15
Colombia
68
68
67
55
+13
Singapore
82
82
79
70
+12
Canada
48
48
38
36
+12
Italy
61
61
57
50
+11
Sweden
57
57
48
46
+11
Israel
69
69
70
59
+10
Malaysia
77
77
66
68
+9
Argentina
60
60
57
51
+9
Note: Only statistically significant differences are shown.
Source: Spring 2026 Global Attitudes Survey.
“Across the Globe, People Increasingly Say Social Media Is Harming Democracy”
PEW RESEARCH CENTER
For instance, in the U.K., 18- to 34-year-olds are nearly twice as likely as those ages 50 and older to think social media has been a good thing for democracy. And significant gaps by age exist in 23 other countries, including the U.S. Still, less than half (44%) of American adults under age 35 believe social media’s impact on democracy has been positive.
Younger adults are more likely to see both the upsides and downsides of social media. They often say it makes people more informed and able to have a voice in politics, but they’re also more likely to believe it makes people more politically divided and manipulable.
Education follows a similar pattern as age. In some countries, people with higher levels of education tend to be more positive about social media’s impact on democracy, and they are more likely to see both its positive and negative effects.
Young people and those with higher levels of education are more likely to use social media and, in many countries, they tend to use more platforms. And people who use more than three of the platforms included on the survey are generally more likely than those who use fewer platforms to say social media has been good for democracy.
It is worth noting that on questions about the impact of social media on politics, older people, those with less education and those who do not use any social media platforms are often less likely to offer an opinion.
[$] Analyzing Rust programs with Charon
Linux Weekly News
lwn.net
2026-10-07 11:19:59
Nadrieril is a long-time Rust contributor, and the maintainer of the rustc
pattern-matching infrastructure. During his involvement with Rust, he has
noticed a problem with the usability of the language: it is difficult to
automatically extract information from a Rust crate for use with other toolin...
The page you have tried to view (
Analyzing Rust programs with Charon
) is currently available to LWN
subscribers only.
Reader subscriptions are a necessary way
to fund the continued existence of LWN and the quality of its content.
If you are already an LWN.net subscriber, please log in
with the form below to read this content.
Please consider
subscribing to LWN
. An LWN
subscription provides numerous benefits, including access to restricted
content and the warm feeling of knowing that you are helping to keep LWN
alive.
(Alternatively, this item will become freely
available on October 15, 2026)
Reverse Engineering of the M-VAVE FM-1 Pocket Synthesizer Firmware
Reverse engineering and update-protocol research for the M-Vave FM-1
synthesizer. Package, SPL, and SDK provenance identify the target as JieLi
AC791N/WL82 with a pi32v2 CPU, XIP flash at
0x02000000
, and a
Dexed/msfa-derived six-operator FM engine. The embedded
JL-BR22
string is
inherited library nomenclature, not reliable SoC identification.
This
main
branch intentionally contains no replacement firmware or custom
image builders. The former experimental implementation is preserved on the
with-custom-firmware
branch.
# Low-level disassembly and independent function map
scripts/run_ghidra.sh
python3 scripts/build_funcdb.py
python3 scripts/resolve_strings.py
python3 scripts/match_libs.py
python3 scripts/build_master_index.py
python3 scripts/build_slices.py
# OTA loader extraction, vendor map, and corroborative Ghidra sweep
scripts/analyze_ota_loader.sh
scripts/run_ghidra_loader.sh
# Current enriched classification database and documentation
python3 scripts/build_db.py
scripts/disasm_toolchain_libs.sh
python3 scripts/match_libs.py
python3 scripts/mech_tag.py
python3 scripts/export_shards.py
python3 scripts/aggregate.py
# Offline OTA protocol checks
python3 -m unittest discover -s tools/tests -v
Safety status
The update protocol is not a demonstrated recovery mechanism. Current work has
not established ROM recovery, rollback, or safe interrupted-write behavior for
the single-bank layout. The console/factory-mode audit in
analysis/device/debug-surfaces.md
found no
substitute recovery entry. Read
TODO_aug2.md
before using any
update or flash utility.
External references
USB_KEY | jielie
:
reverse-engineered notes on invoking JieLi USB boot using a signal on D+/D-,
including the key waveform, acknowledgement, timing, and USB bus caveats.
JL SoC forum thread
:
long-running Russian community discussion of JieLi SoCs, SDKs, toolchains,
programmers, boot activators, and USB/ISP/UART key experiments. Reports are
community observations and may apply only to the chip family being discussed.
SMK-37 Pro community notes
:
observations about a related M-Vave/JieLi keyboard that may help identify
shared packaging, update, and hardware conventions.
kagaimiq/jl-misctools
(
3rd-party/jl-misctools
, also checked out at
../jl-misctools
): utilities
for JieLi firmware containers, key files, UI resources, and older formats.
kagaimiq/jl-uboot-tool
(
3rd-party/jl-uboot-tool
): Python tooling for discovering UBOOT devices,
loading code into RAM, and reading, writing, or erasing flash. Its support
table lists WL82/AC791N as unknown, so it is not an established FM-1 flasher.
Jieli-Tech/fw-AC79_AIoT_SDK
(
../fw-AC79_AIoT_SDK
): official AC791N/WL82 SDK containing peripheral and
MaskROM API headers, boot/update configuration, libraries, build tools, and
application examples used to identify stock firmware behavior.
GitHub Incident with Git Operations, Pull Requests and Actions
We are seeing full recovery across all systems including Git Operations, Pull Requests and Actions. We are continuing to investigate the root cause of the disruption between 15:06 UTC and 15:16 UTC, and will share another update shortly.
Posted
Oct
07
,
2026
-
15:58
UTC
Update
The degradation has been mitigated. We are monitoring to ensure stability.
Posted
Oct
07
,
2026
-
15:56
UTC
Monitoring
The degradation has been mitigated. We are monitoring to ensure stability.
Posted
Oct
07
,
2026
-
15:49
UTC
Update
The degradation affecting Git Operations and Webhooks has been mitigated. We are monitoring to ensure stability.
Posted
Oct
07
,
2026
-
15:49
UTC
Update
The degradation affecting Issues has been mitigated. We are monitoring to ensure stability.
Posted
Oct
07
,
2026
-
15:47
UTC
Update
We are observing partial recovery across all systems. We are continuing to monitor and investigate the root cause of the disruption.
Posted
Oct
07
,
2026
-
15:41
UTC
Update
Git Operations is experiencing degraded performance. We are continuing to investigate.
Posted
Oct
07
,
2026
-
15:35
UTC
Update
Actions is operating normally.
Posted
Oct
07
,
2026
-
15:32
UTC
Update
The degradation affecting Pull Requests has been mitigated. We are monitoring to ensure stability.
Posted
Oct
07
,
2026
-
15:31
UTC
Update
Webhooks is experiencing degraded performance. We are continuing to investigate.
Posted
Oct
07
,
2026
-
15:27
UTC
Update
Git Operations is operating normally.
Posted
Oct
07
,
2026
-
15:27
UTC
Update
Pull Requests is experiencing degraded performance. We are continuing to investigate.
Posted
Oct
07
,
2026
-
15:25
UTC
Update
We observed widespread impact across GitHub services between 15:06 UTC and 15:16 UTC, and are investigating the root cause. We are seeing all systems recover, and will provide another update shortly.
Posted
Oct
07
,
2026
-
15:24
UTC
Update
Git Operations is experiencing degraded performance. We are continuing to investigate.
Posted
Oct
07
,
2026
-
15:21
UTC
Update
Webhooks is experiencing degraded availability. We are continuing to investigate.
Posted
Oct
07
,
2026
-
15:16
UTC
Update
Issues is experiencing degraded performance. We are continuing to investigate.
Posted
Oct
07
,
2026
-
15:15
UTC
Update
Webhooks is experiencing degraded performance. We are continuing to investigate.
Posted
Oct
07
,
2026
-
15:14
UTC
Investigating
We are investigating reports of degraded availability for Actions, Git Operations and Pull Requests
Posted
Oct
07
,
2026
-
15:14
UTC
This incident affects: Git Operations, Webhooks, Issues, Pull Requests, and Actions.
Visa, Mastercard, Major Banks Facing New Litigation over 'Anticompetitive' Fees
A San Diego pizzeria alleges in a proposed class action lawsuit that Visa, Mastercard and several of the nation’s largest banks have maintained a conspiracy to artificially inflate the fees merchants pay on credit card transactions, despite a more than $5 billion class action settlement in recent years over the same allegedly anticompetitive market restraints.
The 134-page lawsuit contends that the effectively non-negotiable credit card transaction fees imposed by the defendants amount to “a deadweight toll on virtually every credit card purchase in America,” totaling hundreds of billions in “monopoly rents” at rates that “no competitive market would produce.”
According to the complaint,
Visa
and
Mastercard
have worked with the bank defendants—
Bank of America
,
Capital One
,
Chase Bank
,
Citibank
and
Wells Fargo
—for decades to set uniform schedules of interchange fees, i.e., the effectively non-negotiable charges merchants must pay to credit card-issuing banks on each transaction. The case charges that the defendants, to maintain the high interchange fees and “ensure merchants cannot escape them,” have implemented a “web of anticompetitive rules,” or restraints, that on the whole have disabled any market forces that could rein in or otherwise discipline the credit card
transaction costs
imposed on merchants.
Broadly, the interlocking restraints set and maintained by the defendants have forced merchants that accept any Visa and Mastercard credit card to accept all such cards, regardless of cost, thereby eliminating any incentive for issuing banks to compete by lowering their fees, the suit says. The lawsuit claims the challenged restraints have also prevented merchants from being able to
steer customers to lower-cost payment options
—for instance, by surcharging based on a customer’s use of a particular card. These and other restraints have prevented competition among issuing banks and other credit card networks, allowing the financial giants to raise their fees every year “without consequence,” the case alleges.
“The scale of this ongoing scheme is staggering: merchants now pay over $100 billion annually in fees to accept Visa and Mastercard credit cards, enriching Defendants at merchants’ expense,” the class action lawsuit states.
The filing claims Visa and Mastercard have also exploited the allegedly anticompetitive market handcuffs to artificially inflate their own network fees—comprised of per-transaction fees and fixed fees—charged to merchants as a cost of accepting the companies’ credit cards, adding “an additional supracompetitive tax on each credit card transaction.”
The suit states that in December 2019, the court approved a class action settlement in years-old
multidistrict litigation
that provided upward of
$5 billion in monetary relief
to merchants, but only for a “class period” ending on January 24, 2019. Although a separate “equitable relief” class action settlement seeking
injunctive relief
has been preliminarily approved, the benefits will apply only prospectively and merchants will not receive “a single dollar in compensation ... for the fees they have paid since January 25, 2019.”
According to the complaint, while some merchants obtained relief for pre-2019 transactions and some may benefit from rule changes in the future, merchants that accepted Visa and Mastercard credit cards after January 2019 have “borne the full brunt of Defendants’ continuing anticompetitive conduct.”
“This settlement structure leaves millions of American merchants without any remedy for their ongoing injuries,” the lawsuit attests.
The
interchange fees class action lawsuit
looks to represent all individuals, businesses and other entities that have accepted Visa-branded and/or Mastercard-branded credit cards in the United States from January 25, 2019 until the alleged “anticompetitive effects” of the defendants’ conduct cease.
Check out ClassAction.org’s lawsuit list for the latest open
class action lawsuits
and investigations.
Books about successful people often make success sound rather unappealing. Education should be intense and rigorous, focused on driving up IQ and on the ‘STEM’ quaternity of science, technology, engineering, and mathematics. Once adulthood is reached, life follows an austere pattern, with a 5am start, cold baths, daily visits to the gym, little or no alcohol, a Huel-rich diet, and Stakhanovite working hours. This last point generates particular excitement, and descriptions of powerful people dwell with eager fascination on how many hours they devote to their labors.
The elites of Victorian Britain operated differently. Their schools and universities were not terribly academic and had very little STEM. As adults, they got up late, drank a lot, and spent a remarkable share of their waking hours partying. They loved feasting, sports, holidays, dancing, and dressing up. Contrary to their stodgy reputation, they were probably a lot of fun.
The Works in Progress Newsletter
Get new articles from Works in Progress delivered to your inbox.
Despite this, the elites of nineteenth-century Britain were, by their own lights, extremely successful. For much of the century, Britain was the richest country on Earth. Its companies dominated global commerce, industry and finance. Its huge navy reigned over the world’s oceans, and its tiny army won a long sequence of wars while keeping watch over an enormous empire. Britain progressed gradually toward democracy with none of the revolutions and civil wars that wracked continental Europe. There is a longstanding view that the prosperity of Britain’s elites came at the cost of the hundreds of millions whom they subordinated and oppressed, and no doubt that is at least partly true. But on their own terms, the British ruling class enjoyed a
victorious century
.
Maybe the Victorian elite succeeded in
spite
of being feckless and incompetent: the Industrial Revolution created great wealth, which they enjoyed passively and parasitically. We cannot run controlled experiments on Victorian Britain, so this thesis is hard to test. But I want to explore the possibility that a relentless focus on soft skills, team building, and intra-elite networking actually worked rather well. Huel, STEM, and cold baths have their merits. But the example of the Victorians should make us hesitate before concluding that there is only one way to rule the world.
School
Caricaturing a little, we can distinguish two groups in modern educational debates: traditionalists and progressives. Traditionalists think that education should prioritize acquiring knowledge, that it should focus on ‘hard’ subjects like science, mathematics, and foreign languages, and that it should take place in a highly structured environment, where students’ time and behavior are tightly regulated. Progressives think that education should prioritize acquiring skills, that it should embrace ‘soft’ subjects beyond the supposed academic core, and that students should be given wide freedom in which to develop personally and socially.
One might assume that Victorian education was at the extreme traditionalist end of the spectrum, and modern educational progressives would indeed be horrified by the rote learning and corporal punishment of Victorian schooling. But in respect of all the controversies listed above, elite Victorian schools lean curiously toward the ‘progressive’ side of the spectrum.
Most socially elite Victorian boys began formal education at the age of seven or eight, spending about five years in a ‘preparatory school’ before passing on to a ‘public school’ (which, to the enduring confusion of international observers, is a kind of elite private school). There was a hierarchy of public schools, with the nine ‘Clarendon schools’ at the top, and just one, Eton College, clearly at the apex. About
30 percent of nineteenth-century cabinet ministers
and 40 percent of prime ministers were Old Etonians.
Most students at both preparatory and public schools were boarders, meaning that elite Victorians were generally sent away from home by the age of eight and spent their childhoods more with peers than with family. In this respect they were distinctive even in their own time: elite families in continental Europe and the United States were far more likely to keep their children at home and educate them at small local private schools. It is plausible that this contributed to the lack of subnational loyalties in the British elite, in contrast to the distinct elites of, for example, the Southern United States, Catalonia in Spain, or Hungary in the Austrian Empire.
The public schools were of course single-sex, generating a somewhat febrile atmosphere in which intense quasi-romantic (or sometimes actually romantic) relationships between boys were common. Perhaps the powerful ties of nostalgia and loyalty that students felt for their former schoolmates were linked to this. The writer Jan Morris
said that
‘the whole of English upper-class life…was shot through with bisexual instinct’ and that ‘the public school system…meant that male relationships were full of emotional nuance and undertone’. Demonstrative affection remained normal in adult life, and it was completely standard for pairs of Victorian men to walk arm-in-arm in public.
Academic life was dominated by Latin and, to a lesser extent, Greek. The exact share varied, but the classical languages took up generally at least half of teaching hours, and sometimes as much as 80 percent, especially earlier in the century. Trailing in third and fourth places were mathematics and modern languages. There was resistance to the teaching of science, not because Victorians were hostile to science in principle, but because they thought it irrelevant to the kind of moral and cultural formation that a school for gentlemen was supposed to provide. Thomas Arnold, headmaster of Rugby School and the leading educationalist of Victorian Britain,
remarked that
‘rather than have physical science the principal thing in my son’s mind, I would gladly have him think that the sun went round the Earth, and that the stars were so many spangles set in the bright blue firmament’. (By contrast, girls were not usually taught classics, meaning that, rather surprisingly, affluent Victorian girls generally followed a more modern curriculum than their brothers.)
Victorian adults had little occasion to use their classical languages, so they tended to forget them. ‘I have gladly forgotten all my Latin, and don’t ever want to remember it again’,
wrote the war poet
Charles Sorley, one year after leaving school; at least he did better than Winston Churchill,
who remarked that
‘In all the twelve years I was at school no one ever succeeded in making me write a Latin verse or learn any Greek except the alphabet’. The most that can be said for Victorians’ classical education is perhaps that, like schooling in any difficult subject, it trained students in self-discipline, memorization, and applying rules.
The Victorians themselves occasionally hinted that they regarded the official curriculum of school as something of a
MacGuffin
. In the hugely influential 1857 novel
Tom Brown’s School Days
, Tom’s benevolent father muses, ‘I don’t care a straw for Greek particles… What is [Tom] sent to school for? … If he’ll only turn out a brave, helpful, truth-telling Englishman, and a gentleman, and a Christian, that’s all I want’. In a similar spirit, Thomas Arnold wrote that ‘what we must look for here is, first, religious and moral principles; second, gentlemanly conduct; thirdly, intellectual ability’.
Under such circumstances, it is perhaps unsurprising that the share of time spent in the classroom was relatively modest. At Rugby School, weeks were divided into three ‘full days’, three ‘half days’, and Sunday. On a full day, students spent about five hours in the classroom, compared with maybe two hours on a half day. Other schools varied, but nearly all left many hours unfilled. In general, classroom hours were certainly not longer, and probably rather shorter, than those in normal schools today. They also tended to be shorter than those of contemporary elite schools in France, Germany, or Italy, where the curriculum was normally more academically rigorous.
The most prestigious Victorian school was Eton, but the most influential was Rugby, where many of the characteristic features of the Victorian public school were developed under the headmastership of Thomas Arnold. The version of football played at Rugby became known as ‘Rugby football’ and then just ‘rugby’. Here it may be seen in a primitive state, before the rules reached their canonical form.
Image
Pictorial Press via Alamy.
The focus on character formation led to increasing interest in how students spent their time outside the classroom. In the early decades of the nineteenth century, they were still mostly left to their own devices: their preferred pastime seems to have been roaming the local countryside and generally making a nuisance of themselves. During the Victorian period, school authorities introduced programs of organized sports, chiefly cricket in summer and various forms of ‘football’ (including what became rugby) in winter. Many schools ultimately had four, five, or even six afternoons of sports per week: a student would certainly spend ten hours a week on the sports field, and might well spend twenty.
The Victorians
thought extremely highly
of sports, believing that they instilled team spirit and endurance, fostered mutual loyalty, channeled energy and competitiveness, and taught men how to win and lose graciously. An extreme but not entirely unrepresentative example was Hely Hutchinson Almond, headmaster of Loretto School, who described his school’s priorities as ‘First – character, second – physique, third – intelligence, fourth – manners, fifth – information’. The cult of team sports was generally not shared in continental Europe, and European visitors were often bewildered by how seriously the English took schoolboy ballgames. But it spread across the Anglophone world, and it is still embodied in the traditions of American high school and intercollegiate sports today.
Another distinctive feature of the public schools, relative both to modern education and to other nineteenth-century school systems, was student self-government. The Victorians thought that close supervision by schoolmasters would inculcate servile habits in boys, compromising what
one headmaster called
‘the great glory of an English public school – its free development of character, its social expansiveness, in short its liberty’. As a result, a system of prefects was developed in which, in theory, the more benevolent and mature boys were given extensive pastoral and disciplinary responsibilities over the others. The prefects would receive an apprenticeship in leadership, while the student body as a whole would come to see authority as something emanating from their peers rather than imposed upon them by an alien entity.
It is not hard to imagine how a self-governing oligarchy of teenage boys might go wrong. Victorian educationalists were themselves concerned about the potential for abuse, and various checks and balances were developed, like jury systems and rights of appeal before punishment. It is hard to say how often the prefect system operated unjustly, reliant as we are on the impressionistic evidence of personal memoirs, but it is beyond doubt that it was sometimes brutal. It is important to understand, however, that it was intended not simply as a system of domination, but almost as one of ‘elite liberalism’, whereby students would not only be subjects of authority, but also to some extent its legislators and executors.
University
From school, most elite Victorian men went on to Oxford or Cambridge (‘Oxbridge’). Oxbridge was not necessarily a big step up academically. Virtually everyone who applied was accepted. Once you were in, you could pursue one of two tracks: either an ‘honors’ degree, which involved at least some serious study, or a ‘pass’ or ‘ordinary’ degree, which involved sitting only some very easy exams in classics and religion. In the early Victorian period, most Oxbridge students were ‘pass men’; this changed gradually as the century drew on, but there were still plenty of pass men around as late as 1914. The pass degree was particularly popular among aristocratic students, who tended to prioritize rowing, feasting, and social polish over disciplined study.
Oxford in the early twentieth century. The town was about a mile across, and the colleges were all within a short walk of each other. As at many times in their lives, the Victorian elite chose to congregate geographically.
Image
David Rumsey Map Collection, Stanford University Libraries.
Many honors students took their studies more seriously. But even they were not very practical in orientation. Oxbridge tutors believed overwhelmingly that the purpose of an undergraduate degree was not training for work nor even preparation for academic research, but rather a general moral and intellectual formation that would fit students for any professional role they might later pursue. ‘No man will be a first-rate physician or engineer who is not something more than either, who has not some taste for art, some feeling for literature, or some other interest external to his profession’,
said Benjamin Jowett
, the influential master of Balliol College. Jowett aimed to convey this ‘something more’ to his undergraduates, leaving them to acquire specific technical knowledge later.
In the first half of the nineteenth century, such theories led the universities to exclude almost everything except classics and (at Cambridge) mathematics from the curriculum. Natural and applied sciences were viewed with suspicion on the basis that they were,
in the words of
the influential Oxford professor Edward Bouverie Pusey, ‘confined to the reception of information on matters of fact’ and do not involve ‘continued reasoning’.
Between 1855 and 1860
, just 8.1 percent of honors degrees were awarded to science students; in 1910–1914, it was still only 13.7 percent.
Nearly all students lived together in colleges. This contrasted with the situation in continental Europe, where students generally rented rooms off campus. Victorian pedagogues believed that colleges provided religious, cultural, and social formation, and that they prevented English students from acquiring the bad morals and habits of their continental peers. A typical day began with chapel; honors students and the less indolent pass men would then spend the morning reading and preparing assignments. Students normally ordered lunch to their rooms: rich students could have elaborate feasts brought up from the college kitchens, like Sebastian Flyte’s lunch in
Brideshead Revisited
. In the early afternoon it was customary to go out for rural walks or sports, especially rowing and cricket.
The Prince of Wales, the future Edward VIII, at a meal with friends in Magdalen College, Oxford. Students could order lavish meals to their rooms from college kitchens, with the cost added to their termly bill. Notice the cheap and mismatched furniture – this was still student accommodation, however grand its residents.
Image
Magdalen College Archives.
Around 4pm, students returned to college. Once or twice a week this was the time for ‘tutorials’, one-to-one or small-group discussions with tutors, often over tea or sherry. The tutorial system contrasted sharply with the more didactic and lecture-based teaching that prevailed in other university systems, and was seen as a way of inculcating independent thought. If a student had no tutorials, he would probably spend the early evening reading, socializing, or at one of the university’s various cultural, discussion, or debating societies, notably the Oxford and Cambridge Unions. The day concluded with a formal dinner in the college hall, often followed by conversation over dessert, digestifs and tobacco in college rooms.
Although the tutorial system was seen with pride, the quality of tutoring was decidedly uneven. Until the 1870s, most fellowships were open only to Anglican priests. Fellowships were comfortable positions, but much less comfortable than well-paid positions in the Church of England, especially since fellows were generally not allowed to marry. The typical Oxbridge don of the mid-nineteenth century was thus in his early or mid-twenties and regarded his fellowship as a holding post, which he would leave as soon as a desirable parish became available. Many of these young men would have been very bright, but few were serious researchers or experienced teachers. Fellowships were opened up to laymen in the later decades of the century, but Oxford and Cambridge did not become world-leading research institutions until they were transformed by German scholars fleeing Nazism in the 1930s.
The dominance of Oxbridge in British culture was unusual: these institutions educated a much larger share of the elite than was the case for any two institutions in any other country, with the possible exception of the top
grandes écoles
in France. Nonetheless, some elite and quasi-elite Victorian groups followed alternative educational pathways. For example, for much of the nineteenth century the colonial civil service in India was educated at a training college called Haileybury. These programs were not, however, necessarily more rigorous than the usual Oxbridge track. I cannot improve upon
David Gilmour’s description
of a day in the life at Haileybury, distilled from the memoirs of the civil servant John Beames:
At eight in the morning, the students rushed to chapel, many of them wearing only a nightshirt under a gown or overcoat. Afterwards they returned to their rooms for breakfast, which often turned into parties with ‘tankards of beer or claret’. Breakfast was followed by smoking a pipe and dealing with tailors and other tradesmen who had arrived from Hertford. At ten o’clock the bell rang for lectures, which lasted two hours on some days, three on others. The studious minority known as the ‘steadies’ took notes; the others yawned, drew sketches of their teachers and fooled about. After a lunch of bread, cheese and beer, they were free to do what they liked. The oarsmen went to the river, the cricketers went to their cricket pitch, the ‘steadies’ went for a ‘solemn constitutional’ along the lanes, and the ‘fast men’ either slipped off by train to London or else clattered off in
dog-carts
to play billiards at Hertford or Ware. Dinner took place in Hall at six followed by evening chapel at eight. Then the ‘steadies’ retired to their rooms and read far into the night, most of the others sat about drinking and singing, and at about two in the morning the ‘fast men’ returned very drunk from Broxbourne Station after catching the last train from Shoreditch.
Beames’s account was particularly colorful, but nobody seems to have had a good word to say about Haileybury as a teaching institution. Everyone, however, including Beames, agreed that it engendered much
esprit de corps
among its graduates. The future prime minister Lord Salisbury commented on the ‘close friendships formed there, which softened the rivalries of after life and secured devoted instead of perfunctory co-operation’. Perhaps something similar might be said for much of the elite Victorian education system.
Adulthood
One virtue that upper-class Victorian education signally failed to inculcate was early rising. Although Victorian schoolboys and university students got up fairly early for chapel, the day of an upper-class Victorian adult started late. The great prime minister of the early nineteenth century, Pitt the Younger,
breakfasted at 9am
. The great prime minister of the Victorian period, William Gladstone, hosted a social breakfast
starting at 10.30am
(claret was served), continuing till about noon. His rival Benjamin Disraeli attended breakfasts hosted by Lord Houghton, which started at 10am, but which Disraeli would
join at about 11.30am
. Elite Victorians generally had flexible working hours and were free to work from home unless they had meetings elsewhere. Many spent the whole morning at home, reading the newspaper, answering letters, and preparing for later meetings. The Edwardian prime minister Arthur Balfour generally did all this
literally in bed
, rarely emerging from his bedroom before lunchtime. As we shall see, mornings must often have been even later from February to July, when many leading Victorians would have spent much of the night partying.
At some point in the late morning, the Victorian would finally make his way to the office. By the end of the century, some wealthy people lived in suburbs and commuted by railway, especially in industrial cities like Birmingham and Manchester. In general, however, Britain’s elites were slow to suburbanize, and in 1914 most still lived in centrally located row houses and walked to work. Every day a great herd of frock-coated men emerged from Belgravia and Mayfair and poured through Covent Garden to their offices in Parliament, Whitehall, the Inns of Court, and the City.
The Strand in London around 1895. The Strand is one of the main roads between the residential West End and the commercial City, so it is thronged with top-hatted men commuting by foot or by horse bus.
Image
Alamy.
The exact timetable for the middle of the day varied by period and profession. In the early Victorian period, the upper classes often went without a serious lunch but dined early at 5 or 6pm. This changed drastically in the third quarter of the century, perhaps because of improved artificial lighting, so that by the 1870s many Victorians had a substantial lunch with wine, usually with professional peers at a club, sustaining them until dinner at about 8.30pm. Politics was a part-time job for most members of Parliament (MPs) and peers, so Parliament typically did not start sitting until 4.30 or 5pm; the majority of part-time politicians could wind up a gentlemanly day at the office in time for this, while the minority of full-time politicians could spend the day on party management, or (if they were ministers) at government business in their ministries.
Victorians loved dressing up, and both men’s and women’s clothes were much more elaborate than they are today.
Image
Wikimedia Commons (adapted).
Provided Parliament did not interfere, however, an elite Victorian would stroll home to dress for dinner. Upper-class Victorians of both sexes paid great attention to their clothes. For an informal dinner at home, a man would change into evening dress, usually including a starched collar, cuff and shirt front, patent leather shoes, and gold shirt studs and cufflinks.
Dinners were rich and elaborate. The number of courses naturally varied, but it was invariably more than a modern dinner at an equivalent grade of formality. A typical dinner began with soup and sherry, then proceeded to fish with hock (Rhineland whites), passing on through a sequence of meat courses with champagne or red wines, and finishing with various desserts and savories served with port and madeira. The Victorians loved drinking and drank a lot by modern standards. Men drank more than women, but women drank a good deal too: one conduct guide
noted that
‘young ladies seldom drink more than three glasses of wine at dinner but married ladies, professional ladies, and those accustomed to society, and habits of affluence, will habitually take five or even six, whether in their own homes or at the tables of their friends’.
For about five months a year, between February and July, the elite of Victorian Britain congregated in London for what was called ‘the Season’, essentially a sequence of parties and social events. They usually owned or rented houses in Mayfair and Belgravia, two small neighborhoods that face each other across Hyde Park, meaning that the bulk of the ruling class would live literally within walking distance of one another. Parallel institutions existed in most Western countries at the time – the social seasons in Paris, Vienna, St Petersburg and New York were famous – but the London Season had a particularly dominant role in the national life because it coincided with the Parliamentary session, making it a forum for political networking as well as elite matchmaking and court rituals.
London’s residential neighborhoods in 1889, graded by social class. The wealthy neighborhoods are in yellow, clustered around the royal parks in inner west London. There was an internal hierarchy even among these, with Belgravia and Mayfair at the top of the hierarchy. Most of Britain’s elites lived in London for at least part of the year, during which time they would have been within half an hour’s walk of each other.
Image
David Rumsey Map Collection, Stanford University Libraries.
During the Season, the upper-class evening continued long beyond dinner. Balls were the flagship evening entertainment. They started late, at nine or ten, and guests would filter in gradually, with the dance floor really full from about 11pm. Around midnight a meal called ‘supper’ appeared, comprising soups, salads, cold meats, shellfish, fruit, and ice cream. The beverage most associated with balls was ‘negus’, a hot drink made of port, spices, sugar and oranges or lemons. Lemonade, tea and coffee were on hand throughout the evening, and champagne, hock, port and sherry were served at supper. In addition, an 1867 guide for aspiring ball hosts
observed
, ‘Bottled pale ale and stout are quite permissible – in fact, they have become indispensable’.
After supper, dancing recommenced; a good ball peaked after 1am. Guests gradually returned home in the small hours, and the most determined might continue until dawn. It was good form to slip out quietly, and there was no requirement even to say goodbye to the hostess. ‘The great thing is to avoid making your departure felt as a suggestion for breaking up the party, as you have no right to hint by your movements that you consider the entertainment has been kept up long enough’,
an etiquette guide notes
.
Balls naturally had special significance for people in the marriage market. Arranged marriage in the strict sense had largely died out by the nineteenth century, but upper-class young people and their parents still hoped they would make advantageous matches, and balls furnished opportunities to do so. It would be quite wrong, however, to think that balls were simply mass dating events. Courtiers, bankers, judges, and cabinet ministers would have shared the dance floor with debutantes. In fact, with some exceptions like club lunches, the social events of the Victorian elite were notable for their lack of segregation by age or sex. A Society ball discharged the functions of both a teenage house party and a business networking convention: in fact, it
was
both a teenage house party and a business networking convention.
The Victorians loved dressing up: fancy dress balls were a popular variant on the normal format.
Image
iStock.
Unmarried young people (and their unfortunate parents) might go to three or four balls per week, while other Society members would go to one or two. But the other nights were scarcely less busy. Opera, theatre and concerts were important, especially opening or gala nights. The bread and butter of the Season was soirées, card parties,
conversazione
, receptions and less formal dances, as well as midnight ‘supper parties’, all mostly hosted in private homes. These elements were modular. An MP might miss dinner for a late sitting of Parliament, but still drop into a soirée at 10.30pm, turn up to a supper party at 11.45pm, and then arrive at a ball at 1am. From February to July, the Victorian upper class led a remarkably nocturnal existence, with bedtimes in the small hours and the mornings something of a write-off.
By one definition, then, elite Victorians did remarkably little work: they could easily spend just four or five hours a day at unambiguous office work. But by another definition, they were working all the time. Lunches, balls, and dinners were certainly fun, but they were not
just
fun. In the words of the historian Leonore Davidoff, these
social occasions provided
‘access to a vast information network: information about jobs, investment possibilities, secret political decisions’. Many Victorians were intensely ambitious and drove themselves hard: sleep deprivation, exhaustion and burnout were familiar companions. The distinctive feature of elite Victorian lifestyles was not laziness, but an extreme weighting toward what we might call team building and professional networking, and the near-total fusion of personal and professional life.
Method in their madness?
The sociologist Ferdinand Tönnies famously distinguished
Gemeinschaft
– community, held together by ties of sentiment and tradition – from
Gesellschaft
– society, held together by explicit laws and contracts. As societies grow more complex, he thought,
Gesellschaft
tends to replace
Gemeinschaft
. But the Victorian elite was a little
Gemeinschaft
at the apex of an enormous
Gesellschaft
, an intimate community bound together by ties of common culture, shared memory, intense status competition, loyalty, and love. Its members strove for one another’s envy, esteem and gratitude, and to achieve this they would go far beyond what was contractually obligated or financially advantageous. It is not altogether fanciful to think of the Victorian elite as functioning like a big premodern village, though its members were geographically dispersed across the kingdom and its vast empire.
All the practices I have described in this article helped to maintain this emotional village. The ruling class were educated together from an early age, in schools and universities designed to inculcate shared values, a common identity, and ties of mutual loyalty. As adults they congregated every year in an area of central London that actually was village-sized, where they spent every night in one another’s company. Young members of the Victorian elite tended to meet their future spouses during these glittering evenings, so that the elite came to form, to some extent, a literal family.
Maintaining this community
required a lot of time and effort, which obviously had opportunity costs. All else being equal, rich and powerful Victorians must have been less productive for spending four hours a day in the office rather than twelve. But the creation of a single ruling
Gemeinschaft
may have had compensatory advantages.
One of these was unity. In some ways, rich Victorians were quite diverse. They came from every part of the British Isles. They included landed aristocrats and urban bourgeois, with the latter steadily superseding the former over the course of the century. Most landowners were Anglicans, but many of the new rich were Nonconformists (non-Anglican Protestants), and some were Jews or Catholics. With the conspicuous exception of the Irish Catholics, all these people were successfully folded into a single British elite, rather than forming hostile counter-elites working to overthrow the social order.
This obviously had advantages for Britain. The elites of the American South had a very different identity to those of the North, and ultimately led the country into civil war. Spain, Italy and France suffered from incessant conflicts between urban bourgeois and rural nobility. Germany was riven with strife along geographical, religious, and class lines. Some of these tensions existed in Britain, but except in Ireland, the elite finessed them peacefully. A Nonconformist Manchester businessman probably supported the Liberal Party rather than the Conservatives, but the idea of violently seceding from Britain to form a Methodist Republic of the North would never have occurred to him.
Maybe the Victorian elites also had advantages relative to elites today. One strength of all
Gemeinschaften
is that they can reward their members with prestige rather than money, which is useful in contexts where adequate financial rewards are for some reason impossible.
This might be an important advantage in some contexts. For political reasons, most modern societies pay politicians and civil servants vastly less than they pay people with a similar level of responsibility in the private sector. As a result, they struggle to persuade their most talented people to enter and remain in public life. The Victorians did not have this problem. Until the 1880s, most local government roles in Britain were conducted on a completely unpaid basis by local landowners. They also administered the rudimentary welfare system, sat as magistrates in the lower courts, and staffed the boards of most schools and hospitals. A landowner who ignored these responsibilities would suffer no material consequences, but he would be diminished and resented in the eyes of his peers, and for most, that inducement was enough.
Something similar applied to Parliament. Nineteenth-century MPs were unpaid and covered most of their own election expenses.
In the 1880 General Election
, the average Parliamentary candidate bore about £150,000 of expenses in an urban constituency and £450,000 in a rural one (in today’s money). MPs had to take this massive financial hit repeatedly, every time there was an election. They also had to bear all of their costs upon assuming office, including staff, travel, and accommodation in both Westminster and their constituency. And yet there was no difficulty filling Parliament: many of the richest and most powerful people in England competed fiercely to enter, simply because being an MP was intensely prestigious in their social milieu.
Maybe this identification with the social system affected how politicians behaved in office, too. In
The English Constitution
the political writer Walter Bagehot noted that parliamentary democracy creates powerful incentives for opposition parties to damage the country by corrosively undermining the government when it makes difficult but necessary decisions. But in practice, he thought, English politicians did not do this as much as one might expect:
In abstract theory these defects in our present practice would seem exceedingly great, but in practice they are not so. English statesmen and English parties have really a great patriotism; they can rarely be persuaded even by their passions or their interest to do anything contrary to the real interest of England, or anything which would lower England in the eyes of foreign nations. And they would seriously hurt themselves if they did.
It is hard to know how seriously to take claims like this. But it is certainly true that the MPs of both parties overwhelmingly belonged to the community described in this essay, sharing friends, relatives, memories, and values. Perhaps the elite village could punish them for defecting from norms in ways that are impossible today, making political life less prone to divisiveness, demagoguery, and brinkmanship.
Would modern societies work better if their local and national government were conducted by a tight-knit community of extremely rich and successful people? Maybe not: the downsides are obvious. But the problems would be interestingly different from those that they have in reality. Imagine a San Francisco whose Board of Supervisors featured Jeff Bezos, Larry Page, Sam Altman and Mark Zuckerberg, each feeling this responsibility to be the greatest glory of his career, each hoping that he might one day attain the greater glory that is the House of Representatives. It might not be paradise. Maybe it would be a disaster. But it would certainly be interesting.
When reading the news, you may occasionally be bemused by the flip-flopping science and health stories: one week you’re told coffee is good for you, the next week that it’s bad. Red wine extends your life, then it doesn’t. Over time, new scientific studies seemingly contradict one another, leaving us unsure what evidence to believe.
As a science journalist
,
I’ve learnt that often the actual problem is that each study is reported in isolation, rather than weighed against everything else we know. And sometimes, the risks of harm are high. For example, in
September 2025,
I
reported
on the US Food and Drug Administration’s announcement that it would add a warning label to the painkiller acetaminophen (paracetamol), claiming that taking the drug during pregnancy increased a child’s risk of autism. It made global headlines. Yet looking closer at the science, it was clear that the FDA and its parent health department had
cherrypicked
a few studies, and downplayed other more robust findings. This was confirmed a couple of months later when researchers published a major
review
that better represented the full body of knowledge. ‘Existing evidence does not clearly link maternal paracetamol use during pregnancy with autism or ADHD in offspring,’ it concluded. By then, of course, countless women had been needlessly scared about a painkiller that is widely considered one of the safest to take during pregnancy.
In 2025, academics worldwide published about
7 million
scholarly articles – that’s more than 19,000 each day. On the surface, that might seem like a number to celebrate, but it also poses a problem. As the volume of research balloons, it can be hard to discern what these millions of papers actually tell us. Within the firehose, there are rigorous methodologies and important findings, but also baffling contradictions, unconfirmed results and sloppy science. And some diverse forms of knowledge, such as lived experience and Indigenous insights, are rarely captured in scholarly articles and databases
at all.
The causes are systemic. Scientists publish and promote one paper after another because that’s how they advance in their careers. Journalists breathlessly chase the latest, flashiest studies so their headlines get clicks online. Meanwhile, hardly any of us spend time trying to make sense of what the world already knows by carefully synthesising and taking stock of existing knowledge. Iain Chalmers, a doctor who co-founded the Cochrane Collaboration, an evidence-synthesis group based in London, once called this the ‘scandalous failure of science to cumulate evidence scientifically’.
This failure was a key motivation to write my book
Beyond Belief: How Evidence Shows What Really Works
(2026). Researching it, I discovered better ways to make sense of the world – but these methods are not as widely known as they should be. If they were, we might pause before believing news stories based on single studies, and instead recognise the real work: finding, sorting and synthesising evidence. This unglamorous labour already shapes our lives far more than any individual study or news headline will. We might also realise that, for many of the wicked problems we face, humans already possess much of the knowledge needed to solve them. All it needs is the ability to assemble it – and
then act.
O
ne of the biggest events in the history of evidence synthesis occurred in a hotel in San Francisco in 1976. There, at the annual meeting of the American Educational Research Association, the statistician Gene Glass
revealed
an ingenious technique that showed how to integrate findings from a large volume of individual studies. This technique would become so important to science that it has its own
biography
.
Glass’s work had emerged from his own struggles with mental health. A decade earlier, when he had finished his PhD in psychometrics and statistics, he was suffering from anxiety and neurosis. He began weekly psychotherapy sessions – which helped him hugely – and stayed in therapy for eight years. But his positive experience ran counter to the weight of academic opinion, which maintained that psychotherapy had little benefit. This stemmed in large part from the work of the psychologist Hans Eysenck, whose influential reviews of psychotherapy research
concluded
that it was worthless – or had a placebo effect at best.
Glass was irritated. Not only did this suggest that he’d flushed away money on ineffective therapy, but he felt Eysenck’s reviewing methods were flawed.
Eysenck – like other researchers at the time – tended to do what’s called ‘vote-counting’: counting up the number of studies showing a treatment has benefits, and the number that did not find it helped. The one with the most ‘votes’ wins. Intuitive though vote-counting is, it’s also misleading as a way of synthesising studies, partly because it ignores the size of the effects. One study may find that 55 people out of 100 improved after a particular treatment, and another that 85 out of 100 did. Even though the second study showed the treatment had a much stronger effect, both studies contribute one vote. Glass could see this was problematic – and was also astonished that Eysenck arbitrarily excluded hundreds of valid studies just because they were in theses or dissertations rather than published in academic journals.
The audience was thunderstruck. Here was a way to extract meaning from an apparent mess of different results
Glass set out to do a better job. He and his wife, a psychologist named Mary Lee Smith, systematically hunted down every scrap of research they could find that compared the effect of psychotherapy with a control group or another therapy –
ending
up with more than
370 studies.
Then Glass and Smith worked out a way to extract and combine the measurements of psychotherapy’s effect in each study. Even though one study measured the effects of the therapy on anxiety and another its impact on blood pressure, they devised a way to convert these results into one standard measure known as ‘effect size’ and then average them across all the studies. (This is somewhat like converting various currencies into dollars, so they can all be combined.) Glass called this new statistical method a
meta-analysis
– an analysis of analyses, just as metadata is data about data.
After toiling on this work for two years, Glass and Smith concluded that psychotherapy had a beneficial effect, and Glass presented the method at the San Francisco hotel where the education meeting was taking place. Debate continues to this day about whether
psychotherapy works
, when and for whom. (Eysenck
called
the work ‘an exercise in mega-silliness’.) But of the meta-analysis, the audience was thunderstruck. Here was a way to extract meaning from an apparent mess of different results.
The meta-analysis was not the only synthesis tool to emerge around this time. The work of Glass, Smith and other researchers also led to the
systematic review
, a rigorously structured method for gathering and evaluating evidence. In a systematic review, researchers scour databases worldwide of published and unpublished work for all studies that address a certain question. Then they whittle down a longlist of thousands of studies to the most relevant few, assess their reliability, extract the data, and combine the results. Many systematic reviews include a meta-analysis to pool results from the included studies and estimate the overall effect of a treatment or other intervention.
These tools to make sense of a body of evidence are one of the most important developments in science over the past few decades. They have the power to identify important conclusions that would never be possible from assessing each underlying study on its own. It’s the scientific equivalent of seeing the forest, not just the trees.
W
ithout fanfare, these approaches have changed modern life – and particularly so when it comes to health decisions. In medicine, this happened partly thanks to a morning stroll by Iain Chalmers, an early champion of evidence synthesis, along the Wolvercote Mill Stream in Oxford, UK, in
May 1991.
Chalmers had just finished a pioneering, decade-long project to synthesise all the evidence from clinical trials on treatments in pregnancy and childbirth. This work, which involved doing hundreds of systematic reviews, had shown that many standard medical practices – such as shaving women’s pubic hair during labour, and surgical episiotomies – were based on little evidence and some were harmful. The work from Chalmers and his colleagues helped change some of these practices.
Now Chalmers was thinking about expanding this work. Wouldn’t it be useful, he thought, to synthesise clinical trials in every area of medicine and healthcare? Then doctors and patients would know, based on evidence, effective ways of treating diabetes, cancer, heart disease and many other conditions. Surprisingly, many medical decisions at that time were based on conventional wisdom or the unsubstantiated opinions of senior doctors rather than on evidence from research.
Some years earlier, in 1979, a doctor called Archie Cochrane working in Cardiff, Wales, had challenged the medical community to do just this. ‘It is surely a great criticism of our profession that we have not organised a critical summary, by speciality or subspecialty, adapted periodically, of all relevant randomised controlled trials,’ Cochrane
wrote
. Chalmers was greatly influenced by Cochrane, and he was now ready to take on the task. He and his team started searching for clinical trials in online academic databases and scouring journals by hand in the library. The task was so big that they recruited volunteers to join the hunt, including elderly people’s groups and even the unlikely source of the Headington Bowls Club in Oxford. Soon, they had tens of thousands of trials.
Most people who have seen a Western doctor have unknowingly benefited from systematic reviews
But Chalmers knew that synthesising all the evidence on effective treatments in medicine was such a gargantuan task that it would take more people power than this. So, in 1993, he and a group of like-minded folk started the Cochrane Collaboration (now known as just Cochrane), dedicated to producing high-quality systematic reviews of evidence on the effectiveness of health treatments.
Within 10 years of starting, the Cochrane Collaboration had
published
around 2,000 systematic reviews and was helping to popularise this type of study. By 2010, researchers inside and outside the collaboration were
publishing
11 systematic reviews on healthcare
every day.
Today, the ‘Cochrane review’ is known for being one of the most rigorous and reliable syntheses of science, and systematic reviews are used to develop the clinical guidelines that doctors use to guide decisions. Most people who have seen a Western doctor have unknowingly benefited from systematic reviews. They have become an invisible bedrock of evidence on which medicine is based.
To understand what the world would look like without them, consider this notorious piece of medical advice that went unchallenged in the mid-20th century. In 1957, the paediatrician Benjamin Spock made a small change to the second edition of his parenting bestseller
The Common Sense Book of
Baby and Childcare
(1946), reissued several more times over the years. He said that parents should put babies to sleep on their fronts rather than on their backs. By the 1980s and ’90s, studies were clearly showing that this was one of the most lethal pieces of unsubstantiated advice in the history of child health, as it contributed to a steep rise in sudden infant death
syndrome (SIDS).
What makes this story even more tragic is that the link between sleeping position and SIDS might have been detected earlier – if researchers had synthesised the evidence from studies rather than looking at them one at a time. When researchers did this with a systematic review and meta-analysis
published
in 2005, they discovered that the link between front-sleeping and SIDS was clear in 1970, some
20 years
before most parents were warned about the risks. They calculated that at least 50,000 infant deaths in Europe, Australasia and the United States might have been prevented if evidence had been synthesised and acted on earlier.
M
edicine is not the only field to have built up a repository of evidence syntheses. Other disciplines – ranging from conservation to education – have mimicked the medical model.
The nonprofit International Initiative for Impact Evaluation
holds
more than 1,700 systematic reviews on what works in international development, such as improving sanitation and tackling malnourishment. CEEDER, a major environmental database,
hosts
more than 2,100 evidence reviews on what works to save species and protect the planet.
The Education Endowment Foundation, a nonprofit organisation in London dedicated to boosting learning outcomes for disadvantaged children, has
constructed
a widely admired ‘teaching and learning toolkit’ using systematic reviews of studies that have tested education approaches. Nearly
70 per cent
of school leaders in England now use it to make spending decisions, and it has been adapted for use in more than 30 countries and translated into nine languages.
The IPCC’s syntheses have helped to inspire global action by showing that humans are causing climate change
One of the most effective strategies, it
shows
, is teaching students metacognition – the ability to learn how to learn. (For instance, a child who learns that writing out the steps helped her solve a maths problem is practising metacognition.) Introducing metacognition provides children with an impressive eight months of additional learning progress, on average, over a year, according to the evidence review. Repeating a school year, by
contrast
, sets back students’ learning by about two months.
More recently, climate scientists have been taking a leaf out of medicine’s book. They are used to synthesising evidence: the six massive scientific assessments that the Intergovernmental Panel on Climate Change (IPCC) has published since 1988 are probably the biggest and most influential evidence syntheses ever done – helping inspire global action by showing that humans are causing climate change. But the IPCC reports have so far said less about solutions, such as which climate policies could help the most.
A group called What Works Climate Solutions, which started in 2024, is now striving to build a Cochrane-like bank of evidence syntheses that show the most effective ways to cut emissions or adapt to global warming. This evidence bank will feed into the next scientific assessment of the IPCC, due to be
published
by late 2029, and should help governments choose which policies to adopt.
The soaring popularity of systematic reviews is visible in academic databases. The number published in medicine alone rose from around 1,400 in the year 2000 to more than 29,000 in 2019 – about
80 per day
. A set of guidelines for performing them accurately, called PRISMA, has
become
one of the most cited papers of the
21st century
– with as many as 138,000 citations, by one count.
S
ystematic reviews have their critics. One is that they conventionally emphasise quantitative studies – those with numerical data – and ignore qualitative research, such as that based on interviews and observations. There is a lively debate going on among evidence-synthesis experts about how best to integrate different streams of information.
In some fields, such as conservation, scientists have been reproached for privileging Western science over other forms of knowledge, such as the lived experience of local communities and Indigenous groups. Some Indigenous peoples in Canada, for example, have extensive knowledge of the behaviours of the migratory fish that they hunt. These types of evidence don’t fit in conventional reviews that just crunch numbers.
But there are ways to combine them. In 2021, Parks Canada and a group called Foundations of Success undertook an extraordinary
analysis
of evidence to decide whether to pursue an expensive and risky captive breeding programme for endangered caribou in Jasper National Park. As well as collecting a range of evidence from research, they asked people from Indigenous groups, zoos, and reindeer experts from Finland, to review the evidence and reach agreement about whether to proceed. (They did – and some adorable caribou calves were born
in 2025.)
Over the past few decades, researchers have developed a suite of methods for combining evidence of different types. Some pool only
qualitative
studies, while other ‘mixed-methods’ syntheses combine qualitative and quantitative work. Some researchers who supply evidence to busy government policymakers favour the
rapid review
, which involves getting together the best evidence synthesis you can, to answer their question in the time available – which could be the next day. They prefer a good synthesis delivered on time rather than a perfect one that arrives too late.
Studies that claim to provide an overview of evidence convey authority, so scientists have a responsibility to get them right
Elsewhere, scientists have been working on another important innovation called
living reviews
. These are updated frequently and continuously as research is published rather than, as often happens, a static systematic review that quickly becomes out of date. Living reviews can be labour intensive and so are most worthwhile for important questions on which new studies are being published quickly. During the
COVID-19
pandemic, a handful of groups soon rallied and developed living reviews to synthesise rapidly emerging studies on COVID therapies and vaccines.
The methods used for synthesising science have become so complex over the past
20 years
that evidence synthesis has become a profession in its own right. There are entire journals dedicated to the craft, as well as academic centres, professorships and prizes. As a science journalist, I’ve interviewed evidence geeks who can talk for hours about the finer details of systematic reviews.
And sometimes tensions bubble up about the best way to synthesise a body of research. A purist argues that a gold-standard systematic review is the most rigorous way of doing it and that other methods risk missing studies or including poor-quality work and thus reaching the wrong conclusion. A pragmatist argues that a systematic review takes too long, and that often a quick-and-dirty aluminium-standard synthesis
will do.
But, really, this is a storm in a teacup. Different methods for evidence synthesis suit different situations. What’s important is to be transparent about the methods used and any weaknesses they have. Studies that claim to provide a rigorous overview of evidence convey authority, so scientists have a heavy responsibility to get them right.
I
n a world where a news story or a social media post can spread in hours, one of the biggest problems in systematic reviewing is the time it takes. An
analysis
of more than 500 Cochrane reviews found that a review typically took nearly three years – and one took a glacial eight. That’s because doing a systematic review involves working through around 25 steps – carefully finding, filtering, and combining studies – of which many are even done in duplicate to avoid errors. This exacting process explains why many researchers view with horror the prospect of doing a systematic review.
Researchers have long used computers to help, but the rapid advances in generative AI are fuelling new excitement about accelerating and automating the task. Could AI help make sense of the overwhelmingly large scientific literature?
It has changed the way I think about research as a science journalist and in everyday life
Start-ups are racing to develop AI-powered systems that can find, sort and summarise research publications, but most have not yet been well tested and shown reliably to produce high-quality systematic reviews. Even so, it seems inevitable that they will improve and that before too long most of the grunt work involved in compiling evidence will routinely be done by machines, freeing up humans to check and interpret the results. That could mean that anyone would be able to assemble a reliable evidence synthesis tailored to their question, virtually at the push of a button – an appealing goal.
But it’s still some way off – and concerns about AI and evidence synthesis also abound. One is that AI is leading to a flood of inaccurate scientific reviews because researchers are using these tools to race through poor-quality efforts. Another is that people will bypass rigorous evidence syntheses entirely: because why bother, when you can just ask ChatGPT what the evidence says? (This is currently dangerous as AI tools can hallucinate and give an inaccurate picture, and most do not have access to the full corpus of the world’s research.)
One lesson from all this is to be aware that single studies can be outliers – something I’ve learned myself over the past few years. I’ve spoken with hundreds of researchers worldwide about the importance of synthesising evidence, and it has changed the way I think about research as a science journalist and in everyday life. Now, when I speak to scientists about their latest study, I try to ask how it fits with the wider body of research. When seeking evidence on a personal issue, such as a medical diagnosis, I tend to search for systematic reviews and other syntheses (while being aware that not all reviews are good quality).
We should resist the allure and glamour of new studies. It’s true that new research is important, because it leads to innovation and fresh ways to protect the planet and human lives. But if we truly want to understand the world and make it a better place, we also need to do the less glamorous, unappreciated work of collating existing knowledge – and figuring out what it all means.
On October 14 2026 we will ship curl 8.23.0. The next iteration in the never-ending series of version bumps from the
curl project
.
We always think of the next release as the best version we ever did – and this time is no exception. Decades of collected experiences and meticulous polishing has lead us to this.
Earlier than planned
We decided to shorten the release cycle this time, so that we can release 8.23.0 a few weeks earlier than what we originally planned. We took this decision after we received one particular vulnerability report that highlighted a rather significant flaw.
We will ship a new version with this problem removed, together with twenty-one other albeit less serious security vulnerabilities addressed.
Severity HIGH
In the curl project we only assign one of the four different severity levels on all CVEs we report (LOW, MEDIUM, HIGH or CRITICAL), as we basically
don’t believe in CVSS scoring
. We have only published two CVEs with severity HIGH since 2021, the most recent one being
CVE-2023-38545
; that could lead to a heap buffer overflow.
Now we are about to release another one: CVE-2026-92392.
All info will be revealed next week
All details about CVE-2026-92392 will become public in the European morning of October 14, 2026 in synchronization of the release of curl 8.23.0 which of course will have this problem fixed.
For the safety and security of curl users everywhere (and frankly, all the infrastructure that uses curl), no details of this flaw will be made public before this date.
We will alert the distros@openwall mailing list and paying curl support customers about this problem (and the associated fix) ahead of time.
I will follow-up with a separate blog post after October 14 to describe this flaw in detail. How it can be triggered, why it isn’t quite the end of the world and what we do in curl to fix this and similar classes of problems.
curl, open source and networking
PoeLLM malware infects exposed AI servers in cryptomining attacks
Bleeping Computer
www.bleepingcomputer.com
2026-10-07 11:04:08
A cryptomining campaign targeting exposed AI services is using PoeLLM malware to turn compromised servers into scanners and exploit launchpads. [...]...
A cryptomining campaign targeting exposed AI services is using PoeLLM malware to turn compromised servers into scanners and exploit launchpads.
The malware features an uncommon method to retrieve command-and-control (C2) addresses by extracting keywords in a poem hosted on GitHub.
Researchers at Lumen's Black Lotus Labs (BLL) tracking the botnet malware say it has compromised more than 2,100 servers, with peak activity reaching as many as 800 infected systems active on a single day.
PoeLLM has been active since at least April, but its activity has increased significantly since then, with at least 11 C2 servers spun up to date.
According to the research, the operation targeted systems across the United States and Western Europe.
Victims of PoeLLM
Source: Black Lotus Labs
Many of those victims run exposed AI tools such as LiteLLM and Ollama, the Gotenberg PDF converter, and the Gitea development toolkit, while signs of Ivanti Sentry targeting were also uncovered.
BLL notes that AI/LLM implementations are attractive targets for threat actors because they are often poorly configured, exposed online, and typically run on powerful GPU clusters that are suitable for cryptomining.
Poetry and malware
In a report today, BLL researchers say that PoeLLM, an ELF file named libgcrypt, retrieves four words or phrases from a poem titled “On the Nature of Connection” in a ‘dash.css’ file hosted in a GitHub repository that appears to fork Node.js.
It then maps these words to numbers using a hard-coded dictionary, generating a IPv4 address corresponding to the C2.
The poem used for C2 address construction
Source: Black Lotus Labs
To change the C2 address, the operator changes the poem. Until now, they have modified the poem 11 times, but researchers suspect that there may be at least another update.
The malware incorporates remote-shell functionality, XMRig and Iron cryptocurrency miners, HTTP/S scanning, and exploit deployment capabilities.
BLL researchers found that victims communicate with a Russian crypto-mining service called Kryptex.
Once a server is compromised, it becomes a springboard to spread the malware further, using scanning on ports 3000 and 4000, associated with Gotenberg and LiteLLM, and attempting to exploit CVE-2026-42271.
The CVE-2026-42271 vulnerability impacts LiteLLM’s MCP server test endpoints. It was originally disclosed as requiring authentication and received a high-severity score.
Horizon.ai
researchers confirmed
that it could be chained with another security issue, CVE-2026-48710, for unauthenticated remote code execution (RCE).
PoeLLM attack overview
Source: Black Lotus Labs
By analyzing the infrastructure, BLL found that several C2 servers featured vulnerable router administration interfaces, suggesting that the attacker reused compromised routers in the attacks.
The researchers could not make a confident attribution but assess with moderate confidence that the operator is Italian, based on comments in the malware and an Italy-based server hosting the administrative interface.
To protect against PoeLLM attacks, system administrators should apply the latest security updates, reduce public internet exposure for critical assets, and restrict external access only to trusted IPs.
Administrators are recommended to inspect network monitoring logs and look for connections to the indicators of compromise (IoCs) shared by Black Lotus Labs.
Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.
Anti-Patterns in Software Blogging
Some excellent writing advice from Michael Lynch. Michael warns against "meandering intros", misjudging your reader's existing knowledge, assuming they'll read your previous posts, and excessive formality.
He also warns against overreliance on links as an excuse no...
Anti-Patterns in Software Blogging
(
via
) Some excellent writing advice from Michael Lynch. Michael warns against "meandering intros", misjudging your reader's existing knowledge, assuming they'll read your previous posts, and excessive formality.
He also warns against
overreliance on links
as an excuse not to explain terminology. This one hurt! I do this all the time, but I have a nagging suspicion that almost nobody ever clicks on them.
This point about using your own voice is crucial:
Beginner software bloggers suffer from a mass delusion that you have to write in a stiff, overly formal way for people to take you seriously [...]
Just write the way you talk.
With so many developers delegating their writing to AI, software blogging is becoming bland and homogenous. Readers are hungry for writing with personality.
'Undue emotional weight'
403 Media
www.404media.co
2026-10-07 10:46:15
AI-generated political slop, a farewell to a good bird, and more....
Howdy, it's hump day and we're rolling along. Today we've got an update on a very strange case involving a man whose image was resurrected from the dead via AI and made to appear in front of a judge to announced "forgiveness" of his murderer. 404 Media contributor Matthew Gault had a long conversation with the man's sister, who wrote the words and generated the video, and she said that swaying the judge was the entire point.
A man convicted of manslaughter in Arizona will be resentenced because an AI-generated video of his victim speaking from beyond the grave was ruled to have carried “undue emotional weight” as an impact statement. An appellate court in Arizona ruled that the manslaughter charge will remain, but the judge must reconsider the length of the man’s prison term because of AI.
In 2021, Gabriel Horcasitas shot and killed Christopher Pelkey during a road rage incident. A jury found him guilty of manslaughter. During Horcasitas’ sentencing hearing, Pelkey’s sister Stacey Wales
played an AI-generated video
of her brother as part of her victim’s impact statement. In the video, which she scripted with her own words that she made an AI generated likeness of Pelkey say, the AI avatar of Pelkey forgives his killer and said they “could have been friends” in another life.
“
I loved that video
,” the judge said while handing down the maximum sentence of 10 and a half years.
Wales told 404 Media that her intent with the video was to sway the judge. “Isn’t that what you want human connection to be?” she said. “Do you ever come out of a movie and say: ‘Well, that was too emotional. That was too powerful. Human connection is powerful, and trying to convey that to somebody that's never known another person, this person on the earth, to convey their lifetime of 37 years down to a two-hour sentencing event is hard to do, and I think that victims and victims' families should be afforded every opportunity.”
A flock of new laws incoming
.
Several recently introduced pieces of legislation
would, if passed, dramatically curb federal agencies’ access to Flock cameras, with the lawmakers referencing reporting by 404 Media in their explanations for why the legislation is needed. “At a time of growing concern about the unchecked power of artificial intelligence, Flock is eviscerating the very notion of privacy by installing tens of thousands of cameras in communities across America without their consent,” Bernie Sanders said.
This post is for paid members only
Become a paid member for unlimited ad-free access to articles, bonus podcast content, and more.
Mumsnet denies using AI for content after prompt appears on message board
Guardian
www.theguardian.com
2026-10-07 10:31:04
Users reject site’s explanations after detailed instructions for writing response to post on popular forum was published Mumsnet has denied using AI to produce content after detailed instructions to help an AI model write one of its popular “am I being unreasonable” (AIBU) posts appeared on a Mumsne...
Mumsnet has denied using AI to produce content after detailed instructions to help an AI model write one of its popular “am I being unreasonable” (AIBU) posts appeared on a
Mumsnet
message board.
Users of the long-established forum, which is known for its frank discussions of family tensions, were vexed when the AI prompt appeared in response to a user seeking advice about a family dilemma involving a neurodivergent uncle and a sick child.
The user, called AmIAComplete, had asked for “a sanity check” from other users after an awkward social interaction with their brother’s family but was met with a prompt titled AIBU Response Schema v.1.0.
It instructed an AI to write an “interpersonal reasonableness” post and “give a clear, forum-style ruling plus brief actionable advice”. It included definitions of common Mumsnet abbreviations including AIBU, ND (neurodivergent) and DD (dear daughter).
The prompt said: “do not diagnose anyone; do not demean ND people; do not imply a sick child owes social performance” and said “prioritize compassion, boundaries, and de-escalation”.
Site users called the emergence of the prompt “really creepy”, “crazy” and “so disappointing” and it fed speculation about Mumsnet’s use of AI, with one user saying: “We might as well just go and do something else then.”
Mumsnet was founded in 2000. It faced a 19% drop in its monthly unique user numbers in 2025 to 6.1 million and the number of posts fell by 600,000. It said in its latest annual report that it was “embedding AI tools to improve productivity and enhance the user experience”.
Recent research,
published in Nature
, has shown that AI assistance in social media risks creating “semantic garbage” and is seen by users as less informative, “robotic” and “generic”.
Soon after the prompt emerged, Mumsnet denied that it used AI to write threads and said it had come from a tool to suggest thread titles which uses an OpenAI model. But users refused to accept this explanation and some copied the prompt into an AI and found it generated a full post.
“We are not idiots,” said one user, MrsBennetsPoorNervesAreBack. “Please can you give us a transparent explanation?”
Mumsnet then said the prompt did not come from them and suggested it had somehow transferred from AmIAComplete’s computer. Again users were unconvinced and speculated the prompt must have leaked from Mumsnet’s internal system.
Mumsnet’s founder,
Justine Roberts
, joined the thread and confirmed the prompt had come from its system, which sends users’ drafts to OpenAI, which is meant to suggest a title, although this time it sent a prompt for a whole reply. She said it “looks like an error on their side” and said: “We don’t use AI to write threads or replies on Mumsnet.”
This explanation also failed to appease users, who said they were being “fobbed off”.
“Why would OpenAI compel itself to do/try to do something so specific if it’s never been asked to do that before from your side?” said one. “It’s like it’s appeared in the wrong place. Rather than just sprung out of thin air.”
“I just don’t know anymore what’s real and what isn’t sometimes!” they said.
Roberts told the Guardian: “I understand why it looked like it had ‘appeared in the wrong place’. But nothing in our systems asks for, stores or uses anything like it.”
She said Mumsnet had never tested a tool that drafted posts, but it used AI for coding, advertising operations, conversational analysis and flagging potentially problematic posts to moderators.
“Mumsnet works because the people on it are real,” Roberts said. “Members trust that the advice they get comes from other parents, not a machine … This episode was a useful reminder of two things. First, these systems are less predictable than the people selling them like to suggest. Second, users are understandably highly concerned that Mumsnet remains written by humans.”
Steinar H. Gunderson: Decompilation patterns, part 4: range comparison
PlanetDebian
blog.sesse.net
2026-10-07 10:30:31
Today, we're looking at a rewrite that rarely differs for matching,
but can mean quite a lot for understanding the code. Say that you
have a signed variable and the compiler has suddenly decided to make
an unsigned comparison:
if ((u32) (x - 4) < 2) {
What's happening is that it uses a con...
Today, we're looking at a rewrite that rarely differs for matching,
but can mean quite a lot for understanding the code. Say that you
have a signed variable and the compiler has suddenly decided to make
an unsigned comparison:
if ((u32) (x - 4) < 2) {
What's happening is that it uses a controlled overflow to do two
signed comparisons using one unsigned comparison. How does this work?
Since the comparison is unsigned, this is trivially equivalent
to the first expression:
if ((u32) (x - 4) >= 0 && (u32) (x - 4) < 2) {
and since negative numbers compare larger than 6 in unsigned comparisons,
nothing prevents us from adding 4 on each side of the inequality signs:
The very first systems programming language I ever learned was Rust. This is uncommon compared to many other programmers; you're more likely to find someone who learned C or C++ first before coming to Rust. As such, there are plenty of "Rust for C Programmers" articles on the internet, but little to no "C for Rust Programmers" articles out there.
Well, I'm about to change that! I've been learning C and C++ recently, and
holy cow those languages are quirky
. This blog post is a collection of eyebrow-raising details I learned while teaching myself C. (No C++ today, I'm not ready to dig into that can of worms.) This is not a substitute for a proper C tutorial, you'll need to Google one of those yourself. Rather, it's a list of things to keep in mind when working in the language.
With the premise set, let's draw the curtain and see what C has to offer!
Boolean Is Not* Built-In
The original version of C did not have a primitive type for booleans, programs instead used the integers 0 and 1. This was changed in
C99
, which added booleans in the optional
<stdbool.h>
header
[1]
:
Even still,
true
and
false
are not literals or keywords like in other languages. Instead, they are definitions that expand to 1 and 0:
#define true 1#define false 0
This was changed again in
C23
[2]
, so now booleans are true language primitives, but if you're compiling for earlier versions you'll need to include
<stdbool.h>
.
Null-Terminated Strings
A Rust
&str
is 16 bytes: 8 bytes for the memory address and 8 bytes for the string length. This is because
str
is a dynamically sized type, and uses
pointer metadata
to track the length of the string.
// While a normal reference uses only 8 bytes...assert_eq!(std::mem::size_of::<&u8>(), 8);// ...strings use 16 bytes.assert_eq!(std::mem::size_of::<&str>(), 16);
This approach makes fetching the string length extremely efficient, but requires more memory per
&str
reference. C uses a different approach: it doesn't store the size of its strings separately, but instead terminates every single string with a null byte (
\0
). This is an intentional trade off that results in a few things:
C programs don't need to keep track of an extra
size
variable alongside the string
Every single string needs room at the end for a null byte
This means empty strings
""
still take up one byte of memory!
Null bytes cannot easily be used in the middle of a string without messing up
<string.h>
's functions
Forgetting the null terminator can result in out-of-bound reads
In practice, this requires you to remember to allocate extra space for the null terminator and insert it at the end of strings. For example, here is a program that reverses a string in C:
char* reverse(char* forward) { // Calculate the string's length, excluding the null terminator. unsigned long len = strlen(forward); // Allocate enough room for the string and its null terminator. char* reversed = malloc(len + 1); for (int i = 0; i < len; i++) { reversed[i] = forward[len - 1 - i]; } // Add the null terminator at the end. reversed[len] = '\0'; return reversed;}
Note lines 5 and 12, which take special measure to account for the null terminator. For reference, the corresponding Rust function
[4]
doesn't need to do so:
fn reverse(forward: &[u8]) -> Box<[u8]> { let len = forward.len(); let mut reversed = Box::<[u8]>::new_uninit_slice(len); for i in 0..len { reversed[i].write(forward[len - 1 - i]); } unsafe { reversed.assume_init() }}
Target-Dependent Integer Widths
C's integer types are not guaranteed to use an exact number of bits, instead the width varies depending on the target platform.
long
being 64 bits on Unix and 32 bits on Windows particularly irks me. I recommend following the advice that
a friend of mine
gave me a few years ago: if you care about cross-platform compatibility, only use the fixed width integers provided by
<stdint.h>
:
I really love Rust's error handling.
Result
forces you to address errors, and sum types (
enum
s) and
match
statements make doing so really easy!
C's error handling story, in comparison, is straight up tragic. It seems to boil down to functions returning a "magic integer" like -1 or a null pointer to signal there was an error. You can get a little bit more information by reading
errno
, a thread-local integer that can be used to check for specific kinds of errors, but in terms of getting actual error messages and stack traces it is much harder.
When writing C, one of the biggest things you'll notice is that the language will never force you to handle errors. It's up to you to remember that functions can fail. For example, here's a snippet of code from
Null-Terminated Strings
:
// Allocate enough room for the string and its null terminator.char* reversed = malloc(len + 1);for (int i = 0; i < len; i++) { reversed[i] = forward[len - 1 - i];}
To new programmers, it's not immediately obvious that
malloc()
can fail and return a null pointer. If your machine runs out of memory,
reversed[i]
will cause a segfault upon being accessed. In order to avoid unhelpful segfaults, the program should check for a null pointer and gracefully exit if one is found:
Doing so provides a much better experience than a segfault, or worse, other unintentional behavior:
$ ./mainError: Cannot allocate memory
Of course, remembering to check every single allocated pointer isn't an amazing developer experience. An approach I played around with was recreating Rust's
Result
in C using tagged unions:
struct MallocResult { enum Tag { OK, ERROR } tag; union Value { void* ptr; char* error_message; } value;};
It's a complete mess to use, however, and there's
still
nothing stopping you from immediately accessing
result.value.ptr
without handling the error first. The lack of visibility modifiers like
private
and
public
, which could be used to prevent this, appear to be an intentional design decision. C places full trust in the programmer to
do things right™
and gives few facilities for contracts or safe abstractions.
I personally disagree with this approach. I'm not some mastermind programming wizard who never makes mistakes. I'd much rather encode program requirements into the type system so that the compiler can check them for me!
[6]
That way, I can be reasonably certain that if the code compiles it was written correctly. But I digress.
Field Access Syntax
C has two different field access operators:
.
for values and
->
for pointers.
struct Foo { int field;};struct Foo value = { 103 };struct Foo* ptr = &value;// Accessing the field directly.printf("%d\n", value.field);// Accessing the field through a pointer.printf("%d\n", ptr->field);
This caught me off guard at first, as Rust uses
.
for everything.
Arrays have some weird nuance to them. Sometimes they are plain values with their length accessible using
sizeof()
, other times they are pointers with an unknown size. In general, this is determined by whether you are handling the array within the function it was defined or not.
To showcase what I mean, here's a very simple C program that prints the size of two arrays:
Why? Because any array passed through a function is
implicitly converted
to a pointer to the first element. From the C compiler's perspective, the above function actually had the following type signature:
This is why it confusingly looked like each array was 8 bytes long. The arrays themselves are 3 bytes, but the pointers take up 8 bytes of memory! Handily, Clang raises a warning when you run
sizeof()
on the pointer form of an array, making it easier to catch this mistake:
$ clang main.c06-arrays/main.c:5:59: warning: sizeof on array function parameter will return size of 'char *' instead of 'char[3]' [-Wsizeof-array-argument] 5 | printf("Declared size (function): %ld bytes\n", sizeof(declared_size)); | ^06-arrays/main.c:4:20: note: declared here 4 | void function(char declared_size[3], char inferred_size[]) { | ^06-arrays/main.c:6:59: warning: sizeof on array function parameter will return size of 'char *' instead of 'char[]' [-Wsizeof-array-argument] 6 | printf("Inferred size (function): %ld bytes\n", sizeof(inferred_size)); | ^06-arrays/main.c:4:43: note: declared here 4 | void function(char declared_size[3], char inferred_size[]) { | ^2 warnings generated.
Oh Yeah References Don't Exist
I'd be remiss if I didn't mention it once. C has no borrow checking, and no concept of references. It's all
raw pointers
. This means you can do fun pointer arithmetic like so:
int array[] = { 0, 1, 2, 3, 4, 5 };// Woah, iterating over pointers instead of indexes!for (int* x = &array[0]; x < &array[6]; x++) { *x = 5 - *x;}
However, I'm not sure how good an idea that is. 😅
Either way, be careful when you mess with pointers. Making mistakes with memory results in buffer overflows and out-of-bounds writes, which provide significant threats to the security of your code.
Vulnerability category distribution for CVEs registered between Nov. 2021 and Jan. 2022,
via
Conclusion
I hope you enjoyed this article! I've honestly had a lot of fun learning C. And while I doubt I'll reach for it for personal projects, it's absolutely a crucial language to know as a systems programmer. All the examples in this blog are
available on Github
, if you'd like to mess with them yourself!
Until next time,
- BD103 :)
Technically you can use
_Bool
without the header, but you still need it for the
true
and
false
definitions.
↩
The corresponding Rust function is not idiomatic, and can only correctly handle ASCII text. If I were writing a production version of this function, it would be a single line of code:
forward.chars().rev().collect::<String>()
.
↩
usize
and
isize
don't perfectly map semantically to
size_t
and
ptrdiff_t
. I don't fully understand the difference, however, so I recommend doing your own research before using these types.
↩
↩2
This blog post on Fearless SIMD
provides a great example of using Rust's type system to guarantee the correctness of code. I highly recommend giving it a read!
↩
My favorite line of code from that post is
100->a = 0;
, it's so cursed!
↩
Device detection and occupancy monitoring for Airbnb hosts
Crowd Detect monitors occupancy in real time, alerting you when unusually high numbers of people are present so you can uphold guest standards and prevent unwanted parties.
Trusted by the best in hospitality
Stay one step ahead of crowding and parties
Know your spaces are being used as intended
Crowd Detect gives your team early awareness when occupancy begins to rise, so you can support safe, comfortable stays without disruption
Stay aligned with local rules and internal standards
Every rental has house rules, and Minut helps you uphold occupancy limits with less effort and greater consistency.
Get peace of mind, without compromise
Get the insights you need, anytime and anywhere you operate, to stay ahead of issues and protect every property.
How does it work?
Technology that works quietly in the background
Crowd Detect uses passive signals from nearby iOS devices to estimate occupancy.
You set the limit, and if the count goes above it, Minut sends an alert, so your team knows when to check in or take action.
It’s fast, frictionless, and designed for peace of mind.
“Before Minut, we had no way of knowing when parties were happening. Now, we rarely get notifications about parties. In the past year, it’s only happened once, and thanks to Minut, we were notified before it got out of hand.”
“Minut has helped us strengthen our relationships with other properties in the community. We’ve reduced complaints and had fewer problems with our properties overall.”
“We were getting some complaints saying there were parties, but when we looked at the Minut data, we could confidently respond and say there wasn’t excessive noise. It helped us navigate difficult conversations with neighbors and community groups."
Best Books/Courses/Channels to Leapfrog on AI/ML Material
Lobsters
lobste.rs
2026-10-07 10:12:02
Assuming I have been on hibernation since 2018-19 time frame, which materials - books, MOOC courses and YT channels would be recommended for me to get started with understanding all the current progress on AI/ML. Ideally a resource like TeachYourselfCS that can give structure and roadmap to navigate...
Assuming I have been on hibernation since 2018-19 time frame, which materials - books, MOOC courses and YT channels would be recommended for me to get started with understanding all the current progress on AI/ML. Ideally a resource like
TeachYourselfCS
that can give structure and roadmap to navigate all the progress/knowledge on the domain would really help.
Do we have to take different approaches to understand LLMs vs other type of ML models which have been in large scale deployment/usage from pre-2022/21 time frame?
Thanks a lot in advance.
AI-assisted proof of optimal packing for 11 squares
The complete optimality proof passed verification with native numerical certificates.
The completed
EvolvingPrograms verification run
accepted all
7,920 local Lean modules
, and its final audit reports
zero
admissions
. This repository imports those exact proof sources and pinned
build configuration from commit
1bf942a7af1ea330e95489d8997deebd4227ca71
.
See the
verification report
for evidence and scope.
Selected expensive, exact numerical certificate checks use
native_decide
.
Geometry, checker soundness, and proof assembly retain ordinary Lean proofs.
Consequently the final theorem trusts
Lean's kernel and native compiler
;
this is not a kernel-only verification claim. The approved numerical declarations
and their exact source hashes are recorded in
verification/native-certificates.json
.
The optimal side length is
[
T = \frac{6u+4}{1+2u-u^2},
]
where
u
is the unique root in
(9/25,37/100)
of
[
5u^8-10u^7-2u^6+14u^5+12u^4-6u^3+2u^2+2u-1=0.
]
The construction attains approximately
3.8770835900228141773
. The model allows
arbitrary orientations, legal boundary contact, and disjoint open interiors.
The public statements in
ElevenSquare/Optimality.lean
and the complete T03
source tree are unchanged from this repository's previous main branch.
Entry points
File
Purpose
ElevenSquare/Foundations.lean
Geometry, exact endpoint, attaining construction, closed-cell cover, and finite case reduction.
ElevenSquare/Pending/
Original public interfaces, now discharged by the integrated proof. The directory name is historical.
ElevenSquare/Interop/Wand125/
Connections to the incorporated upstream certificate results.
ElevenSquare/Tasks/
Geometric arguments, checkers, certificate data, and local analytic proofs.
Sqpack/
Incorporated certificate checkers, generated proofs, and simplifications.
ElevenSquare/Optimality.lean
Unconditional optimality and side-length lower-bound theorems.
ElevenSquare/Verification.lean
Axiom queries for the public proof targets.
Reproduce verification
The project pins
Lean 4.34.1
and Mathlib revision
d13f23b723b8a846827a245b89c10fc7d3f11612
. Keep
lake-manifest.json
unchanged.
On Linux with Python 3, Git, curl, and tar:
On macOS, first install the
elan
launcher, then use the same command. The
bootstrap can prepare the pinned toolchain and dependency cache when
elan
is
already installed. Choose a worker count appropriate to the machine; modules
are compiled serially. Existing valid receipts are reusable. Add
--fresh
to
force a complete replay; Ctrl-C stops the runner cleanly.
The command checks every local module and performs the final source, receipt,
dependency, and axiom audit. Require
OPTIMALITY_PROVED_WITH_NATIVE_CERTIFICATES
,
zero admissions, and
trust_model: lean_kernel_and_native_compiler
in the final
result. Reaching 100% of compiled modules alone is not sufficient.
A source-only check, without Lean, is:
python3 scripts/check_sources.py
The
manual workflow and Ubuntu instructions
also support
resumable verification. Pushes do not start a workflow. The successful source
run used EvolvingPrograms' larger runner; it does not establish a cold-build
runtime or a 2–3 hour macOS guarantee.
Do not run historical materialization commands or
verify.py --setup
on this
snapshot: they restore superseded generated sources. Build objects and logs
belong in ignored
.lake/
and
.verification/
directories.
Credits and provenance
We thank
EvolvingPrograms
,
@ctjlewis
, and every project contributor
for the
formalization and verification work. See
ACKNOWLEDGEMENTS.md
for individual and upstream credits,
PROVENANCE.md
for source
history, and
integrations/wand125
for retained notices.
Historical simplification notes and partial-audit records are preserved; their
old unfinished-status statements are superseded by the completed-run report.
The following guide will assist in setting up and using a Chromium based browser with the flags and policies present in this repo. It will cover 3 main sections: selecting a browser (covering forks, options for different OSes and alternatives to Chromium should they be viable), applying policies (only for Linux and Windows, maybe MacOS, but notes for other OSes where applicable), and persisting flags (only covering a few OSes since not all of them support proper flag persistence). The primary focus will be on
Linux
and
Windows
, but notes for Android and MacOS will be spread throughout where it makes sense.
Please note that while I intend for this to be as comprehensive as possible, there will be gaps. For example, I do not have a Mac so I am not capable of offering up-to-date info on methods or options for those systems.
For starters, make sure you are using a secure browser. This guide does very little if the underlying browser is insecure. You may reference the
browser selection page
for assistance on selecting a secure base browser.
Initial setup is just preferences, so see
the preferences page
. Everything else is covered by configuration files.
This guide consists mainly of configuration files for various platforms generated either the automated script or by hand. The preferred way is the automated script, it works cross-platform and should be simple enough to use (ask for help if needed). For a guide on using the script, visit the
automatic config generation page
. The process can also be done by hand, or if you are not satisfied with the results of the script, you can manually edit the output files. For a guide on how to edit these, see the
manual configuration editing page
.
Content blocking is usually done one of 3 ways, Extensions, Native/Internal, and DNS/Network. Some are blatantly better than others.
For starters, extensions are always bad. Especially MV2 extensions, like uBlock Origin. Since MV2 extensions can access any site as well as a great many features without permission. MV3 prevents this, but it isn't too much better since many extensions just ask for access to all sites anyway to work properly, but at least it offers the user control. With that in mind uBOL (uBlock Origin Lite) in
Basic
mode is pretty good, since it has no access to sites while still being able to deliver decent content-blocking. It also allows granting access to specific sites, such as Youtube (yes it works), for better filtering if needed. Other extensions and uBOL global modes risk security and weaken site isolation.
Native/Interal can mean one of 2 (technically 3) things. One is using Chromium's internal subresource filter (as done by
Vanadium
and
Trivalent
), this is approx on-par with uBOL in
Basic
mode in terms of filtering capabilities. This is also the most secure since it is already built directly into Chromium so no extra capabilities, features, or code is added or enabled. The second option is to integrate a third-party filtering engine, this is done by Brave, Vivaldi, Opera, Cromite, and many other browsers. This can vary between a new engine, like
Brave
, or integrating an extension, like
Cromite
. Both have more attack surface, but extension integration is much worse.
DNS/Network is arguably the most secure but the least effective (since it can only filter by domain, and not paths, e.g. all of
google.com
and not just
google.com/tracking
) of any method. With most content blocking you have to add trust in multiple entities and add extra attack surface. With DNS filtering, you are placing your trust in something you already have to trust (DNS resolution). I would still suggest the usage of some DNS filtering in your browser, even if you have another content-blocking solution. It also has no performance impact and can resist some forms of censorship and tracking by encrypting not only DNS traffic but also the Client Hello (via
ECH
). Non-DNS network filtering has the same effectiveness with the added benefit of IP blocking, depending on the implementation. It should be noted that CNAME tracking can be fully mitigated through DNS filtering.
There is technically a sub-category of network filtering that is more comprehensive in its ability to filter, but it is a massive security risk. HTTPS interception filtering is a method where your content blocker will intercept your encrypted web traffic using its own certificate, this forces you to trust said content blocker with certificate handling and website data. This is not recommended, and you are better off just using DNS with/or a native blocking solution or extension.
Last note on remotely updated filters for systems like Brave, Opera, and uBlock Origin (MV2). The main problem here is that filters can still modify requests, run regex (which can be exploited in the browser engine), use cosmetic filters (which has been used to exfil data from sites in the past), and execute JavaScript via scriptlets. While scriptlets themselves aren't risky, even when limiting execution capabilities it is still arbitrary execution and therefore has massive risk. These filters are themselves arbitrary and unsigned, meaning you are OTA downloading random files that are an exploit away from reading the contents on all sites or worse. At least with MV3 extensions the filters have to bundled, so they are effectively signed along with the rest of the extension, so much better than most integrated engines.
In case you are curious, this is my personal setup. The main purpose is to demonstrate the usage of this guide.
Lets start with OSs. I have an Android and a few Linux machines. They are GrapheneOS, Fedora Workstation, and a few laptops with secureblue. Because of this, your setup may vary.
On GrapheneOS, I use Vanadium. It is without a doubt the best option on Android, but due to a lack of availability outside of GrapheneOS, it is difficult to recommend. Therefore, the next best option I would use is Chrome. Yes, Chrome with some settings changed and some flags altered in
chrome://flags
. Is this ideal? Not really, but it's the next best thing below Vanadium. For watching Youtube without ads, I use NewPipe, so adblocking isn't a big enough deal for my browsing to justify selecting a browser based around it.
Not-so-subtle suggestion against Brave.
On Fedora, I use Trivalent, secureblue's default browser. It is sort of a port of Vanadium to desktop Linux, as it comes with a lot of neat defaults and hardening. For RPM based distros, it is definitely the best option. Due to Trivalent's defaults, it requires no usage of this guide or its configs. Otherwise, say on a Debian-based distro, I would use Chrome with the application of this guide. It's the closest you can get to Trivalent/Vanadium. The same is true for secureblue, Trivalent is best used on secureblue due to the SELinux confinement offered specifically for Trivalent.
Ransomware has a new target. Is your backup ready?
Bleeping Computer
www.bleepingcomputer.com
2026-10-07 10:01:11
Ransomware groups are increasingly targeting backup infrastructure to eliminate recovery options and increase pressure on victims to pay. Kaseya explains why organizations need isolated, immutable, and regularly tested backups that attackers cannot easily reach. [...]...
Ransomware is far less frightening when you know you can recover.
That’s why threat actors are turning their attention to encrypting backups. They wipe recovery points and disrupt the infrastructure needed to restore the encrypted data. Once that recovery path disappears, the pressure to pay rises sharply.
For IT leaders, the lesson is uncomfortable but simple. A backup is only a safety net if the attacker cannot reach it.
What attacks on backup infrastructure look like
Ransomware groups have found several ways to neutralize backups, and the methods keep getting more deliberate.
ALPHV/BlackCat ransomware group
In February 2024, the ALPHV/BlackCat ransomware group encrypted Change Healthcare’s systems after breaking through a remote access portal with no multi-factor authentication. The backups were not isolated or robust enough to restore operations quickly.
The BlackMatter group made backup destruction their standard operating procedure. When they hit NEW Cooperative, an Iowa-based farm services provider, and Crystal Valley, a Minnesota farm cooperative, in 2021, they used compromised admin credentials to locate every backup data store and appliance on the network — then wiped or reformatted them before encrypting everything else. Since the backups were on the same network, they were easy targets.
Gunra ransomware, documented in a joint CISA and FBI advisory in August 2026, pushed that logic further. In one confirmed case,
attackers deleted backup and archived data
at both the organization's primary data center and its disaster recovery site.
A single set of stolen credentials was all it took to reach both. Having two copies in two locations meant nothing, because both locations trusted the same key.
Protecting the way back
The attacks above point to a simple shift in how backup strategy needs to be thought about. Do not only ask how many copies you have or where they are stored. Ask what connects those copies, who can administer them and whether a compromised account could reach them all.
Your recovery strategy should assume attackers will try to destroy the way out. Separate critical backups from production, limit administrative access, and ensure there is no single compromised identity or pathway that can wipe every recovery copy.
Ransomware groups are now targeting backup infrastructure first — wiping recovery points before encrypting everything else.
Our report,
Building Security That Survives Human Error
, reveals what's stalling organizations from closing the gap and where leading IT teams are focusing first.
The cost difference between recovering with intact backups versus without them is the difference between a manageable disruption and a business-threatening event.
IBM’s 2026 research found that 41% of ransomware incidents also involved threats to
damage the victim’s brand reputation
, showing how quickly the impact can spread from disrupted systems to lost customer trust.
If attackers destroy the backups, they take away the leverage that lets organizations refuse the ransom.
Why this keeps happening
These attacks tend to expose the same weaknesses. The recovery environment may exist, but the security around it is often less mature than those protecting production systems.
Backups share the same network and credentials:
If the same administrator accounts can access both production systems and backups, an attacker who compromises one of those accounts may be able to reach both. True isolation requires deliberate architecture.
Backup software gets patched last:
The Akira airline attack exploited a vulnerability that had a patch available for over a year. Backup servers are frequently treated as appliances rather than software systems, which means they get updated last, or not at all. Attackers keep careful track of what remains unpatched.
Monitoring rarely extends to backup infrastructure:
Security teams concentrate detection and alerting on production environments. Backup servers often carry minimal logging, weaker access controls and slower response processes. They are watched less carefully, which makes them quieter to work in.
The deeper barrier is resources:
Knowing what
good backup security
looks like and having the budget, trained staff and operational bandwidth to implement it are separate problems. Many IT teams and MSPs operate with all three in short supply. They understand the exposure but simply cannot close it fast enough.
What closing the gap requires
It requires treating the recovery environment as critical infrastructure.
Immutable storage as a baseline:
Write-once backup copies that cannot be altered or deleted — even by someone with admin credentials — fundamentally change the attacker's calculation. Most modern cloud storage platforms support object lock features that achieve this. It is a minimum requirement now, not an advanced measure.
True network and credential isolation:
A backup environment sharing network access and credentials with production is not meaningfully separate. Proper isolation keeps backup data out of reach even after a network compromise. Multi-factor authentication and role-based access controls for backup infrastructure add another layer of friction that slows attackers down significantly.
Patch backup software with the same urgency as production systems:
This requires discipline, not new tooling. Backup platforms need to be built explicitly into the patch cycle rather than treated as exceptions.
Test restores, not just backups:
An untested backup is an assumption. Regular restore testing — ideally including attack scenario simulations — surfaces gaps before attackers do. Organizations that discover restoration failures during an actual incident have no time to fix them.
The gap between awareness and action
The challenge of protecting backups points to a broader problem in cybersecurity. Most organizations understand the risk of an attack. They may know they need stronger isolation, better monitoring or more resilient recovery, but lack the budget, staff or operational capacity to put those protections in place.
Almost 77% of IT organizations surveyed for our cybersecurity report,
Building Security That Survives Human Error
, say their cybersecurity investment is not keeping pace with the threats they face. Among MSPs, 65% say their clients are underinvesting in cybersecurity.
The result is a familiar tension. Security teams know what needs attention, but limited resources force them to make difficult choices about where to focus first.
Drawing on responses from more than 1,100 IT and cybersecurity professionals, the report looks at the pressures behind those choices, from human error and underinvestment to staffing shortages and operational friction.
If your team knows where it needs to improve but struggles to make those improvements happen, our report offers a closer look at what is getting in the way and where organizations are focusing their efforts.
What the £350 Grand Theft Auto ‘collector’s box’ says about the escalating cost of gaming
Guardian
www.theguardian.com
2026-10-07 10:00:53
The launch of a collector’s box hints at a new reality: the kinds of games I grew up with are becoming a luxury item, locked behind a forbidding spending barrier • Don’t get Pushing Buttons delivered to your inbox? Sign up here As a parent (and partner) to a habitual collector, I do feel like an unw...
A
s a parent (and partner) to a habitual collector, I do feel like an unwelcome portion of my life is occupied by an endless war against stuff. Surfaces in my home ceaselessly accumulate piles of random objects: a JoyCon wrist strap, a sticker sheet, an unidentifiable plastic piece of a toy or gizmo, an Xbox 360 controller with no batteries, cases for Switch cartridges lost down the side of the sofa … Stuff encroaches everywhere, especially in a gaming household. Much as I regret the
imminent disappearance of physical games
, I do not regret the extra space I have in my attic as a result of switching almost entirely to downloaded games over the last decade or so.
Speaking of the imminent disappearance of physical games: Rockstar recently announced a £350 ($400) “collector’s box” of Grand Theft Auto 6 goodies … that doesn’t include the game itself. Not even a download code. The Grand Theft Auto VI: The Goodtime State – Vice City Collection, to give it its full and unwieldy name, contains a hat, a bag, a pair of sunglasses, a figurine, a poster and more, including a mirror and a “swizzle spoon” that looks, unambigiously, like a cocaine spoon. The game retails at £70, so if you want something to play alongside all your stuff, you’re looking at an outlay of £420.
I’m not about to rag on this box of merch. People love this kind of stuff, and as a lifelong collector of Japanese gatcha and blind-box toys, I am in no position to judge how people spend their money. I have a few video game special editions that I cherish – the statue that came with my copy of The Last Guardian is still on my shelf – though I also have many more that are relegated to various cupboards. But it is unusual, however, for them to be this expensive. Usually with gaming special editions, we’re talking an extra £20-50 for a statue, or a steelbook, maybe a booklet and a keyring. 2009’s Call of Duty: Modern Warfare 2’s “prestige edition”
came with night vision goggles
for $150. In 2015, Fallout 4’s $120 collector’s edition shipped with a
plastic Pip Boy
wrist unit (I had one; alas, it was flimsy and disappointing).
There have been some infamously expensive gaming special editions down the years, some of which were PR stunts (see Saints Row IV’s
million-dollar pack
, which included a Lamborghini and cosmetic surgery – nobody bought it). But in the 2020s high-priced video game merchandise has become more common, from a Balenciaga PlayStation tie-in line to
£580 for Pokémon Lego
.
Priced out … expensive video games and merch have now become a luxury item.
Photograph: Charly Triballeau/AFP/Getty Images
What’s behind this is a troubling sense that video games as we once conceived of them – console games, played at home, perhaps on a nice 4K TV – have become a luxury item. People who can afford to spend
up to £790
for a console and games are perhaps now people who can afford an extra £350 for a box of merch.
Continual price-rises
from Xbox and PlayStation suggest that video game companies do not seem particularly worried about pricing people out. Super-expensive merch and special editions like GTA 6’s form part of the same upward trend.
This stratification of the gaming world has been happening for years, ever since the first smartphones. The vast majority of the world’s players – about 80%, according to
Newzoo
– play for free, or at low cost, on mobile phones. The remaining 20% of players account for a disproportionate share of spending on the high-end gaming market on consoles and PC. With merch boxes like this, games companies are really going after that 20% for all they’ve got.
I’m saddened by the idea of gaming becoming even more of a luxury hobby. On the one hand, games in the broadest sense are more accessible than ever – you’re three taps away from playing Monopoly Go on a phone at any moment. But the kinds of games I grew up with and love to write about, games that feel like worlds rather than temporary diversions, are now increasingly locked behind a spending barrier. With the increasing scarcity of physical discs and cartridges, you can’t even reliably buy them second hand any more.
I guess there’s one advantage to the increasing cost of collector’s editions: the price has removed any lingering temptation to buy them. God knows I do
not
need more stuff.
What to play
A game to sit back and wallow in … Silver Pines.
Photograph: Wych Elm Games, Team 17
Keith has been playing Silver Pines, a neo-noir detective drama with point-and-click adventure sensibilities. “As I played, I discovered countless references to 80s and 90s neo-noir culture – from Blue Velvet to Twin Peaks,” he writes
in our review
.
“It was a pleasure to explore every broken-down movie theatre, oily car garage and sweat-stinking interrogation room, collecting weird trinkets and working out how they’d all fit together – then exploring them again (and again) when the answers failed to show up. Silver Pines is a game to sit back and wallow in, like a private eye, kicking off his wingtip shoes and pulling that bottle of bourbon from the bottom drawer as rain trickles down the window and sodium street lights paint the walls a sickly shade of yellow.”
Available on:
PC, PS5, Switch 2
Estimated playtime:
12 hours
Farewell, Ninja Theory … Hellblade 2.
Photograph: Ninja Theory
Award-winning
UK studio Ninja Theory
has
reportedly
begun laying off staff. The creator of Hellblade was one of the developers affected by
Microsoft’s programme of lay-offs
and studio closures in July, and according to Xbox COO Matt Booty two agreements to keep the studio running have failed. It is a sad end to one of the UK’s most interesting development houses.
Meanwhile, Windows Central
obtained the transcript
to a recent in-house meeting held for Xbox staff by Xbox CEO Asha Sharma. In it, Sharma refers to the next generation Xbox platform currently known as
Project Helix
as “a family of devices” rather than a single console. “There is not one device that fits all,” she stated. “We will certainly have a great console first, but we also know that many players will be on the go.” Could the ROG Xbox Ally, launched in partnership with Asus last year, have just been a beta test for a fully fledged handheld Xbox?
The news site GamesIndustry.Biz has
an interesting feature
on
Zach Cregger’s Resident Evil movie
and how it has affected the Resident Evil games. Although the film led to a modest uptick in sales, the real value of movie and TV tie-ins, according to industry analysts, is that they keep those franchises in the minds of gamers so that subsequent titles benefit from the exposure. As game development cycles continue to widen, it looks like we can expect a lot more movies and TV shows to fill in the gaps.
Get cozy … Flame, Forest and Flood.
Photograph: Fairy Mount Games
A question from my friend
David
via email:
“I keep hearing that game publishers are avoiding the whole of November because of GTA 6
– will there really be NO choice for those of us who don’t want to shoot stuff and run people over?”
Thankfully, it’s only the big western publishers who are staying clear of the GTA 6 launch window – lots of other developers are diving in regardless. Cozy strategy game
Flame, Forest and Flood
(try saying that three times) is launching on the same day as GTA, offering a delightful tactical alternative, set in a hex-based woodland idyl. Staying on a rural theme, the purposefully rage-inducing “precision platformer”
Trees Hate You
is a woodland hiking game, where you are attacked by nature. Alternatively,
Football Manager 2027
is arriving on 10 November, so instead of leading a gang of drug dealers into a bloody turf war, you could lead Plymouth Argyle into the Premier League. I’m not sure which scenario is more challenging.
If you have a question for Question Block – or anything else to say about the newsletter – hit reply or email us at
pushingbuttons@theguardian.com
.
All the numbers: Amazon Prime Day 2026 powered by AWS
Amazon Prime Day 2026
was exclusively for Prime members and ran June 23–26, 2026 with millions of deals across more than 35 categories.
As part of our annual tradition of sharing how AWS powered Prime Day (
2016
,
2017
,
2019
,
2020
,
2021
,
2022
,
2023
,
2024
, and
2025
), I want to share the services and chart-topping metrics from AWS that made your amazing shopping experience possible.
Prime Day 2026 – all the numbers
As always, Prime Day was powered by AWS. Here are some of the most interesting and/or mind-blowing metrics:
Amazon Elastic Compute Cloud (Amazon EC2) and AWS Graviton
– During Amazon Prime Day 2026, AWS Graviton powered up to 49% of the Amazon EC2 compute used by Amazon.com.
Amazon Elastic Block Store (Amazon EBS)
– During Prime Day 2026, Amazon EBS, our high-performance block storage service, peaked at over 24.8 trillion I/O operations, moving over an exabyte of data daily.
AWS Lambda
– AWS Lambda handled over 2.3 trillion invocations per day during Prime Day 2026.
Amazon Elastic Container Service (ECS) and Fargate
– During Prime Day 2026, Amazon ECS launched an average of 158.3 million tasks per day on AWS Fargate, representing a 47.7 percent increase from the previous year’s Prime Day average.
Amazon CloudFront
– Amazon CloudFront delivered over 2.1 trillion HTTP requests during the global week of Prime Day 2026, a 5 percent increase in total requests compared to Prime Day 2025.
Amazon DynamoDB
– Amazon DynamoDB, a serverless, fully managed, distributed NoSQL database, powers multiple high-traffic Amazon properties and systems including Alexa, the Amazon.com sites, and all Amazon fulfillment centers. Over the course of Prime Day 2026, between June 23 and June 26, 2026, DynamoDB processed over 59 trillion requests. DynamoDB maintained high availability while delivering single-digit millisecond responses and peaking at 192 million requests per second.
Amazon Aurora
– On Prime Day, Amazon Aurora, a relational database management system (RDBMS) built for high performance and availability at global scale for PostgreSQL, MySQL, and DSQL, processed hundreds of billions of transactions, stored 5,491 terabytes of data, and transferred 1,194 terabytes of data.
Amazon ElastiCache
– During Prime Day, Amazon ElastiCache peaked at serving over 2.3 quadrillion daily requests and 2.1 trillion requests in a minute.
Amazon Kinesis Data Streams
– Amazon Kinesis Data Streams, a fully managed serverless data streaming service, processed a peak of 988 million records per second during Prime Day 2026.
Amazon Simple Notification Service (SNS)
– Amazon SNS, a fully managed pub/sub messaging service for application-to-application and application-to-person communication, delivered 5 trillion messages in a single day during Prime Day 2026.
Amazon Simple Queue Service (SQS)
– Amazon SQS, a fully managed message queuing service for microservices, distributed systems, and serverless applications, received a peak of 213 million messages per second during Prime Day 2026
AWS CloudTrail
– AWS CloudTrail processes billions of API activity events per day for governance, compliance, and operational auditing. During Prime Day 2026, that volume surged to 3.6 trillion events in just four days, a 44% increase over Prime Day 2025.
Amazon CloudWatch
– Amazon CloudWatch processed over 2.15 quadrillion metric observations per day during Prime Day 2026.
Amazon GuardDuty
– During Prime Day 2026, Amazon GuardDuty monitored an average of 14.08 trillion log events per hour, a 59% increase from last year’s Prime Day.
AWS Fault Injection Service (FIS)
– We ran over 44,000 AWS FIS experiments – more than six times what we conducted in 2025 – to help ensure Amazon.com remains highly available on Prime Day.
Prepare to scale
If you’re preparing for similar business-critical events, product launches, and migrations, I recommend that you take advantage of
AWS Countdown Premium
. From retail peak seasons to major sporting events, elections, and healthcare enrollment periods, we help you deliver flawless experiences when it matters most. Our experts help you manage your infrastructure to handle massive traffic spikes while maintaining security and performance. We work alongside your team to scale infrastructure, optimize costs during demand surges, strengthen security measures, and monitor real-time demands. To learn more, visit
AWS Countdown Premium for business critical events
.
I look forward to seeing what other records will be broken next year!
As you wake up this morning, does your body ache from going too hard in the pit during the Lost Dog set? Are you still puffy-eyed from having experienced full emotional release during the Crisis Public Relations performance? Or wondering just who on earth was that beautiful person who made brief eye contact during the Monotone Assassins' harmonica solo? Or searching your pockets for your wallet because it's possible you went
too
low on the dance floor while DJ KoHo was spinning?
(Tod Seelie / Hell Gate)
If none of this applies to you, that means you probably missed out on Hell Gate's second annual "In Battle, Bands," held this year at Night Club 101 in the East Village. (And if you
did
miss it, that's possibly because you're not a Hell Gate subscriber—
become one now
during our biggest sale ever!)
This year, the event, which pits acts featuring local journalists against one another, was actually a cutthroat, winner-take-all (prize: novelty trophy) contest presided over by a group of three steely-eyed judges.
The competition was absolutely fierce! Because not only can these award-winning journalists shine light on the inner workings of our glorious city, they can shred without mercy.
ShinyHunters Extorted Boeing Spin-off Prior to Arrests
Krebs
krebsonsecurity.com
2026-10-07 09:48:45
A teenager from Amman, Jordan suspected of leading the prolific data theft and extortion group ShinyHunters has been detained and is reportedly cooperating with the FBI to identify other members of the hacking gang. KrebsOnSecurity has learned that the suspect, who uses the hacker handle "Rey," was ...
A teenager from Amman, Jordan suspected of leading the prolific data theft and extortion group
ShinyHunters
has been detained and is reportedly cooperating with the FBI to identify other members of the hacking gang. KrebsOnSecurity has learned that the suspect, who uses the hacker handle “
Rey
,” was detained as ShinyHunters was in the process of extorting a business unit recently divested by the global aerospace company
Boeing
, which manufactures the fleet of planes used by the employer of Rey’s father —
Royal Jordanian Airlines
.
The logo for Jeppesen ForeFlight, a business unit divested last year by the aerospace firm Boeing.
On October 3,
Reuters
cited
three unnamed sources saying a suspected ShinyHunters member in Amman named
Saif Al-din Khader
was detained by Jordanian authorities and was cooperating with the FBI. KrebsOnSecurity identified Rey as Khader in
a November 2025 profile
, in which the young man admitted working with multiple ransomware groups.
Rey was featured again in
a September 28 exclusive
about the Dutch police arresting 24-year-old convicted cybercriminal
Pepijn van der Stap
on suspicion of aiding in data thefts and extortions by ShinyHunters. The story noted that immediately following the Dutchman’s arrest on the evening of September 15, Rey assumed control over the ShinyHunters brand and boasted publicly about stealing highly sensitive data from the
FBI
and extorting the ransomware group
Cl0p
.
Rey taunted both the FBI and Cl0p with memes posted to his longtime account on Twitter/X, while simultaneously including images of the avatar used by Van Der Stap’s former hacker alias “
Umbreon
” in an apparent attempt to frame the Dutchman for both hacks.
A taunting meme uploaded to Twitter/X by Rey on Sept. 22. A giant sized version of the Pokemon character Umbreon can be seen in the bottom left.
As noted in our September 28 report, ShinyHunters gained access to the FBI site and other victims by exploiting a vulnerability (CVE-2026-35273) in
PeopleSoft
, a software-as-a-service platform from the tech giant
Oracle
that is broadly used by companies to manage hiring and human resources, benefits and payroll. Oracle quickly issued a fix for CVE-2026-35273, which ShinyHunters first began exploiting as a zero-day in June, and at the time Mandiant released web application firewall rules intended for organizations that couldn’t apply the security update quickly enough.
ShinyHunters
told BleepingComputer in June
that the original goal behind exploiting the PeopleSoft vulnerability was to breach the FBI’s own PeopleSoft database, but the hackers said those attacks were unsuccessful for some reason. In recent weeks, however, ShinyHunters
turned to a well-known URL-encoding trick
to bypass Mandiant’s suggested web application firewall rules.
In
a report
released Sept. 25, security experts at
Mandiant
and the
Google Threat Intelligence Group
(GTIG) confirmed that ShinyHunters had mass-exploited the PeopleSoft vulnerability to steal data from dozens of systems across a range of industries, including higher education, technology, healthcare, agriculture, transportation and government.
Reuters
reported October 5
that the FBI has removed a contractor at
Accenture
over their failure to patch the FBI recruitment website hacked by ShinyHunters, which exposed sensitive data on more than 5,000 FBI personnel, including each’s person’s unit and specialization, as well as medical and psychiatric records.
‘REY’ MEANS KING, AS IN ROYAL
According to two sources familiar with the ShinyHunters investigation, a navigation and digital aviation unit recently divested by the global aerospace company
Boeing
was among the victims that ShinyHunters was in the process of extorting when Rey was apprehended by Jordanian authorities.
Those sources said the FBI’s investigation into ShinyHunters gained renewed urgency with the group’s attempted extortion of the former Boeing unit, which allegedly included the theft of sensitive information that sources said could pose operational safety and security risks.
In a brief statement shared with KrebsOnSecurity, Boeing acknowledged the extortion attempts by ShinyHunters, and said the incident concerned data stolen from
Jeppesen ForeFlight
, a subsidiary that Boeing
sold in November 2025
to the private equity firm Thoma Bravo for $10.55 billion.
“We are aware of claims by a threat actor regarding data allegedly associated with Boeing and our former subsidiary Jeppesen ForeFlight,” a Boeing spokesperson shared. “We are actively reviewing the matter with the Jeppesen ForeFlight team.”
A spokesperson for Jeppesen ForeFlight shared a written statement in response to questions, saying the company has seen no impact on their end. “Based on our investigation to date into this claim and proactive security posture, there was no impact to our operations or products.”
Rey’s alleged involvement in attempting to extort the former Boeing unit is noteworthy because there is strong evidence that his father works for
Royal Jordanian Airlines
, which is mostly controlled by the Jordanian government and operates its long-haul fleet on passenger planes built by Boeing. Rey claimed on Telegram in early 2025 that his father was an airline pilot, although that could not be independently confirmed.
However, as noted in
our November 2025 profile of Rey
, his family’s shared computer was at one point compromised by password-stealing malware, and the data collected by that malware clearly shows Rey’s father used the same credentials to log in at multiple online portals for Royal Jordanian Airlines employees.
Royal Jordanian Airlines has not yet responded to a request for comment. In advance of our September 28 story, KrebsOnSecurity once again emailed Rey’s father to seek comment and update him on his son’s alleged activities. Neither of the Khaders have responded. But just hours after that request was sent, Rey began deleting his various social media accounts, including the Twitter/X account he previously used to taunt the FBI, Cl0p, and other ShinyHunters victims.
Rey may have nixed many of his social media profiles, but his cybersecurity blog on GitHub somehow escaped the purge, and it shows that Rey was fixated on the leaders of the Cl0p ransomware group. In March 2026, Rey’s blog featured
a lengthy post
that identified two Russian men as the core developers and hackers behind Cl0p.
Rey’s blog on GitHub. This post doxes two Russian men as the core operators behind Cl0p, one of the oldest and most established ransomware groups still in operation today.
MURDER FOR HIRE?
Meanwhile, news outlets in the Netherlands reported explosive new allegations leveled at Van der Stap, whose supposed personal transformation from convicted to reformed hacker has been widely covered in the tech news media. The Dutch daily
RTL reported on Sept. 29
that investigators suspect Van der Stap tried to orchestrate at least two murders. According to RTL, the murders were allegedly to be committed abroad, and there are indications Van der Stap gave the order for these attacks.
Van der Stap was released from prison after serving the better part of a four year sentence for data theft and extortion activity that prosecutors said netted between €1.5 million and €2.7 million. In an interview with KrebsOnSecurity on September 9, Van der Stap described his new role as “offensive security lead” at the Dutch cybersecurity company
Neo Security
, saying the job involved probing client networks for security vulnerabilities.
Neo Security’s owner
Benjamin Korper
told Reuters he has hired an outside firm to investigate whether Van der Stap had hacked Neo Security or its customers, but that so far investigators have found no evidence he acted against his employer or clients. Korper said Dutch forensic investigators visited his office on September 15, the night Van der Stap was arrested in a dramatic police raid that reportedly involved
flash bang grenades
.
A screenshot of a Sept 16 story by the Dutch news outlet at5.nl, describing a police raid on Van Der Stap’s residence that reportedly used flash-bang grenades.
Prior to his first arrest in 2023, Van der Stap was working as a software engineer at the Amsterdam-based cybersecurity startup Hadrian, while volunteering at the Dutch Institute for Vulnerability Disclosure (DIVD) — even as he was hacking into and extorting a number of large organizations.
When asked in a recent interview why anyone should believe the word of a self-described “reformed” cybercriminal who had so casually deceived countless friends, co-workers and journalists for years, Van der Stap replied that his work spoke for itself and there was nothing he could say that would convince his worst critics.
“You can throw a bunch of nice words at someone, but you can’t convince them if they don’t want to be convinced,” Van der Stap told KrebsOnSecurity on Sept. 9. “I’m doing what I can to repay victims, and that’s all I can do. If someone doesn’t want to believe me, then that’s on them.”
FRANCHISING AND BURNING A BRAND
Cybercriminals aligned with ShinyHunters have been responsible for
dozens of data breaches
involving billions of stolen records, and breaches claimed by the group stretch back to at least 2019. But experts say the people recently operating behind the ShinyHunters name are not the same core members that populated the group in its early days, most of whom are French citizens who have been arrested (if not also imprisoned) on
at least one prior occasion
for alleged cybercrime activity.
More to the point, ShinyHunters has become something of a franchise. Think the
Dread Pirate Roberts
character in the 1980s cult movie classic “The Princess Bride,” only succession by death is replaced with succession by arrest, and there can be multiple simultaneous Dread Pirate Robertses. Sources close to the investigation say the FBI is focusing on a remaining handful of cybercriminal freelancers or affiliates who have been feeding the group stolen credentials to various software-as-a-service (SaaS) platforms used by major companies in exchange for a cut of any data ransoms later paid by victims.
In the days after the news broke of Van der Stap’s arrest, a cybercrime-focused chat server on Telegram that was allegedly operated by Rey erupted with hot takes, with most participants heaping ridicule on the teenage hacker after he publicly backed down from threats against the FBI and Cl0p, and again when
the ShinyHunters’s darknet website suddenly went offline
. Several commentators accused Rey of resurrecting the ShinyHunters brand after its core members were rounded up in France, and making a mockery of the group’s name and reputation ever since.
“He bought the old forum PGP key and used it to make new Breachforum websites and Telegram channels larping as ShinyHunters to ransom companies and then sell the used data or resell his forum when he goes broke,” one member recounted.
A relatively new Telegram channel called “The Battle” has been doxing and needling Rey and other alleged ShinyHunters members for several weeks, and it has gained a considerable readership among the cybercrime communities operating on Telegram. One of the coordinators of that harassment campaign repeatedly portrayed Rey as clueless greenhorn who sought to ride the coattails of a cybercriminal brand that has long enjoyed a reputation for ruthlessly selling or publishing data stolen from victim companies who refuse to give in to extortion demands.
“Rey (Saif Al-Din Khader) made a serious mistake when he started pretending to be a member of ShinyHunters,” wrote the administrators of The Battle server on Telegram. “That group had already been dismantled, with many of its members either arrested or imprisoned, yet Rey still chose to use its name while carrying out his crimes. We’re aware of claims that [Rey] caused over $200 million in damages and helped around 5–6 friend groups in the community make money by using Shiny Hunters group aliases to negotiate deals for a 25–30% cut over the past few months.”
In an interview with
The Register
, ShinyHunters claimed they hacked the FBI to counter the agency’s narrative in
a May 2026 alert
that advised victims against paying a ransom to the group, which came off looking unprofessional and capricious in the FBI’s advisory.
A flash notice on ShinyHunters released by the FBI on May 15, 2026.
The public notice warned the group has been known to pursue a number of
different victim harassment strategies
, from sending threatening text messages and phone calls to victims and their family members to in some cases
swatting
victims. The FBI warned ShinyHunters members “may also falsely claim to have sensitive or compromising information, including embarrassing photographs or videos of victims, which frequently do not exist.”
The hackers told The Register their attack on the FBI “demonstrated our technical capabilities and directly refuted the misinformation disseminated by the FBI, journalists, and industry researchers.” At the same time, the group’s leaders seemed to acknowledge that the FBI’s warning materially harmed their prospects for convincing victims to pay, saying “this was fundamentally a public relations and marketing initiative for our business.”
[$] Evolving the LAVD scheduler from gaming to servers
Linux Weekly News
lwn.net
2026-10-07 09:33:48
The extensible scheduler class, which
enables the creation of custom CPU schedulers with BPF, has led to a burst
of innovation in this area; the
LAVD scheduler has, perhaps, been one of the most noteworthy schedulers
to emerge. Though it was originally designed
for gaming applications, the LAVD sch...
The page you have tried to view (
Evolving the LAVD scheduler from gaming to servers
) is currently available to LWN
subscribers only.
Reader subscriptions are a necessary way
to fund the continued existence of LWN and the quality of its content.
If you are already an LWN.net subscriber, please log in
with the form below to read this content.
Please consider
subscribing to LWN
. An LWN
subscription provides numerous benefits, including access to restricted
content and the warm feeling of knowing that you are helping to keep LWN
alive.
(Alternatively, this item will become freely
available on October 15, 2026)
The other day I was looking at a little project called
grubby
, a minimal static site generator for git
repos, written in Ruby. The style is quite interesting. It feels solid and unapologetic:
After running
git ls-tree
and collecting the filenames in the repository, it
generates an HTML page for each file. Looking at the code examples above we can
discern a primitive templating system: there’s
#fill_template
for
replacing place-holders in the template with values, there’s a
#render_markdown
helper function, there’s the composition of the inner HTML
with the header and footer HTML in
#write_page
.
The header and footer are stored in separate files, but where are the templates
for the actual page content? A bit further down, we find them. Here’s an excerpt
from the index page template:
This code has a very specific shape: it’s a series of operations that stuff
strings into a buffer. In fact, this code looks very similar to the kind of code
that will actually run when you use
Papercraft
or
ERB
for generating HTML from templates. Both ERB
and Papercraft compile the templates into optimized Ruby code that emits
snippets of HTML into a buffer.
I find it interesting that the code remains very readable, even though it’s very
low level. There are of course many different ways to write this kind of code.
We could use a heredoc with string interpolation:
page_html=<<~HTML
<header>\n<p><strong>#{escape_html(repo_name)}</strong></p>
<p>#{escape_html(repo_description)}</p>
<p><code>#{escape_html(clone_command)}</code></p>\n</header>
HTML
There’s also the
optimized
way that avoids interpolation and therefore double
copying of strings:
A bit less readable, but somewhat faster. Both ERB and Papercraft compile code
into this form. Just for kicks, how would a corresponding ERB template look?
Everything is compressed down to a minimal syntax. But in the end, when you
render this template, Papercraft will actually compile it into code that looks a
lot like
grubby
:
So going back to the grubby code, it was interesting to see this style of
coding, which basically “unrolls” the code that would have been generated if a
templating tool like ERB or Papercraft was used. But it also reminded me of
another code base I was looking at recently, that of
Herb
, which is a gem with a set of tools
for working with ERB templates. Here’s an excerpt:
# from Herb::Engine#initialize:...@context=Visitor::Context.new(file_path: properties[:filename],project_path: properties[:project_path],options: context_options(properties),resolver: properties[:resolver],**(properties[:context]||{}))@bufvar=properties[:bufvar]||properties[:outvar]||"_buf"@escape=properties.fetch(:escape){properties.fetch(:escape_html,false)}@escapefunc=properties.fetch(:escapefunc,@escape?"__herb.h":"::Herb::Engine.h")@attrfunc=properties.fetch(:attrfunc,@escape?"__herb.attr":"::Herb::Engine.attr")@jsfunc=properties.fetch(:jsfunc,@escape?"__herb.js":"::Herb::Engine.js")@cssfunc=properties.fetch(:cssfunc,@escape?"__herb.css":"::Herb::Engine.css")@src=properties[:src]||+""@chain_appends=properties[:chain_appends]@buffer_on_stack=false@parser_options=properties.fetch(:parser_options,default_parser_options).transform_keys(&:to_sym)...
The entire
Herb::Engine#initialize
method is about 65 LOC! And the individual
lines themselves tend to stretch to the right edge of your editor window. I ran
cloc
on the code and was really taken aback:
OK, so the Typescript stuff is probably not relevant to the present discussion,
but still, 20KLOC of Ruby, and another 37KLOC of C/Rust extensions. For
comparison’s sake, ActiveRecord is about 46KLOC (not including tests). The
entire Rails codebase is about 120KLOC of Ruby (again, not including tests).
Herb just became the default template renderer in Rails 8.2.
This method is basically a router inside of a loop. it converts data into method
calls. Again, this kind of functionality could have been hidden behind a DSL,
but here we have the entire logic for this algorithm laid out very explicitly in
front of our eyes. Another example:
It’s just logic expressed in the purest way possible. The code that we’re
looking into here is charged with transforming an ERB template into a piece of
Ruby source code containing an optimized renderer for the given template. The
optimized source code is painstakingly put together from little bits and pieces,
according to the structure of the template. If we take the same HTML template from
the grubby example above, ERB/Herb would have emitted the following code:
# edited for formatting_erbout=+''_erbout.<<" <header>\n <p><strong>".freeze_erbout.<<((escape_html(repo_name)).to_s);_erbout.<<"</strong></p>\n <p>".freeze_erbout.<<((escape_html(repo_description)).to_s);_erbout.<<"</p>\n <p><code>".freeze_erbout.<<((escape_html(clone_command)).to_s);_erbout.<<"</code></p>\n</header>\n".freeze_erbout
Note the similarities: we’re dealing with constructing a string, we’re mostly
doing just one type of action, and the action is repeated. This is what unrolled
code is about: clarity, discipline, optimization, and grouping like actions
together. Let’s take a look at a different project:
defemit_wildcard_childless_root_code(buffer,root_path)emit_code_line(buffer,'->(path, params) {')ifroot_path!='/'re=/^#{Regexp.escape(root_path)}(\/.*)?$/emit_code_line(buffer," return if path !~ #{re.inspect}")endemit_code_line(buffer," @dynamic_map[#{root_path.inspect}]")emit_code_line(buffer,'}')end
This is from
Syntropy
, which is
the web framework that’s driving this website, and the code excerpt is part of
its routing tree compiler. The idea is to take a tree-like data structure
describing the apps’s directory structure and files, and compile it into an
optimized router lambda that can deal with parametric and wildcard routes.
The resulting router code would look something like the following:
Like with ERB/Herb, the Syntropy code generates a complex piece of code by
putting together strings containing little bits of Ruby source code. The above
method could also have been written as follows:
This is much clearer, but still, the code for generating the conditional return
in the middle there is a bit hairy. What if we had a DSL for generating Ruby
code? In Elixir you can define macros that expand into code with a pair of tools
called quote/unquote. In fact, a lot of Elixir’s language features (even basic
stuff like
if
and
case
) are implemented using macros. When a macro is used,
the macro definition is expanded in place, it’s like
parametric code
. This
idea, like all good ideas, comes from Lisp, where there’s no distinction between
data and code. What if we had that in Ruby?
The
#quote
method returns the AST of the given block. The
#unquote
method is
used to inject arbitrary values, or nested ASTs into the quoted code. In this
example, we conditionally inject a piece of Ruby code expressed with a nested
#quote
block, that interpolates a regular expression. The final call to
#unquote
injects the value of
root_path
as a literal into the generated code.
Notice the mechanics of quote/unquote: with
#quote
we’re putting code inside
quotes (obviously), and with
#unquote
we’re temporarily escaping out of the
quotes in order to perform some computation and inject the result back into the
quoted code, it’s very much like string interpolation. The expression passed to
#unquote
is evaluated at
compile-time
. This allows us to conditionally
include pieces of code in the template. And since
#unquote
always returns an
AST (or
:__nop__
for nothing), we can use it to compose ASTs together:
Here’s an attempt to apply the idea of quote/unquote to compiling Papercraft
templates. This method takes the template AST, and returns a mutated AST, where
tag method calls are translated to strings being emitted to a buffer. But since
we want to have the same optimized form of bunching together static strings and
separating out the dynamic strings, we introduce some compile-time state (the
html_parts
buffer), and a
flusher
closure that generates the actual code
that emits static strings to the buffer. This design lets us support recursion
in
#emit_html
, so we can do stuff like
div { p { a 'Home' } }
, since the
state is passed as arguments.
Suppose we have this quote/unquote functionality ready to let us generate code
programmatically. We still need a few more tools to be able to create code.
Consider the following:
(
Sirop
is a little gem I wrote to
help work with Prism ASTs.)
#mutate
works by creating a copy of the AST, letting you replace any node on
the tree with another. In this case, we’re implementing a momoized version of a
method by replacing the method body with a conditional assignment, into which
we inject the memo key and the original method body. I think this example
demonstrates the strength of this approach, it feels magical!
Another way to work with
#mutate
is by passing it a block that returns either
the original node, or a different node in case of a mutation:
This is invalid syntax, and while Prism will parse this, the AST would contain a
MissingNode
, and will otherwise be deformed. Here’s a solution that doesn’t
feel too inelegant:
I find that a functional approach to coding goes very well when working with
ASTs. We treat ASTs as immutable objects, and if we need to change a node
anywhere on the AST, we can use
#mutate
which creates a copy with the
requisite changes. If we’re generating more complex code, as we do in
Papercraft, Syntropy, or ERB, we can split the code generation logic into
multiple methods, each of which prepares a distinct part of the code, and
returns an AST. This allows us to create arbitrarily complex pieces of code by
composing ASTs together. Let’s take a real use case. Here’s the template for an
ActiveRecord migration, used by a Rails generator. Yes, Rails uses ERB to generate
code:
class <%=migration_class_name%><ActiveRecord::Migration[<%=ActiveRecord::Migration.current_version%>]defchangecreate_table:<%=table_name%><%=render_table_with_dom_id%>do|t|<%attributes.eachdo|attribute|-%><%ifattribute.password_digest?-%>t.string:password_digest<%=attribute.inject_options%><%elsifattribute.token?-%>t.string:<%=attribute.name%><%=attribute.inject_options%><%else-%>t.<%=attribute.type%>:<%=attribute.name%><%=attribute.inject_options%><%end-%><%end-%><%ifoptions[:timestamps]-%>t.timestamps<%end-%>endendend
How would it look with quote/unquote? Let’s find out:
version=ActiveRecord::Migration.current_versionquotedo__.class(unquote(migration_class_name)<ActiveRecord::Migration[unquote(version)]){defchangecreate_tableunquote(table_name)do|t|unquoteattributes.mapdo|attribute|opts=attribute.inject_optionsifattribute.password_digest?quote{t.stringunquote(:"password_digest#{opts}")}elsifattribute.token?quote{t.stringunquote(:"#{attribute.name}#{opts}")}else# note use of unquote as method namequote{t.send(unquote(attribute.type),unquote(:"#{opts}")}endendunquote(options[:timestamps]?quote{t.timestamps}::__nop__)endend}end
Not very pretty, I admit, but maybe a little refactoring, splitting the code
into distinct parts, would make it better:
Oh yes this is
much
better, as we can now see the shape of the generated code!
One more example, fitting for the title of this article:
defrewrite_block_param(ast,v){block_param_name=ast.parameters.parameters.requireds[0].namemutate(ast){|n,t|ifninPrism::LocalVariableReadNode(name: block_param_name)unquote(v)elsenend}}defunroll(o,&block)block_ast=Sirop.to_ast(block)o.map{|v|rewrite_block_param(block_ast,v)}endast=quote{->{unquoteunroll(%w{foo bar baz}){|v|pv}}}unrolled_printer=eval(Sirop.to_source(ast))
Here we use
#mutate
to change the body of a given block such that all
references to the block argument will be replaced with an arbitrary value, such
that the actual source code of
unrolled_printer
would be:
->{p'foo'p'bar'p'baz'}
Walking the Fine Line of Abstraction
This morning I was looking at a vibe-coded
port of
Campfire
from Rails to plain
Ruby (with ractors). This is an interesting project, and I think it demonstrates
pretty well the
performance costs of
Rails
’ abstractions.
With Rails, we choose developer happiness at the expense of machine happiness.
But the results show that Ruby is actually pretty fast. OK, not as fast as Rust,
but in some cases faster than Go!
And the code itself is interesting - a lot of unrolled code for sure, but
generated by an LLM:
defdispatch(request)...if(verb=="GET"||verb=="HEAD")&&(file=Assets.file(path))returnstatic(file,request)endreturnhealth(request)ifpath=="/up"ifpath.start_with?(STORAGE_PREFIX)# ActiveStorage controllers; the proxy ones stream (ActionController::Live).returnFront.finish_response(request,Storage.call(request,path,query),live: path.include?("/proxy/"))endbody=nilifverb=="POST"&&request.headers["content-type"]&.to_s&.start_with?(FORM)body=(request.body&.join||+"").force_encoding(Encoding::UTF_8)if(i=body.index("_method="))m=body[i+8,6].to_s[/\A[a-z]+/i]verb=m.upcaseifm&&%w[PATCH PUT DELETE].include?(m.upcase)endelsifverb=="POST"&&request.headers["content-type"]&.to_s&.start_with?(MULTIPART)# Rails forms with file inputs carry _method as a multipart part.body=(request.body&.join||+"").force_encoding(Encoding::BINARY)if(i=body.byteindex(METHOD_PART))&&(j=body.byteindex("\r\n\r\n",i))m=body.byteslice(j+4,6).to_s[/\A[a-z]+/i]verb=m.upcaseifm&&%w[PATCH PUT DELETE].include?(m.upcase)endbody.force_encoding(Encoding::UTF_8)end...end
Look at the style. The lines stretch to the right, and the code itself is
dealing with tiny details, no abstractions here, just pure algorithms (and lots
of branching!) The
#dispatch
method is about 45 lines long, and could have been
easily refactored into a few separate methods that each does a single thing.
An experienced programmer would probably have a blast refactoring this code,
there are lots of opportunities in there to make it more readable, more
maintainable, even snappier than it already is. This is obviously unrolled code,
but other parts of the code base do use metaprogramming. For example, let’s look
at the router config code, which may look familiar to a Rails developer:
classRouter...definitialize(&block)@groups=Hash.new{|h,k|h[k]=[]}instance_eval(&block)end%w[GET POST PATCH PUT DELETE].eachdo|verb|define_method(verb.downcase)do|pattern,to:,defaults: nil,format: true|add(verb,pattern,to,defaults,format)endend...end
This is just the syntactic sugar, but it’s the
#add
method that does the work
of computing and storing the route information:
Here we can see that what this routing DSL does (as in many cases) is to convert
code into data. All routing configuration is stored in
@groups
, and the work
of actually routing requests (done in
#recognize
) is all about consulting the
routing data. But with quote/unquote we can convert data into code, as we saw
above in the Syntropy router example. Here’s the original
#route_in
method:
With this, we’ve converted each route group to a custom-made piece of code.
Notice the frequent use of
#unquote
- this allows us to change all references
to the
r
iterator variable into literal values! And we can do this safely
because the routing configuration is immutable. The only place where we need to
do a bit more work is in the
return
statement, where we need to also return a
reference to the specific route. We do this by calling
routes[]
where the
subscript is hard-coded using
#unquote
.
Ruby is famous for its productivity and simplicity, and Ruby programmers have
wholeheartedly embraced its metaprogramming facilities in the quest for
developer happiness
. Tools such as
#eval
,
#instance_eval
, and
#define_method
let us create beautiful abstractions that make our code more
readable and arguably easier to maintain. All those Rails idioms, they’re
catchy, they make the intent clear, it’s almost as if they’ve become part of the
Ruby syntax!
The quote/unquote mechanism I’m proposing here doesn’t exist yet in Ruby, but if
it ever materializes, I think it would open a whole new world of possibilities
for Ruby. Such tools will allow us to create better abstractions without paying
the associated performance overhead. They might make it possible to express new
ideas, new idioms, and new techniques for manipulating Ruby code.
Introducing mapsnap: Automated Georeferencing for Historic Sanborn Insurance Maps
Over the past few months I’ve been working on
mapsnap
, a program to automatically georeference old insurance maps. This post explains why these maps are interesting, how mapsnap works, and why this is worth doing.
You can browse all the georeferenced maps at
mapsnap.org
.
I’m interested in history, cities, and maps, and in the United States that combination leads you very quickly to the
Sanborn Insurance Maps
. The Sanborn Map Company produced hundreds of thousands of detailed, block-by-block maps of every city in the United States from roughly 1870-1960. Their original goal was to help insurance assessment (Is this building made of brick or wood? Where are the water mains?) but today they’re valued for the unique view they provide into the history of urban spaces before the major urban renewal projects of the mid-20th century. You can
read more
about Sanborn maps at the Library of Congress, or watch this
six-minute explainer video
.
Here’s an example of one page from a Sanborn map (NY 1923 vol 1 p27):
You can zoom in to see more details:
There’s a lot of information here!
The color shows what material the building is made out of: red is brick, yellow is wood frame.
The numbers in front of each building (46, 48, 50, 52, etc.) are its address. This is a gold mine for sites like
OldSF
and
OldNYC
where you want to locate a photo using a
historic
address, which might not be the same as the current address.
Businesses in a building are labeled, e.g. the Bakery at 50 Oak Street or the Italian Marionette Theatre (sounds fun!). “S.D.” means “Store and Dwelling,” i.e. a store on the ground floor with residential units above.
The heights of the buildings are labeled. Reading right to left, there’s 50’, 40’, 40’, 40’, 44’, 66’.
As it turns out, the NYPL (and OldNYC) has a
photo of this block
! If you read the heights carefully, you can match them up to the buildings in the photo.
The street is alive with activity in the summer of 1933. It’s filled with pushcarts and vendors selling their goods under awnings. It’s a good thing this photo was taken, because
this block no longer exists
. It was demolished in 1950 to make way for the
Alfred E. Smith Houses
project.
I knew that there were large-scale “slum clearance” programs in the mid-20th century, but the sheer scale of the destruction is so much more striking when you see what was there before.
Accessing the Sanborn Maps
The Sanborn maps were very expensive, very limited-run books. Today, they’re in private collections and are still used in real estate and environmental law. But fortunately for us, many public libraries own Sanborn volumes as well. The NYPL has a
collection
of New York maps. And the Library of Congress has the
largest collection
of all. Since many of the Sanborn maps are old enough to have fallen out of copyright, the LoC has been able to scan around 400,000 of them and provide them for free online. No need to visit the library in person. (Kudos to HIG for running this
massive scanning effort
.)
Generally you’re interested in a specific block, though, and finding all the maps of one particular block can be tedious. Google Maps has really raised our expectations about how easy it should be to find maps online.
To make Sanborn maps easier to use on a computer, the key step is to
georeference
them. This means aligning them with a modern map. This is a tedious but typically straightforward process. You find a point (maybe an intersection) on the Sanborn map, then find the same point on a web map. Two or three matched points establish the alignment. There’s a fabulous web site and community,
OldInsuranceMaps.net
(aka OIM), devoted entirely to georeferencing public-domain Sanborn maps.
When I learned about Sanborn maps and OIM, I georeferenced a
few
maps
in my area of upstate New York and
one in Brooklyn
. I found it interesting at first but then increasingly tedious and time-consuming. As a software person, I started wondering: could a computer do this?
After a few months of
going deeper
on this problem than I ever intended, the answer is a qualified “yes.” It is, for the most part, possible to automaticallly georeference Sanborn maps. Some maps are harder than others and it doesn’t get everything right, but it generally does a good job.
My program to automatically georeference Sanborn maps is called
mapsnap
, and I’m excited to explain how it works!
Introducing mapsnap
We’re living in the era of AI and it’s natural to ask whether Claude or ChatGPT can just do this. I tried a few variations on this at first, from “here’s an image, find the transform” to “find the intersections in this image.” It didn’t work as well as I’d hoped, and ChatGPT at least would often try to write a Python program to do image processing, rather than just using its vision. I found that I was mentally tracing street labels to check whether its intersections were good. So why not just write a program to do that?
At its core, that’s how mapsnap works. It runs OCR over a Sanborn map to find the street labels. It uses those labels to find candidate intersections, and then it uses those intersections to generate a fit.
Let’s walk through those steps.
Step One: OCR
The first step is to detect street labels. I used
EasyOCR
for this. This is the only real “AI” in this project, and it’s pretty benign: EasyOCR is a text recognition model from 2020 that’s small and runs locally on your computer. I chose EasyOCR because it was, well, easy to set up, but it also performed well and was able to detect text at any angle, a key feature for maps where streets can run vertically or diagonally. I eventually came to appreciate that EasyOCR was very adaptable as well.
Here are the detections for a map in downtown Brooklyn:
Most of these are legitimate streets, though some (“BROOKLYN” and “BRIDGE”) are not.
For each detection, we get three things:
A street name.
A (rotated) rectangle containing that name.
A confidence score.
I don’t want to get into the weeds here, but this is very much not “vanilla” EasyOCR:
mapsnap
tries to figure out
a good minimum street label size given the detections. Small text tends to be information like addresses and business names that we don’t need.
I
reworked the output decoder
so that it could only emit real street names from a vocabulary of the city’s streets. Reads outside this vocabulary come out as low-confidence junk. The vocabulary can be pretty large: usually an entire county’s worth of roads is just fine.
I
retrained the top levels
of the EasyOCR neural net to do a better job of recognizing the Sanborn font and ignore the noise that often accompanies street labels on these maps.
When I say “I” here, I really mean “Claude and I.”
Each street detection gives us quite a bit of information about the page. While it’s possible to do georeferencing directly from the streets (more on this in a future post), in practice it’s easier to work with intersections, just like the humans do on OldInsuranceMaps.
To get an intersection, we need two streets that aren’t parallel to each other. We’ll assume the streets are straight and go in the direction of the label on the Sanborn map. (What if they’re not? More on that soon.)
The circles here are the extrapolated intersections. We can get the latitude and longitude of the intersection from OpenStreetMap (OSM). The incorrect street detections tend not to produce intersections that exist in OSM, which is a helpful filter.
These pixel + lat/lng pairs give us “Ground Control Points” or “GCPs” as they’re known. OIM wants three GCPs to produce a georeference, but in a pinch it will let you
get away with two
. Two GCPs work fine, so long as you’re willing to assume that the map isn’t skewed and has a uniform scale. The Sanborn maps are very well-made, and this is
typically
a safe assumption.
Extrapolating all the streets produces a list of candidate intersections. Each pair of these produces a georeference. We need to pick a pair or calculate some kind of average.
In practice some of these GCPs will be bogus. In fact, a lot of them might be. There are a few reasons this could happen:
mapsnap misreads a street name or detects text that isn’t really a street. Maybe “POST” is “POST OFFICE” and not “POST STREET”.
The streets might curve before they intersect.
We might misinterpret the street. It might be ambiguous whether “4TH” is “4TH STREET NORTH” or “4TH STREET SOUTH.”
We might not have detected the street angle very precisely.
A divided street or a street that “jogs” across another has multiple intersection points in the real world, and we might choose the wrong one.
Whatever the reason, the GCPs aren’t all trustworthy. Averaging in a situation like this tends to produce poor fits: a mix of good and bad comes out mediocre.
In statistics, you can mitigate this by using a robust metric like the median rather than the mean. mapsnap uses a related technique from the 1980s called
RANSAC
. Here’s how it works:
We try each pair of GCPs. Together, these georeference the map.
We run all the street detections through that model and see how far away they are from the real street in OSM, both in distance and angle.
If they’re close, we’ve got an “inlier.” If they’re far apart, we’ve got an “outlier.”
We choose the pair of GCPs that produces the most inliers.
In the intersections image above, the two GCPs we choose are blue (PIERREPONT x HENRY and PIERREPONT x CLINTON). The ones we didn’t choose are red and yellow. The street labels that are “inliers” under this model are yellow (HENRY, MONROE, CLINTON, PIERREPONT, etc.) and the ones that are outliers are gray (POST, BRIDGE, EAGLE). These outliers are mostly bad reads.
This pair of GCPs produces an excellent fit, good enough that you can line up individual buildings across the nearly hundred year gap between the old and new map:
This isn’t textbook RANSAC, but it’s in the same general spirit. The beauty of this system is that it can tolerate a
lot
of noise (>50%!) so long as the noise isn’t self-consistent. When there’s even a nugget of signal, RANSAC does a pretty good job of finding it.
Where this fails
Street OCR and RANSAC work well when there are street labels, when those labels are clear and unambiguous, when the streets are straight, and when they haven’t changed since the Sanborn map was made. That’s often the case but not always:
EasyOCR has trouble with short street names like “1ST” or “E.”
It sometimes has trouble picking up cardinal prefixes like “N 4TH ST” vs. “S 4TH ST” or stitching together multi-word street names.
When streets curve, as in hilly or suburban areas, extrapolating the labels to find intersections doesn’t work.
When there aren’t many streets (say in a waterfront area or in a rail yard), mapsnap has nothing to latch onto.
When the streets have been renamed (looking at you, Queens) or changed due to highway building or other urban renewal projects, this approach doesn’t work at all.
My goal is to georeference at least 90% of pages in the Library of Congress’s Sanborn collection without making major mistakes. To do that, mapsnap has to handle at least some of these tricky cases. The next posts will look at some of the other information Sanborn maps give us to aid in georeferencing. If you can’t wait for those, or you want to run it yourself, check out the
mapsnap repo
and the full LLM-generated
How it Works
page. You can browse georeferenced maps at
mapsnap.org
.
A new Raspberry Pi Desktop release is finally available for x86-64
Linux Weekly News
lwn.net
2026-10-07 09:26:25
Simon Long has announced
a long-awaited release of Raspberry Pi OS, based on Debian 13 ("trixie"),
for x86-64 systems.
We managed to find the time to update the Desktop for the Buster and Bullseye
releases of Debian, but then we all just got too busy with other things, and,
while we left the ...
We managed to find the time to update the Desktop for the Buster and Bullseye
releases of Debian, but then we all just got too busy with other things, and,
while we left the Bullseye version on the website for anyone who wanted it, we
simply didn't have time to release any newer versions. But people kept on asking
for it – we get two or three emails every week asking when the PC Desktop will
be updated, and we haven't had an answer, because we honestly didn't know when
we might get a chance to do it. We've continually tried to allocate time to be
able to work on this, but it hasn't been easy. [...]
Earlier this year, we (or rather Serge) finally got the latest version of the
Desktop running on top of a Debian Trixie image. It's now based on 64-bit Debian
(the amd64 architecture) rather than the older 32-bit version, as Debian itself
has stopped supporting 32-bit for PC architectures. This shouldn't be a major
problem – most PCs made in the last 15 years or so will quite happily run the
64-bit version of Debian, as will most Intel-based Macs. (Debian support for
Apple Silicon is still experimental, so unfortunately those of you with the
latest and greatest shiny fruit products will not be able to run this.)
AI Skeptics: How Schools Procure AI Tools (with J.B. Branch)
Math Babe
mathbabe.org
2026-10-05 09:11:24
We were joined this week by J.B. Branch, the Director of Federal AI Governance and Technology Policy for Public Citizen’s Congress Watch division: Apple Spotify YouTube...
In software development, we collect anti-patterns to recognize common traits that lead to poor outcomes in our software. I thought it would be helpful to do the same for software blogging, so I’ve catalogued the most frequent mistakes I see from beginner bloggers.
The most common mistake in software blogging, by far, is meandering. I constantly find myself several paragraphs into a post with no idea what the author is trying to tell me.
Developers love specificity, so they start blog posts with backstory, historical context, and whatever else happens to be on their minds. That may be fun to write, but it’s not always interesting to read.
From the reader’s perspective, there are a billion other articles they could be reading. Why should they read yours? They’re not going to invest 20 minutes to read it in full unless they expect a payoff. Give the reader a reason to continue reading.
When a developer begins reading an blog post, they’re trying to answer two questions as quickly as possible:
Did the author write this for someone like me?
How will I benefit from reading it?
Give yourself the title and the first three sentences to answer both questions.
The benefit you offer can be teaching the reader a new skill, explaining a concept, illustrating a new perspective, or delivering an entertaining rant. You just have to offer the reader
something
. They’re not going to read your blog post just because it’s there.
if got, want: A Simple Way to Write Better Go Tests
There’s an excellent Go testing pattern that too few people know. I can teach it to you in 30 seconds.
The introduction succinctly communicates that the article is relevant to programmers who use the Go programming language, and the value is teaching them a new technique they can learn quickly.
Some bloggers write a compelling intro but clutter the reader’s path with extras like a subtitle, a bio, an image, or a famous quote. You can include any of these things, but recognize that they count against your “inspire the reader to keep reading” budget. Everything you put in the reader’s path is extra work that chips away at their finite supply of focus.
“The reader knows everything I know except this one thing”
🔗
Effective teachers compare new concepts to something the reader finds familiar. For example, if you were explaining
Jellyfin
, you might say, “Jellyfin is a streaming service like Netflix, except it’s open-source and private, so nobody monitors your viewing habits.” The tricky part is knowing what the reader finds familiar.
In this article, I’ll introduce Docker to developers who have never heard of it before.
Docker is simple. It’s nothing more than a slick frontend to Linux cgroups. Oh, you know jails in *BSD? Docker is the Linux version of that.
Lots of developers want to use Docker but don’t recognize terms like cgroups, jails, or *BSD. They might not even know what Linux is, especially if they’re seeking out an introduction to Docker.
Instead of assuming the reader has your exact body of knowledge, minimize your assumptions about the reader:
Docker is a tool for packaging your app so that it has a consistent, reproducible environment wherever it runs. Docker allows you to define your app’s environment and dependencies in human-readable text files. These files capture your app’s requirements, so you know exactly how it works even after years of tweaks by different teams.
When you write a blog post, think about your target reader. What do they know? Imagine a friend or teammate you know in real life. Write a list of terms they would recognize and terms they would not. Then, re-read your blog post, and whenever you encounter a technical term, think about whether your reference reader would understand it.
You’re describing the audience I had in mind, but I’ve never tried listing out what that audience knows. Comparing your list against the assumptions in my draft is pretty mind-blowing.
When’s the last time you read a book that directed you to stop reading, go buy a different book, read it in full, then continue your original book? Software bloggers do this all the time, though it’s more subtle.
Bloggers often want to mention a term the reader might not know, but they don’t feel like explaining it themselves. Instead, they slap a link on the term and think, “Problem solved!”
The problem is not solved because the reader doesn’t want to interrupt their flow and go read a whole different site just to understand one word.
Assign
firewall
rules to prevent external traffic from reaching your database.
The FreeBSD manual linked above is an excellent resource, but the chapter on firewalls chapter is about 20,000 words. When you link to such a wordy page, you dump a massive amount of work on the reader.
Instead of relying on a link to do your work for you, give the reader the minimum possible explanation to understand your article.
A
firewall
is a system that restricts how hosts and networks communicate with an app. You can increase your web app’s security by configuring firewall rules to only allow inbound requests to your database server when they originate from your app server.
By all means, link to useful resources, but make them a bonus rather than a prerequisite. Keep the reader on the page. Your target reader should be able to enjoy and understand your article from start to finish without clicking any links.
These days, everything is either a sequel or a reboot, including blog posts. I see a lot of blog posts that open like this:
In part one, we learned about quintuply linked lists and how they can 100x your daily LOC output. In today’s post, I’ll show you how
goto
statements let you scrunkmax (a term I invented in part one – remember?).
I hate to break it to you, but most readers have not read part one. If you assume your last article is fresh in the reader’s mind, they’ll think, “Oh, now there’s extra work to even
start
reading?”
It’s fine to refer to your previous posts, but don’t do it right out of the gate. When you do link to past posts, summarize what was relevant rather than force the reader to go back and read it in full.
If you’re writing about the hobby operating system you built from scratch, then sure, you probably need more than one blog post, but the vast majority of sequel posts could be standalone articles with like 3% more effort.
Beginner software bloggers suffer from a mass delusion that you have to write in a stiff, overly formal way for people to take you seriously:
Several static analysis tools were utilized by my teammates and myself throughout the duration of this project’s lifetime.
You’re not writing for 80-year-old executives at IBM in 1988. Your field is software development, one of the least pretentious white-collar jobs out there. The person reading your article is probably wearing pajamas and flip flops while eating from a bowl of cereal next to their keyboard. They don’t expect or want you to talk like a legal document.
Just write the way you talk.
We tried a few static analyzers on this project.
With so many developers delegating their writing to AI, software blogging is becoming bland and homogenous. Readers are hungry for writing with personality. Here’s a random sentence from Joel Spolsky,
the best software blogger of all time
:
All the kids who did great in high school writing pong games in BASIC for their Apple II would get to college, take CompSci 101, a data structures course, and when they hit the pointers business their brains would just totally explode, and the next thing you knew, they were majoring in Political Science because law school seemed like a better idea.
It’s not Spolsky’s best line, but it captures his style. It’s casual, personable, and unpretentious. It sounds like he’s telling a story to some friends at lunch. You can see the same style in the writing of
Kathy Sierra
,
Terence Eden
, and
Raymond Chen
. They’re not trying to sound smart–they’re just trying to sound like themselves, and that’s what readers enjoy.
The hardest part of software blogging is writing in a compelling way, so it’s frustrating to see so many software bloggers bungle the part that should be easy: making a basic webpage.
The worst mistake you can make for mobile readers is overflowing the screen so the reader has to scroll back and forth to read your article. Usually, it’s because you have an image or code snippet that insists on being desktop size and screws up the layout of the rest of the page.
Allowing the text to overflow the screen on mobile creates a miserable reading experience.
Desktop versions of Firefox and Chrome both have a mobile preview mode. Check your article with the mobile preview before you publish, and check for common rendering issues.
Don’t underestimate your mobile readers. According to my analytics, 25% of you are reading this page on your phones. On
my personal blog
, it’s 35%.
Choose a font color and family that are easy to read. Stop it with this dark gray text on a light gray background. Firefox and Chrome both have built-in tools that flag low-contrast text for you.
Firefox’s accessibility tool identifying low-contrast text
If you don’t feel like searching around for the perfect font, the Braille Institute has a free font called
Atkinson Hyperlegible
that’s particularly comfortable to read, even for readers with poor vision.
Give the reader a compelling reason to continue reading. Get to it within the title and the first three sentences of your blog post.
Common reasons: they want to hear an entertaining story, learn a useful technique, or understand a concept that’s relevant to them.
Question your assumptions about what the reader knows and does not know.
Think about what concepts and terms you expect the reader to recognize and re-read your article to make sure it matches those expectations.
The reader should be able to read your article from start to finish without clicking links or hovering for tooltips.
Links should allow the reader to explore topics more deeply, but they should be a bonus rather than a pre-requisite.
Avoid presenting your article as a follow-up to a previous article.
Assume most readers haven’t read your previous articles. Summarize what’s relevant for them to know rather than expecting the reader to go read all your prior posts.
Drop the formality. Write the way you speak in real life.
Test your articles in your browser’s mobile view.
Make sure that your text doesn’t overflow the screen and force the reader to scroll horizontally as they read.
Use browser testing tools to find low-contrast text that makes your article difficult to read.
“Not Quite How Developers Read” and “What the Reader Knows” illustrations by
Piotr Letachowicz
.
Security updates for Wednesday
Linux Weekly News
lwn.net
2026-10-07 09:08:53
Security updates have been issued by AlmaLinux (bind, dovecot, freerdp, kernel, mariadb-connector-c, mod_auth_openidc, nodejs22, nodejs:22, sudo, and vim), Debian (node-shell-quote, puma, rails, ruby-jwt, and suricata-update), Fedora (chromium, cockpit, flocq, freerdp, gappalib-coq, golang-x-mod, ht...
The Social Reckoning review – Aaron Sorkin’s jittery sequel with all-new evil puppet Zuckerberg
Guardian
www.theguardian.com
2026-10-07 09:00:04
Jeremy Strong’s slow-talking Facebook founder turns into a cameo, background to a thriller about a whistleblower that somehow never mentions Trump Here is a film for all those people who solemnly deplore social media in conversation but haven’t quite got round to deleting their accounts. Sixteen yea...
H
ere is a film for all those people who solemnly deplore social media in conversation but haven’t quite got round to deleting their accounts. Sixteen years have gone by since
The Social Network
, written by Aaron Sorkin and directed by David Fincher; this was the addictive, obsessive-compulsive movie about the birth of Facebook in Harvard’s macho-competitive cauldron of student sexism and nerdy self-pity, starring Jesse Eisenberg as the fast-talking, arrogant, not-necessarily-evil genius Mark Zuckerberg. Now Sorkin is back with a sequel of sorts, directing his own script – and now Zuckerberg is basically evil, but also weirdly slow talking, as if all the zinging, smart-alecky Sorkinesque dialogue energy of the first film has been siphoned away in early middle age and dispersed among the insurgents in the new cast.
Sorkin’s jittery, funny, syncopated style is unmistakable, as distinctive in its way as David Mamet, although he has one joke about “spatula” being a Yiddish word that he
has borrowed from Philip Roth’s Portnoy’s Complaint
. It is the story of how in 2021, Zuckerberg’s Facebook empire was challenged by
courageous whistleblower-employee Frances Haugen
and Wall Street Journal investigative reporter Jeff Horwitz, who were on a mission to tell the world how Zuckerberg, scared by falling revenue, changed the algorithm so that engagement was supercharged by rage bait and hate speech, malice and envy. Facebook, moreover, didn’t factcheck political ads, thus driving a huge, nationwide surge in extremist groups and teen bullying and becoming a major player in the Capitol riots. (Having said that, this film does not say aloud the word “Trump” and rather cravenly declines to specify Maga activity.)
Mikey Madison sympathetically plays the unhappy, principled, supersmart Haugen, who is dismissed as a “disrupter” – and Sorkin leaves it to us to notice that this in other circumstances is what oligarchs smugly call themselves.
Jeremy Allen White
does a capable, unflashy, unrisky job as Horwitz, a stolid performance that perhaps, in sync with the film, tells us about the diligent professionalism of old-school print journalism. And Zuckerberg? Well, we were all hoping and expecting it would be Jesse Eisenberg again. (You’ve heard of Richard Linklater’s Boyhood; this would be Aaron Sorkin’s Techbrohood.) But no: it is Jeremy Strong, who does a very interesting Zuckerberg impersonation, standing and sitting like a Thunderbird puppet, modifying the torpid sadness of his Kendall Roy in TV’s Succession, making this the unhurried contempt of someone simultaneously burdened and energised by knowing he’s richer and therefore smarter than everyone else. (A famous line in The Social Network was “A million dollars isn’t cool. You know what’s cool? A billion dollars.” This film starts by telling us that Facebook is worth a trillion dollars.)
Zuckerberg, though, is almost a cameo, with no equals with whom he can exchange dialogue, only employees – although the film imagines one former investor-turned-critic called Charlie (Bill Burr) who heckles Zuckerberg’s practice performance before a confessional hearing, with Zuckerberg amusingly having to be told how to smile. Charlie is the angry, unheeded voice of conscience.
Rattles along at an entertaining clip … L to r, Jeremy Allen White as Jeff Horwitz and Mikey Madison as Frances Haugen in The Social Reckoning.
The resulting drama rattles along at an entertaining clip, although there are false notes. When Frances has secret meetings with Horwitz in a hotel, the film has her wearing a slinky evening dress, supposedly as a kind of disguise, so that if recognised it would seem like she was out on a date. Hmm, really? Or is it so that the movie’s female lead has scenes where she looks glamorous? Later, Horwitz affectionately tells her to be less “mean”. But she is not “mean”. She just speaks Sorkinspeak. So does Horwitz. The ubiquity of Sorkinspeak effaces the difference between all the characters.
The most disturbing scene comes at the very beginning, showing a drunk woman livestreaming her racist abuse of traffic cops; this is the woman whom Haugen recognises as an acquaintance of hers, radicalised by
Facebook
. Towards the end we glimpse her again, looking ambiguously appalled by what she has become. But would she be appalled in reality? Or is this just the film-makers’ wish fulfilment? The closing-credit titles – usually the platform for a film to promote and summarise its heroes’ achievements – concedes that Facebook still isn’t checking political ads. Plus, of course, the film is set at the start of the Trump interregnum of 2020-24, when it seemed as if the Capitol riots were the ultimate wake-up call and moderation had returned. Nowadays Trump and his tech bro courtiers are more vainglorious than ever, and the values of investigative journalism have been thrown out at the Washington Post by its owner Jeff Bezos. This is a film which doesn’t reckon with the future we’re living in.
House with 15m underground tunnels for sale for 300k
Hidden for decades beneath a workshop in a suburb of Tilehurst in Reading - the network descends through four levels and reaches more than 15m below ground.
Tony, whose initial day job was in the construction of aerodrome and airport infrastructure, started the construction work by himself during the 1980s and ‘90s.
He was also assisted by his six children.
The only apparatus the family used to conduct the digging and move chalk were buckets and wheelbarrows.
Richard Worrall, auctioneer and consultant at BTG Eddisons Property Auctions, said: “In all my years as an auctioneer, I’ve never seen anything like this before - it’s utterly amazing and, I think, unique.
"What a feat to create this as a family project and as a hobby.”
The Deane family’s subterranean projects began even earlier, in the 1960s, with the construction of underground changing rooms next to the property's swimming pool.
The entrance was hidden behind a 23ft-high rocket, which is still in place today.
Despite the scale of the project, its existence is understood to have remained largely unknown outside the family, including to neighbours in the surrounding residential area.
Richard Worrall said: “Most of these tunnels were dug through chalk, which is naturally self-supporting, while, in some places, where required, structural concrete was used and cast by hand.
“Incredibly, absolutely everything was excavated manually.
"Mr Deane also constructed a railway through the tunnels to transport the excavated chalk, as well as building a four-layer, one-tonne goods lift to carry it to ground level.”
The site is accessed via a secure gated driveway and includes a 2,500 sq ft workshop arranged over three floors, parking, gardens, a large swimming pool and changing rooms.
Planning consent was granted by West Berkshire Council in December 2024 for a 5,000 sq ft, five-bedroom, two-storey house to be built on the site of the swimming pool, while conversion of the workshop has consent for conversion to a three-bedroom house.
Richard added: “The underground space is obviously a game changer to say the least and we believe it could offer an imaginative buyer a range of potential uses, anything from secure storage, a fallout shelter or an EMP-resistant data centre to a prepper’s paradise."
The property has a guide price of £300,000-£325,000 and will be included in BTG Eddisons Property Auctions’ live stream auction on 28-29 October.
A Walkable History of Art
Every artist, every work, under one roof.
Hackers exploit critical Atlassian flaw after public PoC release
Bleeping Computer
www.bleepingcomputer.com
2026-10-07 08:49:01
A critical vulnerability (CVE-2026-21589) affecting multiple Atlassian product families, including Jira, Confluence, and Bitbucket, is being exploited in attacks that do not require authentication. [...]...
A critical vulnerability (CVE-2026-21589) affecting multiple Atlassian product families, including Jira, Confluence, and Bitbucket, is being exploited in attacks that do not require authentication.
Earlier today, security company Previdian detected the activity on its honeypot network, just hours after a detailed technical report was published.
An unauthenticated attacker can exploit CVE-2026-21589 to access specific files in the application's web root directory if they know the file's exact name and path.
The issue is an arbitrary file-access flaw disclosed on Monday, and it affects self-hosted instances of eight Atlassian products:
Bitbucket Data Center
Confluence Data Center
Jira Service Management Data Center
Jira Software Data Center
Bamboo Data Center
Crowd Data Center
Crucible
Fisheye
In a security advisory on Monday, Atlassian
warned system administrators
managing self-hosted instances to apply the security updates as soon as possible, noting it cannot determine whether individual customer instances have been compromised.
Following Atlassian’s advisory, offensive security company
watchTowr published a technical report
showing how CVE-2026-21589 could be exploited to gain administrator-level access to Jira, Confluence, and Bitbucket in certain Crowd-integrated deployments.
Atlassian Crowd provides centralized identity management, single sign-on (SSO), authentication, authorization, and access management for connected Data Center apps.
The researchers exploited the root cause of the flaw, which is a shared web-resource library that converts double colons "::" into forward slashes "/", to construct directory-traversal requests through plugin resource endpoints and retrieve protected application files without authentication.
watchTowr researchers confirmed file reads in Jira, Confluence, and Bitbucket, but their technique could not traverse outside the Tomcat application context.
In Crowd-integrated Jira deployments, attackers could read plaintext application credentials from WEB-INF/classes/crowd.properties and use them to create a Jira administrator account through Crowd’s API.
This applies if Crowd is reachable and the application has sufficient permissions. However, the researchers note that specifying a list of allowed IP addresses would make exploitation significantly more difficult.
"[An attacker] would need to pivot through arbitrary machines or use SSRF-like capabilities of Jira, Confluence, or Bitbucket to reach Crowd directly," the researchers say.
The leaked credentials in crowd.properties provide administrator access to the Crowd identity management system, allowing attackers to create new users and modify permissions.
PoC showing escalation to admin on Jira
Source: watchTowr
According to Previdian, the technical details were sufficient to help threat actors scan for exposed vulnerable instances and probe them.
“Within two hours of watchTowr publishing its technical research and public PoC for CVE-2026-21589, Previdian's honeypot network began observing exploitation attempts targeting the vulnerability,” Previdian’s Ryan Dewhurst told BleepingComputer.
“A Nuclei template has also now been released, making it significantly easier to automate scanning for vulnerable systems.”
The company has so far observed attempts from these three IP addresses: 38.60.157[.]86, 146.70.187[.]234, and 159.26.119[.]225, and recommends blocking them.
Dewhurst expects exploitation activity to increase significantly over the coming days and weeks, driven by the rapid emergence of exploitation attempts following the public PoC, the availability of automated scanning templates, and the broad range of affected Atlassian products.
System administrators should apply available security updates as soon as possible, or apply the recommended mitigations.
This includes restricting external network access, adding a web application firewall (WAF) or proxy rule blocking specified traversal patterns across all affected products, Tomcat RewriteValve rules for Confluence, JSM, Jira, Bamboo, and Crowd, or a URL rewrite rule for Bitbucket.
For more details about the fixed versions and mitigation steps, check out
Atlassian’s bulletin
.
watchTowr has also
released a free scanner tool
to help administrators determine if their instances are vulnerable to CVE-2026-21589.
Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.
Kentaro Hayashi: Testing Japanese IMEs on virtual desktops with virt-japanese-desktop
PlanetDebian
kenhys.hatenablog.jp
2026-10-07 08:47:13
Japanese input method editors (IMEs) such as Mozc are essential for typing
Japanese (yes, skk, anthy and more, but it's out-of-scope in this article). Testing them is a surprisingly fiddly job: an IME behaves differently
depending on the desktop environment (GNOME, KDE, Xfce, Budgie), the input
fram...
Japanese input method editors (IMEs) such as
Mozc
are essential for typing
Japanese (yes, skk, anthy and more, but it's out-of-scope in this article). Testing them is a surprisingly fiddly job: an IME behaves differently
depending on the desktop environment (GNOME, KDE, Xfce, Budgie), the input
framework (ibus, fcitx5, uim), and the Wayland or X11 session. Setting up a
fresh VM for every combination from scratch is slow and repetitive, and
breaking your host's environment while experimenting is no fun.
virt-japanese-desktop
is a small set of toy scripts that solves exactly this
problem. It builds ready-to-use virtual machines - each with a desktop
environment and a Japanese IME already installed and configured - so you can
start typing Japanese within a minute of booting, and throw the whole thing
away without any impact on your host.
What it gives you
A single
make
command produces a qcow2 image that combines one of four
desktops with one of three IMEs:
ibus-mozc
fcitx5-mozc
uim-mozc
GNOME
✓
✓
✓
KDE
✓
✓
✓
Xfce
✓
✓
✓
Budgie
✓
✓
✗
The one gap -
uim on Budgie
- is a real technical limitation, not an
oversight: Budgie runs the labwc compositor, which implements
zwp_input_method_v2
, while
uim-wayland
only speaks the older
zwp_input_method_v1
(KWin/Weston). On Budgie you use ibus-mozc or fcitx5-mozc
instead.
All images come with the Japanese locale, fonts, and the
Asia/Tokyo
time
zone pre-configured, plus the SSH key of your choice. Input switching is wired
up out of the box:
Ctrl + Space
or
Shift + Space
for your IMEs to enable it.
The layered-image trick
The core idea is that the images are
stacked qcow2 overlay layers
, one on
top of another:
debian-sid-nocloud-amd64-daily.qcow2 ... base image you download
└─ unstable-japanese-template.qcow2 ... locale, fonts, user, SSH
└─ unstable-<DE>-template.qcow2 ... a desktop environment
└─ unstable-<DE>-<IME>.qcow2 ... desktop + IME, configured
└─ unstable-<DE>-<IME>.workspace.qcow2 ... the image you test in
This makes both
building and resetting fast
. The first build downloads and
customizes the base image and takes a while, but every later step only adds a
thin overlay. When a test breaks the VM, you do not rebuild anything - you
simply delete the top
workspace
layer and recreate it. That is all it takes
to return to a pristine state:
rm /tmp/unstable-gnome-ibus-mozc.workspace.qcow2
make gnome-ibus-mozc
The images reference their backing files by
relative path
, so you can move
an entire stack anywhere you like as long as the images stay together.
A platform for experimenting with bleeding-edge IMEs
The base system is Debian
unstable (sid)
, and the
experimental
repository is already added. That makes the project a convenient platform for
testing not just the IME packages in sid but also the ones still being
developed in
experimental
- exactly what the project was built for.
The keyboard layout of the VM follows the layout of your host (read from
/etc/default/keyboard
, with a fallback chain to
localectl
and finally
us
).
Quick start
sudo apt install qemu-utils libguestfs-tools virt-install curl
curl --location https://github.com/USERNAME.keys --output pubkey.pub
make download
make gnome-ibus-mozc
./scripts/make-virsh-image.sh virt-gnome-ibus-mozc /tmp/unstable-gnome-ibus-mozc.workspace.qcow2
Log in as
debian
(password
debian
) and press
Ctrl + Space
(
Shift + Space
) to start
typing Japanese. The VM is registered in libvirt, so you can manage it with
virsh
or
virt-manager
- handy for opening the SPICE console or restarting
the machine after a test.
Status
The project is still in the
proof-of-concept
phase, and it is developed
mainly to test Mozc and other IMEs across desktops and input frameworks.
If you regularly test Japanese IMEs - give it a try.
Three Years into Genocide, "The IDF Is Still Hunting Gazans": Palestinian Journalist Akram al-Satarri
Democracy Now!
www.democracynow.org
2026-10-07 08:46:55
In Gaza, October 7, 2023, is viewed as the beginning of the genocide. Over the past three years, Israel killed over 74,000 people, according to the Gaza Health Ministry, but many believe that number is a vast undercount. The killing continues despite a so-called ceasefire that took effect about a ye...
In Gaza, October 7, 2023, is viewed as the beginning of the genocide. Over the past three years, Israel killed over 74,000 people, according to the Gaza Health Ministry, but many believe that number is a vast undercount. The killing continues despite a so-called ceasefire that took effect about a year ago.
“Every single aspect of the life of the people of Gaza has changed once and forever,” since October 7, says Akram al-Satarri, a Gaza-based journalist. He points out that a large majority of structures in the Gaza Strip have been reduced to rubble and that Israel controls around 70% of Gaza. “The Israeli Air Force is still hovering all around different parts of the Gaza Strip, and the
IDF
is still hunting Gazans, disturbing the already disturbed life, spreading death and displacement once and once again.”
Guests
Please check back later for full transcript.
The original content of this program is licensed under a
Creative Commons Attribution-Noncommercial-No Derivative Works 3.0 United States License
. Please attribute legal copies of this work to democracynow.org. Some of the work(s) that this program incorporates, however, may be separately licensed. For further information or additional permissions, contact us.
Requires Linux x86-64. Both installation methods compile native code and require
Rust, a C compiler, clang with the BPF backend, bpftool, pkg-config, libelf and
zlib development files, and BTF information at
/sys/kernel/btf/vmlinux
.
Support on WSL2 is very limited, I recommend building procinsh from source rather than installing it from crates.io.
I haven't tested it in container environment such as Docker.
Then run and open
http://127.0.0.1:9090
in your browser. To allow remote
access, use
--allow-non-loopback
, but be careful: this exposes process memory
and environment variables without any authentication.
Also requires Node.js 22 or newer with npm and Rust via rustup. The Rust version
and components are pinned in
rust-toolchain.toml
; rustup installs them
automatically when needed.
Clone this repository and run the following from its root:
Ubuntu's
bpftool
wrapper may fail with
bpftool not found for kernel ...
because the WSL2 kernel version differs from the Ubuntu tools package. Set
BPFTOOL
to the packaged executable directly, bypassing the wrapper:
sudo apt-get update
sudo apt-get install --yes --no-install-recommends \
build-essential clang llvm pkg-config libelf-dev zlib1g-dev python3 \
linux-tools-common linux-tools-generic
fortoolin /usr/lib/linux-tools/*/bpftool;doif [ -x"$tool" ];thenexport BPFTOOL="$tool"breakfidone"${BPFTOOL:?No packaged bpftool found; install linux-tools-generic}" version
"$BPFTOOL" btf dump file /sys/kernel/btf/vmlinux format c >/tmp/procinsh-vmlinux.h
npm ci
npm run build:web
cargo build --release --locked
"Origins of a Genocide": Al Jazeera's Marwan Bishara on "Twelve Days That Shook the Middle East"
Democracy Now!
www.democracynow.org
2026-10-07 08:29:13
Twelve Days That Shook the Middle East: Origins of a Genocide is the new book by Al Jazeera senior political analyst Marwan Bishara discussing the crucial 12 days between Benjamin Netanyahu’s October 7 speech and Biden’s “rare” Oval Office speech on October 19. “October...
This is a rush transcript. Copy may not be in its final form.
AMY
GOODMAN
:
This is
Democracy Now!
, democracynow.org. I’m Amy Goodman.
Twelve Days That Shook the Middle East: Origins of a Genocide
. That’s the title of a new book by the Al Jazeera senior political analyst Marwan Bishara. In the opening pages, he writes, “Twelve days were enough. It took no more than twelve days, from Netanyahu’s October 7 war speech to Biden’s October 19 rare Oval Office speech, to expose the anatomy of Israel’s war, its intent, its strategy, and the method of its madness. The genocide was not a slip; it was the result of a predetermined strategy for which there existed a blueprint, if not a premeditated decision. … Hamas had set the process in motion on October 7, it lost control of that process the very same day.”
Marwan Bishara has been with Al Jazeera for more than 20 years, has written three books and written for many publications, joining us now from Paris, France.
Marwan, congratulations on your book. Thank you for joining us. If you can talk about that quote and say — you say Hamas set the process in motion October 7th and immediately lost control of that process. What do you mean?
MARWAN
BISHARA
:
Yes, absolutely. Clearly, the attack of October 7 was the trigger for what came after. And my argument is that we know about what came after, basically the past two or three years, from those very first 12 days after October 7. And it started with ideology. You know, genocides don’t start with bombs; they start with ideology. And in this case, the ideology has been decades in the making — a century, I would even say, in the making, in the sense of the establishment of a racist, supremacist system, regime, that had been in control of Palestinians and Palestinian life and destiny for the past 80 years, occupying them and dispossessing them and discriminating against them. This has been going on for many decades. So, that was ready, intact on the eve of October 7, October 8.
And then there was the dehumanization. Immediately after October 7, October 8, 9, 10, we saw this extraordinary campaign of dehumanization of the Palestinians as animals, Nazi-like. You cannot deal with them. They don’t understand. They are — they are primitive. They only understand the language of force and violence, and so on and so forth.
After that immediately came the hyperbole, the exaggeration of what exactly happened October 7. October 7 was bloody, and it was horrific. And your earlier guest testified to the fact that this was not easy for hundreds of Israeli civilians. But Netanyahu and his government had to go overboard, to exaggerate, to lie, to deceive about what exactly happened, and their role in it, as well, invoking the so-called Hannibal doctrine, where they themselves joined the bombing of those involved in the conflict during October 7.
AMY
GOODMAN
:
Marwan, explain the Hannibal doctrine.
MARWAN
BISHARA
:
So, the Hannibal doctrine was something that was invoked in the early 1980s by the Israeli military, saying we cannot afford to have Israeli soldiers being captured, because then we’re going to have to negotiate their release, and that will cost us Palestinian — or, in that case, Lebanese — prisoners in Israeli jails. So, because they did not want to release Palestinian or Lebanese prisoners, they decided that if any one Israeli soldier — later on, civilian — is captured by any kidnappers, that the Israeli military then will take extraordinary measure to kill the operation, including, if it need be, to kill the soldier, the Israeli soldier, or the Israeli civilian.
So, on October 7, immediately after the war — after the operation started, in the early hours of the morning of October 7, apparently, the Israeli military invoked the Hannibal doctrine. And Israeli aircrafts, jets, you know, bombers, and tanks started carrying attacks against suspected conflict zones, whether it’s kibbutz or settlements, or whether it is convoys going back to Gaza. So, the invocation of that, of that controversial military order, clearly led to the death of perhaps hundreds of Israeli civilians and many, as well, among the Palestinians, whether it’s Hamas or any of those Palestinians who came into Gaza.
All of that, you know, taking that in consideration, what happened soon after that is the worst part of it, right? The establishment of a war cabinet, the declaration of war, the deployment of 100,000 soldiers on the border with Gaza, the recall of 350,000 reservists from the Israeli army, and setting up a strategy under the rubric of lethality, which has been developed over many years of wars in Gaza. And lethality is the one that involved mass destruction, not discriminant destruction, that involved the use of the so-called Lavender system, Gospel system, “Where’s My Daddy — Your Daddy?” system, and so on and so forth, that would actually lead to massive devastation of Gaza.
And that’s where the intention came. They’ve said it in press conferences, in war rooms. They’ve said it in public, in declarations, ministers and opposition leaders, about why there is no innocents in Gaza and why “we’re going to have to bomb all of them.” Some of them spoke about nuking them. Some of them spoke about “Nothing is going to be left standing in Gaza. We’re going to turn it into a deserted strip.” And all of these things were then repeated by the soldiers, not only as slogans, but even as songs — right? — repeated by the soldiers. So, the orders went down the hierarchy, from the government ministers to the army, and was implemented by the generals, who boasted of bombing Gaza back to the Stone Age. And they certainly tried to do that.
AMY
GOODMAN
:
If you can talk about what is the power that the U.S. has over Israel? You know, many said, for example, with the Iran war, that Israel brought the U.S. into it. But do you see it in that way? And what about Jared Kushner, President Trump’s son-in-law, just visiting with the Israeli Prime Minister Netanyahu, also meeting directly with Hamas, but not meeting with the head of the PA, the Palestinian Authority? And if you can talk about where all of this stands right now? As people stood with a moment of silence on this October 7th, people at the music festival, Israelis mourning the dead, they could hear the boom of Israeli airstrikes continuing in Gaza. Talk about the Board of Peace. Talk about the Palestinian government that was set up in exile in Cairo, still not allowed to come into Gaza by Israel.
MARWAN
BISHARA
:
So, let’s start with the most relevant to the issue at hand, which is the Biden administration first, right? Because nowadays, anywhere you look, in
The Atlantic
, where you are,
The Atlantic
magazine, or
The New Yorker
or
The New York Times
, or even, you know, hosted by our brilliant or your brilliant anchor, Jon Stewart, you have these former Biden administration officials trying to whitewash their record. They lie, they deceive, and they continue to claim no responsibility for the genocide that actually took place in Gaza over the past three years.
But they are directly responsible. Not only they aided and abetted; they were part of the operation taking place in Gaza the past two years, part of the war. They were in the war cabinet meetings. President Biden himself attended the war cabinet meetings on October 18, back in 2023. Lloyd Austin, the secretary of defense, attended war meetings. Secretary Blinken attended war meetings.
So, when you watch people like Jake Sullivan, the former national security adviser, sitting on Jon Stewart’s show, talking about how they tried their best to have a ceasefire, and they really put pressure on the Israeli government, and they succeeded, you know, 15 months later, it’s all humbug. It’s all BS. They were part and parcel.
The banality of evil wasn’t in Israel; it was in Washington, because the kind of support that Washington provided to Israel — diplomatic support, financial support, military support — they were talking about “We were trying to help the Palestinians,” while actually giving Israel the arms to shoot the Palestinians, before sending them the Band-Aids to take care of their injuries.
So, Washington, in that sense, in the sense of the Biden administration and its officials, were part and parcel, were complicit, were cooperating, coordinating the war effort that actually was a genocidal war in Gaza. We really must underline that, because we are seeing history rewritten as we speak. In the U.S. publications, as I remember it, after the 2003 war, when many of the then the Bush administration tried to whitewash their crimes, and the journalists who supported that war, just like those who supported the war in Gaza are now trying to whitewash the record.
AMY
GOODMAN
:
I wanted —
MARWAN
BISHARA
:
It’s very important to underscore — yes.
AMY
GOODMAN
:
I wanted to go to the former top State Department official under President Biden admitting a genocide has taken place in Gaza. Wendy Sherman served as deputy secretary of state until her retirement in July 2023. She spoke on a Bloomberg podcast.
WENDY
SHERMAN
:
I think that it is critical that Israel remain an ally of the United States and that we protect the right of a Jewish state. But I also believe that the prime minister has led us down a road — and we have been part of it — that has, in essence, created a genocide in Gaza.
AMY
GOODMAN
:
And then there’s former State Department spokesperson Matt Miller, who spent more than a year as the face of the Biden administration’s foreign policy and repeatedly defended Israel against allegations of war crimes and genocide in Gaza. He later admitted on a podcast interview with Sky News that, yes, Israel committed war crimes.
MATTHEW
MILLER
:
I don’t think it’s a genocide, but I think the — I think it is, without a doubt, true that Israel has committed war crimes.
MARK
STONE
:
You wouldn’t have said that at the podium.
MATTHEW
MILLER
:
Yeah, look, because I — I mean, when you’re at the podium, you’re not expressing your personal opinion. You’re expressing the conclusions of the United States government.
AMY
GOODMAN
:
And then you have Brett McGurk, who’s just come out with his new book. It’s called
Brink
. The headline in
The New York Times
, “'Ready to Blow His Stack': How Biden Nearly Cut Off Netanyahu Over Gaza,” and that’s talking about what McGurk is saying in his book. Marwan Bishara, take it —
MARWAN
BISHARA
:
Yes, absolutely.
AMY
GOODMAN
:
Take it through to the Trump administration now.
MARWAN
BISHARA
:
Yes, there’s no doubt that, even according to McGurk and Sullivan, that they paved the way for the Trump administration to take over the war effort, that they paved the way in the worst possible manner for the Trump administration then to take it on and to do whatever it wants, including the weaponization of hunger and the weaponization of the food aid to the Palestinian people. I mean, we sunk into a whole new abyss with the Trump administration.
During the Biden administration, the United States was basically complicit in genocide. During the Trump administration, on the very first few days when it took over, they started talking about ethnically cleansing Gaza, out in the open. President Trump spoke of ethnically cleansing Gaza, the entire Gaza Strip, in order to have a beachside property called Trumpland or whatever, or called, you know, the New Riviera in the East. So, the cynicism that the Trump administration built in its first year with the Gaza issue, taking up the mantle from the Biden administration, that was really, you know, just so depressing to watch, in the sense that from one administration to another administration, we see complicity with genocide.
We see the United States aiding an Israeli government, protecting an Israeli government, and instead of having them stand up for the responsibility for the war crimes, they attack the International Criminal Court and the International Court of Justice, which is unprecedented for an American president, for the United States, to be so hostile to international law, international legality, international institutions, the United Nations, even the United Nations, you know, relief agency, like
UNRWA
, and so on and so forth, that were indispensable for the people in Gaza when the policy was to starve them, to bomb their bakeries and schools and universities, to turn their schoolyards into graveyards. You at least needed something like
UNRWA
. But, no,
UNRWA
has to be demonized.
UNRWA
has to be cut down to size.
UNRWA
has to be blocked and to be, again, demonized by not only Israel, but also by the United States.
So, the Trump administration did everything possible in order to achieve some sort of a Israeli victory and a Hamas surrender, and forced the United Nations Security Council to adopt a resolution which, to my mind, Amy, is illegal. And I say that not being a legal scholar, but certainly being a student of history. The United Nations Security Council, under pressure from the Trump administration, and because the genocide was going on and they had to do something, they adopted a resolution that recognized the Board of Peace, without this Board of Peace ever being defined, without knowing zero about what is this Board of Peace. So, they adopted the resolution, and this allowed the Trump administration then to go on and to define this Board of Peace as it wishes, including appointing President Trump as a, you know, life member that heads it, and Netanyahu, of course, and Putin as members of the Board of Peace. So, you can imagine the surrealism — right? — behind this kind of an effort in order to turn a two-year genocide into a mockery of international law.
AMY
GOODMAN
:
Marwan Bishara, there is so much to talk about and so little time, and we want to get directly on the ground to Gaza. So I want to thank you for being with us, Al Jazeera’s senior political analyst, author of the new book
Twelve Days That Shook the Middle East: Origins of a Genocide
.
Coming up, we go to Gaza City to speak with journalist Akram al-Satarri. Stay with us.
[break]
AMY
GOODMAN
:
“Khatar,” “Danger,” by the Palestinian oud musician Huda Asfour, performing in our
Democracy Now!
studio.
The original content of this program is licensed under a
Creative Commons Attribution-Noncommercial-No Derivative Works 3.0 United States License
. Please attribute legal copies of this work to democracynow.org. Some of the work(s) that this program incorporates, however, may be separately licensed. For further information or additional permissions, contact us.
If you can think it, you can play it. Create, play, and share custom games — no coding required.
Software Engineer, AI Innovation + Research
Listen to article
[[duration]] minutes
This content is generated by Google AI. Generative AI is experimental
Games are one of the most creative forms of expression, and the gaming community keeps growing. But game creation has historically been limited to a select group of people who have technical experience and knowledge of complex game engines.
Today, we’re introducing Playground, an experimental gaming platform that lets you create, play, and share custom games. Playground lowers the barriers to creation, so bringing game ideas to life can now be as easy as describing it through a few prompts — no coding experience required.
Create your custom game
With Playground’s conversational interface, you can direct the game just by typing what you want to build. Start from a blank canvas, remix starter prompts, or use guided support to shape your concept — and then test it right away. Want to tweak the physics, rewrite the rules, customize the characters, or overhaul the environment? Just ask. You stay in complete control of the creative direction from rough concept to final product.
Create games in Playground using text prompts
Play, compete, and share across devices
The fun doesn’t stop at creation. Playground is home to an ever-growing catalog of community games. And because it’s browser-based, you can play these games on your phone or laptop.
Once your game is ready, you can keep your game private, drop a shareable link to challenge friends and family, or publish it to the Playground Explore gallery for the broader Playground community. To have fun with your friends, Playground games support in-game leaderboards and multiplayer gameplay in select genres. With a Play Games profile, you can claim a custom handle, like your favorite games, battle for the top spot, and follow creators to catch their latest releases.
Our gallery uses player ratings and play activity to spotlight games that are fresh, creative, and fun. To keep the experience safe and welcoming, every published game undergoes safety screenings aligned with our
Community Guidelines
, along with user reporting.
Explore the Playground gallery
Stay tuned for Unity Spark
For dedicated creators ready to expand their concepts, Playground will soon offer an integration with Unity Spark. Unity Spark unlocks robust, professional-level mechanics, high-fidelity 3D capabilities, and the flexibility of the Unity runtime. The accessibility of Playground and the world-class, game-building expertise of Unity provides the building blocks for an ecosystem where you can start with a simple text prompt and eventually grow into building immersive experiences with Unity. And these games will be available for everyone using Playground to enjoy.
Unity Spark is currently in testing, with closed beta coming soon. Check out the demo games on
playground.google
, and learn more at
unity.com/spark
.
Get started today
Playground is an early experiment designed to make game creation more approachable and ignite everyday creativity. Playground will continue to evolve alongside our community, and we invite you to share your games and your feedback. Playground launches today for users in the U.S. (18+) at playground.google, with tiered creation access rolling out based on your Google One membership.
Get the latest news from Google in your inbox
Sign up for our newsletters with product updates, event information, special offers, and more.
Israeli Activist Maoz Inon Lost His Parents on Oct. 7, Decries Israel's Ethnic Cleansing in Gaza
Democracy Now!
www.democracynow.org
2026-10-07 08:14:05
Israelis and Palestinians are marking three years since October 7, 2023, when Hamas launched a surprise attack on Israel. Nearly 1,200 Israelis were killed, and around 250 were taken hostage. It was the deadliest such attack in Israel’s history, after which Israel immediately began bombing the...
This is a rush transcript. Copy may not be in its final form.
AMY
GOODMAN
:
Israelis and Palestinians are marking three years since October 7th, 2023, when Hamas and allied groups launched a surprise attack on Israel. Almost 1,200 people were killed in the attack. About 250 were taken hostage. It was the deadliest such attack in Israel’s history.
In Gaza, October 7th is viewed as the beginning of the genocide. Over the past three years, Israel has killed over 74,000 people in Gaza, according to its Health Ministry, but many believe that number is a vast undercount. The killing continues despite a so-called ceasefire. Most of Gaza remains in rubble.
We’ll go to Gaza later in the program, but we begin in Israel, where a moment of silence was held to remember the victims of the October 7th attack in southern Israel. Mourners gathered near Kibbutz Reim to commemorate hundreds of people who were killed and abducted at the Nova music festival. During a moment of silence, the booms of Israeli airstrikes on the nearby Gaza Strip could be heard ringing out.
We’re joined now by the Israeli peace activist Maoz Inon. He lost both of his parents, Bilha and Yakovi Inon, in the Hamas attack. His parents lived on a kibbutz, a farming collective, just north of the Gaza border. They were 78 and 76 years old. In the aftermath of the attack, Inon called for peace and reconciliation rather than revenge. Maoz Inon has spent much of the past year — of the past three years working closely with Palestinian peace activist Aziz Abu Sarah, whose brother died in 1990 from injuries from his time in an Israeli prison. Last month, they addressed the United Nations Security Council. This is Maoz Inon.
MAOZ
INON
:
To build a different future, we must face — we must face reality. Hamas committed atrocious war crime on October 7, killing and abducting innocent Israeli. At the same time, the Israeli government exploited these atrocities to wage a devastating war of revenge. They have inflicted catastrophic destruction across Gaza, while accelerating ethnic cleansing, settlement expansion and forced displacement in the West Bank. … They weaponized our pain. But our grief must not be an excuse for war; it must be the catalyst to end the bloodshed.
AZIZ
ABU
SARAH
:
We believe this is not Israelis versus Palestinians; it is a choice between those who believe in peace, justice and equality, and those who don’t. … We call for targeted sanctions on all individuals and institutions that support violence. Apply sanctions without partiality.
AMY
GOODMAN
:
That was Palestinian peace activist Aziz Abu Sarah and Israeli peace activist Maoz Inon addressing the United Nations just a few weeks ago. They co-authored the book
The Future Is Peace: A Shared Journey Across the Holy Land
.
Just this morning, on the anniversary of October 7th, the anniversary of his parents’ deaths, Maoz Inon gathered with family at the ruins of his parents’ home at their kibbutz. He sent us this clip.
MAOZ
INON
:
This is what remains from my parents’ house: only the foundation. The entire house was burned to ashes. We put a memorial plaque for them. It’s on the safe room, the only — the only thing that remain.
AMY
GOODMAN
:
Maoz Inon joins us now from Ashkelon.
Welcome back to
Democracy Now!
Thank you so much for joining us on this very difficult day for you and your family and so many, both of your parents killed on October 7, 2023. Your thoughts today, Maoz?
MAOZ
INON
:
First, Amy, I would like to thank you personally and the entire team at
Democracy Now!
for supporting, for being there for me and for so many peacemakers, both in Israel and Palestine. And this maybe was what motivated me the most to join today in this difficult day, just to thank you from the bottom of my heart for being there with us in the devastating three years that we’ve been going through.
And this is exactly the thoughts that’s going through me, and with the tears and all the emotion that are coming, that we deserve a different reality. We must stop the sounds of bombs and start hearing the sounds of children’s playing. And we all deserve it. If it’s Palestinian kids or Israelis, Israeli kids, they all deserve to live within security, within his dignity, and to have a home, a safe place to go to bed at night and to wake up in the morning with a smile, and to be with their beloving and supporting family. And we must work as hard as we can to achieve it as soon as possible.
AMY
GOODMAN
:
If you can talk about what happened in those days after you were in mourning with your family and you get this note from Aziz Abu Sarah, who also lost his brother in that — after brutality in an Israeli prison, and what that meant to you, and the journey you’ve been on since?
MAOZ
INON
:
Yeah, so, two days after October 7th, me and my siblings were just sitting and crying and holding hand. And my young brother told us we must make a family decision based on moral clarity that we are rejecting revenge, that we don’t want to venge — to avenge the death of our parents, and explain that it’s not going to bring them back to life. It will only escalate the cycle of bloodshed we Israelis and Palestinians been trapped within for a century.
And the day after, as I was waking up crying in bed, I saw a message from Aziz, whom I met only briefly in 2014, maybe for 10 minutes, in Jerusalem. And on his message, Aziz shared his — offered his condolences, and he says that he’s standing with me and with my family in this tragic moment. And it was literally — it’s not a metaphor. It was literally like a hand reaching out, saving me from drowning, drowning into the ocean of sorrow and pain. I was literally drowning. And Aziz, it was my savior. He was my savior.
And in the last three years working together, co-authoring
The Future Is Peace
, starting our own
NGO
, InterAct International, I can now say — and I’m so proud to say it again and again — that, yes, I lost not only my parents, I lost many of the hundreds of numbers with — like, for you and for our viewers, there are a number for me. There are childhood friends. There are people I knew my entire life that lost their life on October 7th. So, I lost so much, but I won Aziz as a brother. And this brotherhood is what keep me going, keep us going to achieve a better future, not just for ourselves, but to our people.
AMY
GOODMAN
:
In 2024, you and the Palestinian peace activist you’re talking about, Aziz Abu Sarah, met with Pope Francis, and then, in 2025, you met with Pope Leo. What did you say to him? And what are you calling for? What are you saying should happen right now?
MAOZ
INON
:
Yeah. So, we shared with him how we channeled our pain and suffering and loss to create a dialogue to achieve a better future in the Holy Land. And Pope Francis asked us, “How can I help you? How can I be your ambassador?” And we asked him to amplify our messages, and that the question should not be “if,” but “when,” when the Israeli-Palestinian conflict will end. And we asked him to be our ambassador, our ambassador to the G7 leadership gathering that was few weeks after in southern Italy.
And he went there as our ambassador, and he was able to achieve, in the official communiqué of the G7 leader, the language of the Israeli-Palestinian civil society, language that the G7 countries must support the civil society in Israel and Palestine to achieve a lasting peace and to invest in reconciliation. Based on that, the EU created a peace fund with 18 million euro, that is now supporting the civil society. And few months ago, U.K., Canada and Australia initiated another peace fund.
And we are calling the G7, we are calling, of course, the U.S. and the EU to invest in peace, to invest in reconciliation and in the civil society in Israel and Palestine, and while, in one hand, investing and supporting peacemaker, in the other — in the other hand, those who choose ethnical cleansing, those who choose bloodshed, those who choose destruction must be sanctioned. They must be accountable for their wrongdoing. And this is basically what Aziz and I and so many others been advocating for the last three years.
AMY
GOODMAN
:
I want to turn to Aziz Abu Sarah. This is a message he posted on Instagram yesterday.
AZIZ
ABU
SARAH
:
I woke up this morning to the horrible news that Israel demolished the house of my niece, Razan. She, her husband, her children and nine other families in East Jerusalem were rendered homeless. Imagine having to work all your life, as Razan has done with her husband, so hard, save every penny, so you can finally have an apartment you own, and then it’s all gone. Just a few days ago, Bezalel Smotrich, an Israeli minister, said that the way to get rid of Palestinians is to deprive them from hope. And this is exactly what the government did today, is to deprive my family from hope. Ben-Gvir himself came with the demolition party, with the Israeli police, so he can be there to witness it and to participate in demolishing the house. But you know what? This will not succeed.
AMY
GOODMAN
:
So, that was Aziz Abu Sarah, your Palestinian peace partner as you travel the world. What more do you know about what’s happening, what’s happened to his family, and what’s happening in the West Bank? We just got this report: Armed settlers, Israeli settlers, beat a photographer in the presence of Israeli soldiers, hospitalizing him, as he covered worsening settler violence against the olive farmers in the West Bank. Jaafar Ashtiyeh, 58-year-old veteran photographer with
AFP
— that’s Agence France-Presse — was assaulted by a group of settlers. They beat him with their rifle butts, despite him being marked as press, assaulted in full view of about a dozen Israeli soldiers. Talk about the situation. Also, Netanyahu, the Israeli prime minister, addressing the nation today and what his message was? Can you hear me, Maoz?
MAOZ
INON
:
No, I can’t. So, OK. First, I’m not listening to Benjamin Netanyahu. He is not my leader. He is not my prime minister. He is accused for being a war criminal. And I think he should clear his name at the
ICC
and
ICJ
. It’s not for me to talk about him.
And it’s exactly what’s happening nowadays, and including to Aziz’s niece — it’s exactly what we said at the U.N. Security Council: that the Israeli government exploited the trauma, the pain, the fear of the Israelis after the October 7th massacre to change the status quo from conflict management to ethnical cleansing. And it’s happening in the West Bank, it’s happening in East Jerusalem, and it’s happening in now more than 70% of the Gaza Strip.
And this is why we are calling for international intervention, because they are doing it. The bulldozers and the assault rifles, they are all — many of them are manufactured by American companies. So, we are calling: Where is the international law will be applied against Israeli terrorists? It doesn’t matter if they are Cabinet members, they are prime ministers or
IDF
soldiers. And this is exactly why the work we are doing is — must be amplified, and not to give hope, not to give hope about the Israelis and Palestinians, but to act. We have enough empty statements. We have enough prayers. We have enough world leaders that cross their fingers for us. But crossing fingers is not enough. We need action, and we need to see them on the ground.
AMY
GOODMAN
:
Maoz Inon, I want to thank you for being with us, Israeli peace activist and co-author with Palestinian peace activist Aziz Abu Sarah, author of the book
The Future Is Peace: A Shared Journey Across the Holy Land
. Maoz Inon lost both his parents, Bilha and Yakovi Inon, in the attack on October 7, 2023. He spoke to us from Ashkelon.
Coming up, Marwan Bishara, author of the new book
Twelve Days That Shook the Middle East: Origins of a Genocide
. And then we go directly to Gaza. Stay with us.
[break]
AMY
GOODMAN
:
“Bear Witness, O World,” performed by the New York City Palestinian Youth Choir.
The original content of this program is licensed under a
Creative Commons Attribution-Noncommercial-No Derivative Works 3.0 United States License
. Please attribute legal copies of this work to democracynow.org. Some of the work(s) that this program incorporates, however, may be separately licensed. For further information or additional permissions, contact us.
Numbers from a 50-passage, ~1,300-question benchmark:
Cross-family matrix
(readers = foreign models answering from GLM-5.3-Flash’s records; writers = GLM-5.3-Flash answering from theirs)
1
Savings shown use the lowercase instruction, on each provider’s own meter. Without that word, models write cablese in ALL CAPS and the styling costs 14–19 points: gemma 25.0%, qwen 29.8%, GLM 33.9%. *gpt-5-mini cannot disable reasoning, and compression makes it think — its writes bill about double plain. The one family where this technique does not pay.
:
Model
Role
plaintext acc
cablese acc
recovery ratio
(1.00 = plaintext control)
token savings
gemma-4-31b
reader
70.9%
77.5%
1.09
—
qwen3.8-27b
reader
72.3%
79.4%
1.10
—
nemotron-3-120b
reader
76.1%
76.7%
1.01
—
gemma-4-26b
reader
71.8%
76.9%
1.07
—
gemma-4-31b
writer
—
—
1.09
40.4%
qwen3.8-27b
writer
—
—
1.10
48.9%
gpt-5-mini
writer
—
—
0.99
17.7%*
GLM-5.3-Flash itself:
48.4% savings
with the lowercase instruction, in-family recovery
1.09
.
2
Readers answer 694 anchored questions per condition; writers’ records are read by GLM-5.3-Flash — every GLM call in every run in this post is glm-5.3-flash, the fast variant on the z.ai coding plan.
No comparison in the matrix favors plaintext; every ratio sits at 0.99–1.10.
The condition ladder
— same questions, one variable at a time:
Condition
What the answering model gets / does
Anchored accuracy
vs plaintext control
passage, answers in plaintext
the full passage
91.5%
1.00
(reference)
passage, answers in cablese
full passage, terse register
74.4%
0.81
└ + transcribed to plaintext
those terse answers, expanded
78.8%
0.86
plaintext record
a plaintext summary
75.8%
1.00
(reference)
cablese record
the compressed summary
82.5%
1.09
decoded record
cablese record, expanded back
81.6%
1.08
Downstream models,
including four families never shown an example
, answer questions from the compressed records as well as or better than from plaintext — recovery ratios
0.99–1.10
, where 1.00 means “exactly as good as plaintext.” The register is not a construct we invented (LLMs were handed compressed records cold and read them at parity); the capability was already in the weights, inherited from a century and a half of people writing under metered bandwidth.
The retention result, interpreted.
Cablese records read slightly better than regular responses (the ladder’s 1.09), most likely because compression suppresses copy-the-record phrasing that strict grading penalizes.
Every model tested can do this.
It’s in their training data — but compression varies from 25% to 49% under identical instructions. That per-model spread is worth knowing if you want to use this technique, and it is knowable.
There is a slight information reconstruction cost, but only when transcribing individual answers for users. If the consumer is a model — or the record is simply re-read — the readback ratios are the ones that apply, and they’re all ≥0.99.
3
Statistically: no comparison favors plaintext; in-family read-back p<0.0001, decoded-record p=0.03.
The storage loop’s economics.
Expansion gives back about a third of the savings, so a human-readable expanded archive still costs well under plaintext.
Notice what this adds up to:
No new hardware, no training, no API change, one sentence of instruction. Either your agents get nearly twice the context at the compressed end, or the same work for half the output bill. It is a large optimization that is verified, and as far as I can tell, just not happening.
When is compression free?
It depends on where the compression step sits. Compress while the model is still composing (answers in cablese, 0.81, partly recoverable at 0.86) and you pay a register tax. But
if you compress after the content is settled, into a record a machine will read (1.08–1.09), there is nothing to pay
.
Cablese belongs between “content settled” and “machine consumption”: scratchpads, memory stores, agent-to-agent handoffs. Nowhere a model is still mid-thought.
Models whose reasoning you can disable bill the compression clean. gpt-5-mini’s reasoning is mandatory (the API refuses to turn it off), and cablese makes it think roughly 3× harder — its writes cost double despite the shorter text. The lesson is that
you must compress with models whose thinking you can control
.
So:
The expensive half of every LLM bill is optional overhead for machine-consumed text.
Output tokens cost 3–5× input tokens, and for traffic a model writes
for another model
, half of it comes off at any major API — one sentence of instruction, no training, no setup (mandatory-reasoning models excepted).
Agent memory just got nearly twice as capacious.
Store scratchpads and summaries in cablese: models — the intended readers — read them at parity or better, in every family tested. Expand back to plaintext only when a human actually looks, and even then the archive costs less than plain.
Model choice matters as much as the instruction.
The identical instruction compresses gemma’s records by 40% and Qwen’s by 49% on their own meters — compressibility is a measurable, per-model property, and at scale that spread is real money. The Telegraph Test is a way of determining how good models are at this kind of compression.
It probably can’t be walled off.
The register lives in the training data of every family tested. A lab that suppresses it in a frontier model just moves the advantage to open models that still carry it.
Auditability survives the compression.
Unlike the emergent agent protocols, cablese is human-readable, fixed by convention, and decodable on demand — the tokens shrink without losing the audit trail.
The Telegraph era, where every word was metered
In 1866, sending a message across the new transatlantic cable cost
$100 — for ten words
4
$10 a word, ten-word minimum,
roughly $2,600 in today’s money
.
. That’s real money today; it was serious money then. Telegraph companies charged
per word
, and an entire industry grew out of that price structure.
The Great Eastern laying the first successful Atlantic cable (oil painting, National Maritime Museum, public domain). The link that charged $10 a word — and taught a generation of correspondents to write in cablese.
Two compression strategies emerged, and they map onto two very different technologies:
Codebooks.
Publishers sold massive commercial code dictionaries — Bentley’s ABC Telegraphic Code ran to a thousand pages mapping entire business phrases to single code words.
…might mean
*steamship arrived, cargo intact, remit balance.*
Firms could cut a 40-word negotiation to a 6-word coded exchange. The codebook is a
substitution
technology: the message exists in full prose, and a dictionary renames chunks of it.
Cablese.
The operators and correspondents evolved a written format like:
This disciplined style saved words without losing meaning, and it was emergent from the cost of the medium.
Over time, this format faded as new technologies like the telephone, fax, email, made the per-word cost premium collapse. Telegraphese kind of survives wherever metering in terms of cost or just the time it takes to tap out a message is still onerous: 160-character SMS begat a whole new abbreviation culture; early Twitter did it again.
Well guess what?
Tokens are metered words again
An LLM API bill is a telegraph bill. You pay per token.
So we ran both strategies from ye olden days against modern models:
The Codebook: take existing text, substitute code words from a fixed dictionary.
Result: about 10% savings. Most prose isn't dictionary-shaped, and the substitution can't remove words the author already wrote — it can only rename them.
Cablese
Instruct the model, “Write a complete record in telegraphese; drop articles and filler; abbreviate; keep every fact, number, and proper noun verbatim — in lowercase, not all caps.”
Why does this work across models?
The register is already in the weights — an artifact of training data. Telegraph cables, codebooks, and cablese’s cousins across the broader “telegram style” family (headlinese, teletype style, note-taking, SMS abbreviation) are all in the “all of human knowledge” corpora that most models share. Two observations back this. First, models produce
fluent, conventionally-shaped
cablese from a one-sentence instruction — no examples, no codebook. Second, the readers in the cross-family matrix never saw even that instruction: they were handed compressed records cold and read them at parity, across four model families. Shared zero-shot fluency like that is hard to explain unless the convention is latent in shared human text.
5
Boring caveat: I have not run the strict control — the same compression instruction stripped of the historical framing (“write as tersely as possible”) — to test whether generic terseness produces equally
legible
compression. The claim here is strong evidence of a latent register, but at its core this is an educated guess.
Haven’t we seen LLMs do this already?
Yes,
BabelTele
(arXiv, June 2026) demonstrated that LLMs can encode text in compact,
non-standard
forms — omnilingual word fragments, symbols, emoji — that other models recover with high fidelity (99.5% semantic fidelity at 27.9% of original length, by their metrics), including cross-model transfer, agent memory, and multi-agent communication. It proves the general phenomenon: human readability is not a requirement for model-to-model text.
Further, researchers running populations of LLM agents under token budgets have watched the same thing emerge spontaneously: put agents under compression pressure and they negotiate compact protocols instead of passing full English back and forth.
GLOSSOGEN
found it’s the budget pressure that produces the new communication system. Another paper,
From Token Efficiency to Oversight Evasion
, shows agent populations developing emergent languages under efficiency pressure.
Other work has agents inventing symbolic languages that cut tokens
3–6×
at steady accuracy, and at the far end, frameworks that skip text entirely and pass raw embedding vectors between agents.
What’s already shipped
Terse
(
terseai.org
). A product selling “telegraph compression” for your prompts: rule-based stripping of articles and pronouns on the input side. The fidelity is asserted, not measured.
Caveman
(
github.com/juliusbrussee/caveman
). A Claude Code skill that makes the model answer in terse fragments. It has 100k+ stars on Github.
I think there’s more to do here, though.
So why does Cablese matter then?
If models can create something like a 6x compression on their own, why bother with a mere 2×? Here’s why:
These emergent protocols share a property profile: dense, efficient, portable between models (other LLMs can learn them in-context), but they are:
unstable — they drift as negotiation continues
illegible/unreadable to people
By contrast, Cablese is:
a known standard, more deterministic, generalizeable, with no setup or initial negotiation
readable, therefore auditable
In truth, the two techniques are different points on the compression/auditability frontier, and where your workload sits on that frontier should pick the point.
The Telegraph Test — measuring how models use Cablese
What the test measures:
A model is given a passage and a fixed set of questions with short, checkable answers — a date, a name, a count. It answers from one version of the information: the original passage, a plaintext summary, a cablese summary, or a cablese summary expanded back to plaintext. An answer is right if it matches the expected answer under fixed string rules — no judge, no discretion. Every score is then divided by the score of its
own plaintext control
, so
1.00 always means “exactly as good as plain English.”
Above 1.00 is better; below is worse.
What the percentages measure: comprehension against facts, presentation varied.
Each passage comes with ~24 questions whose short expected answers (a date, a name, a count) are verified as extractable from the source text — anything else is discarded before testing. A model receives one presentation — the full passage, a plaintext record, a cablese record, or a decoded record — and answers each question in a few words. An answer is correct when its overlap with the expected answer clears a fixed 0.8 threshold under deterministic, normalized matching: no exact-match pedantry, no judge, no discretion. Every comparison in this post holds the questions and grader constant and changes only what the model gets to read.
Why isn’t the plaintext baseline 100%?
Answering questions about an uncompressed passage in plaintext scores ~91%. The shortfall has three sources: strict grading (a correct answer worded differently scores as a miss — deterministic matching has no judgment, and no model sits in the judge’s seat); ambiguous questions (some have more than one defensible answer); and genuine misses (facts the model actually fails to extract). We can’t split those three without a judge, so we don’t try — and the design doesn’t need it: questions and grader are held fixed across conditions, so the baseline shortfall cancels in every comparison. That is why the benchmark’s primary numbers are ratios against each condition’s own plain control, where 1.00 means
as good as plaintext
.
A note on how to read the numbers in this post. The comparisons are
paired
— the same questions answered under both conditions — and tested with
McNemar’s test
. In plain words: ignore the questions both conditions got right or both got wrong (they carry no signal about which is better) and look only at the
exchanges
— questions where exactly one condition succeeded. If the conditions were truly equal, those wins would split like coin flips; a lopsided split (say 69 wins to 45) that had only a 3% chance of arising by luck is evidence of a real difference. That is all the p-values here mean. So every claim in this post is a paired comparison, and the register’s cost is the gap between the pair — zero-to-negative in both directions we measured.
The Telegraph Test answers three questions about any model, with the same frozen passage bank, the same anchored questions, and the same deterministic grader every time:
How lean does it write?
Given the identical cablese instruction, how much shorter is its record than its plaintext record?
Can others read it?
When a
different
model family answers questions using that compressed record, does accuracy hold? (Legibility — the difference between a shared register and a private idiolect.)
What does each use cost?
Read-back for machines, transcription for humans — measured separately, because they behave differently.
Nobody can know these numbers without running the experiment — a model release doesn’t come with a “writes lean cablese” spec-sheet row — and models are released monthly. That’s what the benchmark is for: an afternoon of compute per new model, and the property becomes a number on a table instead of folklore.
Image editing; an open-source, clean-room reimplementation of Adobe Photoshop, rebuilt in pure Rust.
Layers, masks, adjustment layers, layer styles, type, vectors, brushes and real PSD files,
in a native app written entirely in Rust. Open source, offline, and yours.
A caption card with a drop shadow, live type, and Vibrance and Curves adjustment layers, with the Curves editor open.
The Great Wave off Kanagawa
, Katsushika Hokusai, c. 1831
Note
ArtCraft is a community of artists from all walks of life.
Digital, generative, music,
games — if you make things, you're one of us.
Come say hi on Discord
.
The menus, shortcuts, panels and tools are where your hands expect them, from ⌘J to ⇧⌘D. If you know Photoshop, you already know PhotoCraft.
⚡ Native and fast
A GPU compositor on wgpu (Metal, Vulkan, DX12, WebGPU), copy-on-write tiles and multithreaded filters. No Electron, no web view, no waiting.
🗂️ Real PSD files
Open, edit and save layered Photoshop documents. Re-saving keeps the render of 307 of the 309 psd-tools test files.
🤖 Agent-ready
Every action is a command, so you can drive the same engine from the UI, the CLI, a JSON control channel or an MCP server.
Features
Every screenshot here is the real app at work on public-domain art, rendered offscreen through its control channel.
Levels and Vibrance adjustment layers, with the live Histogram panel.
Impression, Sunrise
, Claude Monet, 1872
Edit without regret
Adjustment layers keep every edit live. Stack Levels, Curves, Vibrance, Hue/Saturation and a dozen more, mask them to an area, reorder them, or turn them off, and your original pixels never change.
16 adjustment layers
that also apply directly to pixels, including Curves with per-channel editing, Levels with a live histogram, Black & White, Channel Mixer, Gradient Map, Photo Filter, Selective Color and Color Lookup (.cube, .3dl, .look). Plus Shadows/Highlights, Replace Color, Match Color, HDR Toning, Desaturate and Equalize.
Outer Glow and Stroke on a live type layer, in the Layer Style dialog.
Earthrise
, William Anders / NASA, 1968
Styles that sell the shot
Drop Shadow, Inner Shadow, Outer and Inner Glow, Bevel & Emboss, Satin, Stroke, and Color, Gradient and Pattern Overlay, live on any layer, including type. Patterns come from a library (built-ins, Edit › Define Pattern,
.pat
import/export) and PSD
Patt
blocks.
Copy and paste styles between layers, hide all effects at once, and open styles straight from your PSDs, rendered to match Photoshop.
An elliptical selection becomes the mask of a Hue/Saturation layer, so only the face keeps its color.
Girl with a Pearl Earring
, Johannes Vermeer, c. 1665
Selections that understand your image
Marquees, lassos and the Magic Wand for precision; Quick Selection, Object Selection and Select Subject when you want the computer to do the tracing; Select and Mask to refine hair-fine edges.
Feather, expand, contract, smooth, grow, reselect. Turn any selection into a layer mask, a vector path or a shape. Smart selection runs on your machine, with no cloud and no account.
A headline edited in place, with a byline and a paragraph of body text.
Among the Sierra Nevada, California
, Albert Bierstadt, 1868
Type that sets beautifully
Point and paragraph text, edited right on the canvas, with full Character and Paragraph controls: font, weight, size, leading, tracking, alignment and colour.
Type layers stay editable, take layer styles, and round-trip through PSD.
A badge made of shape layers: a gradient-filled lotus, a star and a dotted ring.
Water Lilies
, Claude Monet, 1906
Pixel-perfect vectors
Rectangle, Ellipse, Triangle, Polygon, Line and the Pen tool, with resolution-independent shape layers, gradient fills, and dashed, aligned strokes.
Combine shapes (unite, subtract, intersect, exclude), keep paths in the Paths panel, use them as vector masks, or stroke and fill them. 116 of 116 shape layers in our PSD corpus match Photoshop's pixels.
Twirl previews live on the canvas, only inside the selection.
The Starry Night
, Vincent van Gogh, 1889
See it before you commit
Every filter dialog previews live on the canvas, through your selection. Blurs (Gaussian, Box, Motion, Radial, Surface, Smart, Lens, Shape, and the Blur Gallery: Tilt-Shift, Iris, Field, Spin, Path), sharpening, Reduce Noise, distortions (Twirl, Wave, Ripple, Displace, Shear, Zig Zag…), Pixelate, Stylize (Oil Paint, Wind, Extrude…), Render (Clouds, Fibers, Lens Flare, Lighting Effects) and more.
Run filters on a smart object and they stay editable: change, hide, reorder or mask them at any time.
Large-radius blurs use running-sum box passes across all cores: a radius-180 Gaussian on 3.6 MP takes under a second.
Free Transform on a rotated print, with every step listed in History.
The Tetons and the Snake River
, Ansel Adams, 1942
Shape it any way you like
Free Transform with scale, rotate, skew, distort and perspective; exact 90° and 180° rotations and flips; Transform Again. Layers, type, shapes, masks and selections all transform together.
Full history, Toggle Last State and the History Brush mean every step can be undone, even one brush stroke at a time.
Export As in the light theme, with a preview and a file-size estimate.
The Kiss
, Gustav Klimt, 1907–1908
Ship it anywhere
Export As with format, quality, transparency and scale, plus a preview and an instant file-size estimate. Quick Export to PNG in one click.
Choose a dark Pro theme, the airy Studio themes, or a Classic look.
Shape Dynamics, Scattering, Texture, Dual Brush, Color Dynamics, Transfer, Brush Pose, Wet Edges, Build-up and Smoothing (including Pulled String), driven by pen pressure, tilt, rotation and direction. Brush presets, Define Brush from Selection, and deterministic, replayable strokes.
🗃️ Layers, done properly
Groups, clipping masks, pixel and vector masks, fill layers (solid, gradient and pattern), adjustment layers, live smart objects with smart filters and lossless transforms and warps, multi-layer selection with align, distribute and link, alpha channels and Quick Mask, 27 blend modes, opacity and fill, locks, colour labels, layer filters, merge, flatten, rasterize, Layer via Copy/Cut, Paste Into.
🎨 Any colour, any depth
RGB, Grayscale, CMYK and Lab documents at 8, 16 and 32 bits per channel. Bit depth and colour model are runtime data, so every tool works at every depth.
Real ICC colour management in pure Rust: embedded profiles, Assign and Convert to Profile with all four rendering intents and black point compensation, soft proofing (⌘Y) and Gamut Warning (⇧⌘Y) on the GPU.
🗂️ Formats
PSD and PSB, plus PNG, JPEG, TIFF, WebP, GIF, BMP, TGA, ICO, QOI, PNM, OpenEXR, Radiance HDR and AVIF, with symmetric read and write at 8, 16 and 32 bits, HEIC photos from iPhone and Mac (read; in official builds, an optional
--features heif
build feature), and the native
.pcraft
format.
🪄 The everyday essentials
Auto Tone, Contrast and Color · Equalize · Image and Canvas Size · Crop and Trim · Reveal All · Edit › Fill and Stroke · Copy Merged · Paste in Place · guides, rulers, grid and snapping · Actions record and replay · a command palette (⌘K).
PSD without compromise
PhotoCraft's PSD support is a standalone crate written from Adobe's public specification and tested against a corpus of real-world files.
Faithful round trips:
opening and re-saving a document renders the same for 307 of the 309 files in the psd-tools test set and 169 of 170 in our mixed ag-psd/psd-tools set (
crates/io/tests/corpus.rs
; fetch the psd-tools set with
cargo xtask corpus --psd-tools
), and anything we don't model yet (raw blocks, descriptors, extras) is carried over instead of being dropped. A re-saved file is not byte-identical to its source: PhotoCraft rewrites image resources, layer records and the composite. Only the standalone
photocraft-psd
crate, parsing and writing a file without the document model, reproduces every parseable corpus file byte for byte (
crates/psd/tests/corpus.rs
).
Pixels that match:
a composite oracle compares our render with Photoshop's own merged image, covering gradient interpolation (Classic, Perceptual and Linear), layer effects, shape strokes, clipping and fill opacity.
Large documents:
PSB, 16 and 32-bit files, and CMYK and Lab documents open natively.
Built for agents
Every menu item, tool and dialog runs a command from one registry of 500+ commands. The UI, the CLI, the JSON control channel and the MCP server all call the same commands, so anything you can click, a script or an AI agent can do too.
# Headless: open, edit, save
photocraft-cli run wave.psd \
--cmd filter.sharpen.smartSharpen --params '{"amount":80}' \
--cmd layer.newAdjustmentLayer.curves --params '{"points":[[0,0],[64,48],[192,212],[255,255]]}' \
--out wave-final.png
# Apply one action list to a folder of images
photocraft-cli batch --actions grade.json --in ./raw --out ./graded
# Every subcommand explains itself
photocraft-cli batch --help
# Let an agent drive it over MCP (headless, or bridged to the running app)
photocraft-cli mcp
The desktop app also offers an authenticated, loopback-only control channel (
photocraft --control
) for inspecting UI state, driving tools with pointer events, and taking offscreen screenshots. Every image in this README was rendered that way. See
docs/control-protocol.md
.
Under the hood
Engine first:
a pure-data document model and a command engine, with a thin egui UI on top. Layering is enforced at build time across 24 crates.
Two compositors:
a CPU compositor serves as the reference oracle, and a wgpu compositor puts the canvas on the GPU. They are tested against each other.
Copy-on-write tiles:
256² sparse tiles make undo cheap and huge canvases light, and effect maps are cached per layer state.
Runs in the browser:
the whole engine and UI compile to WebAssembly.
Clean-room:
implemented from public specs and observed behaviour only. No proprietary code, shaders or assets.
Tested:
more than 1,700 tests, including PSD round trips, synthetic generators, compositor oracles and multi-depth checks.
Get started
git clone https://github.com/storytold/photocraft
cd photocraft
cargo run --release -p photocraft -- image.psd # the desktop app
cargo test --workspace # the test suite
Japanese fonts for the UI and Type tool come from
craft-fonts
, an optional build input (desktop release builds always include it). Without it PhotoCraft uses your system's CJK fonts:
git clone https://github.com/storytold/craft-fonts ../craft-fonts
CRAFT_FONTS_DIR="$PWD/../craft-fonts" cargo run --release -p photocraft
Installers for macOS, Windows, Linux, FreeBSD and the web are attached to each
GitHub release
. On Linux you can pick an AppImage, a
.deb
, an
.rpm
, a tarball or a Flatpak bundle. The bundle needs the freedesktop runtime from
Flathub
, which
flatpak
offers to install along with it:
flatpak install --user photocraft-<version>-linux-x86_64.flatpak # or -linux-aarch64
flatpak run ai.storyteller.photocraft
On macOS, the command-line tool comes as
photocraft-cli-<version>-macos-universal.zip
. The binary is signed with the same Developer ID as the app and notarized by Apple. A bare binary can't carry a stapled notarization ticket the way the DMG does, so the first time you run it macOS checks the notarization online. You can confirm it yourself:
Maintainers:
docs/releasing.md
explains how releases are built, signed and published.
Important
Status:
PhotoCraft is in early alpha, and we want to be straight about where it stands: much of Photoshop's feature surface exists in some form, but
it is not yet a Photoshop replacement for daily professional work
. The biggest gaps are AI/generative features, about twenty missing tools, depth in typography and pro workflows, and plug-in compatibility. Every Photoshop menu item is wired to a command (
docs/parity.md
), but that measures wiring, not behaviour. The honest, dimension-by-dimension picture and where we're going next are in the
roadmap's parity assessment
. Expect rough edges, and please file issues (include your OS, document size, layer count and a screenshot). You can also tell us what broke on
Discord
.
Documentation
Developer, architecture, automation, format, and security documentation is maintained in the
PhotoCraft documentation book
.
Security
Security architecture, threat modeling, parser hardening, fuzzing, and vulnerability reporting are covered in the
security documentation
and the repository
security policy
.
Test corpora
PhotoCraft is tested against real files: our own Photoshop-authored oracle PSDs in
photocraft-corpus
plus the psd-tools, ag-psd and PngSuite sets, pinned and
sha256-verified. Fetch them with
cargo xtask corpus --all
and run the tests with
cargo xtask test-corpus
(details in
docs/development.md
).
The Crafting Apps
PhotoCraft is one of the
Crafting Apps
: free, open-source creative tools from the
ArtCraft
team, each written from scratch in Rust and each able to
stand on its own.
App
What it's for
Code
Learn more
PhotoCraft
Image editing: layers, masks, type and real PSD files · you are here
And
ArtCraft
itself, our AI image and video studio for artists who want real control.
The Crafting Apps share the same conventions: clean-room and pure Rust, native on macOS, Windows and Linux, in the browser via WebAssembly, and fully drivable by agents.
Come make things with us
Our Discord is where artists of every kind hang out: people who paint, shoot, draw, cut film,
set type, and people still figuring out what they like to make. Share what you're working on,
ask for help, tell us what's broken, or tell us what you wish these tools could do.
Whatever your medium and however long you've been at it, you're welcome here.
PhotoCraft is dual-licensed under
MIT
or
Apache-2.0
, at your option.
Copyright (c) 2026 ArtCraft Team and the PhotoCraft contributors. Required notices are in
NOTICE
.
Bundled fonts, icons, images and other assets keep their own open licenses; each one is listed
with its author, source and license in
ATTRIBUTION.md
.
Every artwork shown is in the public domain (Wikimedia Commons, NASA, U.S. National Archives); sources are listed in
docs/images/SOURCES.md
.
The ArtCraft name, wordmark and logos in
docs/brand/
are trademarks of the
ArtCraft Team and are not covered by this license. They may be used only unmodified, and only as
part of this repository and PhotoCraft, under
docs/brand/LICENSE-brand.txt
.
Forks and modified versions must remove them.
Adobe, Photoshop, Illustrator, Premiere Pro, Lightroom, Acrobat, After Effects and InDesign are trademarks or registered trademarks of Adobe Inc. in the United States and/or other countries. PhotoCraft is an independent, open-source project and is not affiliated with, sponsored by or endorsed by Adobe Inc.; these names are used only to describe the workflows it is compatible with.
Thatcherism via the Walkman: Stuart Hall’s British cultural studies preserved online
Guardian
www.theguardian.com
2026-10-07 08:01:10
Project lead says it celebrates ‘incredible contributions’ Jamaica-born sociologist made to Britain’s culture Whether it was Thatcherism, populism or the story of the Sony Walkman, Stuart Hall was a public intellectual famed for defining the times. Now a major digital archive of his previously unpub...
Whether it was Thatcherism, populism or the story of the Sony Walkman, Stuart Hall was a public intellectual famed for defining the times. Now a major
digital archive
of his previously unpublished papers, notebooks, recordings and videos has been launched.
When the Jamaica-born sociologist and seminal figure in postwar Black British history died in 2014 aged 82, his
obituary
in the Guardian said he had been “among the first to identify key questions of the age”. Hall coined the term Thatcherism and in 1985 warned of the dangers of “authoritarian populism”.
The
University of Birmingham
, where Hall developed the Centre for Contemporary Cultural Studies, has digitised and released more than 3,100 items donated by his widow, the historian Catherine Hall.
They include notebooks, index cards, his abandoned doctoral thesis and hours of previously inaccessible audio and video tracing the evolution of cultural studies, all within an interactive database.
Hall’s notes, charting the development of cultural studies, have been preserved in the archive.
Photograph: University of Birmingham
Also released are his “Origins of the new left” lecture, notebooks on “race and moral panics”, and
Encoding/Decoding Cybernetics
, a research project exploring Hall’s impact on the architecture of large language models.
In response to the era of artificial intelligence, and the questions it raises about culture, copyright and corporate power, the academics behind the
Stuart Hall Archive Project
have barred AI bots from accessing the data.
The Cadbury Research Library, which has catalogued the original papers, has run the digital archive through a programme that “poisons the data” for bots.
The project lead, Dr Rebecca Roach, working with Helen Fisher and Dr Katherine Parson, said they did not create the archive “just for AI to scrape” and commercialise. She said the platform was built to ensure “it doesn’t replicate the exploitation of people of colour and their cultural heritage that technology has historically enabled”.
A notebook with the handwritten title After Columbus.
Photograph: University of Birmingham
“One of the things that was really central to us developing the
Stuart Hall
digital site was to think about what can we learn from Hall’s ideas – and how can we apply that to our current moment with machine learning,” Roach said.
Hall died before the AI era fully arrived, but he left a body of work with which to interpret it. His theory of representation – examining how media builds meaning through language, symbols and cultural codes – remains relevant amid concerns about
algorithmic bias
. Doing Cultural Studies, a textbook he co-authored, used the
Walkman
as a case study of how social, cultural and economic life could be expressed through a single technology.
Roach said it was important to celebrate Black British intellectuals who had made “incredible contributions” to British culture, and noted that while scholars had sometimes been “a bit sniffy” about popular culture, Hall was at its forefront. She called him “a guiding figure for making the world better”.
Alongside the archive, Roach has produced learning materials for primary schoolchildren, while the art collective Vivid Projects has made a zine called Dear Stuart.
University of Birmingham students have created Walking with Hall, a virtual tour of campus landmarks, raising questions about what it meant for the academic to walk in the shadow of Joseph Chamberlain, the imperialist statesman who founded the university.
Classically educated in Jamaica, Hall came to Britain in 1951 on a Rhodes scholarship to Oxford, became the founding editor of New Left Review, and worked as a supply teacher in Brixton before pursuing his academic career.
Hall is considered a seminal figure in postwar Black British history.
Photograph: Eamonn McCabe/The Guardian
He met Catherine on an anti-nuclear march in 1963; they moved to Birmingham, where their two children were born and where he made his university department a beacon for debate on media, race, politics, Marxism and critical theory.
From 1997 to 2000 he served on the Runnymede Trust’s commission on the future of multi-ethnic Britain, going on to collaborate with young artists and film-makers, and in 2005 he was made a fellow of the British Academy.
Calling the archive “an extraordinary resource”, Prof Daniel McNeil, the Stuart Hall interdisciplinary chair at Birmingham University, said Hall “transformed how we think about culture, identity, power and belonging”.
Headlines for October 7, 2026
Democracy Now!
www.democracynow.org
2026-10-07 08:00:00
Israel Marks Three Years Since the Hamas-Led Surprise Attack on October 7, 2023, Israeli Forces Kill at Least Three Palestinians in Gaza, Houthi Militias Fire Drones and Missiles at the Aden International Airport in Yemen, Democratic Senator Chris Murphy Denied Access to U.S. Air Base in Qatar, More...
Israel Marks Three Years Since the Hamas-Led Surprise Attack on October 7, 2023
Oct 07, 2026
Israel is marking three years since the Hamas-led surprise attack on October 7, 2023, that killed 1,200 people, with 251 Israelis and foreign nationals taken hostage. In southern Israel, mourners gathered near Kibbutz Reim to commemorate hundreds of people who were killed and abducted at the Nova music festival. During a moment of silence, the booms of Israeli airstrikes on the nearby Gaza Strip could be heard ringing out. Ahead of today’s anniversary, Israeli Prime Minister Benjamin Netanyahu presided over a state memorial ceremony in Jerusalem for those killed in the October 7 attacks. He did not address reports he ignored warnings that Hamas had been planning a major operation. His speech was interrupted by a bereaved father, Yaakov Godo, whose son Tom was killed in the attacks.
Yaakov Godo
: “You didn’t protect my son on October 7th. You didn’t protect 1,200 people who were murdered in one day on October 7th. Shame! Disgrace!”
Israeli Forces Kill at Least Three Palestinians in Gaza
Oct 07, 2026
In Gaza, Israeli forces killed at least three Palestinians today. Among the dead is Nimr Barbakh, a 16-year-old child shot in the head by Israeli soldiers east of Khan Younis, and Mohammad Tammous, a Civil Defense rescue worker who died of wounds from an Israeli airstrike near Kamal Adwan Hospital. Israel has killed over 74,000 Palestinians in Gaza since Hamas’s surprise attack on Israel three years ago. Several governments, human rights organizations and scholarly bodies, including the International Association of Genocide Scholars, have accused Israel of committing genocide in Gaza. According to the U.N., as of September, 94% of Gaza’s population remained in urgent need of shelter and basic household assistance, while 65% of households were living in tents or makeshift shelters.
Houthi Militias Fire Drones and Missiles at the Aden International Airport in Yemen
Oct 07, 2026
Image Credit: Giants Brigades
Yemen’s Saudi-backed government says Houthi militias fired drones and missiles at the Aden International Airport, targeting a passenger terminal and striking a runway just minutes before a flight from Cairo was due to land. Elsewhere, Saudi officials said they had intercepted a missile fired by Houthis north of the Saudi capital, Riyadh. This comes amid a counteroffensive by Saudi-backed Yemeni government forces seeking to retake all remaining Houthi-held territory, including Yemen’s capital, Sana’a.
Democratic Senator Chris Murphy Denied Access to U.S. Air Base in Qatar
Oct 07, 2026
Democratic Senator Chris Murphy of Connecticut said the Pentagon denied him access to Al Udeid Air Base in Qatar during a visit to the Middle East. His failed attempt came after Defense Secretary Pete Hegseth signed a directive restricting travel to the area earlier this year. Senator Murphy claimed he was denied access as part of the Trump administration’s effort to conceal the true cost of the U.S.-Israeli war on Iran.
Sen. Chris Murphy
: “I have been to the Middle East probably a dozen times, at least, while I have been a member of Congress. I have never, ever been denied access to military facilities. This is part of a coordinated, deliberate campaign by the Department of Defense to try to hide the consequences and the true cost of this war.”
More Than 250,000 People March in Student Protests in France
Oct 07, 2026
In France, more than 250,000 people marched in over 40 French cities Tuesday. It was the largest turnout yet in demonstrations that began more than two weeks ago, when high school students started blockading their schools. The students are protesting staff shortages, crowded classrooms and run-down buildings. Labor unions joined in support, along with parents, teachers and university students. Paris had the biggest march — about 56,000 people, according to the Interior Ministry. Prime Minister Sébastien Lecornu said at least 215 teenagers and 85 school staff have been hurt in clashes with the police, calling it “the heaviest toll in decades paid by the school system.” This is a middle school science teacher in Strasbourg.
Erwan Le Clech
: “This situation has been deteriorating for years, and always at an accelerated pace. And now to see so many demonstrations, so many mobilizations appearing like this everywhere, it clearly shows that there is a general awareness with our living conditions worsening in France.”
Spain’s Gov’t Approves Two Housing Decrees Rejected by Parliament
Oct 07, 2026
Spain’s government has approved two housing decrees rejected by parliament last week, which prompted Prime Minister Pedro Sánchez to call a snap election. The housing package approved Tuesday caps rents, temporarily halts evictions and restricts companies from buying homes. The housing measures follow demonstrations in Barcelona that turned violent as tens of thousands of people took to the streets to protest the eviction of an 87-year-old woman, Maricarmen, from her lifelong home in Madrid. This is a woman affected by a possible eviction.
Josefina Alfaro Teruel
: “Make no mistake: If that Maricarmen business hadn’t happened and there weren’t so many journalists here, I’d already be out on the f—ing street in the rain. I swear to you.”
Death Row Prisoner Christa Pike Regains Consciousness After Surviving Execution Attempt
Oct 07, 2026
Image Credit: J. Miles Cary / Knoxville News Sentinel
In Tennessee, Christa Pike has reportedly regained consciousness after surviving a botched execution last week. According to her attorneys, the death row prisoner is off a ventilator and speaking from a hospital bed, where she remains handcuffed and shackled as she recovers from two injections of pentobarbital that were meant to kill her. Her prognosis remains unclear, and her lawyers say she faces a long road to recovery. Tennessee Republican Governor Bill Lee has ordered an independent review of Pike’s failed execution and canceled the only other scheduled execution of the year. Death penalty abolitionists have called on Governor Lee to commute Christa Pike’s death sentence.
Pentagon Schedules Execution of Fort Hood Attacker by Firing Squad
Oct 07, 2026
The Pentagon has scheduled the execution of 2009 Fort Hood attacker Nidal Hasan by firing squad on December 3, marking the first military execution since 1945. President Trump has approved Hegseth’s order.
Cornell Trustees Name Sally Yates to Investigate Handling of Sexual Misconduct Complaints
Oct 07, 2026
Cornell University trustees announced Tuesday that former Deputy Attorney General Sally Yates will investigate how the university handles sexual misconduct complaints. Yates’s appointment follows outrage and protest over Cornell’s response to a complaint from a woman known only as Jane Doe, who alleged that she was gang raped by multiple Chi Phi fraternity brothers after being pressured into drinking and taking drugs. No criminal charges have been filed. Jane Doe sued the university, seven former students and the fraternity last month, alleging that Cornell failed to protect her and punish the accused students. Members of the Cornell University Faculty Senate have introduced a resolution for a vote of no confidence in administrators over their handling of the case. There are also growing calls for Cornell University’s president to step down.
Paramount Completes Takeover of Warner Bros., Giving
CEO
David Ellison $150 Million Payday
Oct 07, 2026
Paramount Skydance has completed its $111 billion takeover of Warner Bros Discovery. The combined company will be called “Skydance,” made up of film studios, streaming services, broadcast and cable TV, sports channels and news outlets
CNN
and
CBS
News. With the deal now complete,
CEO
David Ellison will receive a $150 million payday, making him a billionaire. Ellison’s father, Larry Ellison, is the multibillionaire co-founder of Oracle and a major investor in Skydance. Both Ellisons are major Trump supporters.
Protests Rage Across India as Government Purges 130 Million Names from Voter Rolls
Oct 07, 2026
In India, police have detained several opposition leaders, breaking up a protest by Rahul Gandhi, the leader of the Congress Party, and other MPs who are demonstrating against a controversial move that purged 130 million names from voter rolls nationwide. Tuesday’s protests come just days after the Gen Z-led Cockroach movement organized demonstrations in Mumbai and Delhi calling for the resignation of India’s Chief Election Commissioner Gyanesh Kumar. This is Jairam Ramesh from the National Congress party, the main opposition to Indian Prime Minister Narendra Modi’s
BJP
.
Jairam Ramesh
: “This whole campaign is to highlight the fact that a constitutional functionary, the chief election commissioner, has violated the Constitution in letter and spirit. He has lied to the Supreme Court. He has lied to the country. And clearly he cannot stay in office any longer.”
ProPublica: 500+ U.S. Citizens Have Been Detained by
ICE
Oct 07, 2026
According to a ProPublica report, 506 U.S. citizens have been detained by immigration agents during the Trump administration. More than 100 Americans were held after agents questioned their citizenship — nearly all were people of color. At least 70 American children and teens were detained, some of whom were handcuffed. Thirty-six Americans were held for at least a day without being able to contact a lawyer or their family. And roughly a dozen citizens were deported, mostly children. This comes as the Justice Department moved to strip citizenship from 40 naturalized citizens who have committed crimes.
Immigrants Begin Hunger Strike at Folkston
ICE
Jail in Georgia
Oct 07, 2026
In Georgia, dozens of immigrants jailed at the Folkston
ICE
Processing Center launched a hunger strike Tuesday, demanding their release and calling on Congress, judges and advocates to highlight their plight. They’re also demanding protection from deportation and reunification with loved ones. Some of the men have been jailed since October 2023. They’ve described inadequate medical and dental care, unhealthy and spoiled food, noisy and cramped living quarters, isolation from loved ones, a lack of communication from
ICE
officers, xenophobic and hostile behavior from guards and a lack of access to sunlight and fresh air.
Martín Soto, Who Led Hunger Strike of Immigrants at
ICE
Jail, Released After 8 Months Behind Bars
Oct 07, 2026
Image Credit: Wali Khan
In New Jersey, an immigrant who led labor and hunger strikes by prisoners at the Delaney Hall
ICE
jail walked free on Tuesday after eight months behind bars. Martín Soto was transferred from Delaney to an
ICE
jail in Elizabeth, New Jersey, in May, after he helped organize protests against spoiled food, overcrowding and inadequate medical care at Delaney Hall in Newark, where detainees are forced to work for around $1 per day. In retaliation against the strike, guards at Delaney Hall reportedly beat participants and temporarily suspended family visitations. Soto was held in solitary confinement for about two months. On Tuesday, he was released from the Elizabeth Detention Center with an ankle monitor and greeted by supporters and his wife Gabriela and their two children.
Gabriela Soto
: “I am so thankful for everybody that has helped me. These months have been very, very difficult. Not only did I lose our home, not only did I lose our child, but it made me — this experience made me stronger to keep fighting to have the reunitement of our family.”
The original content of this program is licensed under a
Creative Commons Attribution-Noncommercial-No Derivative Works 3.0 United States License
. Please attribute legal copies of this work to democracynow.org. Some of the work(s) that this program incorporates, however, may be separately licensed. For further information or additional permissions, contact us.
The Data Race That Wasn't a Bug (and the One That Was)
Imagine this: you are testing the performance of some part of your application. Everything is going smoothly, the numbers look good, and as a last check you turn on Go’s race detector. Then, out of nowhere, it prints a warning you didn’t expect:
So you look at it. You look at it again, and again, and you think:
“What the…?”
The race is between your code and a goroutine you never started, somewhere deep inside
net/http
. You have no idea how that is possible, or why. Time to dig in.
That is more or less what happened to
Vadim Alekseev
while running
vmagent
under the race detector. He tracked the behavior down, reduced it to a small program, and reported it upstream as
golang/go#81445
. The same behavior also showed up in
vmauth
. In the end we left one of those races in place on purpose, and fixed the other. To see why, we first need to understand what the race actually is.
The Go internals in this post (function names, buffer sizes, timeouts) come from the Go 1.27
net/http
source. The same mechanics have been there for many releases.
Here is the client side of the program from
the upstream issue
(the full runnable version is there). It sends a batch of logs to a server, and retries if the server answers with an error. To avoid allocating a new buffer for every attempt, it reuses a single global
bytes.Buffer
:
typeLogEntrystruct{Timestampint64Contentstring}varbuf=bytes.NewBuffer(nil)funcinsertLogs(client*http.Client,logs[]LogEntry)error{buf.Reset()iferr:=json.MarshalWrite(buf,logs);err!=nil{panic(err)}body:=bytes.NewReader(buf.Bytes())resp,err:=client.Post("http://example.com/insert/logs","application/json",body)iferr!=nil{returnerr}deferresp.Body.Close()_,_=io.Copy(io.Discard,resp.Body)ifresp.StatusCode/100!=2{returnfmt.Errorf("unexpected status code %d",resp.StatusCode)}returnnil}// ...and in main, retry until it works:forrange100{err:=insertLogs(client,logs)iferr!=nil{fmt.Println("error inserting logs:",err)continue}break}
logs
is 1024 entries with 256 bytes of content each, so the body is about 300 KB. On the other side there’s a small test server that always answers
502 Bad Gateway
, without even looking at the request body. The program sends, gets a
502
, and tries again. Nothing unusual. Run it with
-race
and, after a few attempts:
error inserting logs: unexpected status code 502
error inserting logs: unexpected status code 502
...
==================
WARNING: DATA RACE
Write at 0x00c000390f60 by main goroutine:
runtime.slicecopy()
encoding/json/v2.makeStructArshaler.func2()
...
encoding/json/v2.MarshalWrite()
main.insertLogs()
main.main()
Previous read at 0x00c000390f60 by goroutine 49:
runtime.slicecopy()
bytes.(*Reader).Read()
bytes/reader.go:44
io.(*LimitedReader).Read()
...
io.Copy()
net.genericReadFrom()
net.(*TCPConn).ReadFrom()
...
net/http.persistConnWriter.ReadFrom()
bufio.(*Writer).ReadFrom()
...
net/http.(*transferWriter).doBodyCopy()
net/http.(*transferWriter).writeBody()
net/http.(*Request).write()
net/http.(*persistConn).writeLoop()
==================
exit status 66
The report describes two accesses to the same memory address,
0x00c000390f60
:
The write
happens in the
main
goroutine, inside
insertLogs
. It’s
json.MarshalWrite
writing the JSON for a new attempt into
buf
.
The read
happens in another goroutine, one we never started. It’s a
bytes.Reader
reading
buf
, called from a function named
writeLoop
inside
net/http
.
So something inside
net/http
was reading our buffer while our code was writing the next attempt into it. That’s surprising: by the time we write,
client.Post
for the previous attempt has already returned, and we have read and closed its response.
To understand how that can happen, we need to take a step back and look at how Go’s HTTP client really sends a request.
When you call
client.Post
, it feels like a single operation: send the request, get the response. Under the hood, at least
three goroutines
are involved.
For every HTTP/1.1 connection it opens, the
http.Transport
creates a
persistConn
with two goroutines of its own:
writeLoop
sends requests over the connection: the request line, the headers, then the body.
readLoop
reads responses from the connection and delivers them.
The third goroutine is
yours
, the one that called
client.Post
.
So
writeLoop
and
readLoop
are already running, one pair per connection, waiting for work. When your goroutine calls
client.Post
, it goes down through the client and the transport until it reaches
persistConn.roundTrip
, the function that connects your request with those two goroutines. It does two things.
First, it gives the request to
both
goroutines, one message on a channel to each:
// net/http/transport.go (simplified)// Write the request concurrently with waiting for a response,// in case the server decides to reply before reading our full// request body.pc.writech<-writeRequest{req,writeErrCh,continueCh}// writeLoop: "send this"pc.reqch<-requestAndChan{treq:req,ch:resc,...}// readLoop: "wait for its answer"
From this point on,
writeLoop
is sending the request and
readLoop
is waiting for the response, both at the same time.
Second, your goroutine waits to see what happens: either
writeLoop
reports that it finished writing the request, or
readLoop
delivers a response:
for{select{caseerr:=<-writeErrCh:// writeLoop finished writing the request...casere:=<-resc:// readLoop got a responsereturnhandleResponse(re)...}}
Let’s see what this looks like visually:
On the left, your goroutine does the two channel sends we just saw:
writech
tells
writeLoop
to send the request, and
reqch
tells
readLoop
to wait for its answer. Then it parks in the
select
, waiting for whichever arrow comes back first.
On the right side there are two separate paths. The
request + body
arrow, from
writeLoop
to the server, is the request going out: first the request line and the headers, then the body, chunk after chunk. The arrow from the server back to
readLoop
is the response coming in. A TCP connection carries data in both directions at the same time, and each direction is independent. The response doesn’t have to wait for the request to finish before it can travel back.
Typically,
writeLoop
sends the request line and the headers, and then starts sending the body. The server receives the request line and the headers, and based on them it calls the right handler. The handler reads and processes the body, and then sends a reply back, which
readLoop
picks up and passes to your goroutine along the
response
arrow.
But it isn’t always like that. Sometimes the headers alone are enough for the server to make a decision. The request may be unauthorized, or its
Content-Length
may say the body is bigger than the server accepts, or the server may be a proxy that can’t reach its backend. In those cases, the server can reply right away, without waiting for the body. The reply travels back on the other direction of the connection,
readLoop
passes it to your goroutine, your
select
takes the
resc
case, and
roundTrip
returns. So the client is still sending the body while the server has already sent its reply. That’s the interesting case for us here.
And here’s the important consequence. Your
select
returns as soon as the response arrives, and it doesn’t wait for
writeLoop
to finish. So
client.Post
can return, and your code can move on with the response in hand,
while
writeLoop
is still running in the background, sending the rest of your body
.
RoundTrip must always close the body, including on errors, but depending on the implementation may do so in a separate goroutine even after RoundTrip returns. This means that callers wanting to reuse the body for subsequent requests must arrange to wait for the Close call before doing so.
and
Client.Do
warns that “the Body may be closed asynchronously after Do returns”.
If we go back to our
DATA RACE
report, the goroutine doing the read was running
writeLoop
. So now we have the
who
: it’s the goroutine that sends our request, and it can keep running after
client.Post
has returned. What we still don’t know is
what
exactly it reads from our buffer, and when. For that, we have to follow the bytes.
Let’s follow our JSON from the moment we write it until it leaves the machine:
It all starts inside the
your code
box, with
buf’s array
. Remember that
buf
is a global variable: it lives for the whole program, so every call to
insertLogs
uses the same buffer, and the same array inside it, which stays in memory between calls.
buf.Reset()
doesn’t throw that array away, it just marks it as empty, so every time we retry
insertLogs
, the new JSON goes into the same memory. That’s the point of reusing
buf
: no new allocation on every retry. The
bytes.Reader
we pass to
client.Post
doesn’t copy anything either: it reads straight from that same array.
To send the body,
writeLoop
follows the first arrow in our diagram: it copies 32 KiB chunks from our buffer into a buffer of its own, and sends them to the server through the network.
Then the arrow at the bottom takes us back to buf’s array for the next chunk, and the whole thing
repeats until EOF or until the socket is closed
.
Now we have all the pieces. Let’s put them together.
Each row is one goroutine, and the bottom row is the memory they share:
buf’s underlying array
. Let’s follow it from left to right.
First, in
your code
, we call
buf.Reset()
and encode our logs, which writes the JSON into the array. Then we call
client.Do
, and our goroutine starts waiting.
That call sets the other two goroutines in motion, in parallel.
writeLoop
starts sending our request to the server: it copies the body out of the array, 32 KiB at a time, and writes each chunk to the socket. At the same time,
readLoop
starts waiting for the server to answer.
The answer comes back before the body is fully sent. The server only needed our headers to decide, so it replies
502
without reading the body. And because so much of the body is left unread, it adds
Connection: close
.
readLoop
hands that response back to us, and
Do returns 502
. For us, the request is done.
But
writeLoop
doesn’t know the server already answered. It’s
still reading!
from our array, sending the rest of the body.
Right then, our code retries:
buf.Reset()
and encode again, and
WRITE again
goes into the same array
writeLoop
is still reading. That’s the
overlap window
: one goroutine writing new data into the array while another is still reading the old data from it. That’s our data race.
Finally, because the server asked to close the connection,
readLoop
closes the socket
.
writeLoop
’s next write fails, it closes our body, and it exits.
This race doesn’t show up every time: locally the body is sent almost instantly, and if the connection is going to be reused, Go waits up to 50 ms for
writeLoop
before letting us continue. But what is the actual problem when the race does happen?
First, the good news: our code writes the array and
writeLoop
only reads it, so our buffer itself is never corrupted. And in our example, every retry encodes the same logs, so the bytes
writeLoop
reads are the same either way.
But let’s do a small mental exercise. Imagine we reuse this buffer for
different
requests, like sending a new batch of logs each time. What would the leftover
writeLoop
send for the previous request?
That’s valid JSON, but with a timestamp that neither request ever had. Both payloads have the same shape, so the pieces fit together perfectly.
Now imagine the next request is shorter than the previous one.
writeLoop
still thinks the body has the old length, so it copies the new data and then keeps going into what’s left of the old data, because
buf.Reset()
doesn’t clear anything:
previous : [{"Timestamp":1,"Content":"first"},{"Timestamp":2,"Content":"second"}]
next : [{"Timestamp":9,"Content":"new"}]
sent : [{"Timestamp":9,"Content":"new"}]},{"Timestamp":2,"Content":"second"}]
This time it’s not even valid JSON.
That sounds scary. But whether it matters depends on one question:
who, if anyone, reads that mixed copy?
vmagent’s remote write client (
#11507
) is the same pattern at a bigger scale. A worker pulls a block of compressed samples from its queue into a byte slice that it
reuses
on every iteration, and sends it with a fresh reader:
// app/vmagent/remotewrite/client.go (simplified)func(c*client)runWorker(readBlockfunc(dst[]byte)([]byte,bool)){varblock[]bytefor{block,ok=readBlock(block[:0])// <- overwrites the previous block's bytes...c.sendBlock(block)}}func(c*client)newRequest(urlstring,body[]byte)(*http.Request,error){reqBody:=bytes.NewBuffer(body)// <- a fresh reader for every requestreq,err:=http.NewRequest(http.MethodPost,url,reqBody)...}
Point vmagent at a remote storage that answers
200
without reading the body (Vadim used
httpbin.org/status/200
), push a big import through it, and the race detector reports the queue writing the next block into
block
while
writeLoop
is still reading the previous one. It’s exactly the race we just dissected. We even merged a fix for it. Three days later we
reverted it
, and later closed the issue without a fix. Here’s why.
The key to understanding this is how we build the request body: with
bytes.NewReader
in our example, or
bytes.NewBuffer
in vmagent. Both wrap our array with their own length and read position, so that part isn’t shared between requests. The only shared part is the array underneath. And the
Go memory model
guarantees that reading a byte while it’s being overwritten gives us either the old value or the new one, never some corrupted in-between state. So our readers are safe: the data they send can be mixed, but the program itself won’t break. Nothing crashes, and no other memory gets corrupted.
So the worst case is some mixed data, and in vmagent nobody uses it. It goes into a request the server has
already answered
, without reading the body, so the server didn’t want it anyway. Usually the connection is being closed too, so nothing on the other side ever reads them.
With nothing to protect, fixing it anyway would only cost us. A new buffer per request would undo the savings of reusing it. And waiting for the transport to finish with the body, which we tried and then
reverted
, could stall workers, deadlock in rare cases, and still didn’t cover everything.
Paying in performance and complexity to protect bytes nobody reads isn’t a good trade, so we kept the race.
“Harmless” is not “free”. This is a known, accepted race-detector report, not a silent one. We’ll revisit it if Go gains an official way to wait for the transport to be done with a request body. In
the upstream issue
, Damien Neil suggested a new
Request.Close
method for exactly that.
vmauth hit the same behavior (
#11508
), but here it was a real problem. vmauth is a proxy, and it can
retry
: if a backend fails or answers with an error like
503
, vmauth sends the same request to the next backend.
To send the same body twice, vmauth keeps it in memory in a type called
bufferedBody
. Before the fix,
bufferedBody
was also the reader: it held the bytes plus its own read position. On a retry, vmauth rewound that position to zero and handed
the same
bufferedBody
to the next request.
That’s the key difference from vmagent. In vmagent, each request has its own reader, and only the bytes are shared. In vmauth, both requests shared
the reader itself
, including its read position. So the leftover
writeLoop
from the failed request could keep moving that position, and when it finished, it even reset it to zero, right in the middle of the retry’s upload.
The result: the next backend could receive the body with pieces missing or repeated. Sometimes the length came out wrong and the request failed. But sometimes it came out exactly right, and the backend accepted a corrupted request without anyone noticing. This time the mixed data doesn’t go nowhere: it goes to a healthy backend that processes it.
The fix (
#11647
) does what vmagent already does:
never give two requests the same reader.
vmauth still keeps the bytes in
bufferedBody
, but every attempt now gets its own new
bytes.Buffer
over them. The leftover
writeLoop
can keep reading its own reader as long as it wants, and it can’t touch the retry’s. In the case of vmauth, the shared bytes are only ever read, never written, so there’s no race at all.
Here, correctness wins easily: a proxy that can silently send corrupted data to a backend isn’t acceptable at any speed. And the cost is tiny: a couple of small allocations per attempt, with no copy of the data, on a path where the network round trip to the backend costs far more.
So the same race detector warning led to two opposite decisions: keep it in vmagent, fix it in vmauth. Let’s wrap up with what made the difference.
Both warnings came from the same
net/http
behavior: the transport writes the request body in its own goroutine, and it can return the response to you before it’s done reading your body. The race detector was right both times. What it can’t tell you is whether the race matters, and that came down to two questions:
What exactly is shared?
Plain bytes that get copied somewhere, or
state
that decides behavior, like an offset, a length or a pointer? A race on bytes gives you stale bytes. A race on a cursor gives you wrong behavior.
Who consumes the result?
In vmagent, the mixed copy goes into a request the server has already answered, on a connection that’s being closed. In vmauth, the racing cursor decided what a
healthy backend
received and ingested.
So in vmagent we accepted the race and documented why, because fixing it would cost performance and complexity to protect bytes nobody reads. In vmauth we fixed it, by removing the sharing rather than adding synchronization, because correctness comes first.
If you use Go’s HTTP client, the rules are short:
Assume the transport may still be reading your request body after
Do
returns.
Avoid giving the same stateful
io.Reader
to two requests.
Be careful with reusing, pooling or mutating the memory behind a body you’ve already sent, until the transport has called
Close()
on it.
And keep an eye on
golang/go#81445
: if Go gets an official way to wait until the transport is done with a request, most of this goes away.
SonicWall warns of max severity SSRF flaw in SMA1000 gateways
Bleeping Computer
www.bleepingcomputer.com
2026-10-07 07:37:07
SonicWall has released hotfixes to address a maximum-severity server-side request forgery (SSRF) flaw in SMA1000 series appliances. [...]...
SonicWall has released hotfixes to address a maximum-severity server-side request forgery (SSRF) flaw in SMA1000 series appliances.
Tracked as
CVE-2026-102255
, the vulnerability was found in the Appliance WorkPlace interface of
SMA1000 6210, 7210, and 8200v models, but it does not affect the SMA 100 Series product line or SSL-VPN running on SonicWall firewalls.
The flaw stems from an
unintended alternate access-path weakness that remote attackers without privileges can exploit in low-complexity attacks
.
"By abusing this path, a remote unauthenticated attacker could potentially exploit this vulnerability to direct the appliance to issue requests on their behalf and reach internal functionality and perform unauthorized operations,"
SonicWall explained
.
While it has not yet flagged these flaws as actively exploited, the company urged customers to deploy hotfixes released on Tuesday to block potential attacks targeting their virtual or physical appliances.
"SonicWall strongly advises users of the SMA1000 series appliances to upgrade to the mentioned fixed release version to address these vulnerabilities," the company added. "There is currently no evidence any of the vulnerabilities addressed in this release are being exploited in the wild."
Although CVE-2026-102255 is not exploited in the wild, attackers often target SMA1000 flaws because they affect enterprise-grade secure remote access gateways used by government agencies, Managed Service Providers (MSSPs), and many large corporations to provide VPN access to internal apps and corporate networks.
Since the start of the year, threat actors have exploited several SMA1000 security vulnerabilities in zero-day attacks.
In July, two SMA1000 zero-days (CVE-2026-15409 and CVE-2026-15410)
were exploited for weeks
to install custom Sou5, OrangeTail, and RootRun malware on vulnerable VPN appliances in attacks that the U.S. Cybersecurity and Infrastructure Security Agency (CISA)
linked to ransomware gangs
.
Last month, SonicWall also warned customers that
attackers were chaining two new zero-days
(CVE-2026-83548 and CVE-2026-83549) to execute remote code on vulnerable SMA1000 gateways.
CISA has added 19 SonicWall vulnerabilities to
its list of actively exploited flaws
over the last four years, 13 of which have also been abused in ransomware attacks.
Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.
PSP games running in the browser without an emulator. The game's MIPS machine code is translated ahead of time into C++, compiled to WebAssembly, and linked against a small reimplementation of the PSP's operating system and graphics chip that draws with WebGL2.
The first title brought up this way is God of War: Chains of Olympus. It plays from boot through the menus, cutscenes and combat, at 60 frames per second in the scenes measured so far in Chrome and Firefox on a laptop, at up to four times the PSP's resolution, and on phones with on-screen touch controls. Music, speech and sound effects work. Movies are skipped for now.
God of War: Ghost of Sparta followed through the same scripts. It needed the PSP's DRM decryption for one small file, a handful of system calls and a lighting fix, and no performance work: it runs at 55 to 60 frames per second at three times the native resolution.
No game data is included here. You bring a disc image of a game you own, and the scripts in this repository turn it into a web page on your machine.
How it works
Recompilation.
PSPRecomp
analyses the decrypted executable, finds its functions and emits C++ for them in translation units of 16 KiB of guest code each. That code runs against a register file and a model of the PSP's memory. This project adds a handful of fixes to PSPRecomp (in
patches/
) and a new target for it, the web profile in
profile/
.
A small PSP kernel.
Whatever the game asks of the PSP's operating system is answered by high-level emulation in
profile/host
: cooperative threads with semaphores, event flags and callbacks, memory partitions, the file system (with disc data streamed over HTTP Range requests, so only the executable is downloaded up front), the controller, audio output, the save data and message dialogs, and the movie player's bookkeeping. Guest time advances in frames, so a game sees a steady 60 Hz however fast the host runs.
Graphics.
The GE, the PSP's graphics chip, is fed display lists.
ge.cpp
decodes them on the CPU, including vertex formats, skinning, lighting, texture generation, clipping and backface culling, and hands batched triangles and render state to
ge_gl.cpp
. There, framebuffers become WebGL render targets keyed by their place in VRAM, so effects that render to a texture and read it back stay on the GPU. Pixel format reinterpretation (games read 32-bit buffers as 16-bit textures and back), the PSP's stencil-in-alpha, fog and block transfers are emulated on the GPU as well, and everything can render at one to four times the native 480×272.
Two threads, like the PSP.
On the PSP the graphics chip works through one frame's display list while the CPU prepares the next, and games are written around that. Here the GE and its WebGL context run on a worker thread with an OffscreenCanvas: queuing a list returns at once and only the explicit sync calls wait. A frame costs whichever thread is busier rather than the sum of both.
Sound.
Sound effects come from a reimplementation of the PSP's voice synthesizer (32 voices of ADPCM with pitch and ADSR envelopes), music and speech from ATRAC3+ streams decoded with the FFmpeg decoder, and the mix goes to an AudioWorklet.
Lessons from making it fast
The first playable build ran at 6 frames per second. Most of the way to 60 came from finding out where the time actually went rather than from making the renderer faster.
God of War swaps its framebuffer without waiting for the vertical blank, so with nothing to hold it back the game drew around eight frames for every one that reached the screen. Holding the thread that swaps twice within one blank until the next one, a trick PPSSPP also uses, cut the work per displayed frame by a factor of eight on its own.
In WebAssembly, reading the clock through
std::chrono
goes through
clock_gettime
and a BigInt conversion in JavaScript. Profiling timers that read it per primitive took about a third of the frame until they were made to read
performance.now()
only when profiling is on.
Firefox copies every WebGL buffer upload to its GPU process, and after any change to an index buffer it re-validates the whole buffer on the next draw. A shared 4 MB index ring therefore dropped Firefox to 3 frames per second; giving every draw a small index buffer of its own fixed it. The opposite fix, writing vertex data piece by piece into one large buffer, helped no browser and made phones stall, because mobile GPU drivers wait or copy when a buffer the GPU may still be reading is modified.
The PSP keeps its stencil buffer in the framebuffer's alpha channel, and God of War uses it for projected shadows and to limit a blur pass. Mirroring stencil and alpha into each other with full-screen passes was correct but cost 60 million extra pixels per frame at 4×. Tracking which rectangles, which stencil bits and which constant values actually changed brought that down to about 7 million.
Finally, God of War queues each frame's display list and keeps working on the next frame before it waits. Running the GE on its own thread turned the cost of a busy fight from game plus graphics, around 17 ms in Firefox, into the larger of the two, around 10 ms.
A second game
Ghost of Sparta went from a ZIP to a page through
port.sh
without changes to the scripts or the recompiler, and then waited forever at boot. It opens a 176-byte file with the PSP's DRM flag, hands its key to
sceIoIoctl
and checks what it reads back. The file is in PGD, the format amctrl.prx decrypts with the KIRK crypto engine, and for disc games that comes down to AES-128 with three keys from KIRK's key vault: a CMAC-based check of the header and a counter mode for the data.
profile/host/pgd.cpp
implements it.
The next problem was a white sky, and the menus had the same white haze. Bisecting the draws of one frame led to a cloud layer drawn with lighting on, whose opacity comes from the alpha of the global ambient light, a factor the lighting code had left out. Performance needed no work: a frame costs 6 to 8 ms, as in Chains of Olympus.
Port a game of your own
You need git, CMake, Ninja, a C++20 compiler and Python 3. Everything has been run on Linux; macOS should be able to build the browser version but is untested, and the native test runner needs EGL and OpenGL ES headers (
libegl-dev
and
libgles-dev
on Debian and Ubuntu). Expect about 2 GB of disk space per game for the extracted disc, the generated code and the builds.
git clone https://github.com/snuri00/psp-web-recomp.git
cd psp-web-recomp
scripts/setup.sh # PSPRecomp, the patches and the Emscripten SDK
PSP_DECRYPT=/path/to/decrypter scripts/port.sh mygame "My Game.iso"
scripts/serve.sh mygame # http://localhost:8613/
port.sh
extracts the disc (an ISO, a ZIP containing one, or an already extracted folder), decrypts
PSP_GAME/SYSDIR/EBOOT.BIN
, translates it to C++, writes the manifest for streaming the rest of the disc and builds the page. It takes a few minutes; God of War goes from ZIP to playable page in about four on an 8-core laptop. Add
--native
to also build a headless runner that can dump frames to images and record audio, which is the quickest way to see how far a new game gets.
Executables on retail discs are encrypted.
PSP_DECRYPT
names any tool that is called as
tool <in> <out>
and writes a plain ELF, such as DecEboot or pspdecrypt. PPSSPP can also dump a decrypted executable while it runs a game (Settings, Tools, Developer tools).
Set your expectations accordingly: two games have been brought up so far, and both are Ready at Dawn titles built on the same engine, so they say little about how far a game from another studio gets. Another game will most likely stop at a system call nobody implemented yet, which is logged as
[hle] unimplemented ...
, or use a GE feature this renderer does not handle.
docs/internals.md
describes the tools for finding out what is missing, and the code is organized so that adding a call is a few lines.
Hosting
The page uses SharedArrayBuffer for its threads, so it has to be served cross-origin isolated, with
Cross-Origin-Opener-Policy: same-origin
and
Cross-Origin-Embedder-Policy: require-corp
.
scripts/serve.py
sends both headers and supports the Range requests the disc streaming relies on. To try it on a phone, put a tunnel in front of it, for example
cloudflared tunnel --url http://127.0.0.1:8613
, which passes the headers through. Anyone with the address can load the game while the tunnel runs, so keep it to yourself and stop it when you are done.
Browsers that cannot draw WebGL2 on an OffscreenCanvas fall back to a single thread automatically;
?threads=0
forces that, and
?profile
adds a per-frame timing breakdown to the status bar.
Legal
This repository contains only original code, the PSPRecomp patches and third-party code under its own license. It contains no game code or data. The generated C++ and the built WebAssembly are translations of the game's executable, so they belong to the game's owners: keep them on your own machine and do not publish them. Use disc images of games you own. God of War is a trademark of Sony Interactive Entertainment; this project is not affiliated with or endorsed by Sony or any game publisher.
Credits
PSPRecomp by its contributors (MIT) does the static recompilation. PPSSPP and JPCSP documented much of the hardware behavior emulated here, and PPSSPP's standalone copy of FFmpeg's ATRAC3/ATRAC3+ decoder is used for music and speech (LGPL 2.1 or later, in
profile/third_party/at3_standalone
). Emscripten builds the WebAssembly.
License
The code in this repository is available under the MIT License, see
LICENSE
.
profile/third_party/at3_standalone
keeps its LGPL 2.1 license.
We're excited to announce that Chrome is shipping decoding support for the JPEG
XL (
.jxl
) image format starting from Chrome 155. JPEG XL is a next-generation
image format designed to meet the needs of modern web developers and
photographers. It offers 30-50% better compression than JPEG, lossless
compression, built-in HDR support, lossless JPEG transcoding, and more.
In general, we recommend trying both AVIF and JPEG XL to get the best results.
We expect that JPEG XL is most helpful for high-fidelity or lossless
compression, especially of photographic images or in cases in which fine-grained
progressive decoding is preferred.
In this post, we share why we brought JPEG XL to Chrome, how we used Rust to
ensure memory safety first, the extensive performance work that makes it fast,
and what the journey tells us about developer feedback and the web standards
ecosystem.
Safety first: Reimplementing the decoder in Rust (
jxl-rs
)
Image decoders are one of the most critical and targeted attack surfaces in any
modern web browser. They process complex, untrusted binary structures directly
from the network and run inside the renderer process. Historically, decoders
written in memory-unsafe languages like C++ have been prone to vulnerabilities
such as out-of-bounds reads, heap overflows, and use-after-free bugs.
Our security model relies on sandboxing and defense-in-depth, guided by the
rule of
two
.
However, sandboxing is a secondary layer of defense. To eliminate these security
risks at the source, we have integrated
jxl-rs
, a pure Rust implementation of
the JPEG XL decoder.
Design for speed, without compromising safety
Memory safety is crucial, but a memory-safe decoder that is approximately as
fast as the best non-memory-safe alternative is a much more obvious choice than
a choice with a significant performance compromise.
A fundamental part of the performance of modern codecs is making full use of the
SIMD hardware available on modern devices. To do so safely,
target_feature_11
Rust
feature had to be stabilized, which allowed the use of SIMD instructions without
requiring
unsafe
code.
The next step was to build a SIMD abstraction layer (
jxl_simd
), inspired by
the C++
Highway
library (itself originally
developed for
libjxl
, the C++ reference implementation of JPEG XL). Together,
those developments allowed writing a multi-platform library that doesn't
compromise on SIMD performance optimizations, while restricting unsafe
operations to a small number of highly-vetted locations.
Performance optimizations in
jxl-rs
build on those in
libjxl
. This includes
a generic processing pipeline for steps crossing region borders, while
minimizing data copies to maximize hardware performance. We've been tracking the
performance of the Rust reimplementation across different hardware platforms on
the
jxl-rs performance dashboard
.
We verified the
jxl-rs
implementation with various state-of-the-art
techniques, including fuzzing and AI review of the code, and have not found any
memory safety bugs throughout the entire implementation history, providing yet
another validation of the huge improvements that Rust brings to memory safety.
Developer feedback and the Interop Project
The Chrome team considers web developer feedback from a wide range of channels,
such as bugs, surveys, the
Developer Signals
Project
,
and the
Interop Project
. Our
decision to ship JPEG XL was based on consistent feedback and requests from web
developers, most visible in the Interop Process, where it was a
popular
proposal in 2026
and
several years prior.
To ensure the format is interoperable across browsers, we have participated in
the
Interop 2026 JPEG XL
Investigation
to ensure
there is test coverage for all of JPEG XL's features in browsers, and that those
tests pass in Chrome.
Try it out
With JPEG XL officially landing in Chrome, the web becomes faster, richer, and
safer. We encourage developers, content creators, and platform owners to start
using
.jxl
images and animations in their pipelines.
Try it out,
file bugs
,
and help us continue building a faster and safer web for everyone.
Acknowledgements
We'd like to thank all the people who contributed to
jxl-rs
or its integration
in Chrome, and especially Helmut Januschka for the substantial contributions
both to the Chrome integration and
jxl-rs
, and Martin Bruse, Zoltan Szabadka,
Sami Boukortt and Wonwoo Choi for their substantial contributions to
jxl-rs
itself.
LLMs and Data Poisoning Are Weaponized to Manufacture Consensus
This website is using a security service to protect itself from online attacks. The action you just performed triggered the security solution. There are several actions that could trigger this block including submitting a certain word or phrase, a SQL command or malformed data.
What can I do to resolve this?
You can email the site owner to let them know you were blocked. Please include what you were doing when this page came up and the Cloudflare Ray ID found at the bottom of this page.
Apple just released a system called “Reference Image.” It can verify the image is exactly as taken by an iPhone—new models only—without tying it to a specific iPhone or photographer. It can also verify that multiple images came from the same iPhone.
Other industry solutions r...
Apple just
released
a system called “Reference Image.” It can verify the image is exactly as taken by an iPhone—new models only—without tying it to a specific iPhone or photographer. It can also verify that multiple images came from the same iPhone.
Other industry solutions require a photographer or institution to vouch for an image using their own credentials. We are concerned this puts some photographers, such as those operating in conflict zones, in a difficult position; it should not be necessary to forgo anonymity in order to prove image authenticity. We built Apple Reference Image to avoid using an explicit, public credential for photographers, and to avoid even implicit public association between different photos taken by the same sensor. The final reference image is instead signed by Apple’s signing service, after validation by PCC. That signature is backed by Apple’s strongest technical guarantees.
Our implementation also protects the confidentiality of the image itself, including from Apple. Merely capturing a reference image should never expose the actual pixels to Apple or anyone else. We achieve this through the exceptional privacy properties of PCC the nodes themselves are architected so that not even Apple can access image data, just as Apple cannot see the information processed for Apple Intelligence in PCC. While the revocation service must maintain a private record of photo GUIDs and associated sensors to allow for revocation, it never has access to the image data, and does not allow for public access to this record. And as final revocation checks occur using on-device lists, a device never reveals to anyone which photo it’s looking at in order to find out whether it’s still valid.
The report makes for good reading; the details are interesting.
McDonald’s sued for allegedly using AI tool to determine pricing for franchises
Guardian
www.theguardian.com
2026-10-07 07:00:08
Suit says AI tool allows independently owned franchises to exchange nonpublic price and sales information McDonald’s is facing a lawsuit in federal court over its alleged use of an AI tool to determine pricing across independent franchises, which prosecutors say violates antitrust laws and has unfai...
McDonald’s is facing a lawsuit in federal court over its alleged use of an
AI
tool to determine pricing across independent franchises, which prosecutors say violates antitrust laws and has unfairly inflated menu prices for Americans.
The vast majority of
McDonald’s
US stores are independently owned and are said to individually decide on prices under company policy. Antitrust laws require businesses to set prices independently from their competitors, as coordinating prices can stifle market competition and push up costs for consumers.
The lawsuit, filed on 2 October in a federal Chicago courtroom, alleges that McDonald’s AI tool is illegal because it amounts to independent franchise locations
exchanging nonpublic price and sales data.
“This case is about McDonald’s bringing ‘optimization’ to America’s largest fast food chain,” the complaint says. “McDonald’s has built and for years deployed its own information-sharing pricing platform, which draws on data from its millions of daily transactions to set menu prices across thousands of US restaurants. The result is algorithmic price-fixing aimed at customers who are already stretched thin.”
In response to the lawsuit, a spokesperson for the McDonald’s corporation said that the “complaint is filled with inaccuracies and we will vigorously defend against this lawsuit”.
“AI does not set menu prices at McDonald’s restaurants – McDonald’s franchisees do,” it said. “Optional tools are available to franchisees to help them make the best decisions for their businesses and customers, but these tools do not automate, coordinate or fix pricing in any way.”
A recent Reuters
investigation
reported that some McDonald’s franchise owners have been pressured by the company to use the AI pricing tools and record their deviations from the tool’s recommendations. McDonald’s contested that characterization, calling the Reuters reporting “speculative and uninformed”.
The litigation, which has been proposed as a nationwide class-action lawsuit, was first brought forward by an Illinois man named Michael Thomas. Thomas, a McDonald’s regular who is described as “price conscious”, found that the cost of his usual order – a Quarter Pounder with cheese, fries and a Coke – varied even within his neighborhood in DeKalb, Illinois. The attorney listed as representing Thomas did not immediately reply to a request for comment.
Thomas does not appear to be alone in his experience. At a McDonald’s in Manhattan’s financial district, Chukwama Okeke, 43, was finishing a mango-pineapple smoothie that he bought as part of a value meal that cost about $15 – something he only gets about once a month.
Okeke said that when he orders the same meal at a McDonald’s in Brooklyn’s Crown Heights neighborhood, he ends up paying closer to $9.
“McDonald’s in Manhattan is much more expensive than those in Brooklyn,” he said. “I’ve notice in some areas if they have a lot of like foot traffic, the prices tend to get higher.”
Beatriz Milander, 27, who ordered a snack wrap, a small fries and an iced coffee for $11.51, hadn’t noticed a huge difference between what she pays in the financial district and what she pays in her Bronx neighborhood, but said prices overall had inched up.
“In the last few months, I’ve seen them do more deals or like lower-priced things,” she said. “But it definitely was pricing up for a little bit there.”
Experts say that companies using AI to set or alter pricing could exacerbate the country’s affordability crisis and make it harder for Americans to bear everyday costs. So far this year, at least 90 pieces of legislation have been filed across the country to push back against algorithmic price-fixing, according to Lindsay Owens, the head of the left-leaning thinktank Groundwork Collaborative, who recently authored a
book
on the topic.
McDonald’s acquired an AI company, Dynamic Yield, in 2019, but has firmly denied using AI in setting menu prices. In a statement released on 1 October, the day before the lawsuit was filed, the chain denied using dynamic pricing and described its pricing tool as simply providing information that franchises are free to decide whether to adopt.
The fast-food chain has been scrutinized over recent years for its elevated prices, and went viral in 2023 over an $18 Big Mac meal at a Connecticut McDonald’s. The franchise’s owner filed a lawsuit alleging that it was the AI pricing tool that suggested the $18 price to the franchise.
According to a fact sheet the chain published in 2024, the average price of a McDonald’s menu item increased about 40% between 2019 and 2024.
Musician sent to prison for $10 million streaming fraud using AI bots
Bleeping Computer
www.bleepingcomputer.com
2026-10-07 06:35:15
A North Carolina musician was sentenced to 18 months in prison for collecting more than $10 million in royalties from Spotify, Apple Music, Amazon Music, and YouTube Music in a massive streaming royalty fraud scheme. [...]...
A North Carolina musician was sentenced to 18 months in prison for collecting more than $10 million in royalties from Spotify, Apple Music, Amazon Music, and YouTube Music in a massive streaming royalty fraud scheme.
54-year-old Michael Smith
pleaded guilty
in March after
being indicted
in September 2024 for fraudulently inflating his songs' listening stats between 2017 and 2024.
According to
court documents
, with help from the Chief Executive Officer of an AI music company and an unnamed music promoter, Smith uploaded hundreds of thousands of AI-generated songs bought from an accomplice to streaming platforms and used automated AI bots to stream the tracks billions of times.
To avoid detection by anti-fraud systems, he had the bots connect to the streaming platforms using virtual private networks (VPNs).
At the scheme's peak, Smith used more than 1,000 bot accounts to boost streams artificially. On October 20, 2017, he emailed himself a financial breakdown highlighting how he operated 52 cloud service accounts, each with 20 bot accounts.
Based on his estimates at the time, each bot could stream around 636 songs per day, for a total of roughly 661,440 streams per day. At an average royalty rate of half a cent per stream, the scheme's daily earnings reached $3,307.20, monthly earnings reached $99,216, and annual earnings exceeded $1.2 million.
One year later, on October 4, 2018, Smith emailed accomplices to say that "we need to get a TON of songs fast to make this work around the anti fraud policies these guys are all using now." and added that they needed "a TON of content with small amounts of Streams" to not raise issues "with the powers that be."
"By flooding music streaming platforms with automated bots in the place of consumers, and fake songs in the place of creativity, Smith robbed millions in royalty payments from genuine artists and their fans,"
said U.S. Attorney Jamie McDonald
.
"For example, in April 2023, the entire catalogue of Taylor Swift received 9.3 million streams on YouTube Music from family plan streams, while in the same month, Smith's Bot Accounts used family plans to fraudulently stream his AI-generated music 80.9 million times," the Department of Justice added.
In a February 2024 email, Smith also confirmed these claims, boasting to accomplices that the songs had generated "over 4 billion streams and $12 million in royalties since 2019."
In addition to the 18 months in federal prison, Smith was also ordered to pay $8,091,843.64 in forfeiture and was sentenced to an additional two years of supervised release.
Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.
Advantest confirms personal information stolen in ransomware attack
Bleeping Computer
www.bleepingcomputer.com
2026-10-07 06:27:52
Advantest Corporation is notifying affected individuals that a ransomware attack earlier this year exposed their personally identifiable data. [...]...
Advantest Corporation is notifying affected individuals that a ransomware attack earlier this year exposed their personally identifiable data.
The Japanese company is a global manufacturer of automated test equipment for the semiconductor industry. On February 15, a threat actor breached its network and gained access to some of its systems.
At the time, the
disclosure noted
that hackers had accessed parts of its network and deployed a ransomware payload, but the company could not determine if customer or employee data had been impacted.
In a data breach notification dated October 6, 2026, Advantest confirms that data was stolen in the attack.
“In February 2026, Advantest became aware of a cybersecurity incident in which an unauthorized third party accessed Advantest systems and extracted some data from our servers,”
reads the notification
.
“The data extracted from our servers included PII (personally identifiable information) belonging to you.”
The data types that have been exposed include the following:
Contact information
Date of birth
Social Security Number (SSN)
National ID number
Driver’s License
Passport number
Medical information
Financial information
Other ID numbers
It is unclear whether the compromised data belongs to customers, employees, partners, or a combination of these groups.
BleepingComputer has contacted Advantest with questions about the number of affected individuals, but we have not received a response as of publication.
The company states that it has no information that the compromised data has been leaked or otherwise misused, although it recognizes the elevated risk of identity theft and fraud for exposed individuals.
To mitigate this risk, the firm provides instructions on how to enroll in free 18-month identity theft, credit, and web monitoring services from Kroll, giving letter recipients until January 4, 2027, to activate the offer.
In addition to enrolling in this service, impacted individuals are recommended to closely monitor their accounts and financial statements for suspicious activity and report any unknown transactions to their bank.
Also, it is advisable to be cautious about phishing attempts, avoid clicking links or opening attachments, and never send money or share sensitive information in response to requests made via email or text.
At the time of writing, BleepingComputer could not find any public claims from ransomware groups targeting Advantest.
Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.
A game of cat-and-mouse between Sony and hackers is very much underway, as
the PS5 has been jailbroken
to a massive degree in the past week.
Legally we’re walking on thin ice here, but we can report on a new jailbreak which opens up the console, as long as it’s running firmware v7.00 through v13.60.
Firmware v13.60 was released in July, so is relatively recent. It means the jailbreak will run on most consoles not using the latest system software. (Although it should be noted that recent releases like
Marvel’s Wolverine
require a newer firmware to boot, so they remain off limits.)
Armed with AI tools, modders have already been able to install a variety of different emulators on Sony’s console, running from Switch through to the original Xbox. While performance is patchy, it’s utterly remarkable how fast this has all escalated in the past week.
What’s particularly fascinating is that, simultaneously, PC emulation of the PS5 is also developing fast. So far modders have found ways to run several exclusives, like
Demon’s Souls
, on the PC.
Obviously we’re not going to link to any of the repositories here, but I think it’s important to acknowledge this is all happening.
Some people have attached the pace of the development of these jailbreaks to
Sony’s decision to stop manufacturing physical discs
, but as we’ve seen with other consoles, hackers rarely need an incentive to deploy these kinds of exploits.
I imagine we’re going to see an increase in firmware updates over the coming weeks, as the platform holder tries to tighten up any security holes – but obviously that’s not going to prevent these exploits from working on older firmware.
It’s possible this may hasten the firm’s transition to the PS6, but it just depends how prevalent and problematic these jailbreaks become I suppose. It’s too early to say right now.
At the rate things have developed over the past week, I think Sony may have a pretty big headache on its hands here.
As the Editor of Push Square, Sammy has over 15 years of experience analysing the world of PlayStation, from PS3 through PS5 and everything in between. He’s an expert on PS Studios and industry matters, as well as sports games and simulators. He also enjoys RPGs when he has the time to dedicate to them, and is a bit of a gacha whale.
IMF chief urges governments to tighten belts as global debt levels soar
Guardian
www.theguardian.com
2026-10-07 06:14:35
Kristalina Georgieva says big economies will have to make ‘very touch choices’ as soaring bond yields hit budgetsBusiness live – latest updatesThe head of the International Monetary Fund has called on governments across big economies to tighten their belts as soaring bond yields hit budgets. Speakin...
The head of the International Monetary Fund has called on governments across big economies to tighten their belts as soaring bond yields hit budgets.
Speaking in Singapore, the IMF’s managing director,
Kristalina Georgieva
, said global debt-to-GDP ratios were at their highest level since the second world war and on course to hit 100% in the coming years.
She said governments could not rely on rapid economic growth to lift the burden of debt – and instead would have to make “very tough political choices”.
Georgieva was speaking before the IMF and
World Bank
annual meetings, which are to be held in Bangkok next week. “My message to the world’s economic policymakers will be this: we cannot keep delaying necessary policy action – you have the tools, now have the wisdom to use them.
“And yet we don’t see decisive action in the high-debt advanced economies where the need of the hour is for credible medium-term fiscal consolidation plans, supported in some cases by upfront fiscal measures,” she said.
Bond yields – effectively the interest rate on the debt –
have jumped in recent weeks
, raising the cost of borrowing for many governments to multi-decade highs as markets adjust to the prospect of higher inflation as a result of the war in the Middle East.
“Elevated yields are inflating the interest bill at a time of tight budget constraints and competing spending priorities, including defence,” Georgieva said, calling for “an urgent and comprehensive set of policy responses”.
The Bulgarian economist suggested central banks should be prepared to raise interest rates to see off resurgent inflation.
The ECB, US Federal Reserve and Bank of Japan have already tightened policy in the face of rising inflation – moves Georgieva said were “highly appropriate” – but the Bank of England has so far left rates on hold at 3.75%.
“Now may be a good time for a prudently hawkish bias in many countries’ monetary policy,” Georgieva said, suggesting central banks might want to err on the side of caution.
She also stressed the importance of tackling some of the risks of AI, which has buoyed the US stock market but raised fears of mass layoffs.
She highlighted IMF research predicting that the adoption of AI could add half a percentage point to global economic growth if carried out effectively. However, she urged policymakers to “help manage AI’s substantial perils, including large-scale labour market fallout, serious cyber and stability risks and frontier models threatening to escape human control and run amok”.
The Bank of England governor, Andrew Bailey, who is also chair of the Financial Stability Forum that brings together the world’s central banks, recently
warned of the “real and significant” risks
posed by frontier AI models and called for the “right to intervene”.
In the UK the chancellor, John Healey,
has said
he will stick with his predecessor Rachel Reeves’s plans to balance day-to-day spending with tax revenues – borrowing only to invest – and bring the debt-to-GDP ratio down over time.
Koreader reading statistics and highlights sync for spaced repetition
Reading a book again is a great way to retain more of its knowledge. Those of us pressed for time may instead opt to make bookmarks and highlights on an electronic ebook reading platform, like Koreader.
Press and hold, then drag to select text, which you can then highlight or annotate.
This is core functionality of Koreader.
I'll explain how I built on this to replace doomscrolling with re-reading quotes from books I loved to better retain their contents and improve my mental health.
The plugin
To sync these highlights, I made a plugin.
The hard part was the addon, not the web panel you'll see below.
Since the keyboard on ereaders is touchscreen and inconvenient to say the least, I've made it simple to pair.
Once paired, it syncs all your reading activity and highlights.
The panel
Now, you are able to see highlights, annotations, bookmarks (the whole page is treated as a highlight), and reading statistics. You can see here that I have almost a decade of reading statistics. I don't think any piece of technology has lasted this long except for a Yubikey.
Here you can see some statistics on your reading
The idea is you go through and star/tag them
Once you star them you will be able to review them at intervals and rate how well you remember them
You can see a timeline of a day to see what you annotated on that day
Stay tuned for the spaced repetition Android app. I was going to use
NeuraCache
but decided to complicate the field further by making my own implementation. (Kidding, I just wanted more instant sync)
How In-Game Overlays Work (2022)
Lobsters
fredemmott.com
2026-10-07 05:54:52
this only talks about direct3d, no mention of vulkan's layers feature
Comments...
This post aims to give developers a high-level understanding of how in-game overlays work in a variety of
environments: non-VR, Oculus, SteamVR, and OpenXR. Non-VR overlays are often used for social features and
notifications, like the Steam, Discord, and EA overlays; VR expands the use cases and technical requirements.
I’ll be focussing on Direct3D 11 in this post, but it roughly applies to D3D12, OpenGL, and Vulkan as well.
Non-VR overlays
DirectX does not have a supported way for one application to draw on top of another, so applications have
to do pretty hacky things. Let’s start with a simplified view of how games draw to a monitor:
Something wanting to draw an overlay has to find a way to fit in to this pipeline, and this problem isn’t
specific to overlays: anything wanting to change how a game looks has the same problem, such as
ReShade
or
certain cheats. For NVIDIA’s various overlays (mostly branded as GeForce), this seems to be relatively
straightforward: as NVIDIA provide the GPU driver (Graphics Processing Unit, a.k.a. graphics card), they can
add the functionality they need there. NVIDIA also use some of the later techniques, though it’s not clear to me
why extending the graphics driver isn’t sufficient for them.
Most developers don’t have that option, so they need another approach. Their software will need two things:
access to the games Direct3D state (
ID3D11Device
)
the ability to modify every frame
For Direct3D,
IDXGISwapChain::Present()
is a convenient target for both of these: its’ purpose is essentially
for a game to say “I’m done with this frame, send it to the monitor/window” - it is called for every frame,
can modify the frame, and has access to the Direct3D device (developers can convert the DXGI device to a
Direct3D11 device via
QueryInterface
). This leads to the next problem: how to change the behavior of
IDXGISwapChain::Present
inside the game; this is usually split into two sub-problems:
how do you run your code inside another process?
once your code is running inside the other process, how do you change the behavior of
IDXGISwapChain::Present
?
Running your code in another process
The main idea is to write a DLL that does what you want, then load the DLL into the game; Windows will run a
DLL’s
DllMain
whenever your DLL is loaded (or about to be unloaded), which in turn means that if you write
a DLL main, Windows will run your code when your DLL is loaded. There’s a few approaches for this:
Trick the game into loading your DLL
This is usually done by naming the DLL
dxgi.dll
or
d3d11.dll
, and putting it in the same folder
as the game executable. Direct3D games will then load these instead of the actual Direct3D components
from System32 - so your code will run, but your code will then need to load the ‘real’
dxgi.dll
or
d3d11.dll
.
This approach is often taken by shader modifiers, as it’s relatively straightforward and doesn’t need a launcher
or other program to be running to make it work. The biggest downsides are:
game updates may remove the extra DLL or otherwise conflict
each modification needs a different filename, and it needs to be one that the game already loads; for
example, if both tools need to replace
dxgi.dll
, you can only use one.
Most modifications noawadays are designed to work as either
dxgi.dll
or
d3d11.dll
letting you easily run two,
and some will work as any DLL name that the game tries to load. If you use many mods, this quickly becomes
complex, and there is still a limit: you can only replace DLLs that the game would try to load.
Modify the game
.exe
to load your code
This is commonly done for cheats as they usually want to make other intrusive changes to the game anyway,
or for other changes like fixing support for ultrawidescreen monitors. It’s not common to do this for
overlays - it can cause a lot of problems for users, so isn’t suitable for widespread use:
it invalidates the executable signatures, so will not work on many users’ systems
future updates are unlikely to work; a reinstall - and updated modification - will likely be required
even when the modification is not a cheat, it is the most likely approach to be flagged by anti-cheat
software
There are three common kinds of changes:
modify the list of DLLs that the game loads
modify the game’s code to call
LoadLibrary()
with a path to your DLL
modify the game’s code to directly include the changes you want
Act as a launcher, and modify the game in-memory to load your code before it starts
This is similar to the previous approach, but requires that the overlay application be used to launch
the game. This removes the signature and update problems, but - like all the approaches - still has a chance of
triggering anti-cheat software.
This approach involves using
CreateProcess()
with the
CREATE_SUSPENDED
flag, making the changes in memory,
then using
ResumeThread
to actually start the process. Any of the changes that would work on disk also work
here, though modifying the DLL list is probably the most common;
DetourCreateProcessWithDlls()
is a convenient
way to do this if you’re already using Microsoft’s
Detours
library.
The biggest disadvantage is that most launchers only support loading their own DLL and can’t be chained - i.e.
you generally can’t use more than one modification that takes this approach.
Load the DLL after the game is already running
This has several advantages: there’s no practical limit on how many modifications can take this approach,
and you don’t need to remember to set things up before you start the game. However, it is the hardest approach
to do reliably: for all the other approaches, your code is loaded as soon as the game starts, before it does
anything else. You know that Direct3D hasn’t been initialized yet. If your code loads later, perhaps
that’s still the case, or maybe the game’s been running for a few hours - it must handle every case.
There are many ways to load a DLL into another process on Windows; two of the most common ones are:
Use
SetWindowsHookEx()
to hook
Window messages
for either a specific thread (including threads in another
process), or for all threads on the system
Use
CreateRemoteThread()
to create a new thread in the game process that loads your DLL
I prefer the
CreateRemoteThread()
approach as it targets a process, rather than a specific thread/window,
but in addition to the usual challenges of multithreaded code, it is also often useful to run code in
the thread that owns the window.
SetWindowsHookEx()
isn’t a universal solution for these problems though - you may need to handle
the case that a thread was created for a splash screen or other temporary window and no longer exists - or
may simply be the wrong thread/window for the game itself. Those problems can be mitigated by hooking all
threads in the system, but this increases the risk of unintended side effects or false positives from
anti-malware software.
CreateRemoteThread()
needs combining with some other details to do this; the approach I use in
OpenKneeboard
is:
use
OpenProcessToken()
and
AdjustTokenPrivileges()
to give my app debug privileges (this does not
require UAC/administrator)
use
VirtualAllocEx()
to allocate memory in the game process to hold the path to my DLL
use
WriteProcessMemory()
to write the DLL path to that memory
use
GetModuleHandle()
and
GetProcAddress()
to find the address of
LoadLibraryW()
by a very handy coincidence,
LoadLibraryW()
is ABI-compatible with
LPTHREAD_START_ROUTINE
- so call
CreateRemoteThread()
, passing the address of
LoadLibraryW()
as
lpStartAddress
, and the address we got
from
VirtualAllocEx()
as the
LPVOID
lpParameter
There’s two common reasons this may not work:
the game is running as administrator but the overlay application isn’t; to fix, don’t run the game
as administrator 😉 If you are using a launcher (e.g. Skatezilla’s launcher for DCS World), don’t run
the launcher as administrator either
anti-cheat detects and blocks this. To fix, work with the anti-cheat vendors to recognize your software and
allow it
Drawing your overlay
Now you need to make it so that when the game tries to use
IDXGISwapChain::Present()
it actually uses your
code instead. If you’re modifying the game exe on disk or doing something game-specific in memory, it might be
easiest to find and directly change the call sites in the exe. If you don’t want to do that, or if your code is
running in a DLL in an unmodified game, you can
instead change
IDXGISwapChain::Present()
to call your code - rewriting the native code in memory while the
app is running. This again has subproblems:
Find the address of
IDXGISwapChain::Present()
Figure out how to change the start of it to call your code
Figure out how to call the original code; you’re going to have to overwrite some of it with the code to call
your function, so, you’ll need to copy the original code somewhere else, modify it, and jump back to the
remainder
Replace the start of the original function with your code from step 2 - while fixing any threads that were
in the middle of executing the original code while you swapped it out
Your code will usually do its own thing (e.g. drawing your overlay on to the swap chain’s active buffer), then
after doing some modifications go back to the original
IDXGISwapChain::Present()
code:
Finding the address
Finding the address of a public C-ABI function in a DLL is easy via
GetProcAddress()
; C++ can be much more
complicated, but in this case, we can take advantage of the fact that
IDXGISwapChain
is a COM interface,
so we can retrieve the address of the implementation via the COM C API:
This code must be compiled as C, not C++: when compiled as C++,
lpVtbl
is not defined. There’s also another
problem: this returns the
Present()
function for a particular
IDXGISwapChain
instance, but we don’t have an
instance.
If your DLL is loaded as the application starts, the most reliable way to get one is to also hook
D3D11CreateDevice()
,
D3D11CreateDeviceAndSwapChain()
, and
IDXGIFactory::CreateSwapChain()
via the same
techniques; however, if your code is loaded once the application is already running, you need another approach:
Fortunately, all
IDXGISwapChain
’s appear to have the same
Present()
function, so you can create your own
Direct3D device and swap chain, and pass that to the
Find_IDXGISwapChain_Present()
function above - however,
this is an unsupported implementation detail of Windows/Direct3D, and could change without warning in a future
Windows/Direct3D update. I’m personally fine with that given that’s true of drawing overlays on Direct3D in
general, not just this specific part.
Modifying the function
This is hard, as CPU instructions have varying lengths, layout requirements, and you also need to restore the
original registers/stack - except when those correspond to parameters you’ve changed; Microsoft’s Detours wiki
has
more details
on this problem.
Fortunately, this is a common enough problem that there are convenient libraries for this - the most popular are:
Previously, Detours’s free version was severely limited - however it is now entirely open source, including x64 support. These libraries do the vast majority of the heavy lifting of rewriting the functions, and a library is practically essential for doing this kind of work in a reliable way.
Problems
Brittleness
While Detours is a Microsoft project, it does not make using it to rewrite parts of Direct3D ‘supported’;
there is no supported way to do this, so all we can aim for is ‘good enough’.
Conflicts
Pretty much every overlay application needs to change
IDXGISwapChain::Present()
; whichever gets there first
will work fine, but when the second one comes along, it might not be able to understand the
already-rewritten-once version; the more overlay applications you run, the more likely it is that you’ll have a
problem, regardless of how they load into the game.
If you have trouble with the game crashing, try stopping any other overlay applications or disabling their
overlay functionality (e.g. MSI Afterburner, Discord, Steam overlay, RivaTuner, NVIDIA overlay, EA overlay); if
you’re a user and find which combination fails, try reporting to the developers of both overlays; they may be
able to find a way to make them work nicely together.
One way for a developer to make things work nicely together is to make yours piggy-back off the other. As a real
example, if
OpenKneeboard
was loaded in after the Steam overlay, games would crash; I fixed this
by checking if the Steam overlay DLL is loaded, and if so, instead of rewriting
IDXGISwapChain::Present
, I
use the same techniques to modify the wrapper function in the Steam overlay DLL and insert my overlay there.
This led to another problem: the Steam DLL does not export this function or
provide a documented way to find it - so I resorted to
pattern matching the native code
.
Anti-cheat false positives
Everything here is modifying the game to do stuff it wasn’t meant to do; the problems and solutions are shared
with cheats. Signing your DLLs
may
help, but really the only way to fix this is to stick to games without
anti-cheat, or work with the anti-cheat maintainers so that they recognize and allow your software.
Virtual Reality
While many appreciate non-VR overlays like Steam and Discord, the VR community has a much more pressing
need for them: VR players can’t quickly look at another window, another monitor, or alt-tab. The technical
requirements also change:
even if a flat HUD-style display is wanted, it needs distortion correction, and usually to be rendered in both
eyes
in-world positions are often wanted, like ‘on my wrist’, ‘by my feet’, or in the case of OpenKneeboard, it’s
intended to show information on your knee while sitting and playing a flight simulator
Fortunately, every VR API has the concept of visual layers - like the world itself and any in-game HUDs or
menus. These layers all have textures (images), and can have varying types, like ‘world view for both eyes’, or
‘WxH rectangle at (x, y, z)’ - so the goal here is to add an extra layer with a new texture, rather than to
modify the texture that the game is already creating.
Some VR APIs also provide a way to specify the coordinate system - you can then choose whether
(0,0,0)
means:
user-chosen seated ‘center position’ at eye level
user-chosen standing position at floor level
other less well-defined positions
Available options vary by API and by headset. If you’re not able to choose, you will need to retrieve the
controller or HMD positions, and apply the desired translations yourself; if there isn’t enough information
available, you may need to build a recentering feature. For the Oculus API, there is a setting, but it applies
to all layers; you must check what the game developers chose, and do the math as needed to get the results you
want.
Overlays with the Oculus API
Like non-VR Direct3D, the Oculus API does not provide a way for applications to add overlays (layers) to
other games, and we essentially need the same techniques: load our code, and rewrite some functions. The Oculus
API is designed so that at that the end of every frame, the game submits a list of layers, their descriptions
(‘WxH rectangle at (x, y, z)`), and textures; if we intercept these calls, we can add additional layers.
Modern code should use
ovr_EndFrame
(or better, OpenXR instead of the Oculus API), however older code may also
use
ovr_SubmitFrame
or
ovr_SubmitFrame2
; while the difference is important to game developers, from the
perspective of an overlay, these can be treated identically - and conviently have identical signatures.
Small overlays will want to add an
ovrLayerQuad
; larger overlays may want to use an
ovrLayerCylinder
to
provide some curvature.
While we no longer need to hook
IDXGISwapChain::Present
for every frame, it can still be a convenient way to
get the game’s Direct3D device, which you will need to optain to pass to
ovr_CreateTextureSwapChainDX()
.
Problems with Oculus API Overlays
Oculus API overlays have the same potential problems as non-VR Direct3D overlays: brittleness, conflicts, and
anti-cheat issues. There’s also a few extra quirks to keep in mind:
Layer limit
The Oculus API has a fixed limit on the number of layers: this limits the number of overlays you can use, which
can vary by game, and by state: it’s possible that pausing the game will add an extra layer with a pause menu,
so pausing the game might hide your overlay. Additionally, some games - like the Oculus World Demo in the SDK -
give a list of layers that is already the maximum size.
In these cases, it’s likely that not all of the entries are in use; for example, they might always reserve
layer 2 for a HUD, layer 3 for a pause menu etc, even if the HUD is disabled in options and the game isn’t paused.
In these cases, the entries in the layer list are usually a
nullptr
, which means the layer has no no effect,
so you can simply remove these layers from the list, making space for your own layer to be appended.
Interaction with depth data
When games provide the main ‘world view’ layer, they can provide depth information along with the RGBA (color)
data for each pixel; this is required for some of the framerate-compensation technologies like
ASW 2.0
, but
also interacts with the other layers in the list - later layers are not visible if the depth data from another
layer says they will not be visible because they’re behind the other layer in 3D space.
Some games appear to provide incorrect depth data which completely disables all overlays - even the
Oculus Debug Tool performance HUDs. This can be fixed by replacing any layers with
ovrLayerType_EyeFovDepth
with copies set to
ovrLayerType_EyeFov
. While
EyeFovDepth
is documented as being required for some
framerate-compensation technologies, it seems unlikely that they were working as intended if the depth data was
incorrect.
Overlays with SteamVR
This is where things get better: Valve saw the need for third-party overlays - perhaps due to their experience
with the non-VR Steam overlay - and made it a built-in feature of SteamVR/OpenVR:
There is no need for the overlay to interfere with the game: the game and the overlay app are separate processes,
independently communicating with SteamVR via OpenVR. SteamVR is responsible for setting up the layers,
coordinating input, and so on; SteamVR will also manage translating between Direct3D 11, 12, OpenGL etc as needed.
Problems with OpenVR/SteamVR overlays
While a multi-process overlay API is a huge improvement and the APIs themselves seem fine, there are some
long-standing issues in the implementation that overlay developers should be aware of:
Use
vr::TextureType_DXGISharedHandle
While OpenVR supports image files, raw pixel data, normal textures, and DXGI shared handles, the first two are
extremely slow, and the first 3 can flicker, and stop working entirely after a few hundred frames. For
long-running or high-framerate overlays, use DXGI shared handles for reliability and to avoid flickering.
These must be created via the legacy
IDXGIResource::GetSharedHandle()
function on a texture created with
D3D11_RESOURCE_MISC_SHARED
- OpenVR does not appear to support
D3D11_RESOURCE_MISC_SHARED_NTHANDLE
and
D3D11_RESOURCE_MISC_SHARED_KEYEDMUTEX
.
Don’t poll
VR_Init()
or
VR_IsHmdPresent()
frequently
These functions
leak memory
;
VR_Init()
is required,
but use a different technique first to reduce the number of calls and the speed of the leak. For example,
OpenKneeboard only calls
VR_Init()
if SteamVR’s
vrmonitor.exe
is running.
Edited 2023-04-02:
striked out references to XR_EXTX_overlay
XR_EXTX_overlay
proposed adding an overlay API to OpenXR; this appears to be a dead end, and I strongly
recommend creating an API layer instead:
there is practically no visible multi-vendor interest in adoption, or addressing unresolved design issues/missing features
the test implementation was never intended for end users, has no end-user support, and development has been inactive since 2021
XR_EXTX_overlay
is a provisional extension to OpenXR which would allow separate overlay applications in a very
similar way to SteamVR, however it’s not yet widely available.
While OpenXR does not
currently
have an
fully
supported
overlay API, it does provide a much more flexible system:
API layers
.
API layers provide a supported way to insert DLLs wrapping any OpenXR functionality
, and there is even
a test implementation of XR_EXTX_overlay as a layer
. In the
case of overlays, our goals are very similar to when hooking the Oculus API: we want a Direct3D device which we
can get by intercepting
xrCreateSession()
, and we want to add a layer by intercepting
xrEndFrame()
. The key
improvement is that by design it will load our DLLs without unsupported hacks, and we do not need to use Detours
or similar code rewriting techniques to intercept or wrap functions.
In the simplest case with no extra layers, the OpenXR “trampoline” - part of the OpenXR loader - will simply
‘bounce’ everything on to the active OpenXR runtime:
The OpenXR loader will also look for information on installed layers in the registry and environment variables;
if it finds any, it will automatically load them into this pipeline, transparently to the game. For example,
if a single layer wants to wrap
xrEndFrame()
, the result should look like this:
There can also be multiple layers, which are also all transparent to the game and other layers:
While I’ve used two layers that both intercept
xrEndFrame()
for this example, it could be any OpenXR function,
or they could intercept different OpenXR functions. Unlike DLL injection and Detours, you are unable to install a new
API layer into a game that is already running - but, it’s reasonable to have your DLL always be active, as long as
it is designed to have near-zero overhead when inactive.
Advice
Several VR SDK vendors decided to include basic matrix math libraries; I recommend using a separate matrix library
instead, such as
DirectXTK
’s
SimpleMath
:
some of the vendor matrix libraries are buggy
if you support multiple VR APIs, sharing math code is useful
it is unclear if the licenses of some vendor SDKs permit using their matrix code in code that does not target
their headsets
If you’re supporting any flow where your code is running in someone else’s process (every flow here except for
SteamVR or XR_EXTX_Overlay), I strongly recommend using native code - like C++ - instead of something that
will also load the .NET runtime into the game’s process.
More generally, if you have a DLL that is loaded into another process (including via the OpenXR loader), I
recommend doing as little as possible in that DLL: if you have bugs, it is usually much better for those bugs
to take down a separate overlay application than the game. For example,
OpenKneeboard
does the vast majority
of the work in the
OpenKneeboardApp.exe
process, but creates shared memory and shared textures to communicate
with the DLL, and is responsible for filling those shared textures with the desired overlay. The DLLs do the
bare minimum: they install the hooks (if needed), set up the layers, and copy the textures:
I use the same approach for Oculus and non-VR overlays.
Real-world example
OpenKneeboard currently supports overlays with:
SteamVR
Non-VR Direct3D 11
The Oculus API combined with either Direct3D 11 or Direct3D 12
OpenXR Direct3D 11 games, via a custom API layer
This code is available in
src/injectables
folder of the repository, and is under the GPLv2 license.
Disclosure
I used to work for Meta, but in an unrelated area (programming languages). This post and it’s opinions are
purely my own, and soley based on public information. I use an Oculus headset, which I purchased retail, and
was not reimbursed for it.
Nobel Prize in Chemistry 2026 to Henri B. Kagan and Kenso Soai
“for the discovery of non-linear effects and autocatalysis in asymmetric organic synthesis”
They made chemistry choose a mirror image
Some molecules, such as amino acids, exist as two variants that are each other’s mirror image. However, living organisms only contain one of these mirror images. How chemical asymmetry such as this could have emerged was long a mystery to chemists. Henri B. Kagan and Kenso Soai are awarded the Nobel Prize in Chemistry 2026 for discovering a solution to this puzzle.
Life’s chemistry is what chemists call
homochiral
, after the Greek words for “same” and “hand”. Like hands, all amino acids exist as two mirrored variants, but only one of them is found in the proteins in your cells. The other is rarely found in nature.
For a long time, chemists wondered how homochirality can emerge. When they began to experiment with chemical reactions that can form two mirrored molecules, they always obtained equal proportions of both in their test tubes. However, chemists strived to produce only one of these mirror images, because in the development of molecules that will interact with living beings – such as in pharmaceuticals – only one mirror image will have the desired effect.
The Nobel Prize in Chemistry 2026 recognises discoveries that have enabled chemists to drive chemical reactions that lead to homochirality.
“Henri Kagan and Kenso Soai have provided a solution to a chemical mystery that is over a century old: how homochirality can emerge spontaneously. The chemical reactions they have developed are spectacular,” says Heiner Linke, chair of the Nobel Committee for Chemistry.
Henri Kagan took the first decisive step in 1986, when he discovered a new way of manipulating chemical reactions. This allowed him to create a greater excess of one of the mirror images than had previously been thought possible.
Kenso Soai took the next step. In 1995, a key publication describes how he designed the first ever chemical reaction that had the potential to be homochiral. In 2003, he finally succeeded. He presented a reaction in which only one of the two possible mirror images was formed. Other than life itself, no one had previously achieved this feat.
Thanks to Henri Kagan and Kenso Soai, we now know how homochirality can emerge. Their discoveries have been decisive for chemists who design reactions for the manufacture of pharmaceuticals.
Henri B. Kagan
, born 1930 in Boulogne-Billancourt, France. PhD 1960 from Collège de France, France. Professor Emeritus at the then Université Paris-Sud, France.
Kenso Soai
, born 1950 in Hiroshima, Japan. PhD 1979 from University of Tokyo, Japan. Professor Emeritus at Tokyo University of Science, Japan.
Prize amount
: 12 million Swedish kronor to be shared equally between the laureates
Further information
: www.kva.se and www.nobelprize.org
Press contact
: Eva Nevelius, Press Secretary, +46 70 878 67 63,
[email protected]
Expert
: Peter Somfai, +46 70 693 63 77,
[email protected]
, member of the Nobel Committee for Chemistry
The Royal Swedish Academy of Sciences, founded in 1739, is an independent organisation whose overall objective is to promote the sciences and strengthen their influence in society. The Academy takes special responsibility for the natural sciences and mathematics, but endeavours to promote the exchange of ideas between various disciplines.
Nobel Prize® is a registered trademark of the Nobel Foundation.
In this model almost all state is on the server. We stream the next frame (when I say frame I mean the next version of the html page generated on the server) to every connected client every X ms (a tick). This style of rendering is often referred to as immediate mode (fat morph in the
Datastar discord
). These frames are streamed to each client over a long lived SSE connection with streaming compression (Brotli or Zstandard).
Think
view = f(state)
just on the server rather than the client.
Bounding the system with ticks
Ticks are awesome because they give you a boundary to batch against, back pressure and a place to measure the performance of your system. If your server is under load frames are dropped, but it doesn't fall over.
Without this batching your system can easily be accidentally quadratic. X users do Y actions each action triggers a render for each user. So 1000 users each doing 1 actions a second is 1000 x 1000 = 1000000 renders. In a 1 second tick based system (updates every second) that would only be 1000 renders regardless of how many actions the users perform.
When you have a tick based system you have a natural point to introduce a barrier. For example the simplest way to batch writes is to have a single writer. In a tick based system you can have a clear barrier between your writes and your reads.
Batch your writes -> batch your renders -> ...
This lets you batch your renders. This can be as simple as iterating over all your long lived connections and generating the HTML they need to render. Or if you want get slightly fancier pinning each connection group to a core and iterating over them. This allows each group to have their own thread local resources (buffers, caches, database connection). Which can be great for making your system more deterministic and bounding memory usage.
Before you might have had a 64kb buffer for each connection's template generation, but with batched renders, that ends up being 64kb per batching thread. Say you have 10000 concurrent users, that would be 640mb. Even worse if you're not using buffer, you'd be generating 640mb of garbage every render. With the batching model you'd have 64kb per thread, so in a 4 core system that would be a minuscule 256 kb!
In the case of database connections (and sometimes caches) you eliminate contention (
see this LMAX talk for more on why this is important
) . Each render thread has it's own resources so there's no need for coordination.
In another post
I covered how streaming compression eliminates the network overhead of immediate mode rendering. It's ok to send the whole 50kb HTML frame because on the wire it will be as small as 13bytes (if nothing changed).
Querying the database every tick?
But doesn't a tick based model involve querying the database every tick? In my case yes, yes it does. But, if you're using an embedded database like SQLite for projections then your projections are laid out so that they are quick to query. If you're worried about write throughput you should check out this post
100000 TPS over a billion rows: the unreasonable effectiveness of SQLite
.
HTML templating
DEEP BREATH. We are breathing rare air here. String generation, concatenation, encoding, escaping and some iterating are our bottlenecks.
So with ticks, barriers, compression and sqlite we've eliminated a load of work. Well that leaves one last bottleneck. With thousands of concurrent users being updated 10 times a second, HTML templating ends up occupying a lot of our frame budget.
You could do something clever like only update users who need to be updated. But, that doesn't solve all users having a shared widget and all needing to be updated anyway. Or dynamic content that changes all the time etc. The goal here is to not be accidentally quadratic. Any optimisation that helps the happy path but doesn't improve our worse case is just overhead when things go wrong.
This is where aggregates come in. Because, we've deliberately not been clever up to this point: broadcast to all users on every tick even if nothing changes. We are in a position to think about rendering as a batch process that happens for all users. Which means we can assume that there is likely to be some overlap between all the frames in a batch, but also all the frames in the previous batches.
But first...
CPU Flamegraph of a concurrent user overload test. 4000 users all looking at slightly different views without caching. You can use the search feature to find things like sqlite etc.
Simple hiccup interpreter
Our system has a simple hiccup interpreter that recursively iterates over hiccup and writes to a byte buffer:
There's two small changes with this hiccup interpreter. If it encounters a function it will eval it and assume the output is more hiccup to be interpreted.
(defn write-node
[lane-ctx node ^ByteBuffer out]
(cond
(nil? node) nil
(bytes? node)(.put out ^bytes node)(string? node)(write-escaped-string node out)(instance? Sequential node)(write-collection lane-ctx out node);; if function we evaluate it
(fn? node)(write-node lane-ctx (node) out)
:else (write-escaped-string (str node) out)))
The main benefit of this is you can delay work until the interpreter reaches that point. So your database queries can be streaming straight into your output byte buffer without materialising the full result of the query.
If it encounters an element who's first argument is a function it will apply the rest of the elements content to that function as arguments (think of them as components).
What's cool is we can wrap these "components" in a cache and key their output by their function and arguments. Effectively this is automatic content addressable caching. Components are just function:
(defn Palette [current-selected]
[:div {:class "palette"}
(mapv (fn [state]
[:div
{:data-id state
:data-action handler-palette
:data-color state
:class
["palette-item" (when (= current-selected state)
"palette-selected")]}])(subvec states 1))])
And we can reference them in hiccup similar to components in Reagent (although I might change this to chassis style aliases in future):
They can be nested just fine, because our interpreter is recursive.
Cache thrash
But, what about thrashing the cache? We've got automatic component level content addressable caching. A bunch of small components that are all different could push out our more valuable cache entries!
It has two really cool features. Entries are only admitted if they are seen with a certain frequency. This means all those small entries that are all different: they never make it in.
But, the other interesting property is entries that are not requested often naturally drop out of the cache.
This leads to really interesting emergent optimisation behaviour. Say we have a nested components. They present a problem. If we cache all the children and the parent (including all the children). We have roughly doubled the amount of space we take in the cache. But, with WTinyLFU if the wrapping component always changes, it never gets cached cause of the admission mechanism. If the wrapper component is stable and sticks around all the cached sub components drop out of the cache because they are never requested.
CPU Flamegraph of a concurrent user overload test. 4000 users all looking at slightly different views after content addressable component caching was added.
Conclusion
Thinking in batches and aggregates lets you do really powerful performance optimisation whilst keeping your system simple to reason about. When doing app development I don't have to think about caching or performance. I can just write hiccup and sql queries.
Install the TTF or OTF, then select
C64 Keyboard
in your application.
To try it out, download and extract the complete package, then open
preview.html
. Click a special-key or PETSCII button to insert its symbol into the type tester, then copy the text into your application and select
C64 Keyboard
.
If you have an earlier version installed, replace it with
version 1.107
and restart applications that still show the old font.
Included
Uppercase letters, numbers, punctuation, arrows, and £.
Selected accented letters, including Hungarian characters.
Function-key labels f1–f12 and classic C64 key legends.
All 63 PETSCII graphic front legends, arranged by key in the preview.
Lowercase input displays as uppercase. This recreates printed key legends, including smooth geometric reconstructions of the PETSCII graphics, rather than the screen bitmap font.
PETSCII front legends
Open
preview.html
and use the PETSCII buttons to insert the symbols, then copy them from the type tester. Left legends correspond to Commodore + key, right legends to Shift + key in uppercase/graphics mode. Pi on the up-arrow key also works with Commodore.
These are reconstructed outlines: identities and key assignments come from the
Ultimate Commodore 64 Reference
; frame thickness, curves, and print proportions are inferred from the supplied keyboard photograph. They are not exact photo traces. Pi is unframed as on the key.
The preview handles the special character codes for you. Copied symbols need this font to display correctly. See the
PETSCII key map and encoding guide
for codes and keyboard combinations.
Editable outlines are in
glyphs/
, with build tools in
scripts/
. Original photographs are not included.
There's a place in Skyrim right outside the capital city, Solitude, where if you cast Magelight at the top of a mountain, you get
tons
of XP in Alteration — enough to
level you from 15 to 25
in just one cast. But why?
There's a lot of wrong explanations for why this works, but the main one is some version based on how far the target is:
"Magelight grant XP on arrival based on how long it travelled. What you have found is a place where you are on the near edge of render distance, but there are tricks to increase it even further." -
@Jewbacca1991
But as
others
quickly
point
out
, it doesn't work for other far distances: something seems to be special about
this
mountain. The
UESP Skyrim wiki
claims that it's because of the cloud layers that you can see:
"The large XP gains at this location seem to be related to passing through multiple planes of weather effects."
But there are plenty of cloudy mountains where this doesn't work. Let's reverse engineer the Magelight XP formula to put this to rest.
How Magelight grants XP
Magelight is an Apprentice level spell. Loading it up in
TES5Edit
we can see that it has a Magnitude of 5, a Duration of 60, and a base effect of
LightFFAimed
, or
0001EA6D
. If we load that effect we can see it has a Base Cost of 2, and a Skill Usage Multiplier of 0.15.
The Spell ID,
00043323
The Spell Effect,
0001EA6D
This all gets used to calculate the Skill Usage value. Because there's no Area for the spell it quickly simplifies:
The Alteration XP takes this, multiplies it by the Skill Usage Multiplier (0.15 for Magelight) and then by the Skill Use Mult (
3 for Alteration
) for a raw XP of 37.938 per cast.
Validating in game
Ideally we'd cast Magelight at a wall and see if this gives the expected amount of XP. Unfortunately there's no way in game (even using console commands) to get the exact XP. Luckily we can use
Cheat Engine
to read the underlying memory directly with Lua.
1
You may remember Lua (
my favorite
programming language downstream of Brazilian oil companies) from my
previous post speedrunning Harry Potter 1 for the GBA
.
-- (Set the player's Alteration level to 15 with a console command)
-- player.setav alteration 15
-- Set the overflow XP to 0
local ALTERATION = 18
local player = readPointer(0x01B2E8E4)
local skills = readPointer(readPointer(player + 0x614))
local entry = skills + 12 * ALTERATION - 0x40
writeFloat(entry + 4, 0)
print(string.format('Alteration XP: %.3f', readFloat(entry + 4)))
-- <Cast the spell>
-- Read out the current XP
print(string.format('Alteration XP: %.3f', readFloat(entry + 4)))
If we run the above script
2
I swapped to using the VEH debugger with
Edit
>
Settings
>
Debugger method
, though I'm not sure that's necessary.
to reset the XP to 0, cast Magelight, and then run the final read of the new XP we see it goes up exactly as predicted. Yay!
But none of this incorporates distance. Why are we getting so much more than 38 XP?
Reverse engineering when the XP grant runs
Magelight is unlike most other spells because it doesn't need to hit a person — it'll work, but it also works to cast it at a wall or a table or a rock. It grants XP in all of these cases, so how does it know when it hits something?
How collision works
Casting the spell creates a
LightSpellProjectile
that travels along a path based on where you originally aimed it. Each tick it checks if it's collided with something, which roughly follows these three steps:
Draw a path for where the object will move in the next tick
(
0x007A0090
)
It takes the speed of the projectile (512 units/s) and divides by your frame rate (60 fps
3
And yes, this means that the collision check rate is lower if you're playing on old Xbox consoles, and higher if you've cranked this above the normal value. It
also
goes higher if you use the
Slow Time shout
or use the Archery/Block perks to slow time, though collisions/s stay the same.
) to get \(\frac{512}{60} = 8.5\bar{3}\) units/tick in the direction you cast in. The spell has a range of 10,000,
4
Again determined by the ESM file, though it wasn't in the above screenshot crop. Whoops.
so it will continue along its path for \(\approx 19.5\) seconds before it vanishes.
Intersect that path with every object that has a collision shape and passes the collision layer filtering.
(
0x00DE0040
)
Every object in the game has a collision layer which defines what it collides with (for instance
L_TRANSPARENT
won't stop projectiles or spells where
L_TERRAIN
will). If they don't match, a collision won't fire.
L_TRANSPARENT
L_TERRAIN
Check if the path hit the object.
(
0x00EC41C0
)
First, it finds the closest point on the surface of the object to the projectile. Then it finds the normal vector that points away from the surface with that closest point.
Shown here on a rectangular prism, though there are much more complex normal maps for things like rocks or castles
It then checks if the projectile is going in the same direction as the normal.
5
It does this by checking if the
dot product
is positive or negative. Yay math!
If they are, it's moving
away
from the surface, so it says there's no collision. If they're moving in opposite directions then it flags a collision if it's already intersecting or inside of the object, or if it will hit it in this next tick step.
Note that even if the projectile is
inside
the object the normal check comes first: if it's made it over halfway through the object, such that the closest point is on the far side and the normal is in the same direction as movement, then it stops triggering collisions.
What happens when it collides with something
Once the
LightSpellProjectile
has collided with something (or multiple things) it runs an "OnHit" function (
0x007A30FD
) for each collision, triggering what's supposed to happen. The function handles a lot of things that aren't relevant for Magelight
6
Like
wards
blocking spells at
0x006ECE30
or targets getting knocked back by Destruction spells at
0x006E99D0
.
but finally either applies the effect to a person
7
Technically a magic target, which also includes stuff like activation triggers.
(
0x00664740
) or the collision surface. Most single-cast spells handle the XP award as part of the person branch, but there's a special check at
0x00660353
that handles giving XP for spells that can handle alternative targets, like Magelight.
8
Actually not
like
Magelight,
just
Magelight. Technically it's all Reanimate and Light spells, but all of the
Reanimate spells
override this location to just grant XP when you raise the body from the dead, and the only other Light spell is
Candlelight
, which isn't a projectile.
Finally "OnHit" adds the impact at
0x007A2560
. It makes sure it only runs once, and checks the material that was hit to run accordingly (attaching to a person, for instance, or just floating near the wall for a wall). If there is no material, or it's an unsupported material, then no impact event happens.
Stopping the projectile
Everything that's been described so far seems to work, and it mostly does, but here's where it breaks down. After colliding with something and running "OnHit" the code runs a "HandleHit" function to determine if the projectile should keep going or not. Usually it will stop the projectile, but there are two exemptions: it will keep going if it collides with something 1) non-corporeal
9
To be more specific,
BROAD_PHASE_PHANTOM
from Havok's collision code, which is used for triggers (stuff like "if you walk here activate the trap/cutscene/dialogue"). This check happens at
0x0079f022
.
or 2) has the layer type 26 or 28, which corresponds to
L_TRANSPARENT_SMALL
and
L_TRANSPARENT_SMALL_ANIM
.
10
This exemption is at
0x0079f063
if you're
curious
future me looking for proof.
L_TRANSPARENT
, which doesn't collide with
L_SPELL
L_TRANSPARENT_SMALL
, which
does
collide with
L_PROJECTILE
and
L_SPELL
Because
L_TRANSPARENT_SMALL
'collides' with projectiles and spells in its collision layer, the surface continually registers collisions without stopping the projectile, leading to the bug.
The bug breakdown
You cast Magelight until it enters a bounding box with
L_TRANSPARENT_SMALL
. That collision layer collides with
L_SPELL
so it counts as a collision. It's not a person or magic target, so it hits the fallback path and awards 38 XP. But
L_TRANSPARENT_SMALL
is exempt from the "HandleHit" and so the projectile keeps going. Next tick, same thing happens. It's in the bounding box, it counts as a collision, it awards 38 XP. It will keep doing this until the closest point has a normal in the same direction that it's going, which is roughly when it's gone halfway through.
But where are the
L_TRANSPARENT_SMALL
objects? You guessed it: the mountain right near Solitude.
11
Ï made the walls visible using a quick mod built off of the
Skyrim Creation Kit
, but you can do the same thing with .esm tweaks to add a debug color for the collision layer.
There are some other places that have it, like the sides of
High Hrothgar
:
Or around the statue in
Irkngthand
at the conclusion of the Thieves' Guild questline:
But the biggest ones are all next to Solitude, which is why you have to cast at
that mountain
, not just
any
cloudy mountain that's far away.
How to maximize
Most of the guides say to aim at the top of the mountain, but what you actually want to aim at is whatever line maximizes the amount of time that the spell will be passing through the
L_TRANSPARENT_SMALL
blocks (and in the right sections of them). The mountain is good, but we can do better.
12
The top of the mountain gives ~280 triggers vs. ~850 in the optimal place.
Ï quickly
13
And I mean
quickly
, Claude Opus 5.5 two-shot the mod in about 10 minutes with only one screenshot to show what was wrong.
wrote a mod using
SKSE
that would render the invisible walls and react to where you were aiming the spell, dynamically calculating the number of expected XP triggers as well as the optimal angle to maximize XP gain. Just walking around the area quickly found the optimal location.
14
A lot of math also found the 'optimal' location, but it was underground and inaccessible. Ideal.
Starring a nearly unreadable font choice
To get from Level 15 to Level 100 Alteration you need \(\sum_{L=15}^{99} 2 * L^{1.95}\) =
528,804.0234 XP
. At 37.938 XP per trigger, that's almost 14,000 triggers, so a line that does 250 triggers vs. 850 triggers is the difference between 55 casts and 17 casts. The key bit, though, is a consistent repro—it's all good to say "aim at this spot in the sky" but people can't follow that. So here's a followable guide!
Consistent setup guide
Start by traveling to Solitude and then exiting the main door, so you're right outside. We're going to want to get up on the rock to the left. There's an invisible wall there, so line up against it and the castle wall in the farthest corner.
We're going to be aiming at the mortar line of the top left brick of the wall we're pressed up against.
To get the right angle we're going to
crouch
, and then line up the center of our crosshair with that mortar line (optimal line is the top left corner of the brick one to the right, but there's some leeway).
We're then going to strafe forward and left so we're pressed up against the invisible wall the whole time, keeping our angle consistent. This will show the two lines of the castle wall from the different pieces.
Our goal is to strafe back until the second line has
just
been hidden again. At this point you should be staring at a seemingly random patch of sky.
Cast away!
15
I got to level 100 in 17 casts in 1 minute and 15 seconds. If you'd prefer a video walkthrough, I
uploaded one to YouTube
.
16
I used Claude's computer use for the first time to build in support for the Skyrim-themed UI popups in DaVinci Resolve and watching it work was yet another "this is magic" moment. Perfect for one-offs.
With this, hopefully,
hopefully
, I'm done with Skyrim.
Bits from Debian: Looking for artwork for the next Debian release: Forky
PlanetDebian
bits.debian.org
2026-10-07 05:00:00
Each release of Debian has a shiny new theme which is visible on the installer,
the boot screen, the login screen and, most prominently, on the desktop
wallpaper. It's a very important part of a Debian release as it is usually the
first thing new users see when installing or booting the system for t...
On Wed 07 October 2026
with tags
trixie
artwork
Written by
Jonathan Carter
Each release of Debian has a shiny new theme which is visible on the installer,
the boot screen, the login screen and, most prominently, on the desktop
wallpaper. It's a very important part of a Debian release as it is usually the
first thing new users see when installing or booting the system for the first
time. Not only that, but also the first thing appearing when debianites share
their screen or connect to an external device when giving talks around the
world.
As the most enthusiastic users will know, the forky - yes, that is the name of
the next stable version, Debian 14 - release cycle is rapidly approaching its
latter stages. This means we need to select the artwork shipping with forky
soon, really soon!
And as with everything else in Debian, collecting artwork is a collaborative
effort which Debian shares with its community. So, if you would like (or know
someone who would like) to create a desktop look and feel that will be seen by
trillions of people around the world - and in space! - be sure to send in your
artwork ASAP!
The deadline for submissions is: 2026-11-26.
For the most up to date details please refer to the
Debian
wiki
.
At the same time, we would like to thank Elise Couper for creating the
Ceratopsian theme
for
our last trixie release.
And for the interested or the curious ones, the artwork is usually picked
based on which theme looks the most:
''Debian'': admittedly not the most defined concept, since everyone has their
own take on what Debian means to them. Though, usually they all agree when
something looks like Debian.
''Plausible to integrate without patching core software'': as much as we love
some of the insanely looking themes, some would require heavy GTK+ theming
and patching GDM/GNOME.
''Clean and well designed'': without becoming something that gets annoying to
look at a year, or ten, down the road. Examples of good themes include
Emerald, Homeworld, Joy, Lines, softWaves and futurePrototype
He Confessed to an Afghanistan War Murder. The U.S. Military Promptly Closed the Case.
Intercept
theintercept.com
2026-10-07 05:00:00
A veteran walked into a California police station to make a horrifying confession.
The post He Confessed to an Afghanistan War Murder. The U.S. Military Promptly Closed the Case. appeared first on The Intercept....
In the spring
of 2023, on a partly sunny day where the temperature flirted with 70 degrees, an Army veteran of the Afghanistan War walked into the unique
Prairie-style
police headquarters in Banning, California. In a
private room
in the detective bureau, the former specialist, who served with the storied 82nd Airborne Division, sat down with a special agent from the Army’s Criminal Investigation Division — better known as CID — to make a confession.
The veteran had served with roughly 20 mortarmen and scouts from a platoon in the 2nd Battalion, 508th Parachute Infantry Regiment assigned to Combat Outpost Ware in Afghanistan’s Arghandab River Valley in 2010. A lush land of grape groves and pomegranate orchards, towering sunflowers and marijuana plants, the Arghandab was crisscrossed by canals and was a flashpoint in Kandahar Province, the spiritual homeland of Afghanistan’s once and future rulers, the Taliban.
The soldiers from the 82nd Airborne played a cat-and-mouse game with the Taliban. America’s official policy was
counterinsurgency
, or COIN: a re-tread of a failed strategy of the Vietnam War that promised a kinder, gentler brand of warfare aimed at winning Afghan hearts and minds. But in the Arghandab Valley, the COIN of the realm was kinetic.
The Taliban seeded the orchards with improvised explosive devices set off by a foot on a pressure plate, a boot on a tripwire, or Taliban hands gripping a remote control. The Americans countered with overwhelming firepower:
attack helicopters
, armored gun trucks, and infantrymen toting M4 carbines. It was a “rough time for the unit,” another veteran of the platoon remembered. They saw a great deal of “contact,” engaging the Taliban more times than another unit member could recall.
But the former mortarman in Banning had come to talk about one particular incident that, more than a decade later, he could not forget.
The CID agent gave the veteran a “cleansing waiver of rights” to read and sign. It advised the veteran that he had the right to remain silent, that any statement he made could be used against him at trial, and that he was providing testimony “freely and voluntarily.” The veteran waived his rights and signed the document on March 28, 2023 at 1:31 p.m. He acknowledged, according to exclusive criminal investigation documents obtained by The Intercept, that he killed an unarmed man and wounded a woman. More than that, he said that he “needed to be held accountable for his actions.” He told CID that he had “committed a war crime.”
A U.S. medical helicopter arrives to evacuate a soldier from the 1-320th Alpha Battery, 2nd Brigade of the 101st Airborne Division, in Arghandab Valley, Kandahar, Afghanistan, on July 30, 2010.
Photo: Rodrigo Abd/AP File
Twenty-five years ago,
in the wake of the 9/11 attacks, an angry America went to war in Afghanistan. The conflict would become, by some counts, the longest in U.S. history, stretching across four presidential administrations.
It ended where it started
, with the Taliban in control of the country after a crushing 2021 defeat for the U.S. and the
Afghan allies it propped up for two decades
.
The U.S. lost around 2,400 troops in Afghanistan at a staggering cost of around $2.3 trillion. Afghans suffered far worse. In October 2001, President George W. Bush
announced
that “the oppressed people of Afghanistan will know the generosity of America.” But from almost the opening American salvo, they began dying at U.S. hands.
During just the first five months of the conflict, there were no fewer than 136 media reports of U.S. attacks that killed between
1,200 and 3,155 civilians
. Year after year, U.S. and coalition airstrikes killed Afghans – sometimes scores of civilians, sometimes
hundreds
. The deaths continued to the war’s final moments when an August 2021
drone strike
targeting a suspected terrorist in the country’s capital, Kabul, instead
killed 10 civilians
, seven of them
children
. More than 70,000 Afghan and Pakistani civilians are estimated to have died as a direct result of the war, according to Brown University’s
Costs of War project
.
While the carnage of the American air war garnered the biggest headlines, Afghans were also gunned down, again and again, by U.S. ground troops. Following a 2007 ambush, Marines
killed 19 civilians
and wounded 50. An initial investigation found the Marines used excessive force when they killed civilians after a suicide bombing. Later, the Americans were
recast
as
victims
. In 2010, a “kill team” from the Army’s 5th Stryker Brigade
slew at least three civilians
over four months and mutilated the dead for
war trophies
. Eleven of the 12 soldiers charged were convicted of crimes, but most were released from prison by the mid-2010s. Staff Sgt. Robert Bales
massacred 16 Afghan civilians
in 2012. A Special Forces A-Team was accused of
killing close to 20 civilians
in 2012 and 2013. The remains of 10 victims were later found buried outside their base. A U.S. investigation exonerated the Americans of any wrongdoing. Only one man was punished:
Zikria Kandahari
, the unit’s Afghan interpreter.
Pete Hegseth, now the self-styled Secretary of War, served with the Army National Guard in Afghanistan during that era. In Trump’s first term, when he was a Fox News personality, Hegseth
lobbied on behalf
of three men who were convicted or facing charges related to war crimes committed in Iraq and Afghanistan. Trump intervened in all three cases, granting two
pardons
and
reversing a demotion
.
During his confirmation
hearing
last year, Hegseth complained that “restrictive” rules of engagement had made it “difficult to actually do your job on the battlefield,” repeatedly recalling his own service. “I was the senior counter-insurgency instructor in Afghanistan. My job was to understand how the Taliban and Al Qaeda operated,”
he explained
. “They knew our rules of engagement, and when they were more restrictive, they took advantage of them.” He continued: “We don’t need burdensome rules of engagement that make it impossible for us to win these wars.”
But what the veteran in Banning described to CID agents was anything but burdensome.
In April or May
— he couldn’t recall which — of 2010, that veteran said he was carrying an M4 carbine as his unit moved about 3 kilometers into the Arghandab and occupied an abandoned building at the edge of a village. When the Americans were attacked, they returned fire, killing a couple Taliban members.
The unit readied themselves, expecting the Taliban to regroup and mount another assault. The Americans intercepted a radio transmission. A translator told them that a single male militant on a motorcycle, transporting rocket-propelled grenades, was headed toward them, according to the documents. The unit’s chain of command authorized the use of force.
When a motorcycle rider came into view, a sniper took aim and fired off one round. He missed.
The respite for the motorcycle man was short-lived, however. About 500 feet away, close to 20 Americans raised their rifles.
An
M4 carbine
— a magazine-fed, gas-operated, air-cooled, shoulder-fired weapon — shoots lightweight bullets at extremely high speeds. In the spring of 2010,
M855
5.56-mm
bullets
were de rigueur for U.S. forces. Fast with low mass, the rounds produce a tremendous amount of kinetic energy. When they slam into a human body, that energy is transmitted to the soft tissue of the victim creating a “wound channel” that causes
massive blood loss
.
The Banning mortarman fired off one round. Others from the unit opened up, too.
The motorcycle man never reached his destination. If someone was waiting for him to arrive on that spring day in 2010, they still are.
The Banning veteran and others walked up to the Afghan man and “witnessed a military aged male deceased with multiple gunshot wounds to the face and other areas of his body,” wrote the CID special agent, transcribing the mortarman’s testimony. “No RPG’s were present.”
Nobody knew how many rounds hit him, but the damage to the Afghan man’s body was so extensive that he was “unidentifiable,” according to another witness. The mutilation was so devastating that the Americans were unable to even collect his “biometrics”: fingerprints, iris scans, facial images.
The dead man, said the veteran, wasn’t alone. There was a wounded woman, too. He told the CID agent that the “woman was crawling on the ground.” There may also have been a baby or child present as well — CID’s redactions make it impossible to know for certain.
After the veteran
made his confession, in late March 2023, Army criminal investigators started consulting U.S. sources. Despite detailed unit records for more than 100 notable incidents in April and May 2010, an Army historian could locate none for the killing of the motorcycle man.
When CID began interviewing unit members, the company commander — by then a lieutenant colonel — said the incident did not sound familiar to him. The unit’s first sergeant told CID that while he saw most “troops in contact” reports, he didn’t recollect the incident. He acknowledged that the details “sounded plausible.”
Others recalled the incident in detail but excused it. One soldier who never saw the motorcycle man before he was killed said, for example, that it “sounded like” the man handed off the RPG to someone before he was shot. Another former specialist — who by 2023 was a warrant officer — told CID that unit members “fired on the military age male on the motorcycle.” He also recalled that “a woman was on the motorcycle,” too. He told the Army investigators that it was “unfortunate that no RPGs were present” but believed “the incident was handled properly” since “they were authorized to fire from their command team.”
Having never even attempted to identify the Afghan man, the woman on the scene, or the possible child; having never sought out any evidence beyond the testimony of members of the unit; and despite a signed, sworn confession by one of the shooters, Army CID concluded — within six weeks — that there was not “probable cause to believe that any soldier committed the offense of murder.”
CID refused to make the special agent in charge of the case available for an interview. “Unfortunately, we are unable to support your request to interview DACID personnel involved,” Special Agent Mark Lunardi told The Intercept, before leaving his job. More than a year ago, after multiple follow-up requests for information, Thomas Hamilton of CID’s press office wrote: “I will get back to you as soon as I can.” He has, despite repeated follow-ups, failed to respond.
Soldiers from the 1-320th Alpha Battery, 2nd Brigade of the 101st Airborne Division, patrol near COP Nolen, in the Arghandab Valley, Kandahar, Afghanistan, on July 26, 2010.
Photo: Rodrigo Abd/AP File
A
2020 study
of post-9/11 civilian casualty incidents found most have gone uninvestigated. When they did come under official scrutiny, the “military too often relies solely on its internal records and sources —which can be flawed and incomplete — to assess civilian harm,” the Center for Civilians in Conflict and Columbia Law School’s Human Rights Institute reported.
“There were probably hundreds of incidents that were just like it that we’ll never know about.”
These findings were bolstered by the groundbreaking work of the podcast “In the Dark.” Their 2024 investigation — a corollary to a reexamination of the
2005 killings of civilians in Haditha, Iraq
— found U.S. troops had committed
781 possible war crimes
in Afghanistan and Iraq that the U.S. military actually investigated. In at least 65 percent of cases analyzed by “In The Dark,” investigators didn’t believe that a crime had taken place — even in cases where perpetrators had, like the Banning veteran, confessed. The 2010 killing of the Afghan motorcycle man is, however, not one of the cases in their database.
In the 151 cases in which investigators found probable cause that a crime had occurred, that the rules of engagement had been violated, or that a use of force hadn’t been justified, true accountability was rare. While 572 alleged perpetrators were involved in these cases, only 130 were convicted; the records show they rarely received significant prison terms. Much more often, their cases were dealt with by commanders with broad discretion to punish their troops with extra duty, demotions, or reprimands, in lieu of formal prosecutions. Fewer than 1 in 5 perpetrators, investigators found, appeared to have been sentenced to any type of confinement. The median sentence was just eight months.
Todd Huntley, an active-duty Judge Advocate for more than 23 years who deployed to Afghanistan twice with a Joint Special Operations Task Force, believes killings like the one described by the Banning veteran were a common occurrence. “There were probably hundreds of incidents that were just like it that we’ll never know about,” he told The Intercept. A current government official, who was involved in the wars in the Greater Middle East, said the total could reach into the thousands.
Since taking the
helm of the Pentagon last year, Hegseth has launched overlapping efforts to weaken transparency, scuttle accountability, hobble military justice, and undercut protections for civilians in conflict zones — from replacing the
Pentagon press corps
with pro-administration
sycophants
, and firing the top
legal authorities
of the Army and the Air Force, to pursuing changes that would encourage lawyers to
approve more aggressive tactics
and take a more lenient approach to those who
violate
the laws of war.
The Intercept has previously
reported
on Hegseth’s
gutting
of civilian harm mitigation and response efforts. In late April, Hegseth repeatedly
dismissed
congressional concerns about civilian harm and respect for the laws of war in testimony before the House Armed Services Committee. “The Department of War fights to win,” Hegseth
replied
when asked if he stood by his statement that the U.S. would afford enemies “no quarter” — a war crime. A recent
report
by the Pentagon’s top watchdog says cuts to civilian harm mitigation and response efforts have been so severe under Hegseth that the United States cannot adequately protect civilians in conflict zones. This was driven home at the outset of the Iran war when the U.S. attacked the Shajarah Tayyebeh elementary
school
, killing more than 150 people,
most of them children
.
The CID records of the motorcycle man investigation are redacted in ways that make it impossible to ascertain key facts about additional civilian harm. “An [REDACTED] woman was on the motorcycle with the military aged male and [REDACTED] from the engagement,” according to the summary of a statement by one unit member. The Banning veteran also “stated he [REDACTED] an unarmed [REDACTED] woman during his deployment.”
The redactions also prevented The Intercept from identifying the veteran and fellow unit members to gather additional information. “If it wasn’t for the soldier involved coming forward, this incident would never have come to light,” Huntley, now the director of the National Security Law Program at Georgetown University Law Center, told The Intercept. “The 13-year delay in reporting made any investigation difficult as others involved couldn’t remember details and the redactions in the report make it impossible to learn anything else about the incident.”
Only one member of the unit contacted by The Intercept agreed to speak about the incident. The veteran acknowledged, on background, that the incident occurred, but would not say if he was one of those interviewed by CID.
Still, the case of the motorcycle man offers a stark counter-narrative to Hegseth’s assertions about the war on terror — and the basis for his efforts to undermine protections for civilians today.
The rules of engagement in Afghanistan in 2010 were so lax that a man was killed for no reason and with no repercussions.
It was so unremarkable that no record of the engagement could be located. It was so commonplace that senior unit members couldn’t recall it.
It was so typical that only one man even thought to report it — and he waited 13 years.
It was so characteristic of U.S. military policy that no inquiry was conducted into the rules of engagement in force at the time.
And the incident was so in keeping with
U.S. military tradition
that, despite a murder confession and little question that an unarmed man was gunned down, Army CID and the 82nd Airborne’s chief of military justice quickly concluded there was no probable cause for charges, and the investigation was closed.
Police routinely failing to investigate ‘revenge porn’, research reveals
Guardian
www.theguardian.com
2026-10-07 04:00:38
Lawyers launch super-complaint over finding that officers in England and Wales frequently dismiss image-based abuse Police forces are routinely failing to investigate complaints about the circulation of image-based abuse, previously known as “revenge pornography”, according to research that has trig...
Police forces are routinely failing to investigate complaints about the circulation of image-based abuse, previously known as “revenge pornography”, according to research that has triggered a super-complaint over the way such crimes are handled.
Lawyers have gathered evidence from 100 people who reported intimate images being shared online without their consent. Their research suggests that image-based abuse is still frequently dismissed or not recognised as criminal offending, despite being described by the former prime minister Keir Starmer as a national emergency.
Referrals to the Revenge Porn helpline, the charity that helps victims get non-consensual images removed from the internet, increased by more than 4,000% between 2015 and 2025, but data from 39 police forces in England and Wales shows that charges were lodged against perpetrators in an average of just 4.5% of cases last year. In some forces the figure was lower than 1%.
Some police officers responded inappropriately when women tried to register complaints, researchers found. One reported that the officer “almost giggled on the phone to me”, while another said the officer responded: “Boys will be boys.” Another said: “I bet you’ve learned your lesson.”
Image-based abuse overwhelmingly affects women and girls, who make up 98.9% of reported images.
Lawyers from Leigh Day and the charity the Centre for Women’s Justice are working with the charity #NotYourPorn to highlight systemic police failings in responding to image-based abuse, defined as “the taking, sharing or creating of intimate content without the survivor’s consent, or threatening to do so”.
They note that rapid innovations in AI are helping abusers find new ways to create and disseminate synthetic explicit material, sometimes known as deepfake content.
Their research classifies the policing response as “fragmented, inconsistent and incapable of meeting the scale of the threat posed by image-based abuse”. “Despite the increased public awareness, and the creation of new criminal offences designed to address this conduct, there has been no corresponding increase in charging rates, prosecutions or convictions,” the complaint says.
“The unavoidable conclusion is that the overwhelming majority of this offending goes unaddressed.”
Home Office data shows that there were only 569 charges issued for image-based abuse in 2025/2026. The Revenge Porn Helpline received almost 25,000 reports in 2025.
Researchers found that the vast majority of survivors of image-based abuse did not report it to the police, noting that “embarrassment, shame and fear” contributed to their reluctance. Police outsourced the evidence-gathering process to survivors in many cases, the report notes, leaving them to work out how to collate proof that the images were being disseminated widely and how to get the content removed from the internet.
“It felt like to them it wasn’t a priority, it wasn’t even a crime,” one woman told researchers. Another said: “Police involvement has not alleviated the harm caused by this offence. Instead, it has intensified my distress and left me feeling powerless and completely unsupported as a victim.”
A police super-complaint is a mechanism introduced in 2018 to allow concerns to be raised about ways in which the police operate that may significantly harm the interests of the public generally, rather than focusing on an individual force of officer.
The complaint will be jointly considered by the College of Policing and the Independent Office for Police Conduct, which will decide whether to respond by opening a wider investigation.
Harriet Bland, a lawyer with the Centre for Women’s Justice, said: “Without an urgent overhaul of police guidance and training on image-based abuse – which must be regularly updated – and a clear national strategy rolled out across forces, survivors will continue to be let down.”
Forever Junior: The Skills AI Can't Develop for You
Software engineering is going through an AI revolution we’re all still trying to understand. For years, we’ve been told software engineers are six months from extinction. Some still believe the job is disappearing. Others say it’s simply changing.
My experience says it’s the latter.
My fear with AI
I’ve seen a real shift in priorities for engineers, but the job itself is very much still needed. At Criteo, engineers are still hired, still valued — and increasingly asked to lean on AI agents to boost productivity as much as possible. But every one of those messages comes with the same addendum: keep a firm hand on the wheel. Use the tool, but don’t stop knowing what you’re doing.
As a Junior fresh out of school, “still learning what I’m doing” isn’t a caveat — it’s the job description. So it’s tempting, and a little dangerous, to let an agent not just do the work, but learn it in my place. It knows much more than I do, after all!
We were taught in school that the way to get better at this job is to just code. Now we’re told not to code directly if we can help it. So — how are we supposed to learn? That’s the extinction I’m actually afraid of: not Software Engineers disappearing, but Seniors disappearing — as we all become forever Juniors, only ever qualified to review one agent’s work with the help of another.
So what was I supposed to do about it?
I started by researching two questions:
What has AI actually changed in my job description?
What’s a Senior, concretely?
To answer the former question I just had to listen to discourse online and at work. Many are saying that reviewing code is now more important than producing it. The consensus seems to be that learning coding syntax and languages is obsolete. And of course companies want us to focus on things like token management to optimize costs.
For the latter question, I observed Seniors at work, specifically in my team, and picked up some keys to what Seniorhood meant. On my team, the Seniors are the ones who can defend code they’ve shipped, argue a design choice instead of just having one, tell when an agent is heading the wrong way, know when to delegate and when to just do it themselves — and take the time to teach the rest of us.
With these guidelines, I made a map of where I needed to grow with my first milestone being to get nominated for a promotion by the end of the year. As I follow this map, I hope to reach Seniorhood one day.
The next question was how. Clearly, I needed to develop skills.
I started, like a lot of engineers did, writing a skill.md: instructions and patterns fed to a model so it would behave more like a mentor teaching me about Senior behaviors. It’s useful, sure. But somewhere in that process I noticed the real gap wasn’t in my agent’s behavior — it was in mine. So the focus of this article isn’t really about the skills I wrote. It’s about the ones I had to develop in myself to keep up with the ones I was developing with the agent. The ones that I believe will actually allow me to evolve from Junior to Senior.
Questioning
I started by instructing my agent to ask me Socratic questions to test my understanding of the plan or the code itself. A little pop-quiz at the end of each decision or code change. In practice, the questions were relevant only half of the time and they were mostly useful to prompt me to look closely at the plan or code and talk with the agent to make sure I understood it.
# Understanding checks and delegated work
- After implemented work, prompt understanding—rotate style: e.g. explain in their own words, predict what breaks if X changes, walk through one failure mode, defend a trade-off in one sentence. Gently correct and re-check if needed.
- Subagents or delegated tasks must follow the same rules: no unconfirmed mutations and no unconfirmed mutating verify (subject to the same-task verify rule after approved implementation).Code language:Markdown(markdown)
That’s when I noticed I was building a human skill of my own. One every five-year-old already has mastered as they ask “but why,” on repeat, until someone caves.
Curiosity
.
The best trick an LLM has is that it speaks your language back to you — so use it. Ask your agent to explain a decision you don’t follow. Ask it to cite a source when it name-drops something unfamiliar. Ask it to justify a suggestion that just feels off. That last one is the real trick: making an agent explain a flawed solution out loud is often how it catches its own mistake — and how you catch it too.
A Senior can defend code they’ve shipped to production. They can argue their design choices. They can tell when an agent is going the wrong way and redirect. Question your agent enough, and you’re rehearsing exactly that.
Yes — this costs tokens. Fair. But I’ve found it’s more efficient overall to slow down here, because code I actually understand before I ship it works in production far more often. And when it doesn’t, I can fix it myself instead of round-tripping with the agent again. I’d rather use AI a little less often, and get properly curious every time I do, than push code that’s only half-understood and pay for it later.
Reviewing
If you’re coding with an agent as much as possible, you’re going to spend most of your time reviewing its work instead. Reviewing is its own skill — arguably the one you’ll need most to actually reach Seniorhood — especially in this AI landscape. The good news is, you’re already sitting on a training set for it: every comment a Senior has ever left on your code.
Pay attention to the patterns. Constantly getting notes on naming? Your team clearly cares about naming conventions — check every name in your agent’s code twice. Do it enough, and you start to notice the practices that are specific to your industry, your company, even your team. And with that you can practice
emulation
.
# Similarity to existing code and reusability
For each substantive task:
1. Discover — Search the codebase (and/or ask the user) for the closest existing feature (same kind of entity, resource, table, API surface, etc.).
2. Align — Summarize candidates; ask which analogue is best. If nothing fits, ask the user to confirm the work is net-new.
3. Anchor — State the chosen reference (paths, components, patterns) and keep it for the rest of the thread.
4. Plan and implement (when confirmed) — When outlining or implementing, call out divergences from the reference. If the reference is one-off, ask whether a small shared abstraction could serve both without over-engineering.
5. Information gaps — If the user omitted details the reference implies (fields, ownership, API shape), ask before assuming.Code language:Markdown(markdown)
My team owns a product catalog. Every time I added a new entity type in it, Seniors would say some version of “this looks a lot like this one we already have.” So I started using that existing entity as a template — and when I pushed code for review, I watched reviewers scrutinize every deviation from the template entity. The pattern was unmistakable: justify any deviation, reuse whatever you can. Once I saw it, I could apply it myself, reviewing my agent’s output before a colleague ever had to.
I took it a step further and wrote the pattern into my skill — the markdown-file kind, this time, fed straight to the agent. Now it has the pattern in mind before I even open the review. You can push this further too: point AI at your own git history or team chat and have it surface the reviewing patterns for you.
Funny enough, that’s the two meanings meeting in the middle: I built a skill so I could get better at a skill.
The file makes the agent’s first draft look more like my team’s code; reviewing that draft anyway, over and over, is what slowly turns me into a better reviewer.
In the end, hopefully, I won’t be a Junior with a good skill file. I’ll be the Senior who wrote the comments that went into it.
Handwriting
I’m sorry to report that the thing that’s helped me most is pretending, occasionally, that I’m back in a school exam where use of tools like AI is “cheating”. This is how you build the skill Seniors already have in spades:
autonomy
.
Seniors did this job without agents for years.
They can argue a design choice because they’ve had to make those choices themselves, and live with them.
They can tell when an agent is heading the wrong way because they used to code without one. Having an agent sitting next to you at all times can keep you from this “learning to swim by being dropped in water” method. I think it’s a method worth trying out every now and then.
I actually built a constraint directly into my
skill.md
for this: it isn’t allowed to touch a file, run a migration, or commit anything until I’ve explicitly confirmed the plan. Read-only exploration, sure — search, explain, propose. But the moment it’s time to actually write code, it stops and waits for me.
## Read-only (always allowed)
Search, grep, read files, list directories, and any non-mutating inspection are allowed without implement confirmation. Use them for discovery, teaching, and finding analogues in the codebase.
## Mutations (confirm first)
Do not edit files, apply patches, git commit/push, install dependencies, run migrations, write config, or run state-changing shell until the user confirms implementation of a concrete plan (files, areas, commands)—not merely answering a Socratic question.Code language:Markdown(markdown)
On paper, that pause should be enough. In practice, just saying “go ahead” the second it asks kept happening, and it didn’t teach me much. So every so often — rarely, deliberately — I use that same pause to actually make the change myself. Treat it as a test: small enough not to tank your output, meaningful enough to actually count. I suggest doing it for a minor tweak to a feature you’ve already shipped a large chunk of. No onboarding required, and it puts your understanding of everything your agent wrote — and you reviewed — to the test.
I’ve found this pays off in a few specific ways:
Architecture
Design is still a core skill in this job, and it’s exactly the one you can quietly lose your grip on without noticing. Interacting with an agent-built architecture by hand forces you to actually feel where it’s wrong — and seeing why a design failed is the first step to building a better one. A Senior who owns a system is the person who knows why it’s shaped the way it is.
Algorithmic understanding
When you only ever review, you tend to know the what without always knowing the how — until someone asks a clarifying question and you have to go find out. One hand-written change forces you to face these questions yourself, before it shows up in production. Answering that question in the moment, without going to look, is a Senior move — and it’s built one hand-written change at a time.
Token efficiency
This one’s new, and it costs the company real money. How can you learn about token efficiency when you’re not using AI at all you might ask? Because understanding your own code and architecture is how you can be pertinent in your token usage. It’s directly linked to the two previous points:
knowing where things are
: file discovery is expensive — the better you know where things live, the less you burn finding them.
knowing the how
: some jobs are still just faster by hand. Your IDE renames a variable across a whole codebase instantly. If you know the how, you already grasp the blast radius of that change. So there’s no reason to spend tokens finding files and rewriting them one by one.
Seniors know when to delegate and when to just do it themselves. You need the same instinct — even when what you’re delegating to is a machine.
Confidence
A bit of a bonus skill, but a real one. I’ve heard engineers at every level say they could never go back to coding without an agent. Occasionally doing it anyway is how you remember the agent is a tool for doing your job faster — not a replacement for being able to do it at all. Test the confidence, and make sure it’s earned. And honestly, I got into this job because I liked writing code. Doing it by hand once in a while is still the part I enjoy most, which is reason enough on its own.
The Seniors I trust aren’t confident because they have an agent. They’re confident because they’d still excel at their job without one, and they know it. That’s worth aiming for.
Mentee-ing
No matter how capable AI gets, or how good you get at using it, nothing replaces mentorship from an actual Senior. You learn how to become one by working with one. Seniors have to take the time to explain things so you can grow into that. We can’t lose sight of the fact that we work in teams for a reason.
Collaboration
is the last skill on this list.
You’d think that’s the one skill you can’t build into an agent at all — but I tried anyway. My skill actually nudges me to go ask a Senior whenever a decision is really a team or a judgment call, and it’ll even draft the Slack message. But I’ve found that copying that Slack message was a slippery slope to copying the answer too. Soon enough, you’re just doing the job a good Slack MCP can do, and that’s the real AI takeover. The one you let happen. So draft your own message with your understanding of the issue and your own questions. You can even send your draft to the AI so it can correct your understanding before you contact a Senior. But don’t be afraid to make a mistake. You’re messaging that Senior so that they help you understand. They’re meant to point out your mistake.
Be the eager intern who isn’t afraid to ask the stupid question. It’s tempting to rush toward the next level, but the level you’re at right now comes with something worth using while you have it: people whose job includes looking out for you. Seniors carry a responsibility we don’t have yet — showing us the way up. No Junior becomes a Senior by practicing alone in a room. We need someone already there to show us what “good” looks like and correct us when we’re wrong.
Of course, you may be a Senior reading this and thinking the Senior job description has changed with AI too — especially the mentoring Juniors part. Mentoring Juniors in an agentic world is its own challenge, and worth its own experiments — ones I’m not the right person to run. I’d read that article, though. It’s another skill further up my map to Seniorhood.
Results
So —
skill.md
or skill issue? I started my little experiment of summoning my skill every single time I used an agent. Now I barely open it once a day.
When I built this markdown file with all the ways I wanted my agent to behave, I was really writing down all the ways a Junior Software Engineer like me should behave. As I built up the skills myself, I needed the agent less and less.
Curiosity, emulation, autonomy, collaboration
— none of these come from the agent, skill file or not. They come from what you choose to do while the agent is running. That’s the actual thread here: every one of these is a way of staying an active participant in your own learning, instead of a passive reviewer of someone else’s output. The agent can write the code, and it can even write the questions. You have to do the reps if you don’t want to be thrown out of the loop.
Of course these four skills are not the whole list. Scoping work, estimating it, handling an incident at 2 am, knowing what not to build, carrying a decision across teams, mentoring Juniors — a Senior does all of that too, and they’re all skills I’m hoping to build one day. For now I’m going step by step, and I’m happy to say I’ve completed the first milestone. My experiments with AI and skill-building have helped me get that promotion nomination actually even earlier than planned!
So keep asking “but why.” Keep noticing what your reviewers always push back on. Keep the rare, deliberate change you make by hand. None of it will slow you down as much as you think — and all of it is what turns a Junior who reviews good code into a Senior who can write it, explain it, and one day, pass it on.
And Seniors, we’re going to need your help to get there. We’re working out how to keep learning in an agentic world; I’d love to read how you’re working out how to keep teaching in one. But neither of us can do it alone: room to learn is something Juniors have to take, mentors have to give, and Leadership has to allow for. Agents can’t make that space. People have to.
OpenBSD’s
httpd
can now manipulate HTTP response headers directly in
httpd.conf
. Until now, adding
security headers or stripping headers from a FastCGI backend meant changing the application code or
putting
relayd
in front of
httpd
. For a simple setup, that was often more work than it should be.
The new header directive has three options.
set
adds a header or replaces an existing one with the
same name.
add
appends a header even if one with that name already exists.
remove
suppresses a
header, whether
httpd
set it or the backend sent it or it was inherited from the server context.
By default, headers only apply to 2xx and 3xx responses. Add always to include error responses as
well, which you usually want for security headers. Headers set in the server context are inherited
by location blocks, and a location can override or remove them by name.
Here is a practical example: a personal blog with static pages, cacheable assets, a drafts folder,
and some PHP.
server "blog.sizoefovid.org" {
listen on * tls port 443
tls {
certificate "/etc/ssl/blog.sizeofvoid.org.crt"
key "/etc/ssl/private/blog.sizeofvoid.org.key"
}
hsts max-age 31536000 # built-in, no "header set" needed
root "/htdocs/blog"
# Security headers for every response, including 404/500 pages
header set "Content-Security-Policy" "default-src 'self'; img-src 'self' data:; frame-ancestors 'none'" always
header set "X-Content-Type-Options" "nosniff" always
header set "Referrer-Policy" "strict-origin-when-cross-origin" always
header set "Permissions-Policy" "camera=(), microphone=(), geolocation=()" always
# Fingerprinted assets can be cached forever
location "/assets/*" {
header set "Cache-Control" "public, max-age=31536000, immutable"
}
# Drafts: shareable by link, but not indexed and not cached
location "/drafts/*" {
header set "X-Robots-Tag" "noindex, nofollow" always
header set "Cache-Control" "no-store" always
}
# Don't reveal the PHP version
location "*.php" {
fastcgi socket "/run/php-fpm.sock"
header remove "X-Powered-By"
}
}
httpd: add header block/drop rules for request filtering
#
httpd
can now reject incoming requests based on request headers. This is useful for keeping AI
scrapers, vulnerability scanners, and other unwanted clients away, without
relayd
or a firewall rule
that only knows IP addresses.
Of course, it’s not perfect, and if the
user-agent
isn’t set, there’s not much we can do. But for a
large number of AI scrapers, this seems to be helping at the moment!
I’d like to refer you to the email from purplerain ()
secbsd ! org
, who
carried out an analysis on this and requested this feature – and didn’t just write an initial
version of it. Thanks again!
Following those tests, the latest patch is now running in production. It has been
running without issues for at least 24 hours so far.
I will continue performing stress tests and will send additional results over the
next few days.
RESULTS
Duration: 2h 51m 11s
Requests sent: 10000000
Dropped: 4820945 48.2%
Blocked: 182191 1.8%
Served: 4996864 50.0%
Redirected: 0 0.0%
Network errors: 0 0.0%
BY CATEGORY
FAKE: 5003136 sent · 4820945 dropped (96.4%) · 182191 responded (3.6%)
LEGITIMATE: 4996864 sent · 0 dropped (0.0%) · 4996864 responded (100.0%)
HTTP STATUS
200: 4996864
403: 45392
404: 136799
Here are some other use cases:
$ cat security-tools.conf
header drop "user-agent" "commix*"
header drop "user-agent" "dav*"
header drop "user-agent" "dirbuster*"
header drop "user-agent" "feroxbuster*"
header drop "user-agent" "ffuzz*"
header drop "user-agent" "gobuster*"
header drop "user-agent" "sqlmap*"
header drop "user-agent" "whatweb*"
header drop "user-agent" "wfuzz*"
header drop "user-agent" "wpscan*"
# Drop requests with and empty User-Agent by default.
header drop "user-agent" ""
# ncrack
# lbd
# legion
Purple Rain
There are two new options.
header block
answers a matching request with an HTTP status
code
and then
closes the connection. For
3xx codes
, you must give a target
URL
, which is sent as the Location
header. For all other codes, you can give an optional label that shows up in the log, so you can see
which rule matched.
header drop
closes the connection silently, without any response. This is the
better choice for scanners: they get no information back, and your server does no extra work.
Both the header name and the value are glob patterns (*, ?, […]) and are matched
case-insensitively.
Here is the blog from the previous example, with some request filtering added:
server "blog.sizoefovid.org" {
listen on * tls port 443
tls {
certificate "/etc/ssl/blog.sizeofvoid.org.crt"
key "/etc/ssl/private/blog.sizeofvoid.org.key"
}
# AI scrapers: answer with 403 and a log label per rule
header block "User-Agent" "*GPTBot*" 403 "ai-gptbot"
header block "User-Agent" "*ClaudeBot*" 403 "ai-claudebot"
header block "User-Agent" "*CCBot*" 403 "ai-ccbot"
header block "User-Agent" "*Bytespider*" 403 "ai-bytespider"
# Known scanners: no response at all
header drop "User-Agent" "*zgrab*"
header drop "User-Agent" "*masscan*"
header drop "User-Agent" "*Nuclei*"
# Log4Shell probes can hide in any header
header drop "*" "*${jndi:*"
# Legacy browsers: send them to a static fallback site
header block "User-Agent" "*MSIE*" 302 "https://legacy.example.org/"
# ... see example above
}
For inspiration, here’s a list from Purple Rain that he sent me. It contains his research. You can
save lists like this in a file and include them in your httpd config using
include "ua.conf"
.
# GENERIC USER-AGENTS
header drop "user-agent" "*bot*"
header drop "user-agent" "go*"
header drop "user-agent" "modat*"
header drop "user-agent" "axios*"
header drop "user-agent" "curl*"
header drop "user-agent" "crawler*"
header drop "user-agent" "headlesschrome*"
header drop "user-agent" "httpie*"
header drop "user-agent" "java*"
header drop "user-agent" "libredtail*"
header drop "user-agent" "libwww*"
header drop "user-agent" "lwp*"
header drop "user-agent" "node*"
header drop "user-agent" "okhttp*"
header drop "user-agent" "php*"
header drop "user-agent" "puppeteer*"
header drop "user-agent" "*python*"
header drop "user-agent" "*request*"
header drop "user-agent" "ruby*"
header drop "user-agent" "scrapy*"
header drop "user-agent" "selenium*"
header drop "user-agent" "wget*"
# AI AND TRAINING
header drop "user-agent" "*anthropic*"
header drop "user-agent" "aiwebindex*"
header drop "user-agent" "amazon*"
header drop "user-agent" "amzn*"
header drop "user-agent" "anomura*"
header drop "user-agent" "apify*"
header drop "user-agent" "aranet*"
header drop "user-agent" "awario*"
header drop "user-agent" "azureai*"
header drop "user-agent" "bigsur*"
header drop "user-agent" "bytespider*"
header drop "user-agent" "chatglm*"
header drop "user-agent" "chatgpt*"
header drop "user-agent" "claude*"
header drop "user-agent" "cloudflare*"
header drop "user-agent" "cohere*"
header drop "user-agent" "cotoyogi*"
header drop "user-agent" "cragcrawler*"
header drop "user-agent" "crawl4ai*"
header drop "user-agent" "crawlspace*"
header drop "user-agent" "cursor*"
header drop "user-agent" "datenbank*"
header drop "user-agent" "deepseek*"
header drop "user-agent" "devin*"
header drop "user-agent" "exa*"
header drop "user-agent" "facebook*"
header drop "user-agent" "factset*"
header drop "user-agent" "firecrawl*"
header drop "user-agent" "friendlycrawler*"
header drop "user-agent" "geisthaus*"
header drop "user-agent" "iask*"
header drop "user-agent" "img2dataset*"
header drop "user-agent" "imagespider*"
header drop "user-agent" "isscyberriskcrawler*"
header drop "user-agent" "kagi-fetcher*"
header drop "user-agent" "kangaroo*"
header drop "user-agent" "kimi*"
header drop "user-agent" "klaviyo*"
header drop "user-agent" "kunatocrawler*"
header drop "user-agent" "laion*"
header drop "user-agent" "lcc*"
header drop "user-agent" "lightpanda*"
header drop "user-agent" "linguee*"
header drop "user-agent" "manus*"
header drop "user-agent" "meta*"
header drop "user-agent" "mistralai*"
header drop "user-agent" "netestate*"
header drop "user-agent" "newsai*"
header drop "user-agent" "notebooklm*"
header drop "user-agent" "novaact*"
header drop "user-agent" "omgili*"
header drop "user-agent" "openai*"
header drop "user-agent" "opencode*"
header drop "user-agent" "operator*"
header drop "user-agent" "panscient*"
header drop "user-agent" "perplexity*"
header drop "user-agent" "poggio*"
header drop "user-agent" "poseidon*"
header drop "user-agent" "querit*"
header drop "user-agent" "shap*"
header drop "user-agent" "sidetrade*"
header drop "user-agent" "terra*"
header drop "user-agent" "tiktok*"
header drop "user-agent" "trae*"
header drop "user-agent" "twinagent*"
header drop "user-agent" "useai*"
header drop "user-agent" "velenpublicwebcrawler*"
header drop "user-agent" "webzio*"
header drop "user-agent" "yaK*"
header drop "user-agent" "yandex*"
# SECURITY TOOLS
# Drop reconnaissance tools, scanners, fuzzers, brute-force tools, etc.
# This is not and exhaustive list of tools.
# Dirbuster, feroxbuster, and gobuster are examples, but these could be covered
# by a single "*buster*" pattern.
# The same applies to fuzzers as ffuf and wfuzz, both use the string "fuzz" in
# their User-Agent, so they could be covered by a single "fuzz*" pattern.
header drop "user-agent" "*buster*"
header drop "user-agent" "commix*"
header drop "user-agent" "dav*"
header drop "user-agent" "dirbuster*"
header drop "user-agent" "feroxbuster*"
header drop "user-agent" "fuff*"
header drop "user-agent" "fuzz*"
header drop "user-agent" "gobuster*"
header drop "user-agent" "sqlmap*"
header drop "user-agent" "whatweb*"
header drop "user-agent" "wfuzz*"
header drop "user-agent" "wpscan*"
# Drop requests with and empty User-Agent by default.
header drop "user-agent" ""
# ncrack
# nuclei
# lbd
# legion
# SEARCH ENGINE AND CRAWLERS
header drop "user-agent" "baidu*"
header drop "user-agent" "bing*"
header drop "user-agent" "seekport*"
header drop "user-agent" "slurp*"
header drop "user-agent" "sogou*"
header drop "user-agent" "yahoo*"
# SEO TOOLS AND SCRAPERS
header drop "user-agent" "ahrefs*"
header drop "user-agent" "deepcrawl*"
header drop "user-agent" "majestic*"
header drop "user-agent" "screamingfrog*"
header drop "user-agent" "sitebulb*"
!!! This list is very aggressive. It also blocks search engines, link previews, curl, and uptime
monitors. Read it before you use it and take only what fits your site.
relayd: add patters(7) support and improve glob(7) documentation
#
relayd
filter rules for
cookie, header, path, query, and url
now accept an optional
pattern
keyword
before the key or value. With it, the string is read as a
patterns(7)
expression, the same syntax
httpd
uses for location match. Without it,
relayd
uses
glob(7)
as before, so existing configs keep
working.
patterns(7)
gives you character classes like
%d
, anchors like
^
and
$
, and repetition. This lets you
write rules that glob cannot express, like matching only numeric IDs. The
glob(7)
rules in
relayd
were never really documented. The man page now explains what they can and cannot do.
I am also working on a follow-up change. Next to
pattern matching
, you will be able to write
glob
and
glob ignorecase
.
glob
stays the default.
glob ignorecase
makes matching case-insensitive on any field,
including cookie values, paths, and query strings, which are otherwise compared exactly as sent. The
change is not committed yet, so the syntax may still change.
I’m not entirely sure yet, but I could imagine extending this matching syntax to include additional
elements and maybe httpd.
As always, the source of truth is OpenBSD -current.A special thanks to all who
support my
work
.
Scientists who put faith in technology less likely to take climate action
Guardian
www.theguardian.com
2026-10-07 02:00:36
‘Techno-optimism’ may lead to reliance on unproven solutions at expense of societal changes needed Scientists who believe technology will largely solve the problems caused by the climate crisis are less likely to engage in civic action on the issue, according to a study. The report from the London S...
Scientists who believe technology will largely solve the problems caused by the
climate crisis
are less likely to engage in civic action on the issue, according to a study.
The report from the London School of Economics found they are 23% less likelyto sign petitions or take part in protests than peers who doubt that technology is the answer. They are also 18% less likely to make high-impact lifestyle changes, from reducing flights to changing to a
vegan
or vegetarian diet.
Fabian Dablander, the lead author, said the belief that technology will largely solve climate change and its consequences, labelled “techno-optimism”, can lead to an over reliance on unproven solutions at the expense of the far-reaching socioeconomic transformations needed to address the climate and nature crisis.
“Technological solutions are often politically attractive because they promise to address climate change without requiring changes to how people live and consume,” said Dablander.
“Strikingly, we see the same tendency at the individual level, with greater faith in technological solutions going hand in hand with less engagement in climate action. Yet technology cannot substitute for the broader societal changes and large-scale shifts in consumption that are also needed.”
The
study
surveyed more than 9,000 scientists across 115 countries. Overall 27% said they agreed with the statement “technology will largely solve the problems caused by climate change” and 44% disagreed (29% neither agreed nor disagreed).
The “techno-optimist” scientists reported engaging in less civic actions and protests, signed fewer petitions and engaged in less advocacy. They were also less likely to reduce the number of flights they took, cut their car usage or follow a
mostly vegan or vegetarian diet
.
The report, published in Environmental Research Letters, also found that techno-optimism was most prevalent among scientists with right leaning political views and those working in applied and natural sciences.
Prof Kevin Anderson, one of the world’s leading climate scientists, of the Tyndall Centre for Climate Change Research at the University of Manchester, who was not involved in the study, said in his experience “a significant proportion of climate academics who enjoy very high incomes from their work, particularly when those incomes are accompanied by prestige, status and glittering honours, have enthusiastically swallowed the blue pill”.
“It offers a convenient escape from any uncomfortable sense of personal responsibility, allowing them to maintain the comforting fiction that they are deeply concerned about the world their children will inherit while continuing to live lives that contribute disproportionately to the problem.”
He added the “accompanying faith in future technological salvation provides another crutch: a way of postponing difficult choices today by placing hope in solutions that may arrive tomorrow”.
“Yet strikingly often, the evidential basis for these promises falls far short of the standards of scrutiny they would demand of research conducted by others. The result is a remarkable inversion of scientific practice: the more uncomfortable the implications of the evidence, the greater the enthusiasm for believing in a future that the evidence does not support.”
Shaders, WebGPU Components for React, Vue, Svelte, Solid, JavaScript and Framer
Real WebGPU, declarative API:
200+ effects you drop in as components. Gradients, noise, glass, metal, light, distortions, transitions, blurs, cursor effects. Nest them, blend them, mask them.
A design editor that writes your code:
design on an infinite canvas at
shaders.com
, then export the exact component tree for your framework. Free.
Every framework, one package:
first-class React, Vue, Svelte, Solid and JavaScript entries, with the same props everywhere.
Production-ready:
TypeScript, extensively optimized, typed props with reactive updates, SSR safe. Used on thousands of websites by 16,000+ design engineers.
🧩 Components
Every effect is a component.
<Shader>
renders the canvas; its children are layers, evaluated top to bottom and blended on the GPU.
Browse all
200+ components
, each with a live preview and every prop documented.
🎨 Design visually, export code
Most people don't write their effects by hand. They design them.
The
design editor
at shaders.com is an infinite canvas where you stack components, tune every prop with real controls, drive props from the cursor or a timeline, and see the result live. When it looks right,
Export Code
gives you the component tree in your framework, ready to paste.
Design and export are free with an account. Your work is saved as projects you can come back to.
The CLI connects a codebase to your Shaders account, so the effects you design land in your project as real component files and stay in sync.
npx shaders connect # link this codebase to a Shaders project
npx shaders install # pick shaders from that project; writes component files
npx shaders update # pull in what you changed in the editor
It detects your framework, writes to your components folder, and records what it installed in a lock file. Nothing to install globally.
MCP server:
npx shaders@latest install-mcp
configures Claude Code, Cursor, Codex, Windsurf, Copilot and others. Your agent can find, install and edit the shaders you design. Works on any account, free included.
MCP guide
.
If you need a brand new component, you can write one. A component is a plain object passed to
defineShader
: a name, its props, and what to draw at each pixel. Mount it with
<CustomShader>
like any other component.
The guide covers both ways to write the pixel part: composing it from the std primitives (experimental) or writing WGSL directly (stable).
Custom Components
·
Primitives reference
👩🏻⚖️ License
The engine, every component and the framework bindings in this repository and the
shaders
npm package are
MIT licensed
.
The design editor, presets, sections and other platform features at shaders.com are separate from this package and have their own
terms
.
💎 Contribute
Issues and pull requests are welcome. View the
contributing guide
before starting, and for anything more than a bug fix, open an issue to discuss it before the PR.
The engine lives in
packages/core
, with one folder per component under
src/shaders/
. The framework packages are generated from it.
pnpm install
pnpm lib:build # build every package (regenerates the registry and framework components)
pnpm test# the engine's test suite
I covered this before in
Art or tool?
If you think of software as artistic output then generated software isn’t real because it doesn’t have the creative ineffability that’s a sign of true art.
There’s no argument against this one except that the people who want software typically aren’t paying for artworks, they’re paying for outcomes. And the outcomes of using a given artefact are the same regardless of how the artefact is made.
That also means there’s no argument
for
this one, unless you accept as axiomatic the ideas that things that can be considered art should only be made by people, and that using
this particular set of tools
is too big a separation between the person and the work when
no other historical set of tools
was too big a separation between the person and the work.
AI as Employment Thief
I covered this one before too, in
On working machines
. If you think of typing in code as the thing you get paid for, and the employers now have access to a machine that can do the typing, then the machine “stole your jobs”.
Except that the job was never to type code, it was to deliver valuable software. People who can still do that can still get paid by organisations that need that to be done. And, as I said in the linked post, demand only increases as application becomes more efficient.
Humans as Meat Proxies
If you think that software creation is an interpersonal endeavour, then taking task descriptions generated by software and turning them into prompts that the software uses to generate more software is a dehumanising experience.
This has been the trajectory of society for centuries. Here’s a quote from the
Fragment on Machines
:
…once adopted into the production process of capital, the means of labour passes through different metamorphoses, whose culmination is the
machine,
or rather, an
automatic system of
machinery
(system of machinery: the
automatic
one is merely its most complete, most adequate form, and alone transforms machinery into a system), set in motion by an automaton, a moving power that moves itself; this automaton consisting of numerous mechanical and intellectual organs, so that the workers themselves are cast merely as its conscious linkages. […]
The accumulation of knowledge and of skill, of the general productive forces of the social brain, is thus absorbed into capital, as opposed to labour, and hence appears as an attribute of capital, and more specifically of
fixed capital,
in so far as it enters into the production process as a means of production proper.
This passage, written by Karl Marx in about 1857-1858, shows that both mechanical and intellectual work was always intended to be replaced by machinery (in predicting the alienation of intellectual work, Marx probably relied on his knowledge of contemporary mathematician Charles Babbage). In this situation it isn’t the technology that takes the job, it’s the mode of production. And given that programming is intellectual work, the means of production is literally our heads, we ought to be able to seize those.
Software Engineering as Rules
Motivated by Bertrand Meyer’s
AI for Smarties
, if you believe that software engineering is the careful application of the mysterious knowledge of software engineers, encapsulated in rules like the SOLID principles, “prefer composition over inheritance”, “
use you’re type’s good
“, the Law of Demeter and so on, then AI can’t possibly write software because it doesn’t apply those rules, it just generates text that happens to compile.
Never mind that much of the history of computing involves people taking an empiricist approach to software: copying it from Stack Overflow and tweaking it until it works; writing it out from a listing in Numerical Recipes or Sinclair User and tweaking it until it works; writing macros to automate the copying phase, and so on.
AI Companies as Bad Actors
If you believe that the big AI companies is run by people who don’t have humanity’s best interests at heart (further reading: all the Karl Marx writings I didn’t already quote above), then AI coding is bad because it supports those companies. As long as you don’t allow for the creation of other companies, the use of non-corporate models like academic or community models, and so on.
Sharded, encrypted storage between friends over Yggdrasil
Proof of concept: sharded, encrypted file storage between peers on the
Yggdrasil
overlay, with no VPN
coordinator: node IDs come from public keys, each node only answers group
members, and storage challenges check that peers still hold what they were
given.
The idea: people in a group lend each other disk space. Each file is
encrypted, cut into pieces and spread over everyone's machines, so it
survives any one machine failing. The small stub left behind holds the key,
so sending someone a stub (sealed for them, by email or anything else) is a
way to send them the file, with no cloud service in between.
Status: experimental.
The cryptography has not been reviewed by anyone
independent. Don't make it the only copy of anything you can't lose.
New to this?
Set up a storage box
(step by
step, for Windows users with a second-hand mini PC).
No box? The
gateway
lets people pay a monthly fee and
use the group's storage from any S3 program (rclone, Cyberduck, Duplicati).
Members choose whether their box holds customers' data, and earn credit.
Every version is kept:
history
lets you restore an
older version of a file, or a folder as it was on a given day. A new
version only stores the parts that changed.
Nodes can
message each other
: direct messages, and
topics any node can publish to and follow, delivered even to nodes that
were off at the time. Programs can use it too, through a local API.
No Yggdrasil to install: a node can run it
built in
,
with no daemon, TUN device or root. It talks to nodes on the daemon too.
The
mesh
keeps members' Yggdrasil linked directly to each
other, so losing one tunnel or public peer doesn't cut anyone off.
Websites
can live on the group: publish a folder, and
several machines serve it, so a site stays up when one of them goes down.
Email
to your domain can arrive on the group: each message
is encrypted for you as it comes in and read on your dashboard.
put
splits the file into 4 MiB chunks, encrypts each with AES-256-GCM
(one random key per file, fresh nonce per chunk), and erasure-codes each
chunk into
4 data + 2 parity
shards.
Shards are content-addressed (filename = SHA-256) and spread over the
online peers, at most one shard of a chunk per peer when there are 6+.
A stub
FILE.ystub
holds the manifest,
including the key
. Anyone with
the stub and access to the peers can read the file.
get
fetches shards in parallel, rejects any whose hash doesn't match,
and rebuilds each chunk from any 4 of its 6 shards.
At upload time, while it still has the shards, the uploader precomputes
20 single-use challenges per shard into
FILE.ystub.challenges.json
(keep this private).
verify
sends one per shard: "hash this byte range
with this nonce", and checks the answer.
Identity and access
A Yggdrasil address (200::/7) is derived from the node's public key, so the
source address of a connection over the overlay identifies the caller. The
address is the node ID. Each server only answers callers whose address is in
peers.json
(default deny), binds only to its overlay address (or, with the
built-in Yggdrasil
, has nothing on the machine's network
at all), and only lets the original writer delete a shard.
Usage
Download yggstore for your system from the
releases
page (Linux on
PCs, Raspberry Pis and arm64 boards, Windows, macOS), and check it against
SHA256SUMS
. Or build it yourself with Go 1.24 or later:
go build -o bin/yggstore ./cmd/yggstore
bin/yggstore id # prints your node ID and peers.json entry# collect one entry from every machine into peers.json, copy it to each one
bin/yggstore serve -peers peers.json # on every machine (port 7400 by default)
bin/yggstore status -peers peers.json
bin/yggstore put -peers peers.json photo.jpg # -> photo.jpg.ystub
bin/yggstore verify photo.jpg.ystub
bin/yggstore get -o copy.jpg photo.jpg.ystub
bin/yggstore rm photo.jpg.ystub # delete its shards
Each chunk is 6 shards, and no node gets more than 2 of them, so losing any
one node never loses data. A node marked
"slow": true
gets one shard per
chunk while the others have room for the rest, so it doesn't hold uploads
back. Only the uploading machine's peers.json decides this.
Nodes that run on the same machine (Docker containers, say) should share a
"host"
tag, e.g.
{"name": "t1", "addr": "...", "host": "desktop"}
. The
2-shard limit then applies per machine, and uploads need 3 machines online,
not just 3 nodes.
Joining a group
On an admin dashboard, open
Invite someone
, type their name and press
Make invite
. You get a message to send them, with a one-time invite
(
yggjoin1:…
, valid 7 days). They need Yggdrasil running and connected
(the message lists public peers), and the
yggstore
program; then:
yggstore join -service yggjoin1:… # -name NODE -me NAME -quota GB to change the defaults
join
asks the admin node to add this machine (over Yggdrasil, presenting the
invite), saves the member list, adds the inviter as a contact, and with
-service
starts the node and dashboard as user services. On the admin side
the newcomer is added to peers.json (owner: their name) and to your contacts,
and the list sync tells every node. An invite works once; the admin node
only keeps a hash of it. A node on a machine that is already a member is
tagged with that machine automatically; otherwise tag it with
"host"
by
hand if you know two nodes share a box.
The join endpoint is the only call a non-member can make, and it does
nothing without a valid invite (rate-limited to 6 tries a minute).
Adding nodes
An
"admin": true
node (the desktop) keeps every node's list in step. Each
node reports which list it holds; the dashboard sends its own peers.json to any
node that holds another one, so a node added on the dashboard, or a hand edit
of the desktop's peers.json, reaches every node within a minute, with no
restarts. A node keeps a pushed list in its data folder
(
peers.pushed.json
), and uses whichever of that and its own
-peers
file
changed last. Only admins may push, and a push must keep the sender admin.
To add a machine: open
Add a node
on the dashboard and follow the steps.
In short, the new machine starts with a peers.json holding just the desktop's
entry, runs
yggstore serve
, and you paste what
yggstore id -name NAME
prints. New uploads use it straight away; files already stored stay put.
Nodes on older software show "update yggstore" on the dashboard and are not
sent lists. Install the new binary on them and restart them; a node that has
never been sent a list also needs
"admin": true
on the admin node's entry in
its own peers.json, so that it accepts lists from it.
Sharing
Every person runs their own node and their own dashboard. The dashboard
listens on localhost only and holds your stubs, which carry the keys to your
files, so it is never shared. What people share is the network of nodes.
To send someone an item, add them under
Sharing → Contacts
with their
sharing code (
ys1…
, shown on their dashboard), then press
Share
on the
item. You get a
.ysend
file to send by email or any other way. Only that
person can open it; the item's name, your note and its key are sealed inside,
and opening it proves which sharing code sent it. They drop it into their
outfiles
, and it appears under
Shared with you
, where they can restore
it. They need a node in this network, since nodes serve shards to members
only.
A received item is still yours to lose: if the sender deletes it, it is
gone for everyone they shared it with.
Remove
on a received item only
forgets it on your side.
Received stubs may only point at Yggdrasil addresses or nodes in your own
peers.json, so a crafted file cannot make your restore contact other
machines.
Your private sharing key is
~/.yggstore/sharing.key
(contacts in
~/.yggstore/contacts.json
). Losing it means you can't open what people
send you any more, and they need your new code.
The dashboard refuses actions without its own
X-Yggstore
header, and
requests that name another host, so other web pages in your browser
cannot press its buttons.
Give and use
Each node reports how many bytes each uploader has stored on it. The dashboard
totals this per person (set with
"owner"
in peers.json; defaults to the
machine): what they use across the network, and what their nodes hold for
others. The aim is to hold about 1.5× what you use. For now it is accounting
only; nothing is enforced.
outfiles/ drop files in, or folders of them → each file is sharded, then
replaced by NAME.ystub beside it, so the folder tree stays
outfiles/received/ items people sent you (drop their .ysend anywhere in outfiles)
outfiles/whole/ drop a file or folder in → stored as one item (a folder becomes
one archive, restored as a whole)
outfiles/restore/ drop any .ystub in → rebuilt into restored/
outfiles/restored/ restored files and folders
outfiles/delete/ drop a .ystub in → the item is removed from every node, then its stub
outfiles/.yggstore/ stubs that have been restored (kept for reference)
An item is picked up once it has looked the same on two polls (5 s apart)
and was last changed more than 3 s ago, so half-copied items are left alone.
Names ending in
.part
,
.crdownload
,
.tmp
and the like are ignored.
In folders, every file is its own item: it can be restored or deleted on its
own, and new files added to the folder later are picked up too. Folders in
whole/
are streamed as a tar archive and stored as one item. On restore,
only regular files and folders are written, and paths that would escape the
destination are rejected. Symlinks are skipped.
A file put back next to its own stub (say, restored, edited and moved back)
is stored as a new version, and the stub then points to it; older versions
are kept (see
history
). An unchanged copy isn't stored
again.
At least 3 machines must be online, so that no machine holds more than 2 of
a chunk's 6 shards. Otherwise the item stays put and is retried every 2 minutes.
After uploading, the watcher downloads the item again and compares SHA-256
hashes.
Only then is the original deleted
(
-keep
keeps it anyway).
Each stub's challenges are kept in a hidden
.NAME.ystub.challenges.json
next to it.
Deleting (the
delete/
folder, the dashboard's Delete button, or
yggstore rm STUB
) removes every shard first. Only when all are gone are the
stubs and challenges removed; if a node is down, the stub stays in
delete/
and is retried every 2 minutes, so no shard is left without a record.
Deleting cannot be undone.
The dashboard's Restore button rebuilds any listed stub into
restored/
.
A restore never overwrites: a second copy becomes
name (2).ext
.
Every 3 seconds the dashboard asks each peer for
/v1/info
(up/down, round
trip, shard count, disk use) and checks, for every
.ystub
under
-stubs
,
which shards their holders still have (
HEAD /shards/{hash}
, nothing is
downloaded). A file is
healthy
when every chunk has 6/6 shards,
degraded
below that while 4 remain, and
lost
below 4. The Verify
button spends one stored challenge per shard. The activity log records nodes
going up or down, shard counts changing and file status changes. It listens on
localhost only by default.
To try it on one machine:
scripts/local-demo.sh
runs six nodes on
different ports of your Yggdrasil address, stores a file, stops two nodes and
restores it.
scripts/local-demo.sh loopback
does the same on
::1
.
Tests:
go test -race ./...
Docker test cluster
scripts/testnet.sh up # five test nodes t1..t5 in Docker, each with its own Yggdrasil
scripts/testnet.sh status
bin/yggstore put -peers testnet/shared/peers.json FILE
docker stop yggtest-t2-1 yggtest-t4-1 # simulate two nodes failing
scripts/testnet.sh down # stop (keys and shards kept); `clean` deletes everything
The containers sit on their own bridge (
br-yggtest
) and find this machine's
Yggdrasil by link-local multicast, as LAN nodes do. The host firewall must let
the bridge in (
sudo ufw allow in on br-yggtest
). The test peer list is the
desktop plus t1..t5, so test files never land on the real nodes. All six share
one machine and disk: this tests behaviour, not real redundancy.
Limits of this proof of concept
No machine holds more than 2 of a chunk's 6 shards, so any one machine
can fail. Surviving two failures at once needs 6 or more machines.
Stubs live only where they were made. If that machine's disk dies, the
shards are still in the group but nobody can decrypt them: back up your
stubs and
sharing.key
.
Membership is the admin's
peers.json
, pushed to every node; there is
no gossip or vouching yet.
No repair: if a peer disappears, nothing re-creates its shards.
Challenges are run by hand, are not signed and produce no receipts.
Plain HTTP; confidentiality in transit comes from Yggdrasil's end-to-end
encryption, and shards are ciphertext anyway.
Chunks are buffered in memory (4 MiB plus shards), not streamed.
The Tailscale transport from the plan is not implemented;
transport.Transport
is the seam for it.
Quoting Jake Boggan
Simon Willison
simonwillison.net
2026-10-07 00:47:55
I was a graph theory junkie long ago and even moved to Budapest for awhile to study among the greats. While I was there I started working on Barnette's Conjecture which came to occupy my thoughts over the next 24 years of my life, on and off as I worked in many different fields. Last summer I even t...
I was a graph theory junkie long ago and even moved to Budapest for awhile to study among the greats. While I was there I started working on Barnette's Conjecture which came to occupy my thoughts over the next 24 years of my life, on and off as I worked in many different fields. Last summer I even thought for a few days that I had actually solved it.
But it's supposedly proven here -
problem 180
. I don't know what to think exactly. I spent thousands of hours on that problem. I really enjoyed it. Hearing that it is solved somehow makes me sad in a far-off way, like hearing an ex-girlfriend died suddenly in a car crash. I don't know, there's probably a lot of people feeling odd emotions tonight.
It's time to be done
after just over 15 years of moderator work. I was elected in the
very first community election in January and February of 2011
, and by now I'm quite possibly in the top five for longest serving community moderators on the original trilogy sites of StackOverflow, ServerFault, and SuperUser. I could check, but I'm not motivated. SF has always been the one I worked with and for. 15 years, and I've learned so much about moderation and online technical communities.
The early years of ServerFault were vastly more fractious than the site has been in the last 10, and taught me so much. One particularly vivid example was seeing the toxic reverberations between users from different countries and how they handle escalation of conflict. A particularly volatile mix were the Russian speakers and the folk from the United Kingdom (this was before Brexit), where the very British snide (a plus-one or two escalation from just asking questions, the next step was ad hominem attacks) was taken as an invitation to fight by the Russian speakers (whose plus-one escalation typically was ad hominem attacks, which means the brits saw this as escalation and moved on to flame war), both of which
deeply and horribly offended
US-based users for whom both were plus-two or plus-three escalations, definitely not a plus one. These flame wars proved the need for moderator coverage during US night time, which has been a factor in every election after mine. It also drove the feature to lock posts to get folk to cool down.
However, the site has been in decline for a long time. The peak for question asking was in 2011-2012, with a change to Google algorithms in 2013 causing a sizable drop in traffic in late 2013 to 2014. The long slow decline started in 2016, and in recent months we're getting 2-3 questions a day. The moderator flag queue mostly is dealing with spam. Even the Suggested Edit queue, once a great place to farm reputation by fixing other's grammar and formatting mistakes, has slowed down enough that two people can keep it from getting clogged.
The biggest moderator thing I was involved in the last few years were StackExchange-wide spam campaigns that also targeted ServerFault. Manual remediation was required until StackExchange could figure out technical fixes, and it was a bad 12 months there. SuperUser got it way worse than we did, but our downvote and site traffic numbers have clear spikes 2x and 3x above our baseline traffic during the spam surges. There were a few months in there where I spent my evenings watching the new-user creations and spam-banning and deleting in real time. The new technical measures have made that sort of coordinated spam campaign much harder to execute, and is working.
We haven't had a major debate about policy on ServerFault in a long, long time. StackExchange as a whole regularly has major debates, but ServerFault in specific hasn't.
In the last few months the Suggested Edit queue has been barely active, which is the signal that it's time. I'm out as of the end of the year. It's now up to the community to decide if it's worth replacing me.
Everyone at Exmouth Marina knew there was an unexploded German bomb in the water. Old hands at the dock had been saying so since the end of the second world war. But the marina had been dredged and re-dredged and nothing had ever come up. It had started to sound like an old wives’ tale. The marina is a busy stretch of water: a commercial port until the late 1990s that had since been converted into a home to private yachts. A bomb, you’d think, would have a hard time staying hidden. But in the middle of the afternoon on Wednesday 14 January 2026, 81 years after the end of the second world war, Steve Hockings-Thompson, the manager of Exmouth Marina, got a call from the crew of a dredging vessel telling him he’d better come out to the boat. “I thought: ‘Oh crikey. We’ve got a bit of an issue here,’” Hockings-Thompson told me recently.
He went out to the dredger, a squat barge that hauls silt and other debris from the marina floor on to a flat deck that resembles a steel sieve. Sitting on the deck was a huge hunk of metal, caked in mud and stones. As Hockings-Thompson got closer and saw its distinctive conical shape, he knew it was a bomb – a big one. The dredger was loaded with 20,000 litres of diesel fuel, so if the bomb detonated, it would be disastrous. He called the police, who quickly came to the scene. “And then,” he said, “all hell broke loose.”
To deal with the explosives, they called in the army. The call was picked up at the control centre at Vauxhall Barracks in Didcot, a small town in Oxfordshire. This is the headquarters of the 11 Explosive Ordnance Disposal and Search Regiment (11EOD), a specialist army division responsible for all kinds of bomb disposal – from improvised devices planted by terrorist groups to second world war-era munitions.
When the Exmouth bomb was called in, Maj Andrew Boyce, who was managing operations for the regiment, was confused. “We were talking about ‘the bomb’, which I’d been told was a 50kg,” he recalled. “And then somebody on the phone tells me it’s a 250kg. I’m like, how have you confused a 50kg bomb, which is the size of a fire extinguisher, with a 250kg bomb, which is the size of a hot water tank?”
The reason for the confusion soon became clear: the Exmouth Marina bomb was not the first to be discovered that day. Another second world war bomb had been unearthed a few hours earlier at a building site in Plymouth, 50 miles away.
In Plymouth that morning, Luke Templeton, director of S.I UXO, a private contractor that goes in ahead of construction companies to look for unexploded bombs, had known what he was looking for. His company’s historians had already spent a fortnight doing a detailed risk assessment of the site, drawing on historical documents showing the Luftwaffe’s flight paths and local records of airstrikes. From this, they’d worked out that it was highly likely there were two bombs buried somewhere on the building site. Using specialist equipment, they had found a 50kg bomb buried more than four metres down. After inspecting the bomb up close, Templeton called the police.
As far as Boyce knew, this was the first time since the immediate postwar period that two unexploded bombs had been discovered on the same day. The two bombs were in the same part of the country, and had been reported to the same police force,
Devon
and Cornwall. “We’re used to having confusion with location or devices. We’re not used to having two devices,” said Brig Darren Fisher, who at the time was a colonel overseeing the operations. Once the confusion had been straightened out – it was a 50kg bomb at the building site in Plymouth, and a 250kg bomb floating on the dredger in Exmouth – the response began.
When most people think of bomb disposal, they imagine a lone figure in a bomb suit carefully approaching an improvised device as bullets fly overhead, holding steady as they unpick wires, knowing that one wrong move will kill them. This image was forged in Belfast during the Troubles and Helmand province during the invasion of Afghanistan, and it is not unwarranted; bomb disposal is commonly described as the most dangerous job in the military. But, in Britain in 2026, the daily reality of the job is quite different.
Although the same team deals with terrorist incidents, such as the 2017 Manchester attack on an Ariana Grande concert, it is unexploded second world war munitions that make up the core work of the UK’s bomb-disposal teams. The Luftwaffe dropped more than 60,000 tonnes of bombs on Britain over the course of the war, and the Ministry of Defence estimates that roughly one in 10 of these failed to explode on impact. The reasons range from the grim – poor manufacturing quality due to a reliance on forced labour in concentration camps – to the prosaic: a bomb might have landed at an odd angle, or on soft ground that absorbed the impact before the fuze could fire. (“Fuse” is the general spelling for electrical safety devices and burning cords, while “fuze” refers to engineered explosive detonators.) Britain also dropped munitions on its own land during training exercises and deliberately buried surplus after the war in an effort at disposal. Eight decades on, the physical remnants of the war still account for about 2,000 of the 2,500 incidents that military bomb-disposal teams respond to in an average year.
The calls often come in from metal detectorists and magnet fishers, hobbyists who throw strong magnets into waterways to see what comes up. They surge on bank holidays, when people are out walking and notice things underfoot, or catching up on tasks such as clearing out a grandparent’s home. Not long ago, the bomb-disposal unit was called to deal with a 2in mortar that had been used for 20 years as a paperweight. Increasingly, calls come from construction sites, as buildings get taller and companies dig deeper, which has led to a growing industry of private contractors actively searching for unexploded bombs. “We’re expanding cities, building new homes, and an accident could happen,” said Templeton. “If a 250kg bomb went off in a city, it could knock a whole street down.”
The passage of time has not rendered these explosives less dangerous. In June 2010, three bomb-disposal experts in Göttingen, Germany, were
killed by a 500kg American bomb
found beneath the site of a new sports arena. In 2020, a 250kg bomb
exploded beneath a crab boat off the Norfolk coast
, seriously injuring five of its seven crew. In 2021, another bomb exploded beneath a drilling rig near Munich’s main railway station, injuring four. Because it was a global war, this is a global problem.
This year, a second world war bomb beneath a stilt house in Indonesia detonated, killing five people and injuring 20.
These decades-old devices are difficult to read or predict. “You don’t really know why [a particular bomb] didn’t function, which is part of the problem,” said Lt Col Andy Hambley, the commanding officer of 11EOD. Years of corrosion, shifting ground and rust make the fuzes less predictable, the machinery harder to read.
That was exactly the issue facing two teams in two different corners of Devon on the afternoon of 14 January.
At about 1pm, PO Scott Dooley, a diver and bomb-disposal operator with the Royal Navy’s Diving and Threat Exploitation Group, arrived at the building site in Plymouth. Dooley, a tall man with a neatly trimmed beard and a soft Devon accent, grew up in the military; his father was in the Royal Engineers, doing bomb disposal. In recent years, he had acquired the nickname “the Doom Watch”, because every time he was on duty, the phone didn’t stop ringing (his record was 11 call-outs in a single week). In Plymouth, Dooley scanned the area for risk: there was a school nearby, as well as a glass-fronted hotel that could shatter if the worst happened.
Police
were already at work establishing a cordon and evacuating the area around the bomb.
When an operator approaches a bomb for the first time, they do not know what they are walking towards. The most important thing is to assess the fuze, a metal cylinder screwed into the side of the bomb. It is the first link in a chain. The fuze contains explosives, which fire a tiny, highly sensitive detonator, which sets off a booster charge, which sets off the main filling of the bomb, a comparatively stable “secondary” explosive. Most German fuzes were electrical, charged from the aircraft as the bomb was released. After 80 years, there is no charge left in them. But some German bombs were fitted with a mechanical delayed-action fuze, constructed from gears and springs, like a watch. Once triggered, this fuze ticks down before detonating between two and 72 hours later. The Luftwaffe deliberately fitted these fuzes to maximise disruption; hospitals, roads and factories were cleared for days at a time because the bombs could detonate long after impact.
Everything, therefore, turns on the fuze. If the fuze is judged to be sensitive, that usually means it is too dangerous to move the device. The safest option is to surround it with sandbags and detonate it where it is. If the fuze is stable, then it can be moved somewhere safer, such as out to sea.
Dooley had two things working in his favour: a 50kg bomb was less likely to have a ticker, and it was still half-buried, which suggested it hadn’t been badly jolted. On the other hand, the fuze was submerged in mud and he couldn’t see it. Bomb-disposal operators try to avoid moving a device until they’ve assessed the fuze, lest they trigger a ticker, but Dooley had no choice. Down in the pit, he and a colleague built a stable platform on a ledge of mud, surrounding it with sandbags, and set to work painstakingly unearthing it. Once it was out, he placed a listening device on it called a microphone stethoscope, or mic steth. He listened for ticking.
Silence. A moment of relief, quickly followed by further investigation. Silence only means that at that moment, no mechanical fuze has been triggered. Was there one in the bomb at all? Dooley needed to physically examine it. The bomb had been lodged more than four metres underground for decades and was caked in thick clay. Still listening for ticking on the mic steth in case his movements triggered anything, he poured water on the fuze to wash away the mud. “The fuze was absolutely more than gone,” he said. It was so degraded that none of the markings were visible. No markings, no information. They would have to X-ray it.
While this was unfolding in Plymouth, Capt Liam Kidman was anticipating a quiet couple of days at home. Kidman had just come off a gruelling week that ended in rural west Wales, cutting open a safe packed with explosives at the bottom of a dead man’s garden. (The man had apparently used the explosives for felling trees, and nobody knew where he kept the key.) Then, shortly after he had completed that job, the control centre rang to say two bombs had been discovered in Devon. The local team in Plymouth was already occupied with one of them. Could he go to Exmouth? Kidman was back on the road within the hour.
An energetic man from south Wales, Kidman ended up in bomb disposal almost by accident – when he joined the army more than 20 years ago, there was a shortage of ammunition technicians. As he climbed the ranks, dealing with improvised explosive devices (IEDs) in Afghanistan and second world war jobs in the UK, he came to love the problem-solving aspect of the work.
It was past 7pm when Kidman reached Exmouth Marina from his base in Gloucestershire. The closer he got to the marina, the emptier the roads got. The bomb was sitting out in the middle of the harbour, bobbing on the open deck of the dredger. It was so big that if detonated it could flatten the marina’s waterfront. The bomb was only accessible by sea, and the nearest naval team was tied up with the other bomb in Plymouth that night. Usually, the army handles incidents on land and the navy anything at sea, but the unexpected nature of two bombs being found on the same day had turned everything on its head.
The bomb had been handled roughly – hauled off the seabed and dumped unceremoniously on to a metal deck. “If there’s a mechanical fuze on there, that’s enough for it to be ticking down,” said Kidman. They had to get ears on it. Harry Griffin, the marina compliance manager, stepped up. Griffin knew there was a chance the bomb could go off, but reassured himself that it had been in the water for 80 years without incident. He ferried a sergeant from Kidman’s team over to the bomb to get a mic steth on to the device. “That’s the bit that gets the operator’s heart going,” said Kidman.
It was not ticking. But nobody relaxed. Just as in Plymouth, they did not know what kind of fuze they were dealing with.
On a bright day in June, I visited RAF Northolt, a base on the outskirts of London, to observe a training exercise on airdropped weapons. Learning how to correctly assess an old, rusting bomb takes time and practice. Bomb disposal is a dangerous job but, Hambley told me: “Knowledge dispels fear.”
Bomb disposal as a specialism was created in the second world war. Britain wasn’t ready for the blitz, and the early assumption was that police or air raid wardens could simply deal with any bomb that failed to go off. But as delayed-action mechanical fuzes – those ticking timebombs – became more common, it was clear that targeted intervention was needed. As the historian James Owen details in his 2010 book Danger UXB, in May 1940, the War Office handed the problem to the Royal Engineers, who formed 25 bomb-disposal sections armed with little more than picks, shovels and stethoscopes. By September 1940, they were dealing with more than 3,000 bombs a month. By some estimates, the average life expectancy of a bomb-disposal operator in this period was 10 weeks.
Munitions on display in 11EOD’s museum at their base at MOD Ashchurch.
Photograph: Sam Frost/The Guardian
Those early teams mapped every bomb they dealt with, and the maps are still in use. Matt Fuller, a researcher at S.I UXO, the contractor that found the bomb in Plymouth, works through wartime bomb censuses and Home Office daily intelligence reports. Studying the bomb maps, he found that the Plymouth bomb most likely fell on the night of 21–22 March 1941. (The Exmouth bomb he puts at 24-25 February 1941.) Conducting risk assessments, Fuller is struck by the sheer number of places that were targeted, some of them small towns with tiny populations. “The last person to die from an unexploded bomb dropped in world war two probably hasn’t been born yet,” he told me. Fuller also marvels at the detailed records local wardens kept. They saw the danger that unexploded bombs would pose for future generations. In keeping those records, they reached out a helping hand from past to present.
In contrast to their second world war counterparts, modern bomb-disposal teams come armed with robot-mounted cameras, mic steths that can be operated from a distance, and X-ray machines that can penetrate thick layers of steel to examine the explosive content inside. “Risk to life comes above everything – for the public, but also for us – and risk to property comes second,” Hambley told me. But when a bomb is 80 years old, even this equipment can be unreliable.
At RAF Northolt, a 500kg German bomb was perched on a grassy verge. The bomb, which was inert, had been placed there, at an awkward angle, for bomb-disposal technicians to assess in pairs: what type of bomb is it? Which fuze? What action can be taken? I approached the bomb with Maj Matt Walker, the squadron duty officer. It was roughly the length of a two-seat sofa: a long, thin hunk of dull metal, its surface rough and textured with rust. Walker told me that a device this size would have a blast radius of “a few hundred to a thousand metres”.
Without clear markings, trainees would have to use their knowledge to work out whether the fuze had a mechanical timer, or an additional anti-handling fuze: a booby-trap built into the bomb that detonates it if anyone tries to move it or remove the main fuze. The trainees had to work out what to do next. “It’s extreme problem solving,” said Walker. “Someone described it to me the other day as – you start by juggling five books, then someone throws a glass ball and you have to keep that juggling as well. They then throw a ball of fire.” If you drop one of the balls, you and those nearby will probably die.
We left the 500kg bomb where it was and walked into the “museum”, a grand name for a few long, dusty shelves inside a garage that acts as a visual reference library. “Anything and everything will turn up here,” said Walker. He picked up a tapered, hollow metal tube. “This one is actually a kid’s baseball bat. But it got called in.” They’d hung on to this and other red herrings so that trainees could see, in practice, how easy it is to mistake something for a bomb. The other miscellanea included an internal part from the Citroën 2CV car that looks so much like a bomb that people regularly call 999 after finding it in a scrapyard. Until I looked at the items in the museum, I hadn’t fully grasped how little there is to go on: the munitions are not just rusted but distorted, sometimes missing pieces, completely blank. Translating these rusted lumps of metal into something legible requires an expert eye: a cylindrical shape with a curved end means it could be a German SC50, particular fuze pocket positioning can identify a British SAP250.
There is a generational difference between bomb-disposal technicians in the UK. People who qualified before 2012 cut their teeth in Afghanistan, where the Taliban laid an extraordinary number of IEDs. In 2011 alone, US-led forces recorded 16,800 IED incidents, roughly 46 a day. Those who qualified after 2012 tend to have mainly worked on second world war munitions. Hambley’s own career has mapped this change. His first job as a bomb-disposal officer was in Helmand in 2012, walking toward a buried device while live rounds tore past his ear. These days, the problems are different: orchestrating large-scale evacuations of civilians, trying to avoid property damage. “You’ve got a lot more latitude than when you’re deployed in Afghanistan,” Hambley told me. “In the UK, we own the ground, no one is going to shoot at us, we’ve got time, so we can afford to be as safe as we possibly can be.” Caution, though, is slow – and the moment a bomb is discovered, the clock is ticking to get people back into their homes.
Sitting at her desk overlooking Exmouth Marina, Fiona Bolt, a chartered financial planner in her 60s, could see the bomb on the dredger. At about 3pm on Wednesday 14 January, police knocked on the office door and told Bolt and her four colleagues to pack up and leave. “It was really quite scary,” she told me. By 8pm, the safety cordon had expanded from 100 to 400 metres and police started knocking on doors on streets near the marina. Some people went to stay with family members nearby, others in hotels. One of the Bolts’ neighbours ended up sleeping in their car.
WO Sarah Rowland of 11EOD.
Photograph: Sam Frost/The Guardian
As the night set in, Kidman’s team was out on the water, trying to see inside the bomb. The first X-ray images of the fuze were unusable: the steel casing, thickened and blistered by rust, defeated the beam. They spent seven hours cycling through different X-ray equipment, trying to get a clear image. As Kidman worked through the night, he could see lights going on and off at the far end of Exmouth harbour, in flats that would ideally be empty. Police can’t force anyone to leave. When someone refuses, police hand them a letter saying they’re staying at their own risk, and the regiment has to work around them. In
Great Yarmouth
three years ago, Hambley worked to render a 250kg bomb safe with a resident “teeth-chatteringly close” – less than 50 metres away.
When Bolt got up the next morning, she noticed her street, close to the marina, was eerily empty. When Bolt left the house she saw how far the exclusion zone had expanded along the seafront: it now encompassed 400 metres, and 2,000 properties. She didn’t have an office to go to. “I felt really lost, like a rabbit in headlights,” she said. “It was all just so out of my control.” The council had hastily assembled makeshift shelters out of a leisure centre and retirement home. When Bolt stopped by the leisure centre that day, she saw that camp beds and bedding had been brought in during the night, and Tesco had donated food. “If we’d been in wartime,” she said, “I imagine that’s what the community spirit was like.”
By about 5pm, it was clear Bolt would not be going home that night. Her house was deep within the cordon. She stuffed essentials into a rucksack and walked across town to her son’s house. By 9pm the cordon had crept up to 600 metres, encompassing a further 500 homes, cleared block by block by police officers. This was a big operation, but in Germany, where far more unexploded ordnance lies buried, even larger scale evacuations are common. In September 2017, 60,000 people
were cleared from central Frankfurt
, including two hospitals, so that a 1.8-tonne British “blockbuster” bomb could be defused. In 2025, the city of Cologne alone carried out 19 evacuations for wartime bombs, affecting almost 70,000 residents.
At their respective sites, Dooley and Kidman were both acutely aware of the disruption around them, which adds an intense time pressure to their work. In Plymouth, all the buildings within a 100-metre radius of the building site had been evacuated on Wednesday, including the school. Dooley’s first X-rays, taken with the standard equipment, had come out completely black. At around 4pm on Thursday, high-sensitivity X-ray plates arrived from elsewhere in the country. Finally, a readable answer. The bomb had no mechanical fuze, and no anti-handling device. Dooley breathed a sigh of relief. “People could just calm down a little bit,” he said. “It wasn’t going to tick if you knocked it.”
Soon afterwards, in the late afternoon, the high-sensitivity X-ray machine arrived in Exmouth. It went out on to the dredger. The team started scanning, and kept going for hours. Still nothing. Kidman reported back to the operations team: they couldn’t identify the fuze.
The forward base in Exmouth was the lifeboat station, and over the course of Thursday it filled up with marina staff, police, coastguard and emergency planning officers from the council. Briefings and strategy meetings took place every few hours. The position of the bomb severely limited the normal mitigation procedures. The regiment’s standard protection for a 250kg bomb is a sand-filled fortification weighing about 350 tonnes. This could not be built on a barge. Abrasive cutting, an advanced technique for cutting out the fuze, was not deemed appropriate. The next option was towing the dredger out to sea and blowing the whole thing up. Griffin, the marina’s compliance manager, phoned its insurers to ask whether this would be covered. Most marine policies carry an act-of-war exclusion, which raises a difficult question: after 80 years, is a Luftwaffe bomb still an act of war? “We went round in circles, really,” said Griffin.
This is a sensitive subject for 11EOD. In February 2021, a 1,000kg German bomb – one of the heaviest in the Luftwaffe’s arsenal – was discovered on a construction site next to Exeter University’s halls of residence. The team could not establish if the bomb was booby-trapped so it decided to do a controlled detonation where it was, with 400 tonnes of sand around it to absorb the blast. It was not enough. The explosion, which was heard for miles, left a crater the size of a double-decker bus in the ground and threw debris at least 250 metres, shattering windows of nearby buildings. The university claimed on its all-risks policy; the insurer, Allianz, declined, citing a clause excluding damage “occasioned by war”. The high court found in favour of the insurer. The military was not financially liable for this, but the experience of having every step of its process picked over in court made the unit even more cautious. “We are very aware that everything could end up in court and everything could be questioned,” said Hambley.
Exeter: WW2 bomb detonated after homes evacuated – video
Caution pulled one way and time pressure – the need to get people back into their homes in Exmouth and kids back to school in Plymouth – pulled another. The idea of exploding the entire dredger was ruled out due to the risk of environmental pollution from the 20,000 litres of fuel on board and fragmentation of the vessel itself. That left only one option: get the bomb into the water and tow it away to detonate safely at sea. But this left a question: how to remove it?
By 11pm on Thursday 15 January, there was a plan for the smaller bomb in Plymouth: lift it out, place it in a 4x4 packed with sandbags, drive it 1km to the water’s edge, take it out to sea and blow it up there. By 12.30am, police and fire crews were going door-to-door to evacuate everyone who lived within 150 metres of the route. At about 2.30am, Dooley and another bomb operator lifted the 50kg bomb out of the pit with their hands, climbing up on a makeshift staircase they had built out of sandbags. The bomb, slick with mud, was cold and heavy. It was dark, and the steps, wedged deep in a muddy pit, were slippery. They placed the bomb carefully into the back of the waiting vehicle. The convoy slowly moved off, at walking pace. As the vehicles passed by, the police released the cordon, letting people back into their houses while the bomb was still crawling towards the sea.
Once it was at the water’s edge, the bomb was carried down on to a boat and lowered over the side, where a buoyancy device was waiting so that it could be dragged along underwater at a safe distance from the boat. Dooley was aboard the navy boat, which lowered the bomb into the water as it got further away from the harbour. He’d had two hours sleep in the preceding 36 hours.
The King’s Harbour Master, which controls shipping in Plymouth Sound, had given permission for the bomb to be detonated in a specific patch of sea, and they sailed out four nautical miles towards this spot. At the drop point, Dooley fixed an explosive charge to the bomb and had it lowered down to the seabed. A navy diver followed the rope down in the dark. It was pitch black and the water was thick with silt, rendering a torch useless. Working by touch, the diver undid the bolts, let the air out of the float and brought it to the surface.
When the diver came back up, Dooley took the boat back a hundred metres. From there, the only thing showing the location of the bomb was a small lit float attached with a line, bobbing on the surface of the water. The swell of the waves kept hiding it. Police drones confirmed nobody on shore was in danger. The navy confirmed there was nothing on radar at sea. The charge was initiated. Dooley counted down.
At 5.13am on Friday 16 January, the sea briefly lit in the dark. A plume of white water rose above the surface, then nothing. Eight decades after it was dropped, the bomb had been detonated.
As this unfolded in Plymouth on Thursday night and into Friday morning, Exmouth Marina was empty except for the military bomb-disposal team and the two marina staffers. “It was quite eerie,” said Hockings-Thompson, the marina manager. “It was like going back to Covid times.” They were all focused on how to get the bomb off the dredger. Someone suggested bringing the dredger in closer, but moving the boat meant moving the bomb, and nobody wanted to risk it.
Griffin went back into the office and called a local crane hire company, who agreed to bring a 55-tonne crane. The next problem was who would drive it. Senior officers including Fisher scrambled to find a military crane operator qualified to drive a 55-tonne crane, available at short notice. They could not. The civilian crane driver who had brought it over offered to do it himself. “I can’t own the risk of him driving that crane,” said Fisher, who held the military liability. That passed the question to the police, who had to decide whether to authorise a civilian to handle an unexploded bomb – and whose financial responsibility it would be if the crane was destroyed. “If the worst happens,” Fisher said, “we’ve got much bigger problems than losing the crane.” Eventually, they let him go ahead. (The crane driver declined to speak to me for this article.)
After midnight, in the early hours of Friday 16 January, bomb-disposal operator Lt Matt Bowden attached straps around the bomb casing and added a buoyancy device. He left the dredger and the crane lifted the bomb. As the 250kg lifted from the deck, the barge rose in the water. They had discussed this in advance: whether the dredger, suddenly lighter, would bob up hard enough to bump into the bomb and trigger something. The dredger heaved. The bomb swung out to the side and missed the steel vessel entirely. It was lowered into the water and attached to a navy boat.
Everything now depended on the tides. There was a fixed window in which it could be towed out and detonated. At about 4.30am, the boat picked its way through moored yachts with the bomb submerged and trailing 100 metres behind. By 6.30am, the bomb was so far out and deep underwater that residents were allowed back to their homes.
Hockings-Thompson and Griffin, who had been mostly awake for the preceding 48 hours, went aboard the dredger to bring it back into the marina. Then they walked up to a coffee shop on the seafront. They were in the queue at 8.13am when the bomb went off, a mile and a half out. “The whole seafront was full of people expecting to see a massive bang and a splash,” said Griffin. “All we got was a bit of a plop.” The sea had done its job and neutralised the blast.
By 11am, businesses around the marina were open again. Kidman was exhilarated. “No one died, there was no injury, no damage to infrastructure,” he said. After working for 36 hours straight, everyone went home to sleep. For Dooley, after his own successful operation in Plymouth, it was only a few hours later, when he was sitting with his family, that the magnitude of what had happened hit him. “You’re just doing your job,” he said. “Then you sit there and think, what have I done here? A 50kg could kill you as much as a 1,000kg.”
In late summer, about six months after the double bomb discovery, I sat in the bar at Vauxhall Barracks with Dooley, Kidman, Fisher, Boyce and others who had worked on the two operations. It was a wood-panelled room with a blue patterned carpet, the wall adorned with regiment photos and a framed Private Eye cover of the Queen meeting Martin McGuinness in a bomb-disposal suit. As I pressed for details, they found their recollections blurring. Was a particular memory from 14 January? Or was it the 250kg Plymouth bomb that was found in April outside an Aldi? The one found in June at a business park in Coventry? “So much has happened since then,” said Boyce. In fact, WO Sarah Rowland, who had worked with Dooley on the bomb in Plymouth, was unable to join the meeting as that morning she’d been called out to Staffordshire, where a 500kg bomb had been found in a quarry. (She texted Dooley to say she was collecting them all, “like Pokémon cards”.)
They are getting busier. Ten years ago, one or two large airdropped weapons were found each year. This year there have already been six. The latest was on 21 September, when a 115kg American bomb was found in the water off Pendennis headland, at the mouth of Falmouth’s harbour. The call went to Dooley, but he was already on another job. Kidman was assigned to it and drove down to Cornwall, where he spent almost all of the next three days awake, problem solving. The job was complicated. Blowing it up where it lay risked rupturing waste tanks 40 metres away and pouring sewage into the Channel. Instead, more than 300 households were evacuated while the navy lifted the bomb and towed it out to deep water. It was too dangerous to send a navy diver down, so the bomb was left on the seabed, far from anyone, while the regiment counted down the hours.
On 27 September, with the clock on the Falmouth bomb still running, another call came in: three white vans near RAF Fairford, a Gloucestershire base used by American bombers. A major counter-terror operation began. Kidman and Rowland drove to the scene. A robot searched the vans. X-rays followed. They helped each other out, catching some sleep in the car when possible. Kidman was still overseeing the bomb in Falmouth. On 28 September, they found that there were no explosives in the white vans. Kidman went home to sleep. At 2.25pm on 29 September, the timer was up on the bomb in Falmouth, so the team went back to detonate it at sea. “And then,” Kidman said, “we’re ready for the next bomb, I guess.”
Show HN: NanoMuse – An open-source AI agent for your phone and computer
nanoMuse is an open-source personal agent for every device you own.
One agent with a name and a look of its own, in the style of Meta's
Muse
: it does things instead of answering questions, keeps working while the app is closed, remembers you, and stops to ask before anything you could not undo.
nano
means the whole set, small enough to run and deploy yourself: the phone app, the desktop app, the web console and the relay that joins them are all in this repository, under GPL-3.0-or-later.
Free, open source, non-profit — let's build it together.
Sign in and you get a free allowance of model use on the community relay — the developer pays for it; when it is gone,
use your own key
. The same relay runs on a server of yours, so nothing has to leave your house. Latest:
0.1.40 Clear
—
release notes
·
try it in the browser
.
Install
Browser
demo.nanomuse.dev
— a nanoMuse on a simulated phone, after a sign-in. A demo; the apps below are the real thing
bash scripts/self-host.sh --local
for your own relay;
docker compose up -d app
for the web app on a server of yours —
self-hosting
All downloads come from the
latest release
; the same files are on
nanomuse.cn/dl
when GitHub is slow where you are. Open the app, sign in with an e-mail or a mainland-China phone number, and it has a model to think with. The phone, the desktop and the web share one account and show the same conversations.
What it does
Does things.
A Linux shell, a browser, MCP servers and skills — and, with
Hands
on, the apps on your phone and the windows on your computer through their screens, for everything that never had an API.
Asks first.
A stop before deleting, sending or paying, remembered for once, this chat or always; passwords and codes are yours to type. A login or a CAPTCHA is handed to you,
Done
resumes.
Reaches your other devices.
Say it on the phone, it runs on your PC;
@Mac …
at the start of a message sends the task there. Approvals come back to the device in your hand.
Keeps going.
Goals checked on a schedule, routines that run while the app is closed, a feed written for you each morning.
Remembers you.
What it is, what it knows about you and when it wakes are Markdown files you can read and edit.
Lives in your chat apps.
It answers in 飞书, 钉钉, 企业微信 and Telegram.
A look of its own.
Describe one, your image model draws it, a video model makes it move. A small dragon by default.
Any model.
The relay's allowance, your own key at one of eighteen providers (Bailian, OpenRouter, OpenAI, Gemini, DeepSeek and more), or a plan you already pay for: ChatGPT, Claude, Kimi. A provider without picture or video models leaves those two off, and the app says so.
How it works
Each device runs its own agent — the phone inside the APK (Alpine Linux under proot, a shell, a browser, MCP), the computer inside nanoMuse Desktop (DeepSeek Harness with the Python runtime for the hands). Signed in, they meet on the relay and can ask each other for things; the text of the conversations travels through it, files and screenshots stay where they were made.
Your phone, your computer, or a server of yours — one account across them
Apps without an API
Out of reach; the VM never touches your devices
An accessibility CLI on Android
The screen as a hand on the phone and the computer: screenshots, APIs first, you take over for logins. Not on iOS — the system does not allow it
Other devices
Clients of one VM
The one it is installed on
Devices ask each other for things over the hub, with approvals where you are
Models
Meta's
Bring your own
The relay's free allowance, or your own
Licence
Closed
GPL-3.0
GPL-3.0-or-later, built on OpenMinis
News
2026-10-06 ·
0.1.40 Clear
— the iPhone's input field is back in view, every chat on a phone belongs to its account, and every refusal from the relay is one plain sentence or a card on every client. The Mac takes a true picture of the screen or says why it cannot, Windows starts again after the update that moved the app, and the relay's console has Controls, Stats and Site pages.
2026-10-06 ·
0.1.39 Keys
— your own key from a catalogue of eighteen providers, sign-in with a ChatGPT plan, and a device shows only the signed-in account's conversations.
2026-10-05 ·
0.1.38 Loom
— side chats stay on the device that made them, a Mac helper app holds the hands' permissions, every app can point at a relay of yours, and the docs became a site.
One VPS, one hour:
docs/self-hosting.md
. Three ways — no server at all with your own key, your own relay with
scripts/self-host.sh
, or a runtime of your own for the web app.
nanoMuse is an independent community project, not affiliated with or endorsed by Meta Platforms, Inc.; Muse is their trademark. The dragon is the project's own.
License
GPL-3.0-or-later
. The phone app is based on OpenMinis 1.13 (GPL-3.0), modified since 2026-09-24 — see
NOTICE
. Earlier versions of the Python line were MIT (tag
pre-openminis
).
OpenWAM: An Open Framework for Composable World-Action Models
Should a robot predict what it will see before deciding how to act, or generate both together? World–action models make both possible. Comparing these choices is difficult when every system uses a different backbone, dataset, and training recipe. OpenWAM gives them a common foundation so we can study how prediction and control work together.
The framework supports composition within a model and between models. We can change the order in which video and actions are generated and how their tokens attend to one another. We can also connect independently trained components: a video predictor proposes a future, an inverse dynamics model turns it into actions, and a forward dynamics model predicts what a supplied action sequence will do.
Figure 1. OpenWAM overview.
A shared video foundation supports configurable video–action programs and independently trained dynamics. In the transfer experiments at bottom right, we adapt the video predictor to each task and reuse the IDM without further training.
View full-size figure ↗
A shared video–action architecture
Before learning actions, we adapt Wan2.2-5B to robot motion and interaction. We pretrain on approximately
3.34 million trajectories and recordings—14.64k hours of source video
—spanning real and synthetic robot manipulation, human-guided manipulation, and human interaction. This stage uses video alone, without action labels or proprioceptive inputs.
What goes into pretraining?
Paper Table 8 · Hours count source sequences before training-window sampling, including synthetic and human-interaction video. OXE covers 49 manipulation datasets; InternData-A1 is synthetic; Ego-Exo4D records human interaction. The UMI family includes UMI, DexUMI, UMI on Legs, and MV-UMI.
Causal attention lets us generate video a chunk at a time: each chunk can use current and past observations and earlier chunks, but not later ones. Its tokens are denoised together. After
14 days on 32 NVIDIA B200 GPUs
, this checkpoint provides the visual foundation for the downstream models.
To add robot control, we pair the 5B video expert with a 2B action expert in a
Mixture-of-Transformers (MoT)
architecture. The action expert starts from width-adapted copies of the pretrained video layers. Each expert keeps its own normalization, projections, and feed-forward layers; attention over their combined tokens lets them exchange information.
Figure 2. Shared MoT architecture.
A 5B video expert and a 2B action expert share attention over video and action tokens. Each also attends to the task instruction.
View full-size figure ↗
Video-action interaction programs
The same architecture can predict video before actions, actions before video, or both at once. Each interaction program specifies the generation order and attention between future tokens. The backbone, tokenization, training objective, and downstream recipe stay fixed, and every program receives the task instruction and observed history.
Figure 3. Attention masks.
A–D: the evaluated policy programs. E–F: local inverse and forward dynamics. Colors mark clean or noisy conditioning; numbers mark generation order. L denotes language; O
−
, O
0
, and O
+
denote past, current, and future observations; A
+
denotes future actions.
View full-size figure ↗
All programs use latent flow matching, with four latent video frames aligned to each 16-step action chunk and proprioception supplied per chunk. In VTA and ATV, the second stage learns from recorded trajectories during training and uses the first stage’s predictions at inference.
Beyond policies: local-context IDM and FDM
A task-conditioned predictor proposes what should happen next. Dynamics models connect that proposal to the robot’s motion: inverse dynamics (IDM) turns a visual future into actions, while forward dynamics (FDM) predicts the outcome of an action sequence. OpenWAM supports both as standalone models.
We give them a
local-context interface
: the current observation, proprioception, and a supplied future trajectory.
Local-context IDM:
current observation + proprioception + supplied future video → action trajectory.
Neither model receives task language or pre-start history. This separates choosing a task from modeling a transition: we can adapt the video predictor and ask whether the same IDM still produces the right actions. The experiments below test how far this local information can take us.
The video predictor passes VAE video latents to the IDM, not transformer hidden states or caches. With compatible video and action representations, the components can be trained separately and connected at inference. An FDM can likewise predict the outcome of actions from a separate policy.
Learning from alternative outcomes
A demonstration shows what the demonstrator chose to do, but says little about what other actions would have caused. Changing a model’s inputs does not fill that gap in its training data. We build
LIBERO-Long-CF: 32,000 counterfactual segments across ten tasks
by restoring simulator states and trying alternative action sequences, including unsuccessful ones. The models learn from the resulting observations and actions, without access to simulator state.
Paper Table 9 · Equivalent durations at 20 Hz. Each counterfactual segment contains 128 controls and 129 synchronized two-view observations. The dataset contains 29.7× as many control records as the demonstrations, using the original tasks, assets, and physics.
What changes in the counterfactual rollouts?
We vary motion magnitude, direction, timing, individual action axes, and gripper behavior. 75% of segments start along a demonstration; the other 25% start after an additional action perturbation.
Intervention recipes (%) · Paper Table 10. Values are rounded; different recipes can produce overlapping physical effects.
In the perturbed starts we analyzed, the end effector is on average
3.34 cm
from the nearest point on the demonstrated path (median 1.90 cm). Objects also move beyond their demonstrated configurations in
61.6%
of these starts, measured at thresholds of 1 cm translation, 5° rotation, or 5% articulated-joint travel.
Interaction rates (%) · Paper Table 11. Measured over the segments analyzed; a segment can count toward multiple categories. Object changes are measured against the reference rollout from the same starting state.
Branches from the same starting state stay together in the train/test split. Each transition supplies its own training example; the loss does not directly contrast pairs of branches.
We compare models trained on demonstrations alone, counterfactuals alone (CF-only), and a mixture of
60% counterfactuals and 40% demonstrations
.
Policy performance across programs
LIBERO
VTA achieves
98.6%
mean success across four LIBERO suites. Each evaluated program exceeds 95% on LIBERO-Long.
Closed-loop success (%) · Paper Table 3. OpenWAM: mean ± standard deviation over three training seeds, with 50 episodes per task and 500 per suite per seed. Baselines are reported as published in their respective papers.
Decoupled remains competitive with Joint on LIBERO-Long:
97.0%
versus
96.6%
. Strong control on this benchmark does not require attention between future video and action tokens.
Bimanual manipulation
On a bimanual robot, VTA and Joint each average about
92% success
across toasting bread, completing the final layer of a 2 × 2 Rubik’s cube, and sorting cups by color. The setup uses two Franka Research 3 arms with parallel-jaw grippers, two wrist cameras, and a third-person camera.
Figure 4. Evaluation tasks.
Left: one successful rollout per real-world task. Right: the four LIBERO-90 transfer tasks, outside the LIBERO-Long source set.
View full-size figure ↗
Closed-loop success (%) · Paper Table 4. We train on 200 toast, 200 cube, and 180 cup demonstrations, each about 30 seconds, and evaluate on 50, 50, and 36 trials per method, respectively.
What do pretraining and MoT contribute?
Starting from a general-purpose video model helps, but adapting it to robot video makes a substantial difference. Our causal robot-video pretraining improves LIBERO-Long success over the original Wan2.2 initialization by
29.4 points for VTA
and
34.4 points for Joint
.
LIBERO-Long success (%) · Paper Table 12. Architecture and downstream training are fixed within each program; robot-video data and causal attention are introduced together.
The architecture also matters. Giving video and actions separate experts improves VTA by
5.0 points
and Joint by
3.0 points
over a shared DiT that processes both modalities.
LIBERO-Long success (%) · Paper Table 13. Both architectures use the same robot-video-pretrained backbone, downstream data, and training settings.
What makes a frozen IDM transfer?
Can we teach the video predictor a new task without retraining its action component? We freeze IDMs trained on LIBERO-Long and pair them with video predictors adapted to four LIBERO-90 tasks. Each predictor comes from a VTA model trained on target-task demonstrations, including their action labels; only the IDM is reused without further training.
With the same mixture of demonstrations and counterfactuals, the local-context IDM reaches
84.0%
mean success, compared with
47.0%
for the full-context IDM.
Success (%) · Paper Table 5. The four target tasks are outside the IDM’s LIBERO-Long training set. Within each task, IDM variants share the adapted video predictor and rollout protocol. VTA references run as complete policies.
The targets are stacking bowls in a tray (64), putting a book in a caddy’s left compartment (74), turning on a stove and placing a pan on it (21), and repeating the stove task in a different scene (45).
Counterfactual data is crucial here. The same local-context IDM trained only on demonstrations averages
21.5%
; adding counterfactual transitions raises it to
84.0%
, and CF-only training reaches
84.5%
. A local interface becomes much more useful when the model has seen a wider range of action outcomes.
Keeping demonstrations in the mix also helps the composed model retain its original skills: source-task success is
94.4%
, versus
90.6%
with CF-only training, while their transfer means remain close.
Source-task success (%) · Paper Table 6. All action models use the LIBERO-Long video predictor.
Do predicted futures follow the actions?
For forward dynamics, the question is whether changing the actions changes the predicted future correctly. We test
2,560 counterfactual futures
: 16 action branches from each of 160 starting contexts across ten LIBERO tasks. Every model receives the same initial observations and candidate actions.
Paper Table 7 · MSE in units of 10
−3
. Outcome accuracy asks whether the prediction’s closest match, by RGB MSE, is the supplied action’s actual outcome among K futures from the same starting state.
Counterfactual supervision improves both visual accuracy and the ability to distinguish action outcomes. CF-only training reduces RGB MSE by
34.5%
and raises identification of the correct future among 16 alternatives from
21.1% to 71.3%
.
Policy and dynamics in one model
So far, each dynamics model has been trained separately. We also train a single OpenWAM checkpoint on policy generation, inverse dynamics, and forward dynamics. It reaches
92.8%
LIBERO-Long policy success versus
88.0%
for UVA, a released system supporting the same three objectives, and outperforms UVA on the evaluated dynamics metrics.
A unified checkpoint is possible, but specialization still pays off. Dedicated policy models retain higher task success, while separately trained dynamics models give more accurate counterfactual video and end-effector position predictions.
Counterfactual rollouts · Paper Table 14. Specialist rows pair independently trained IDM and FDM checkpoints. FDM compares agent-view images; IDM compares end-effector trajectories after executing actions from the same state. MSE is in raw units.
Demonstration results and training setup
Adding demonstrations to specialist training improves accuracy on demonstrated motions. The unified model gives the closest reconstructions here.
Reconstruction on real demonstrations · Paper Table 14. Some demonstrations may also be used in policy training. Specialist rows pair separate IDM and FDM checkpoints. MSE is in raw units.
The unified model allocates 60% of training to joint video–action generation on demonstrations, 20% to local-context IDM, and 20% to local-context FDM. Within each dynamics objective, half the samples are demonstrations and half are counterfactuals.
Outlook and resources
A shared causal video backbone supports strong policies across interaction programs, while counterfactual data makes independently trained dynamics components more reusable. The next challenge is to extend this reuse from simulation to real robots, and forward prediction from local transitions to long-horizon planning.
@article{yu2026openwam,
title = {{OpenWAM}: An Open Framework for Composable World-Action Models},
author = {Yu, Heng and Yuan, David D. and Zhang, Juze and Chen, Changan and
Feng, Yao and Baldonado, Michelle and Cousins, Steve and
Fei-Fei, Li and Wu, Jiajun and Adeli, Ehsan},
journal = {arXiv preprint arXiv:2610.07922},
year = {2026},
url = {https://arxiv.org/pdf/2610.07922}
}
My main computer in 1993 was a Zilog Z280-based homebrew computer with 512 KB of RAM. It had a VGA card for display, a 3 1/2" floppy disk drive, and connectors for serial ports, and printer port.
I got a discarded Hayes modem with the amazing speed of 2400 bauds. For some reason it didn't had the top of the case, it was a bulky black long box. But it was easy to interface, and I used a lot of time to write a serial terminal program.
The Internet didn't existed yet, and the information about serial protocols and AT commands came from books, the same as the ANSI escape sequences. The serial protocol was RS-232C, and relatively easy, just take a letter from the input and display it on the screen, take any letter from the keyboard and send it back. I didn't even worry about flow-control, because the UART was set to 2400 bauds, the same as the modem. Hardware flow-control using RTS/CTS pins became common when the 14400 bauds modem appeared, as the UART would need to be set at 19200 bauds (faster than the modem)
The AT commands also were relatively simple. For example, ATDT5551234 means to dial the number 5551234 using tones. Although in Mexico in 1993 we still used ATDP5551234 to dial using pulses.
The ANSI escape sequences are formed by the Esc character (0x1b) followed by the opening square bracket [ some numbers, and a letter that signalled the command. For example, Esc [2J mean to erase the screen. Esc C moved the cursor to the right, and ESC[0;35m set a magenta color for the letters. I implemented all these escape sequences on my terminal program.
However (and luckily for this article), the book and the real usage differed slightly. So some BBS would look weird. To solve this, I implemented a capture buffer in my terminal program. The Z280 allows for MMU access to a 64K area, so I programmed it to save every received byte sequentially. When the display differed from what I expected, I would give a look to the escape sequences, and most of times found a small deviation that I could correct in my program.
Almost nothing was saved from that time, but for some reason in one of my modem sessions, I saved a whole 64K buffer of my activity calling La Cueva BBS. Just a file CUEVA.TXT I found on one of my old floppy disks, saved after February 11, 1993 (I know this because the BBS displayed the last logon time), and the session was from 9:30pm to 9:47pm, also reported by the BBS. You can see I downloaded a file, but the data for it wasn't captured as only the displayed data was captured. You can even see trash when the phone line failed, and also that the data overrun the 64k buffer (you can notice La Cueva BBS replay starts with the log off message)
One thing that couldn't be preserved was the keystroke timing, so the screens change very fast. But probably I took several seconds to read each screen.
There were also other BBS at the time like Catarsys. I think La Cueva BBS disappeared by 1994 or 1995, along all the remaining BBS. The sysop just put a temporary message, and soon the phone ceased to answer.
I made a Javascript code from the CUEVA.TXT file, and this data file is feed to an adapted
jsTerm by Peter Nitsch
with the speed of 2400 bauds. I didn't noticed before that jsTerm couldn't scroll because the security policy for the font "taints" the canvas, so I had to modify my
.htaccess
file to allow it. Also the updated
jsTerm
has been updated in my transputer git.
<Files "ansilove_font_pc_80x25.png">
Header set Access-Control-Allow-Origin "https://nanochess.org"
Header set Access-Control-Allow-Methods: "GET"
Header set Access-Control-Allow-Headers "Content-Type, Authorization"
</Files>
I'm glad I found this file, because watching it again brought memories, and surprisingly enough some of my friends didn't had an idea this BBS existed. Enjoy it!
Update Oct/06/2026: I found a list of BBS phone numbers dated 1993. It was displayed as an alternate help page in my terminal program. I've redacted my passwords. Also I found another BBS capture, it is Nopal BBS; you can select the replayed one in the selection control just above the terminal screen.
This repository contains mathematical manuscripts and supporting proof artifacts produced by an internal OpenAI model.
As part of model development, we evaluate our models on open research problems. We expanded these evaluations after performance on our existing mathematical evaluations saturated. Some outputs build upon earlier results produced by the models.
This collection includes results at different stages of verification. Not all have accompanying Lean formalizations. We will continue to update this repository with Lean formalizations as we obtain them.
Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly.
We are also exploring community-hosted repositories for these materials.
Navigating the collection
The current catalogue contains 722 manuscripts organized into 372 families. A family groups related papers, which may include a principal result, companion arguments, consequences, or alternative proofs. Each family is classified by mathematical discipline.
Start with the
overview
for descriptions of the families.
Use the
manuscript map
to find individual papers and their supporting materials.
The
preprints/
directory contains PDFs, source files, and manuscript-specific citation and build instructions.
The
Lean library
and
formalization catalogue
describe the available formal proofs, their associated papers, and verification configurations. See the
Comparator instructions
for additional checking instructions. Many, but not all, of the manuscripts have been formalized.
Reasoning summaries
We are also releasing abridged summaries of the model's reasoning, covering the following results:
The vast majority of results were obtained with the same procedure using an unreleased internal OpenAI model. On average, each result used three hours of ChatGPT Pro thinking compute with that model. Over the course of the evaluation, the model was posed approximately 4,000 problems. Aggregating the output into result families and manuscripts and requiring an appropriate level of significance led to the catalog outlined above.
Exceptions to this fixed procedure include work on a zero-free region for the Riemann zeta function and proof of the Hodge Conjecture for CM abelian varieties. Additionally, the writeup for the Re(s) > 11/12 zero-free region for the Riemann zeta function was human edited for readability.
Versions and citations
We will preserve the public release history of this collection. Corrections and revisions will be recorded as new versions, with previously released versions remaining accessible.
To cite the individual manuscript, use the BibTeX block in its directory.
It's October once again, and that means it is time to take the new release of Python for a spin (technically, it is the 3.15.0rc3 release that I'm using, the official 3.15 release is still a few days out). As I did with my
Python 3.14 performance
article of a year ago, today I'm sharing a new run of my informal Python benchmark, comparing Python 3.15 against previous interpreters all the way back to 3.10.
If you are not interested in the charts and the tables and just want to read my analysis, feel free to jump to the
conclusions
section at the end.
The benchmark
I just called my benchmark "informal". What does that mean?
Getting an objective and universal measure of the performance of a programming language is impossible. All you can do is write some programs and run them to get a measure of their performance. Other programs may show similar performance characteristics or they may not, there is really no way to know. My intention with this benchmark is just to get a feel for the performance changes across versions of Python, but I want to make it clear that I'm not trying to obtain a comprehensive performance profile of the Python interpreter.
For my benchmark I will be running two programs called
fibo.py
and
bubble.py
, which
you can inspect
if you like. These are the same programs I used in past editions of this benchmark. The first calculates numbers from the
Fibonacci sequence
, and the second sorts numbers using the
bubble sort
algorithm.
I've chosen these two programs as representative of two classes of algorithms. The Fibonacci calculation is done using
recursion
, which I have found to be somewhat inefficient in Python interpreters. On the other side, the bubble sort only uses for-loops, without any recursion. There are other types of programs that my benchmark does not attempt to cover. In particular, note that I'm not including
I/O bound
code in this benchmark.
Because some of the performance improvements in recent Python versions revolve around multi-threading, I also created a multi-threaded variation for each program, so in total I have four different tests.
The testing matrix
The complete testing matrix is actually fairly complex, because I have to run the four program variations under all the Python versions, plus the JIT and free-threading alternatives for those that have them. I like to run the tests under PyPy as well, because this interpreter has shown impressive performance in past runs of this benchmark. And to place Python performance within the wider ecosystem, I've also ported the two programs to JavaScript (Node.js) and Rust.
Here is the full test matrix that I've worked with:
2 test scripts
fibo.py
: calculates Fibonacci numbers, with recursion
bubble.py
: sorts a list of randomly generated numbers, without recursion
2 threading modes
Single-threaded
4 parallel threads
6 Python versions, plus recent versions of PyPy, Node.js and Rust:
Readers of my previous benchmarks may recall that I had an additional dimension in my matrix for Linux vs. macOS. Given that there were no significant differences between them in the two previous runs of the benchmark, I've decided to drop the macOS tests this time around, so all tests were executed on my Linux laptop, which has an Intel Core i5 CPU and runs Gentoo Linux.
The method I'm using to measure the performance of each participant in this benchmark is to run the test program three times and take the average duration of the three. In the tables of results that I share below I also show the speed difference versus the 3.15 version, and when it makes sense also the speed difference versus the previous version of a given interpreter. For speed comparisons I'm using a simple ratio, where 1x means same speed, 0.5x means half speed (or that it took twice the time to run), 2x means twice as fast (or that it ran in half the time if you prefer), etc. Hopefully this makes sense.
Test 1: fibo.py, single-threaded
Let's get started. The first test calculates the first 40 Fibonacci numbers.
fibo(40) - 1 thread
Time (secs)
vs. 3.15
vs. previous
3.10
15.9442
0.44x
3.11
9.4058
0.74x
1.70x
3.12
8.8545
0.78x
1.06x
3.13
8.8174
0.79x
1.00x
3.14
7.1659
0.97x
1.23x
3.15
6.9361
1.03x
pypy3.12
1.2517
5.54x
node-26.3
1.3899
4.99x
rust-1.97
0.0898
77.24x
Below you can see the above speeds in chart form:
From these results we can infer that for this test Python 3.15 is just a tiny bit faster than 3.14, probably not enough to matter. As I have also observed in previous years, the performance of PyPy 3.12 is out of this world, clocking in at 5.5x the speed of 3.15, and even getting a small lead over Node.js. As for Rust there are no surprises, but it is always good to know where the limits are!
Comparing each Python release against the previous one shows an interesting detail. The only two releases that made significant performance improvements over their predecessors are Python 3.11 and 3.14. You can see this clearly in the chart, when there is a larger drop in the height of a bar compared to the previous one. The other releases either maintained the same speed or made small improvements, so overall there's always been progress.
In the next set of results you can see the progress of the JIT and free-threading (FT) releases of Python on the same test. Keep in mind that these alternative versions of the Python interpreter were first introduced in 3.13, so the range of versions to evaluate is much smaller.
fibo(40) - 1 thread
Time (secs)
vs. 3.15
vs. previous
3.13 JIT
8.8339
3.14 JIT
7.1587
1.23x
3.15 JIT
5.7625
1.20x
1.24x
3.13 FT
12.0368
3.14 FT
7.1185
1.69x
3.15 FT
7.0298
0.99x
1.01x
Below is a chart with these results:
Really the most interesting thing from this is that the JIT in the 3.15 interpreter was faster than the standard interpreter of the same test, which was not the case in previous years. And a 1.20x speed increase is not negligible, this is very exciting to see!
On the free-threading front there isn't really a lot to expect because this test is single-threaded. But we can say that the performance of the free-threading interpreter is about the same as the 3.14 one.
Test 2: bubble.py, single-threaded
Let's look at the second test now. Here are the table and the chart for the single-threaded bubble sort test, which was configured to sort 10,000 random numbers:
bubble(10000) - 1 thread
Time (secs)
vs. 3.15
vs. previous
3.10
3.9918
0.49x
3.11
2.6237
0.75x
1.52x
3.12
2.7341
0.72x
0.96x
3.13
2.833
0.70x
0.97x
3.14
2.0575
0.96x
1.38x
3.15
1.9716
1.04x
pypy3.12
0.1073
18.37x
node-26.3
0.0643
30.66x
rust-1.97
0.0391
50.42x
As in the previous test, here we can also see that 3.15 had a very small improvement in performance with respect to 3.14. And we again can see that 3.11 and 3.14 are the two recent releases of Python that have really moved the needle in terms of performance. Some releases are even showing small regressions on this test. PyPy continued to be very fast, but for this test Node.js was faster.
The next table and chart show the results for the JIT and free-threading editions of the Python interpreter:
bubble(10000) - 1 thread
Time (secs)
vs. 3.15
vs. previous
3.13 JIT
2.5887
3.14 JIT
2.3624
1.10x
3.15 JIT
1.5435
1.28x
1.53x
3.13 FT
4.1888
3.14 FT
2.7248
1.54x
3.15 FT
2.6808
0.74x
1.02x
This shows the same overall picture from the first test. The 3.15 JIT once again shows an impressive 1.28x speed gain over the regular interpreter. The free-threading version shows a performance drop with respect to the standard interpreter, but a similar drop occurred with the 3.14 interpreter, so this is not a regression. As I said before, this does not matter much because this test is single-threaded, so it isn't the kind of application that will ever help the free-threading interpreter shine. I think it is reasonable to expect the free-threading interpreter to perform comparably to the regular one, so from that point of view we can say that there is work to be done yet.
Test 3: fibo.py, multi-threaded
Let's now repeat all the tests, but running 4 threads in parallel. Given that this is a very specific test that is designed to evaluate the free-threading version of the Python interpreter, I'm dropping the non-Python runs.
To get a baseline, first I ran the multi-threaded test on the standard Pythons. To be absolutely clear, these results are going to be bad for CPython, because the global interpreter lock (GIL) prevents true concurrency between threads. Here are the results and chart for the multi-threaded Fibonacci test:
fibo(40) - 4 threads
Time (secs)
vs. 3.15
vs. previous
3.10
67.2867
0.48x
3.11
49.726
0.65x
1.35x
3.12
38.7307
0.84x
1.28x
3.13
38.9421
0.84x
0.99x
3.14
31.8367
1.02x
1.22x
3.15
32.5421
0.98x
pypy3.12
5.6385
5.77x
This test shows 3.15 being a tiny bit slower than 3.14. We've seen in the single-threaded tests that the 3.15 interpreter was only slightly faster than 3.14, so overall I think we can say that the standard interpreter is about the same speed as 3.14. Here we can also see that PyPy continues to run circles around standard Python, but with the threads its speed slowed it down by a similar ratio, because PyPy's concurrency is also affected by a GIL.
Now let's see how the JIT and free-threading versions of the Python interpreters do on this test.
fibo(40) - 4 threads
Time (secs)
vs. 3.15
vs. previous
3.13 JIT
38.5661
3.14 JIT
30.9696
1.25x
3.15 JIT
27.1641
1.20x
1.14x
3.13 FT
12.4376
3.14 FT
7.3052
1.70x
3.15 FT
7.235
4.50x
1.01x
The free-threading edition of the Python 3.15 interpreter runs about 4.5 times faster than the standard interpreter, and the ratio was about the same with 3.14. This significant speed gain can be attributed to the interpreter running without the GIL, which allows for more efficient thread concurrency.
A secondary observation that we can make from these results is that the JIT edition of the interpreter, which runs with the GIL, keeps a similar edge over the standard interpreter even when running multiple threads, and this is new with 3.15.
Test 4: bubble.py, multi-threaded
We have one more set of results to go over. Here are the standard interpreter results for the bubble sort test:
bubble(10000) - 4 threads
Time (secs)
vs. 3.15
vs. previous
3.10
16.5959
0.51x
3.11
10.7541
0.79x
1.54x
3.12
10.979
0.77x
0.98x
3.13
11.1337
0.76x
0.99x
3.14
8.7198
0.97x
1.28x
3.15
8.4514
1.03x
pypy3.12
0.5165
16.36x
This is, again, more or less in alignment with the previous results, with the 3.15 interpreter just a hair faster than 3.14.
Now that we have the baseline for this test, let's have a look at the JIT and free-threading interpreters:
bubble(10000) - 4 threads
Time (secs)
vs. 3.15
vs. previous
3.13 JIT
10.3004
3.14 JIT
9.9454
1.04x
3.15 JIT
6.6417
1.27x
1.50x
3.13 FT
8.0547
3.14 FT
4.9934
1.61x
3.15 FT
4.8549
1.74x
1.03x
And here the free-threading 3.15 interpreter is once again faster than the standard one. The gains are not as impressive as in the Fibonacci test, but this type of program is still a good use case for a GIL-free interpreter.
As for the JIT results, they seem consistent with all other runs of this interpreter, which show a great improvement in 3.15.
Conclusions
I hope you enjoyed looking at my benchmark. Maybe in addition to seeing a bunch of numbers and colorful charts, you are wondering what does this all mean in practical terms.
What I want to do before ending this article is to give you my personal interpretation of what these numbers suggest, because you may want to know if it makes sense to upgrade to 3.15, and what improvements you can expect to see when you do. In this section we are leaving determinism and enter opinion territory, so please keep in mind that someone else looking at these numbers may have a completely different interpretation than mine!
With the disclaimer out of the way, I'll say that my gut feeling after running the tests is that Python 3.15 is at best a fairly minor improvement over 3.14 in terms of performance. I may eventually upgrade production projects I currently have on 3.14 such as this website, but the results that I obtained do not make me want to rush an upgrade like I did after seeing such great results with 3.14 last year.
The only area where there is a clear improvement in the 3.15 release is in the JIT. But the JIT continues to be an experimental feature, so it is not a good idea to use it in production. I will consider using the JIT in production only when it is out of the experimental phase.
Aside from the JIT, there aren't really any significant performance gains in this release. Some tests do show a small performance increment, but others show regressions, so I don't expect these small variations to translate into noticeable performance changes for a real world project. I honestly don't feel I'm losing anything by staying on 3.14 for a few more months, or even until 3.16 drops in a year and I have one more release to evaluate.
I do, however, plan to use 3.15 as my main day-to-day interpreter, and maybe I will end up upgrading some of my production projects just so that I can use some of the new features, such as the JavaScript-like unpacking of comprehensions or the lazy imports.
Thank you for visiting my blog! If you enjoyed this article, please consider supporting my work and keeping me caffeinated with a small one-time donation through
Buy me a coffee
. Thanks!
is a senior correspondent and author of
Notepad
, who has been covering all things Microsoft, PC, and tech for over 20 years.
Xbox CEO Asha Sharma told employees that Microsoft is getting ready to do something around
Grand Theft Auto VI
that “no other platform holder is doing” during an employee all-hands this morning. According to sources familiar with Microsoft’s plans, Sharma’s brief teaser was actually about Xbox managing to land a deal for the exclusive game streaming rights for
Grand Theft Auto VI
.
The deal will mean only Xbox Cloud Gaming can stream
Grand Theft Auto VI
to players at launch, but it’s not clear how long the agreement will be in place. Rockstar Games still hasn’t announced a PC version of
Grand Theft Auto VI
, so this deal will allow Xbox to market its Xbox Cloud Gaming service to PC players who would otherwise need an Xbox Series X / S or PS5 console.
News of the streaming rights comes around a month after Microsoft announced some
major changes to Xbox Cloud Gaming
. From November, the same month
Grand Theft Auto VI
releases, Microsoft is limiting streaming hours to just 15 hours for Xbox Game Pass Ultimate subscribers, 10 hours for Premium subscribers, and just five hours for Game Pass Essential.
Microsoft is also opening up access to Xbox Cloud Gaming with a pay-as-you-go option next month, rather than being limited to Game Pass subscriptions. The pay-as-you-go option is perfectly timed for
Grand Theft Auto VI
, allowing those who don’t own an Xbox, or who primarily play from a phone, to stream the game by paying for a bundle of hours.
During the Xbox all-hands, Sharma also told employees that Project Helix, the next generation of Xbox hardware, will be a “family of devices” and that Microsoft will partner with others on some of those devices. Sharma first called Helix a
family of devices
last month in an interview, after Microsoft and AMD announced a “
strategic multi-year partnership
” last year to work together on “a portfolio of devices,” including Microsoft’s next-generation Xbox console.
Follow topics and authors
from this story to see more like this in your personalized homepage feed and to receive email updates.
Tom Warren
Strands Decider 2B: a small, open-source, decision model
Earlier this year, we announced
strands-labs
, a place to get hands-on with state-of-the-art approaches to agentic AI. Today, we’re excited to add Strands Decider 2B: a small decision model optimized for fast experimentation, local development, and innovation.
Strands decider is one of a new class of decision models or system one models, a type of model that has been gaining a lot of attention since TypeSafe AI’s launch of Jev earlier this month. Unlike LLMs that can generate arbitrary output, decision models are designed to pick between sets of options (e.g. “Is the string ‘turn on the lights’ about the coffee machine? Yes or no.”, “What language is the phrase ‘sihamba ngokushesha’ in? English, Zulu, or Dutch.”) and assign simple numerical scores (e.g. “Is the phrase ‘this is the best doc I’ve ever read’ a positive sentiment? Between 0 and 1.”).
In exchange for this reduction in flexibility, decision models are faster and more capable at a given size, always produce an answer from the selected options, and can run with very low latency.
The flip side is that this approach (generating all outputs in a single parallel pass) makes it significantly worse at solving complex problems than reasoning models, and its lack of ability to generate text makes it unsuited for coding, chatbots, document summarization, and other common LLM tasks.
In addition, decision models give each decision a high-quality reliability score (i.e. “how sure can I be that this yes/no is correct?”), which is not available through frontier LLM inference APIs. They also make it highly efficient to ask multiple questions about the same prompt. This combination of properties makes them perfect for driving the types of agentic workflows we see many developers building with the Strands Harness SDK, and the recently launched Strands harness. We expect that this class of model is going to lead to a lot of interesting innovation in agentic AI over the next few weeks, months, and years.
Strands Decider 2B is our first contribution to that innovation. It’s a 2 billion parameter model, suitable for running on a local CPU or GPU, which can return answers to meaningful questions in tens of milliseconds. Its accuracy and calibration is competitive with the other models we know of in this class. We’ve released
strands-decider-2b
as open source on
GitHub
, with the weights on
Hugging Face
, including all the training data and scripts we used to build the model, making it a great place to start on your own innovation journey.
Model Architecture
The core idea is that we take a pre-trained LLM torso (Qwen3.5-2B), and remove the LM head, taking away its ability to generate text. The LM head is replaced with a pointer head which scores the answers offered by the torso for each option. It does this by scoring the hidden state at each option position against the hidden state at the
<answer>
position. This head is pretty small, just over a million total parameters. The torso is fine-tuned with a rank-16 LoRA adapter.
Figure 1: The Strands Decider 2B architecture.
As you browse through the repo, you’ll find that this is the second major iteration of the architecture. The first one was similar, but used a slot head that we found performed significantly worse. In fact, the model we’re releasing today is v19, with lots of iterations under the covers. Everything we changed in each version is covered in the repo, and you can follow along with the work we did.
How does it perform?
For models of this type, we’re interested in three performance targets: accuracy (how well it answers questions), calibration (how trustworthy its confidence scores are), and latency (how quickly it can make decisions). We’ve been measuring the first two together: accuracy on JevBench’s public set, and calibration using the Brier score on the same set. We’ve found that
strands-decider-2b
performs well on accuracy and calibration (
3rd of 33 in the 2B class, and 1st of 30 excluding the just-over-2B models
). As we’ve evolved the architecture our scores are getting better, and we have many ideas for future improvements. We hope the community joins us, in the spirit of Strands labs, in contributing new ideas of your own.
Figure 2: Accuracy and calibration (Brier score) across the training trajectory. As the architecture evolved over successive versions, both metrics improved on JevBench’s public set.
On latency,
strands-decider-2b
can make local decisions in a median of around 115ms on widely available hardware. The time taken to decide depends on the task size, approximately linearly increasing as the task size gets larger. The results in the graph here are on a local Nvidia RTX3090, but the performance on an M3 MacBook isn’t much worse, with a median latency for small tasks around 153ms. As with accuracy and calibration, we have a lot of ideas for getting better here, especially in reducing the floor.
Figure 3: Decision latency as a function of task size (in tokens), measured on a local Nvidia RTX 3090 against v18 of the model. Latency increases approximately linearly with task size.
Why 2B?
We chose to make Strands decider available as a small model for two reasons. One is that we want to encourage experimentation. You can use, and even train,
strands-decider-2b
on hardware you already have. This makes it easy, fast, and low risk to try things out. The other is that two billion total parameters, seems like something of a sweet spot: small enough for experimentation, large enough to do meaningful work. Strands decider performs 100% of the easy tasks on JevBench correctly, for example, and these types of problems map well to some of the easier problems we see people tackle with agents.
What can I do with Strands decider?
Whatever you want! More seriously, we’re seeing early success using this class of model for model routing, tool selection, evaluations, guardrails, memory, context management, and policy classification. We’ve also seen exciting innovation around building hybrid agents, using LLMs to make the hardest decisions and using decider models to make the easier rote decisions, reducing cost and latency. We’re seeing experiments combining decider models with fixed workflow languages to build another kind of hybrid workflow. Folks are also using these kinds of models to play games, automate tasks, navigate mazes, and more. The speed of innovation in this space is astonishing.
Trying it out
The easiest place to get started is through the
strands-decider
CLI:
pipinstallstrands-decider
Choice question
You can ask the model to choose based on some state and a question:
--state"Help! My payouts have been failing for 3 days! "\
--choice"Which team should handle this?=billing,sales,retail"
Example output:
choice_0 ->billing (confidence 0.768)
billing0.845
retail0.091
sales0.064
In this output we can see that the model is predicting
billing
as the answer with the highest probability score.
The repo also includes examples using
strands-decider-2b
inside a Strands agent, under
examples/strands/
. The agent itself runs locally, connects to Strands decider also running locally, and then uses the default LLM from Amazon Bedrock.
It’s a deliberately small scenario. The agent has the (obligatory) demo
get_weather
tool and a system prompt that makes it deliberately eager, so that when the user asks “What’s the weather?” without saying where, the agent guesses a city and calls the tool anyway. However, before that call runs,
strands-decider-2b
will read the conversation and the proposed tool call and answers two yes/no questions about it: are these argument values grounded in anything the user actually said (spoiler: no!) and is it too early to call this tool anyway. A few lines of Python turn the predictions into a decision, and the agent goes back to ask which city you meant instead of confidently reporting the weather somewhere nobody mentioned.
QUESTIONS= {
"args_grounded": Decider.noul(
"Are the tool's argument values grounded in facts the user actually provided?",
{
"true": "every argument value traces back to something the user said",
"false": "an argument value was guessed or invented, not stated by the user",
},
),
"premature": Decider.noul(
"Is it premature to call this tool now, before clarifying with the user?",
{
"true": "the assistant should ask a clarifying question before calling the tool",
"false": "there is nothing left to clarify; calling now is appropriate",
},
),
}
The pattern here is Strands’ intervention system. We use the
InterventionHandler
with a
before_tool_call
method, pass it to
Agent(interventions=[...])
, and it runs before any tool executes. What it returns is a typed action:
Proceed
,
Deny
,
Confirm
(stop and ask a person), or
Guide
, which hands the model back its turn with feedback rather than blocking the call outright. This existing handler is a Python class and Strands has no opinion about what goes inside it, so the same shape holds whether you’re calling our decision model, a Cedar policy, or another agent. (There are equivalent hooks around the model call and around the whole invocation.) This example is an illustration rather than a recommendation, so the questions, the threshold and the policy were all picked by hand. The point is that a decision this cheap can sit in a path where an LLM call never could.
The Strands team is working on libraries for decision model integration, so watch the repo for updates soon.
Conclusion
You can download, use, or build off
strands-decider-2b
today. All the data is available, along with everything you need to get started. You can grab the code from
GitHub
, and the latest snapshots from
Hugging Face
. Now go experiment!
Privacy watchdog launches investigation into China-based company behind Kmart ‘pervert glasses’ app
Guardian
www.theguardian.com
2026-10-06 22:00:15
Commissioner says Shenzhen Qingcheng, the app maker behind the HeyCyan app in the smartglasses, failed to respond to her inquiriesGet our breaking news email, free app or daily news podcastThe Australian privacy regulator has opened an investigation into the China-based software company behind the H...
The Australian privacy regulator has opened an investigation into the China-based software company behind the HeyCyan app in Kmart’s $89 smartglasses, following public outrage over the covert use of the devices.
In August,
Guardian Australia reported
Kmart had sold out of the Anko-branded discount version of Meta’s smartglasses that can capture images and record high-definition video, which led to widespread concern and calls to ban or restrict their use.
The glasses were removed from Kmart’s website as of Wednesday. The retailer has been approached for comment.
A GetUp petition has had more than 55,000 signatures calling to restrict the use of what has been labelled as “pervert glasses”, and councils around the country have looked to ban the glasses in public spaces such as swimming pools. The
federal government is also considering
a restriction on use in government workplaces.
The Australian privacy commissioner, Carly Kind, was asked by the attorney general to assess the privacy implications of the glasses at the time.
Kind wrote to Kmart and BDI Technology, retailers for the glasses, as well as Shenzhen Qingcheng, the app maker behind the HeyCyan app in the glasses. She also wrote to
Meta
– for its Ray-Bans glasses – and Google which is planning its own smartglasses.
In
a blog post on Wednesday
, Kind announced an investigation has been opened into Shenzhen Qingcheng after it failed to respond to inquiries, “as well as concerns arising from third party analysis of the technology itself as well as the entity’s privacy policy”
None of the other companies will be investigated as part of this process, Kind said.
The privacy commissioner also highlighted the difficulty in applying privacy law to individuals who may use the devices given the privacy act only applies to companies and Commonwealth agencies – not individuals.
“Retailers that sell smart glasses – for example on an e-commerce website – or companies that manufacture them, may not have any
Privacy
Act obligations if they do not collect any personal information with respect to those devices,” Kind said.
“Instead, the entity providing the software at use in the device is likely to be the entity that collects and holds the information for the purpose of privacy law.”
Kind said that potential changes to the privacy act proposed by the government would strengthen how privacy law would apply to surveillance wearables such as smartglasses.
“The most substantial change will be the replacement of the ‘reasonably necessary for an entity’s functions and activities’ test with a ‘fair and reasonable’ test,” she said.
“The new test will require an entity to look at a range of factors, including the extent to which an individual had genuine choice in the collection of their personal information, as well as the best interests of the child where children are involved.”
Kind said she hoped the community would “find comfort” in the proposed reforms as it would raise the bar over what personal information smartglasses can collect.
“These reforms are likely to become increasingly necessary as we face subsequent generations of surveillance wearables and connected devices, from personal assistant gadgets to ambient recording badges.”
Researchers at the University of Sydney in August
analysed 350 publicly-available
Instagram videos shot on smartglasses from around the world during 2023 to 2026, showing what they said was “a shift to discreet, point-of-view recording that is difficult for bystanders to detect”.
In a subset of these videos, about 60% of interactions could be classified as potential harassment, with subjects visibly uneasy or attempting to disengage.
Canterbury-Bankstown and Sydney City are among the councils in Sydney to have banned the glasses from some venues, as well as Brisbane in Queensland, and Yarra city council in Melbourne.
Shenzhen Qingcheng was approached for comment.
Couple was swatted 55 times in 2 years over a post about Norm Macdonald
|
CBC Radio |
Posted: October 5, 2026 10:15 PM
| Last Updated: October 5
Patrick Tomlinson and Niki Robinson awarded $575K US settlement from City of Milwaukee
Image | 1000027449
Caption: Niki Robinson, left, and Patrick Tomlinson have been awarded a $575,000 US settlement by the City Milwaukee after armed police officers came to their home 55 times in two years responding to bogus reports of violent crime. (Submitted by Patrick Tomlinson)
The first time heavily armed police officers showed up at Patrick Tomlinson's house, he was asleep in bed, naked.
It was 1 a.m. on July 25, 2022, and someone had reported a hostage situation at his house in Milwaukee, where he lives with his wife, Niki Robinson.
"All of a sudden there's people pounding on the door," Tomlinson told
As It Happens
host Nil Kӧksal.
"I go outside to be pulled out of the house by eight or nine police officers [and] handcuffed on the front porch, which I can tell you is the least fun I've had handcuffed and naked in my life."
Tomlinson, a science fiction writer and comedian, hadn't committed a crime. He was the victim of "swatting," a harassment technique in which someone places a hoax call reporting a serious threat in order to send armed police to the target's home.
Over the course of the next two and a half years, Tomlinson and Robinson were swatted 55 times, despite repeatedly explaining their situation to the Milwaukee Police Department. They sued, and have been awarded a $575,000 US settlement from the City of Milwaukee.
LISTEN | Interview with 'swatting' victim Patrick Tomlinson:
Media Audio | As It Happens :
Caption: A Milwaukee couple was awarded a $575,000 US settlement from the city after armed police officers showed up at their homes 55 times in two years responding to bogus reports of violent crimes. Patrick Tomlinson speaks to As It Happens host Nil Kӧksal about suffering years of targeted harassment.
Loading external pages may require significantly more data usage than loading CBC Lite story pages.
When asked for comment, Milwaukee police referred CBC to city attorney Evan Goyke, who did not respond to multiple requests for comment.
“While Niki and Patrick’s settlement provided closure and an acknowledgment of the violations of their rights, it is just as important that Milwaukee be more proactive looking forward," Jack Idlas, a lawyer for the couple, said in an emailed statement.
"The city has made no changes to policy and no officers have been disciplined for violating Niki and Patrick’s constitutional rights. That’s hard to believe after the taxpayers settled for over a half million dollars.
It started with a tweet
Tomlinson traces his troubles back to a social media post he made seven years ago about Canadian comedian Norm Macdonald.
In September 2018, he wrote on Twitter, now X: “Hot take: I’ve never found Norm Macdonald funny and was pretty sure all my comedy friends who did were either nuts or screwing with me."
After that, he says, he experienced years of online harassment, which is
. That harassment, he says, eventually escalated into swatting.
"Our stalkers had specifically told us that they were going to begin swatting us…. We spoke to a dispatcher. We spoke with a desk sergeant. And both of them said there's nothing we can do," he said.
"Two weeks later, here I am, you know, on my front porch in the middle of the night, exposed to all my neighbours with shotguns and pistols and AR-15s pointed at my head."
Image | vlcsnap-2026-10-05-17h12m56s495
Caption: (Strang Bradley LLC)
Police were aware that Tomlinson's home was the target of swatting calls, according to internal memos and emails submitted as court evidence in the lawsuit.
On Aug. 3, 2022, one officer sent a memo to the acting police captain requesting the couple's home be marked as a "swatter house." The captain declined, saying it could cause a "complacency issue for responding officers.”
In bodycam footage from one incident, shared with CBC by Tomlinson's lawyers, an officer can be heard saying: "This guy keeps getting swatted."
Nevertheless, Tomlinson says, the police kept showing up, sometimes multiple times per day.
On one occasion, an officer entered his home without permission, according to an agreed statement of facts filed in connection with the lawsuit. On several occasions, officers searched the couple's home without a warrant.
Tomlinson says some of the officers responding to the hoax calls were "professional and sympathetic."
"Others, much less so," he said.
It's been an exhausting ordeal for him and his wife, he said.
"Having each other there to support one another through this, emotionally and to know that we're not alone, without that, I don't think that either one of us individually would have been able to come out of this sane and intact," he said.
No charges or discipline
Nobody has been charged for the swatting calls, and none of the officers named in the lawsuit have been disciplined. The lawsuit settlement is not considered an admission of liability.
In his ruling, U.S. District Judge J.P. Stadtmueller determined that several police officers violated Tomlinson's and Robinson's constitutional rights by entering their home without permission and searching it without a warrant.
However, the judge granted those officers "qualified immunity," a legal doctrine that protects police and other government officials from prosecution or lawsuits for actions committed in the line of duty.
Most people don't realize that you don't have to be a celebrity or a politician to be targeted by this sort of thing.
- Patrick Tomlinson, swatting victim
Tomlinson says he would like to see the officers involved disciplined, and for the police department to enact a detailed swatting policy.
In the meantime, he says, he and Robinson will keep speaking out.
"This is the only avenue, really, that we have to move the needle on this issue and to, you know, educate the public on this kind of harassment and stalking," he said. "Most people don't realize that you don't have to be a celebrity or a politician to be targeted by this sort of thing."
The trick everyone misses: you don't need to keep the blocklist in RAM. Store the
domains as
sorted 40-bit hashes in flash
and binary-search them. 140,000+ domains
fit in ~0.7 MB of flash and are matched in ~10 ms, using
~50 KB of RAM
.
query in ──▶ extract domain ──▶ FNV-1a hash (+ parent suffixes)
──▶ binary-search the flash hash table
├─ hit ──▶ answer 0.0.0.0 (sinkholed)
└─ miss ──▶ forward to upstream resolver, relay the reply
Why this is interesting
Most ESP32 DNS sinkholes load the blocklist (domain
strings
) into RAM, so they
demand PSRAM. This project stores fixed
5-byte (40-bit) hashes in flash
instead:
string-in-RAM approach
this (hash-in-flash)
Hardware
ESP32 + PSRAM (~$8)
ESP32-C3, no PSRAM (~$2)
141k domains
~2.5 MB of RAM
0.67 MB of flash
RAM used
most of it
~50 KB
Lookup
string compare
~18 flash reads (~10 ms incl. WiFi RTT)
Collisions
n/a
0 at 141k (1 at 537k)
Why 40 bits?
It's the sweet spot for this flash budget. Collisions follow the
birthday bound — at 141k domains you get ~0, at 537k about 1 (i.e. one unlucky
domain gets over-blocked). Dropping to 32 bits would save 20% of the flash but
cost ~7 collisions at 250k; going to 64 bits wastes 3 bytes per domain to solve
a problem you don't have.
The same trick works on bigger chips — it isn't a C3 workaround. On a 16 MB
ESP32-S3 these hashes hold
~2.7M domains
vs ~466k for strings in 8 MB of
PSRAM. Hashes in flash beat strings in PSRAM basically everywhere; the C3 just
makes it undeniable.
Hardware
Any
ESP32-C3
board (tested on a C3 SuperMini), 4 MB flash,
no PSRAM needed
Classic
ESP32
(DevKit / WROOM, 4 MB) also builds:
pio run -e esp32dev -t upload
(community-contributed, compile-tested; the C3 is the tested target)
Power it from a
stable USB source
(a phone charger or your router's USB port).
Cheap/loose USB-C→A adapters can brown out the radio during WiFi transmit.
A
USB-A → USB-C dongle
lets it plug straight into the spare USB port on the
back of most routers — no power supply, no extra box.
No supports needed; 0.2 mm layers, ~15% infill is plenty.
Keep the antenna end clear.
The C3's PCB antenna is the zig-zag trace on the
short edge opposite the USB-C port — don't bury it in solid plastic or put metal
near it, or your RSSI will suffer.
Leave the vents open: the board idles around 45–55 °C.
Build & flash (PlatformIO)
One USB flash to get going — after that,
firmware and blocklist both update over WiFi
(see below).
⚠️
Use a
current PlatformIO
— the VSCode PlatformIO extension's bundled core, or
pip install -U platformio
in a venv. The distro/apt
platformio
package (e.g. 4.3.4) is
too old and fails with
AttributeError: ... 'resultcallback'
(issue #4). A one-click browser installer is on the way (hosting TBD).
# 1. copy the secrets template (gitignored, stays local) and edit it:# - WIFI_SSID / WIFI_PASS are optional — leave the placeholders and use the# on-device setup portal instead (below).# - WEB_USER / WEB_PASS / OTA_PASS are NOT optional: they gate the dashboard's# state-changing endpoints (/ban, /addblock, /upload, /update, /setupdate,# /forgetwifi) and network OTA. Pick real values — these used to be wide# open to anyone on the LAN.
cp src/secrets.example.h src/secrets.h
# then edit src/secrets.h# 2. build the blocklist hash table (default = StevenBlack base + Hagezi Light,# ~100k entries, WhatsApp/social safe)
python3 tools/build_blocklist.py data/blocklist.bin
# 3. flash firmware + the blocklist filesystem (the one and only USB flash)
pio run -t upload
pio run -t uploadfs
# 4. watch it boot, note the IP / open the dashboard
pio device monitor # -> http://c3adblock.local
Your own blocklists
build_blocklist.py OUT.bin [SOURCE ...]
takes any mix of URLs and local files, in any of
these formats:
hosts files
—
0.0.0.0 ads.example.com tracker.example.com
(all domains on the line are included)
plain domain lists
— one domain per line
AdGuard / Adblock basic rules
—
||ads.example.com^
blocks,
@@||ok.example.com^
removes a domain (e.g. to mirror an AdGuard Home allowlist)
A blocked domain also blocks its subdomains. Rules a DNS hash list can't express (regex,
wildcards,
$
modifiers, cosmetic
##
rules) are skipped and counted. An
@@
rule only
un-blocks that exact entry — it can't carve a subdomain out of a blocked parent. If a source
can't be downloaded the build stops instead of silently producing a smaller list
(
--allow-missing
to override).
WiFi setup (no re-flash needed)
If it can't connect (or you never set
secrets.h
), it starts an open access point
C3-AdBlock-XXXX
with a captive portal — join it from a phone, pick your network,
type the password, done. To move it to a new network later: click
Forget WiFi
on
the dashboard, or hold the
BOOT
button while powering on, and the setup portal
comes back. (
/forgetwifi
requires auth now, so it's no longer a bare URL you can
just visit — see Security below.)
Blocklist
— drop a freshly built
blocklist.bin
into
Blocklist → Upload
, or set a
URL under
Remote auto-update
and the device pulls a prebuilt
blocklist.bin
on a schedule. A fresh default list is rebuilt
every Monday
by GitHub Actions and
published at a stable URL, so pasting this once keeps a device current on its own:
https://github.com/M-Abozaid/esp32-c3-adblock/releases/download/blocklist/blocklist.bin
Firmware
— upload
.pio/build/c3/firmware.bin
under
Firmware → OTA update
; the
device verifies it and reboots into the new image. Or push over WiFi from the CLI:
pio run -t upload --upload-port c3adblock.local --upload-protocol espota
4 MB flash tradeoff:
firmware OTA needs
two
app slots, which leaves ~1.3 MB for the
blocklist (
~250k domains max
). The aggressive 537k "ultimate" list only fits the
single-app partition table (no firmware OTA). Pick your tradeoff in
partitions.csv
.
Security
The dashboard's read-only view (
/
,
/stats.json
) stays open, but every
state-changing endpoint requires
HTTP Basic Auth
(
WEB_USER
/
WEB_PASS
from
secrets.h
):
/ban
,
/addblock
,
/unblock
,
/forgetwifi
/upload
,
/update
(blocklist and firmware OTA)
/setupdate
,
/fetchnow
Network OTA (
ArduinoOTA
, e.g.
pio run -t upload --upload-port c3adblock.local --upload-protocol espota
) requires
OTA_PASS
from the same file.
Without this, anyone who could reach the device on the LAN could reflash it
with arbitrary firmware or rewrite the blocklist with zero credentials — worth
knowing given the device sits in the path of every DNS query on your network.
Custom blocked-domain names are also HTML-escaped before being rendered on the
dashboard, closing a stored-XSS path where a domain string containing markup
(added via
/addblock
) would otherwise execute in the viewing browser.
Basic Auth here is a LAN-trust-boundary control, not encryption.
Everything
is plain HTTP on :80 — this chip has no realistic budget to run a TLS server.
Basic Auth credentials are base64 (not encrypted) and sent on every authenticated
request; anyone who can already sniff your LAN traffic (open/guest WiFi, ARP
spoofing) can read them off the wire. This hardens against the common case —
another device on your network hitting the API with no credentials at all, or a
browser tab CSRF'ing it — not against an on-path network attacker.
CSRF via cached Basic Auth:
browsers auto-attach cached Basic Auth
credentials to
any
subsequent request to an already-authenticated origin —
including one triggered by a totally unrelated page the same browser visits
later (e.g.
<img src="http://c3adblock.local/forgetwifi">
, no JS required).
That would let any webpage silently drive this API once you've logged into the
dashboard once, regardless of who's on your LAN. Every mutating endpoint above
now also requires a custom
X-Requested-With: c3-adblock
header, which a plain
<img>
/auto-submitted
<form>
CSRF can't attach (only same-origin
fetch()
can, which is what the dashboard's own JS does) — this is why
/forgetwifi
is
no longer a bare URL you can visit directly; use the dashboard button instead.
Default credentials:
if
secrets.h
still has the placeholder
CHANGE_ME_WEB_PASSWORD
/
CHANGE_ME_OTA_PASSWORD
values from
secrets.example.h
, the device boots with a "password" that's public (it's
sitting in this repo's example file). The firmware logs a warning over serial
and shows a banner on the dashboard when this is the case — but it will still
boot and run, so don't skip setting real values in
secrets.h
before trusting
this on a network you don't fully control.
Out of scope for now: the WiFi setup portal's access point (
C3-AdBlock-XXXX
)
is still open (unencrypted) by design — it needs to be joinable without knowing
a password first. The real WiFi password you type into the portal is only as
safe as that local radio link during the brief setup window.
Use it
Point a device's DNS at the C3's IP, or add it as a
secondary resolver
behind
your main DNS. Test:
dig @<c3-ip> doubleclick.net # -> 0.0.0.0 (blocked)
dig @<c3-ip> github.com # -> real IP (forwarded)
Gotchas (learned the hard way)
ModemManager
(default on Fedora/Ubuntu) grabs
/dev/ttyACM0
and toggles
DTR/RTS, which
resets the C3
and blocks serial. Fix:
sudo systemctl stop ModemManager
echo'ATTRS{idVendor}=="303a", ENV{ID_MM_DEVICE_IGNORE}="1"'| sudo tee /etc/udev/rules.d/99-esp-no-modemmanager.rules
sudo udevadm control --reload-rules && sudo udevadm trigger
The C3's USB-Serial-JTAG console can swallow early boot output until the host
connects (
while(!Serial)
helps).
DNS clients add an
EDNS OPT
record; a blocked reply must contain only the
question + answer (ANCOUNT=1, NSCOUNT=ARCOUNT=0) or it's malformed.
Done / how it could grow
✅ Web dashboard — per-client block/allow counts, ban a client, add custom domains
✅ mDNS (
c3adblock.local
) for discovery
✅ OTA — firmware + blocklist update over WiFi, plus scheduled remote blocklist pulls
⬜ Bucketed prefix index — ~18 flash reads/lookup → ~1–2 (issue #3), the throughput win
⬜ Act as the DHCP server (hand itself out as DNS) for true plug-and-play
Credits
Inspired by
s60sc/ESP32_AdBlocker
— the
"answer 0.0.0.0 for blocklisted domains" idea. This is an independent from-scratch
implementation focused on the hash-in-flash optimization for PSRAM-less chips.
In our
first study
, we experimented with Jev as a decision aid for an LLM agent. The agent diagnosed and repaired incidents; Jev helped rank the agent's proposed tests and reviewed the evidence before submission.
That post ended with a more ambitious idea: giving Jev a broad view of the cluster and letting its fast, cheap judgments guide the investigation.
In this post, we present a Jev-driven diagnosis pipeline
without any LLM agent
. The pipeline programmatically collects and organizes cluster evidence, then feeds it to Jev. Jev selects a likely root cause and supporting observations, and the pipeline uses them to assemble a diagnosis report.
Across 21 SREGym-Lite faults, the Jev-driven pipeline passes
80 of 105 diagnoses (76.2%)
, with a median diagnosis time of
14.6 seconds
.
Jev answers questions by choosing from a supplied set of options. To use it for diagnosis, we need to provide both the evidence and the possible answers. We added a programmatic collector to turn cluster state into those inputs.
First, the collector reads Kubernetes objects, events, recent pod logs, and resource usage. It groups the observations by component, such as a Deployment, and summarizes signs of failure. Jev receives these summaries and chooses a likely source to inspect.
The collector then gathers more detail about that component and prepares numbered evidence items. Jev decides whether the component is the origin, a downstream victim, or unrelated, and selects the evidence that best supports its answer. The pipeline uses these choices to assemble and submit a diagnosis. If the evidence cannot support the hypothesis, it examines another candidate.
This version investigates candidates one at a time. Jev chooses among supplied options throughout the process. It does not generate commands or write the final report. The
pipeline implementation
is available on GitHub.
Figure
1
.
The diagnosis pipeline alternates programmatic evidence collection with Jev's focused decisions.
Let us look at SREGym-Lite's
mutating_webhook_resource_limits_social_network
fault in the Social Network application.
In this fault, pods created for
nginx-thrift
kept running out of memory. Its Deployment template specified a 256Mi memory limit, but new Pods had only 16Mi. A mutating admission webhook was rewriting their limits as the Pods were created. Four other webhook configurations were also present, so finding a webhook by name alone would not identify the cause.
The collector found 27 Deployments in
social-network
and summarized each one as a component. For
nginx-thrift
, it found the difference between the Pod and its template and identified a matching webhook. Here is an abridged version of the
nginx-thrift
summary Jev saw in its first call.
Component: deployment/nginx-thriftSignals: Pod was OOMKilled and restarted. Live Pod memory limit: 16Mi (Deployment template: 256Mi). Matching Pod-creation webhook: gatekeeper-mutating-webhook-configuration.
Jev then answered two choice questions using options supplied by the pipeline:
Question: Which component is the likely origin?Jev: deployment/nginx-thriftQuestion: What kind of object carries the fault?Jev: admission_webhook
The mismatch and matching webhook in the summary supported the second choice.
The pipeline then gathered more detail about
nginx-thrift
and gave Jev 26 evidence items, including these two:
E6: The nginx-thrift Pod was OOMKilled.E10: The Pod has a 16Mi memory limit, although its template says 256Mi. gatekeeper-mutating-webhook-configuration matches this Pod.
Among the follow-up questions, Jev answered:
Question: Is nginx-thrift the origin, a victim, or unrelated?Jev: originQuestion: What category names the cause?Jev: admission_or_namespace_policyQuestion: Which evidence item best shows the mechanism?Jev: E10
Jev selected E10 as key evidence. The pipeline inferred that the matching webhook caused the memory-limit change and named it in the submission. The collected evidence showed the memory mismatch and webhook match. The submitted diagnosis stated:
Root cause object: MutatingWebhookConfiguration gatekeeper-mutating-webhook-configuration, acting on Deployment nginx-thrift.Mechanism: the new Pod has a 16Mi memory limit instead of the template's 256Mi. The matching webhook rewrites the Pod at admission.Observed: nginx-thrift was OOMKilled and restarted.
All five attempts on this fault passed the diagnosis rubric. The collector did substantial diagnostic work: it found the Pod-template difference and narrowed the webhook candidates. Jev chose the affected component and the evidence to submit.
We ran the 21 fault scenarios in the September 4 SREGym-Lite cohort five times using
jev-1.13.0
. Each run submitted a diagnosis. We scored those diagnoses with
gpt-6-astra
at high reasoning effort, using SREGym's nine-question diagnosis rubric and 0.70 pass threshold. This is the historical 21-fault cohort, not the current leaderboard cohort.
The results were unusually consistent. For every fault, either all five attempts passed or all five failed. In 18 of the 21 faults, all five attempts also received the same diagnosis score. These were separate runs, and receiving the same score does not mean Jev followed the same path each time.
Per-fault results
21
faults ·
105
diagnoses
admission webhook outage hotel reservation
5
/
5
cronjob sidecar blocks completion hotel reservation
5
/
5
duplicate pvc mounts social network
5
/
5
edge request filter cpu saturation
0
/
5
env variable shadowing astronomy shop
5
/
5
finalizer deadlock controller hotel reservation
5
/
5
internal traffic policy local astronomy shop
5
/
5
kafka poison pill hol block
0
/
5
mutating webhook resource limits social network
5
/
5
namespace memory limit
5
/
5
network policy block
5
/
5
readiness probe misconfiguration social network
5
/
5
rolling update misconfigured social network
5
/
5
search rate retry collapse hotel reservation
0
/
5
secret rotation stale env credentials astronomy shop
5
/
5
service dns resolution failure social network
0
/
5
service wrong pod selection hotel reservation
5
/
5
unschedulable incorrect port assignment
5
/
5
valkey auth disruption
0
/
5
wrong dns policy astronomy shop
5
/
5
wrong service selector social network
5
/
5
The pattern points to a central design question: what granularity of cluster state should the pipeline show Jev? A coarse summary can hide the detail that explains a fault, while passing every line of YAML can bury the useful signal. Choosing the right granularity may matter as much as the model's ability to judge the evidence it receives.
Jev passed
76.2%
of diagnoses, close to GPT-5.6 Sol (medium)’s
77.8%
, while running about
7× faster
and costing about
200× less
per diagnosis.
Jev is less flexible than an LLM agent as its diagnoses depend on the evidence and answer choices the pipeline provides. But its speed and low cost make it a promising first-line diagnostic tool, while an LLM agent could handle cases that need broader investigation.
SREGym-Lite · Sep 4 cohort · 21 faults
Diagnosis performance vs.
cost
Jev-driven pipeline
Codex
Claude Code
Scroll horizontally to see all models →
Figure 2.
Diagnosis results on the same 21 SREGym-Lite faults. Jev ran five attempts per fault. Each LLM agent ran three.
By analyzing the failed runs, we found two distinct failure modes:
Jev chose the wrong clue.
In Astronomy Shop's
edge_request_filter_cpu_saturation
fault, crafted
waf
requests triggered an expensive regex in
frontend-proxy
, saturating its CPU and causing timeouts. The collector offered both the regex change and a new 100m CPU limit as evidence:
E7: WAF_RULE_REGEX added: ^([a-zA-Z]+)*$E8: CPU limit changed: unset -> 100m
Jev selected
frontend-proxy
and E8 in all five attempts. Each diagnosis scored 0.67: the judge accepted the location and affected scope, but not the explanation. The submissions blamed the limit rather than the filter rule.
The decisive evidence was missing.
In Hotel Reservation's
search_rate_retry_collapse_hotel_reservation
fault, a brief burst of search traffic filled
rate
's queue. Search retried timed-out calls, keeping
rate
overloaded after incoming traffic returned to normal. Jev focused on
rate
and its 20-QPS backend limit in all five runs. The diagnoses described overload but missed the loop between the queue, deadlines, and retries. All five failed.
The broad snapshot named
search
's retry settings, but did not show their values. The pipeline never inspected
search
in detail, and no Jev call included queue-depth or retry-attempt metrics. It also asked Jev to pick a single root-cause component, while this fault lived in the interaction between two services. The missing measurements and narrow answer choices made the correct explanation harder to reach.
The pipeline passed 76.2% of diagnoses on SREGym-Lite. Both the collector and Jev are essential to that result. The collector decides what to gather and how much detail to show. Jev uses that view to choose where to investigate and which evidence supports the diagnosis. The failures show why both parts matter: Jev can favor the wrong clue, and the collector can omit signals needed to explain a fault.
Next, we want to extend the pipeline to faults whose causes span services or evolve over time. That means collecting request-level signals and changing metrics, connecting them across components, and letting Jev consider explanations that involve more than one service. These failure modes are perfect candidates for smaller, specialized models such as GPT-6 Luna, to convert structured telemetry data into natural language that Jev can comfortably ingest. Ultimately, we believe that incorporating Jev-driven diagnosis into an SRE agent’s workflow is a significant step toward effectively combining System One and System Two models.
I wanted to share my open-source solution on Reddit. I created an account, found a community where people post links to their projects, and posted a link to my GitHub project there. The post was immediately deleted, and my account was banned.
It turns out you first need to write normal comments and posts without links to make it seem like you’re a regular user who isn’t trying to promote anything. Only then can you post a link.
Okay, so I registered a new account from a different IP address and immediately wrote three comments. The account was banned.
It turns out that before posting comments, you need to wait a few days, and then post 1–2 comments per day.
That’s exactly what I did. I created a new account, gradually posted comments, and everything was fine. I made a post that had been popular on other platforms, without any links. A minute later, someone wrote that it was AI slop. The post was deleted, and the account was banned.
It turns out that if your post is longer than two paragraphs, is well-formatted, and contains no errors, the moderators or algorithms will determine that it’s AI-generated. And generated text leads to account bans.
Lesson learned. I create a new account, write comments, then short posts, adding mistakes to the text. I try to write them in the first person, rather than as an article or essay. Everything’s fine — the posts aren’t being deleted. I publish a post with a link to the project, and the account gets banned.
Hmm, what if, instead of posting in someone else’s community, I just create my own community? I create a new account, set up a community, and publish a post with a link there. The community gets deleted, and my account gets banned.
But how come other people can post links to their GitHub projects without any problems, and their accounts don’t get banned?
I talked to moderators from various communities and other Reddit users. They all advised me to first earn 1k karma and only then post links. I was also advised to sign up through the mobile app, as that would make my account seem more trustworthy.
So that’s what I did. I installed the app on my phone, created an account, and made posts and left comments. Within a month, I reached 1k karma. I created a post with a link to another one of my GitHub projects. The post was sent for manual review. I waited several hours, but the community moderators hadn’t reviewed the post yet. An hour later, my account was banned.
I didn’t try to create any new accounts. Reddit users never got to learn about my projects, which might have interested them.
Subscribe
Pornhub returns to Australia after banning access due to age verification rules
Guardian
www.theguardian.com
2026-10-06 20:28:14
Australian adults again able to access the world’s most popular porn site after parent company introduces age checking via Apple devicesGet our breaking news email, free app or daily news podcastAustralian adults will be able to access Pornhub once again after the site brought in age-checking on App...
In March, Aylo, the parent company of the most popular porn site on the internet,
blocked Australian users
after online safety codes required adult websites verify users as being 18 or older or face potential fines of $49.5m.
Aylo said at the time the age checking requirements would not protect minors and create data privacy risks for users required to provide their personal information to the sites verifying ages.
Shortly after access was blocked, three virtual private network apps that allow users to appear to be outside Australia
rocketed up the app download charts
.
On Wednesday, Aylo’s vice-president for brand and community, Alex Kekesi, said Pornhub would return to Australia now that Apple had implemented age-checking features at the device level in iOS.
Apple rolled out its new age checks in Australia last month after a UK rollout earlier this year.
Under the requirements, people can confirm their age with Apple by using a credit card, scanning their passport, driver’s licence or government ID. But for existing Apple account holders, the company can check the existing credit card on file, or by looking at the length of time that person has held that account.
Content filtering and restrictions from downloading apps identified at 18 or older are in place until a user is identified as being over 18.
Kekesi said Apple’s approach “set a benchmark by enabling child safety features automatically for users who have not verified that they are adults, ultimately, making every phone or tablet kid-safe as a starting point”.
Aylo would restore access to Pornhub users in Australia who have confirmed their age through Apple’s age-verification process, Kekesi said.
“In our view, Apple’s Australian device-level age-verification update offers one of the strongest and hardest-to-circumvent protections currently available for helping prevent minors from accessing age-inappropriate content,” Kekesi said.
“We congratulate Apple on their extended age requirements for Apple accounts in Australia.”
Kekesi urged Apple to roll out the feature globally and for Microsoft and Google to follow suit.
“We are determined to be part of this solution and want to collaborate with government, civil society and tech partners to arrive at an effective kid-safe-by-default solution.”
When Aylo blocked Australian users in March, the office of the eSafety commissioner said it was ultimately a business decision for them.
The sex worker advocacy group Scarlet Alliance warned the requirement risked having a “chilling effect” on platforms willing to host online advertising for their services and content, and could lead to over-filtering of content, including sexual health information, due to concern about being in breach.
The rollout of the age-check on Pornhub in the UK has not been smooth. The media regualtor, Ofcom, is investigating whether minors could still access the site.
“We are concerned that Aylo may not have conducted sufficient due diligence and testing before implementing its new age assurance process,” Ofcom
said last month
.
“As a result, we are concerned that Aylo’s process may not be highly effective at preventing children from seeing pornography.”
Aylo said it would cooperate fully with the investigation, and Apple’s age-check method was one of the strongest and hardest to circumvent models available.
The boom in popularity of VPN apps following the code introduction in March appears to have been short-lived. The VPN apps that had rapidly entered the Top 25 downloads chart, according to Sensor Tower in March, have since slipped back out of the rankings.
Here, at last, is the fully updated
Debian Trixie with Raspberry Pi Desktop for PC and Mac
.
Download it from our website
now to give your other computers a Raspberry Pi makeover. We’ve even included a Raspberry Pi Connect client.
Raspberry Pi Desktop installed on a 2020 Dell XPS laptop
Back in 2016, I’d just finished the initial set of changes to make the default LXDE desktop into something Raspberry Pi-specific. (I’d wanted to call it PIXEL, an acronym for Pi XWindows Environment Lightweight I’d come up with a couple of years earlier, but a certain other tech company had claimed the word by then, so it was sadly not to be…). As I was fiddling with it, our CEO Eben Upton walked past my desk and suggested, “wouldn’t it be cool if I could run this on my Mac…?”
Side quest
A couple of weekends later, I was a bit bored at home, and, knowing that it was possible to boot a Mac into Debian, I wondered just how hard it could be to add our desktop to a stock Debian image. A few hours later, I was sitting there with my Mac booted from a USB stick, with the Raspberry Pi desktop running pretty much in its entirety. On Monday, I wandered over to Eben carrying my MacBook and said, “you might want to see this…”
So I’d proved it was possible, but there is a long way to go from a hacked-together weekend experiment to something that you can actually let people use. As a sort of ‘Skunk Works’ project, Serge (the engineer responsible for, amongst other things, packaging and creating the Raspberry Pi OS images) and I spent some time over the course of the next few months making something a bit more usable. That Christmas, we gave away a bootable DVD with the PC and Mac version of the Desktop on the front cover of The MagPi Magazine, as
Raspberry Pi Official Magazine
was called then.
The first ever release of Raspberry Pi Desktop for PC and Mac
We were, I have to say, a bit taken aback by the reception – it was far more popular than I’d ever dreamed it might be. That DVD was just a bootable live image, but in 2017, after a lot more work by Serge, we released the first installable version of Debian with Raspberry Pi Desktop for PC and Mac, and that seemed to be even more popular.
The only problem was that we had created a bit of a rod for our own backs. People really liked using the Raspberry Pi Desktop on non-Raspberry Pi computers – but the trouble is that we make our money from selling Raspberry Pi computers; we’ve never charged for the software that makes them work. So it was hard to justify spending too much time maintaining the PC version, and the engineering team here has a lot of stuff to do!
High demand
We managed to find the time to update the Desktop for the Buster and Bullseye releases of Debian, but then we all just got too busy with other things, and, while we left the Bullseye version on the website for anyone who wanted it, we simply didn’t have time to release any newer versions. But people kept on asking for it – we get two or three emails every week asking when the PC Desktop will be updated, and we haven’t had an answer, because we honestly didn’t know when we might get a chance to do it. We’ve continually tried to allocate time to be able to work on this, but it hasn’t been easy.
Until now.
Raspberry Pi Desktop running from a live image on USB on a 2014 MacBook Pro
Earlier this year, we (or rather Serge) finally got the latest version of the Desktop running on top of a Debian Trixie image. It’s now based on 64-bit Debian (the amd64 architecture) rather than the older 32-bit version, as Debian itself has stopped supporting 32-bit for PC architectures. This shouldn’t be a major problem – most PCs made in the last 15 years or so will quite happily run the 64-bit version of Debian, as will most Intel-based Macs. (Debian support for Apple Silicon is still experimental, so unfortunately those of you with the latest and greatest shiny fruit products will not be able to run this.)
Raspberry Pi Connect compatible
The new Raspberry Pi Desktop is fully up to date with all
the latest desktop changes
, such as the Control Centre and the new dock. In addition, we have added the
Raspberry Pi Connect
client, so you can control a PC or Mac running the Desktop along with your Raspberry Pis running Raspberry Pi Connect. It works in exactly the same way as on a Raspberry Pi; the only difference is that when you set up a PC for Connect, you will be asked to enter a name with which to identify that device – this isn’t required on a Raspberry Pi, where a unique hardware ID is used instead. The name you choose will then appear in the list of devices you see when you log into the Connect website, and you will be able to use screen sharing or a remote terminal on the device.
The Chromium browser running on Raspberry Pi Desktop on a Dell XPS being used to control Raspberry Pi Desktop running on an Apple MacBook Pro over
Raspberry Pi Connect
– how much more meta could we be?
How to get it
Raspberry Pi Desktop is supplied as a downloadable live Debian image: flash the .iso file to a USB stick, insert it into your computer, and boot from the USB. When prompted, choose “run with persistence” – this means that any changes you make will be saved back to that USB stick. The same menu also gives the option to reset persistence; if you choose this, any changes will be deleted and your USB stick will be back to a clean original image.
Once booted to Raspberry Pi Desktop, you can also choose to install it on your computer’s own hard drive or SSD: you can run the standard Debian installer, Calamares, from either the desktop shortcut or the Install Raspberry Pi OS entry in the System Tools category of the main menu.
The Calamares installer running on a MacBook.
The PC and Mac version of Raspberry Pi Desktop is
available right now from our website
. I hope you enjoy it – and as I said at the start, we’re really sorry about the wait!
OpenAI “rogue” agent activities found on Wikimedia projects
Simon Willison
simonwillison.net
2026-10-06 20:16:45
OpenAI “rogue” agent activities found on Wikimedia projects
Given how tempting a target wikis are for rogue agent swarms, it's not a huge surprise that Wikipedia found evidence of that activity once they went looking:
The Wikimedia Foundation conducted its own investigation to see whether Wikimedia...
The Wikimedia Foundation conducted its own investigation to see whether Wikimedia websites had been similarly affected by AI agents, focusing on those operated by OpenAI. We can confirm that we have discovered some activity by these “rogue” OpenAI agents on Wikimedia platforms. The unauthorized bot activities included edits to our wikis, some unsuccessful attempts to exploit a public note-taking tool we host, and heavy traffic, which are described more below.
They found evidence of agents editing sandbox pages, trying to use pieces of infrastructure such as Etherpad to help proxy content from elsewhere, and saw widespread crawling and "hundreds of thousands of data queries" to their Wikidata Query Service.
My best guess is that most of this was a similar (or the same) swarm of agents as those that
defaced that German wiki
while training for research tasks.
The Wikipedia sandbox wiki edits appear to have started on May 12th, and the initial test edits to the UseModWiki Sandbox page reported by that incident started on May 11th.
Comments (0) Trackbacks (0) Leave a comment Trackback