Protocol-aware recovery for consensus-based storage (2018)

Lobsters
www.usenix.org
2026-10-04 16:00:34
Comments...
Original Article
No preview for link for known binary extension (.pdf), Link: https://www.usenix.org/system/files/conference/fast18/fast18-alagappan.pdf.

Hell Gate 2026 Annual Report

hellgate
hellgatenyc.com
2026-10-04 16:00:03
As Hell Gate roars into its fifth year, let's take a minute to talk about how things are going....
Original Article

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to Hell Gate.

Your link has expired.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.

I asked Claude build a physically accurate O'Neill cylinder you can walk around

Hacker News
island-three.gruberbuilds.workers.dev
2026-10-04 15:49:12
Comments...
Original Article

Loading…

Remove and Disable Apple Macos27 AI Models Tool

Hacker News
github.com
2026-10-04 15:42:25
Comments...
Original Article

RemoveMacAI

Turn off Apple Intelligence on macOS 27 and remove its downloaded models.

macOS 27 no longer has a single switch for Apple Intelligence, and its models stay on disk after the features are turned off. RemoveMacAI turns the features off, removes the models and prevents macOS from downloading them again. All changes can be reverted.

RemoveMacAI turning off Apple Intelligence

Demo video

Install

curl -fsSL https://raw.githubusercontent.com/omlahore/RemoveMacAI/main/install.sh | bash

The script downloads the latest release, verifies its SHA-256 checksum and runs it from a temporary directory. Nothing is installed.

Every release is built from its tag by GitHub Actions and carries a build provenance attestation. To check that a download came from this repository's source:

gh attestation verify removemacai-darwin-arm64.tar.gz -R omlahore/RemoveMacAI

With Homebrew:

brew install omlahore/tap/removemacai
removemacai

RemoveMacAI shows the current state and asks for confirmation. It then opens System Settings to install its configuration profile, which macOS requires the user to approve, and removes the models.

Usage

Command Description
removemacai Show the current state, then turn Apple Intelligence off
removemacai status Show each feature and the size of the models on disk
removemacai off --keep <features> Leave the listed features on
removemacai off --dry-run Show the changes without applying them
removemacai revert Undo all changes
removemacai features List the feature names accepted by --keep

To revert with the one-line installer:

curl -fsSL https://raw.githubusercontent.com/omlahore/RemoveMacAI/main/install.sh | bash -s revert

What it changes

Features turned off: Siri (including "Hey Siri" and the menu bar icon), Writing Tools, Genmoji, Image Playground, the ChatGPT extension, summaries in Mail, Messages, Safari, Notes and notifications, Mail smart replies, inline text predictions, Spatial Photos, Photos Clean Up and Xcode predictive code completion.

Models removed: the Apple Intelligence foundation models and the models for image generation and Genmoji, Spatial Photos, Photos Clean Up and Xcode code completion.

Output of removemacai status

How it works

  • A configuration profile applies Apple's restriction keys for Apple Intelligence and forces the settings that have no restriction key.
  • Models are removed through Apple's asset service. System Integrity Protection stays enabled and no files under /System are modified directly.
  • The profile redirects the download of each removed model to a closed local port, so macOS does not download it again.
  • Removing the profile restores the previous settings. macOS downloads the models again when a feature needs them.

RemoveMacAI makes no network requests and collects no data.

FAQ

Does dictation still work? Yes. Dictation is a separate setting, and its speech models are not removed.

Do macOS updates undo the changes? No. The profile, including the download block, persists across updates.

Storage settings still lists Apple Intelligence after the models were deleted. Apple's asset service releases the models right away, but macOS deletes the files on its own schedule. Until then, System Settings > General > Storage keeps counting them under Apple Intelligence.

Why is a process named Siri still running? In macOS 27 the Spotlight window runs as a process named Siri. Some system services also stay loaded; they are protected by System Integrity Protection.

What stops working? The features listed above, apps that use Apple's on-device models (the Foundation Models framework and the Use Model action in Shortcuts), Visual Intelligence and natural-language editing in Calendar.

Requirements

Apple silicon.

macOS Status
27 Supported, tested on 27.0. On 27.0.1, use 0.2.3 or later.
26 and earlier Not supported

Uninstall

Run removemacai revert , then brew uninstall removemacai if it was installed with Homebrew.

Acknowledgements

RemoveMacAI is built on pared , a complete working tool by 4evy that first mapped the asset service, the model sets and several of the settings keys. Its license is in THIRD-PARTY-NOTICES.md .

License

MIT

Homa: The End of TCP for AI Clusters [video]

Hacker News
www.youtube.com
2026-10-04 15:42:25
Comments...

Improper redaction reveals Google Data Center water and electricity usage

Hacker News
www.1011now.com
2026-10-04 15:37:05
Comments...
Original Article

LINCOLN, Neb. (KOLN) - Nebraska data centers are required to turn in an annual report to Nebraska’s Department of Water, Energy, and Environment, but some state statutes are preventing the public from seeing just how much electricity and water data centers are using.

The Google Data Center in Lincoln, named Agate LLC, claimed their electricity usage and water usage were trade secret information, citing Neb. Rev. State §§ 81-1527 ; 84-712.05 and NAC TITLE 115, CH. 2 . In fact, Google did this for all three data centers, including their sites in Omaha and Papillion.

However, using a computer cursor to highlight the redacted text box in the report to DWEE, then copying and pasting it into a separate document reveals that Agate LLC is reporting 52.65 megawatts of electricity during peak electrical demand and 13.299 megagallons of water between cooling towers, evaporative systems and site operations in the last year. That totals out to 13 million gallons of water.

For context, 13 million gallons is enough to fill around 20 Olympic-size swimming pools, and is less than half of what the City of Lincoln reports using on Sept. 29.

The six data centers reporting as of Sept. 30 show a total use of 765 million gallons of water last year, enough water to fill 1,159.93 Olympic-size swimming pools.

More improper redaction shows that the data center using the most water annually is Fireball Group LLC, the Google Data Center in Papillion. Fireball LLC reports 547.88 megagallons for their 2025 Annual Water Consumption.

10/11 filed a public record request Sept. 30 to see redacted information from Agate LLC’s report.

10/11 filed a public record request to Nebraska DWEE on Sept. 30 to see redacted information.

10/11 filed a public record request to Nebraska DWEE on Sept. 30 to see redacted information. (Madison Pitsch | 10/11)

Further improper redaction reveals exactly how much of a 2025 tax refund Google data centers expect.

Google’s reports state that data centers only “utilize or expect to utilize” sales and use tax exemptions under the Nebraska Advantage Act. Reports state no rebates have been received for that program to date, nor have any incentive payments been received, or are expected to be received under the ImagiNE Nebraska Act.

Agate LLC reports expecting a refund of $55,822,472 from 2025 taxes. Fireball LLC is expecting a refund of $39,171,573.39 from 2025 taxes. Westwood Solutions LLC, the Google-owned data center in Omaha, is expecting a refund of $22,558,881 from 2025 taxes.

Documents submitted by Agate LLC show that the gross floor area is 288,530 square feet - about five football fields.

The Department of Water, Energy, and Environment Data Center Task Force is overseeing the execution of Governor Jim Pillen’s July 20 th Executive Order requiring data centers in the state to self-report their impact on Nebraska’s resources; primarily water, power grids and local infrastructure.

Usage reports are due by Sept. 30. Submitted reports can be read here , simply enter “DCR” into the “DEQ Program” field.

Click here to subscribe to our 10/11 NOW daily digest and breaking news alerts delivered straight to your email inbox.

Copyright 2026 KOLN. All rights reserved.

ncdu: NCurses Disk Usage (an updated fork)

Lobsters
github.com
2026-10-04 15:30:36
Comments...
Original Article

ncdu-zig

Description

Ncdu is a disk usage analyzer with an ncurses interface. It is designed to find space hogs on a remote server where you don't have an entire graphical setup available, but it is a useful tool even on regular desktop systems. Ncdu aims to be fast, simple and easy to use, and should be able to run in any minimal POSIX-like environment with ncurses installed.

See the ncdu 2 release announcement for information about the differences between this Zig implementation (2.x) and the C version (1.x).

Requirements

  • Zig 0.17
  • Some sort of POSIX-like OS
  • ncurses
  • libzstd

Install

You can use the Zig build system if you're familiar with that.

There's also a handy Makefile that supports the typical targets, e.g.:

make
sudo make install PREFIX=/usr

Iroh global content discovery

Lobsters
www.iroh.computer
2026-10-04 15:19:42
Comments...
Original Article

What got me excited about IPFS many years ago, briefly after it was announced, was being able to publish a personal website, blog post, or political pamphlet and have it remain available globally as long as enough people are interested in the content. As governments have been more sophisticated in their firewalling methods, this use case for circumvention tools and permissionless global content discovery is still incredibly relevant.

Recent events have added some urgency to this. IPFS shipyard is shutting down . This does not mean that IPFS will stop working, but it does not bode well for the future of the project.

When we had to solve hole punching, we started looking at existing systems and chose the best open source system as an initial starting point for our own implementation. So let's do the same for global content discovery.

There are a number of projects trying to solve this problem. But one project stands above all others: BitTorrent . It just works and has done so for over two decades.

So let's take a look at what makes BitTorrent the current leader in permissionless global content discovery. BitTorrent has a relatively simple protocol for blob transfer and a DHT called Mainline for global content discovery.

The transfer protocol and content discovery are separate systems . In fact, the DHT was developed later than the transfer protocol. BitTorrent was released in 2001 using centralized trackers for content discovery; the Mainline DHT was added in 2005.

BitTorrent works by creating a .torrent file that contains information about the data to be downloaded. The file is encoded using bencode , which is conceptually similar to JSON.

pieces contains the concatenated SHA-1 hashes of all piece length -sized pieces. The transfer protocol downloads blocks of these pieces from different peers. 1

The transfer protocol allows requests for ranges within pieces, but you can only validate a piece after you have downloaded it completely and computed its SHA-1 hash.

I don't want to dwell on this too long, but I think that BLAKE3 verified streaming is a superior replacement for the transfer protocol. We implemented a protocol called iroh-blobs that uses BLAKE3 verified streaming for sharing blobs of data over iroh connections. With BLAKE3, we only need a single root hash, and neither a piece length parameter nor hashes of the individual pieces. This is what allows iroh-blobs to validate any piece of a large blob without needing intermediate hashes.

For a visual explanation of how it works, see my BLAKE3 and Bao deep dive , which covers verified streaming, outboard encoding, and range requests.

We still have some work to do to make multiprovider downloads of single blobs efficient, but the streaming protocol itself with its fine grained validation is superior to the BitTorrent transfer protocol.

Initially BitTorrent used trackers to get providers for a torrent. What the DHT adds is a way to get providers without having to rely on trackers.

This works by computing the SHA-1 hash of the info section, which includes the hashes of all pieces. Then you announce on the DHT that you have content for this hash. The mainline functions for this are announce_peer and get_peers . Each DHT node will store a large set of providers.

Many more modern protocols also use DHTs for content discovery. But Mainline works extremely well compared to many modern alternatives. It is also an extremely large and stable public DHT deployment, so it is unlikely to go away any time soon. You can look at current mainline statistics .

Mainline DHT statistics from IPinfo.

So let's take a look at why.

The mainline protocol takes into account that a DHT operates across an extremely large number of nodes. You can't afford large per-node connection state. Therefore both queries and responses are constrained to fit into single non-fragmented UDP packets.

There are a number of other limitations compared to modern DHT implementations that all serve an important purpose:

  • the data stored for a provider is just the public host:port pair of the provider as seen from the DHT node . There is no user defined information in a provider record. 2

  • For newer extensions like BEP 44 the data is fully self-contained and verifiable, either a tiny piece of data or a signed record, both limited to fit into a typical MTU.

  • you can only store data after first querying the DHT node and returning the short-lived token from its response, proving that at publishing time you can receive packets at your claimed IP address. 3

The result of this minimalism is that mainline lookups typically complete in less than a second.

Here is a real BEP 44 lookup of a pkarr record using the get_mutable example from n0-mainline .

If you are traumatized by DHT lookups taking forever or timing out: it doesn't have to be this way . The mainline DHT shows that millisecond lookups are possible at a global scale.

Mainline for finding blobs providers

Now that we have established why mainline works so well, let's see if there is a way to use it for iroh blobs content discovery.

Mainline provider records are just an IPv4 host:port pair. But iroh connections are dialed by cryptographic identity, currently the Ed25519 public key aka EndpointId . So to use mainline provider records for iroh blobs endpoint discovery, we would need a way to know which EndpointId is currently listening on this host:port .

In most cases this host:port will be behind a NAT, so it is not reachable .

I tried a number of ways to use current mainline mechanisms such as BEP 44 to store this mapping, but currently this is not possible. So we need a tiny extra UDP address index service ( udp-addr-index ) to provide this mapping.

This can be an incredibly simple service. It just stores some tiny arbitrary metadata for each verified UDP host:port pair and is reachable exclusively via a simple single-packet UDP protocol. Since mainline requires UDP, and this is an extension for mainline, we don't need to handle the case where sending UDP packets is not possible.

Writes need a mechanism similar to the mainline or QUIC token mechanism to verify that the remote is actually reachable on the host:port pair. Reads don't need this, we just require a padded query packet to prevent amplification.

The service stores up to 1 KiB of arbitrary metadata. It's a last writer wins map from a live UDP host:port pair to a tiny blob. Mappings expire after some time, so a content provider has to update the mapping at regular intervals as well as when its public UDP host:port pair changes.

The current implementation keeps the map purely in memory. Records have to be updated at regular intervals anyway, and not requiring persistence makes the service much simpler and cheaper to operate. You can run an address index service for millions of iroh endpoints on a single small box with a publicly reachable UDP socket.

The address index service does not depend in any way on iroh. It is infrastructure that could also be useful for non iroh applications using mainline.

In the long term I would love the address index service function to be handled by a BitTorrent Mainline extension. It is simple, generically useful, and not specific to iroh.

For our endpoint discovery purposes we store a record containing the endpoint id, current UDP socket addr, and a timestamp, signed with the endpoint private key. So only the owner of the endpoint key can write a valid signed record, so you can't impersonate other endpoints.

The address index service however does not check anything about the endpoint. What get_peers in conjunction with this service gives us is just a list of candidate iroh endpoints which might serve the content.

Endpoint discovery doesn't have to be for content discovery. We also have an example that shows how to discover peers for a gossip topic .

So now let's take a look at the overall workflow when announcing and discovering content. For both announce and discovery we will need a mainline DHT node. We use the n0_mainline crate just like in the existing iroh-mainline-address-lookup crate.

Announce

We want to associate the UDP socket of n0_mainline with our EndpointId . As outlined above, we use an UDP addr index server for this: we publish a signed record containing our EndpointId via the same UDP socket to the addr index server. n0_mainline has a mechanism to publish arbitrary UDP packets on its socket and to intercept incoming UDP packets. This needs to happen at regular intervals as well as immediately when the public host:port pair changes.

For the announce itself: the mainline keyspace is 20 bytes, usually used to announce SHA-1 hashes of the info section of a .torrent file. We want to announce BLAKE3 hashes, so we first compute the SHA-1 hash of the BLAKE3 hash. Then we use announce_peer to announce that we are providing data for that hash. These announces also need to happen at regular intervals.

You might wonder if SHA-1 is still safe for this purpose. Its collision resistance is broken, but there is no known practical preimage attack. More importantly, the DHT is just a best-effort mechanism for finding candidate providers. We verify downloaded data against the original BLAKE3 hash, so a misleading DHT result can waste time, but cannot make us accept the wrong content.

Discovery

For discovery we first need to find at least one working service that can translate host:port pairs into EndpointId s. We have two mechanisms built in for this: a rendezvous hash that allows address index services to announce themselves, and a BEP 44 record that is a curated list of good address index services. Announce and discovery need to agree on the rendezvous hash or BEP 44 record name to find address index services.

Once we have at least one working address index service, we compute the SHA-1 hash of the BLAKE3 hash we are looking for and call get_peers . The output of get_peers is a list of host:port pairs which we translate to EndpointId s using the address index service.

At this point we get a stream of unverified EndpointId s. We currently do a quick BLAKE3 size query for the content we are looking for to make sure they are live and serve the right content, then hand them over to the iroh-blobs downloader to download the actual content.

Note that the actual connection now uses the EndpointId and iroh's built in address lookup and hole punching to establish a connection. The QUIC connection does not work via the announced UDP host:port pair. That is just a key for the address index service.

All of the above is implemented in the experimental iroh-content-discovery repository. Its workspace contains the address index service, its protocol and client as well as a few extra crates that make use of it.

To see the entire workflow in action, we can run the blobs example . It runs both publish and resolve in one process, but discovery nevertheless goes via mainline and the address index service.

So now we have a mechanism to do content discovery for blobs. This will be useful for tools like sendme and in general for sending around large amounts of data.

But what about permissionless publishing of websites such as a blog? To make this work we need some extra components.

We need a syntax for content-addressed links. We don't go into a giant rabbit hole about how to encode these links. Our content addressed links are always BLAKE3. We need to reserve some global namespace for them, so we reserved a domain blake3.net . A content-addressed link is just https://<hash>.blake3.net , with the hash encoded using zbase32 .

A content-addressed URL with its 52-character z-base-32 BLAKE3 hash and blake3.net namespace labeled. A content-addressed URL with its 52-character z-base-32 BLAKE3 hash and blake3.net namespace labeled.

We don't want to run an actual gateway at blake3.net . That would be a very bad idea for various reasons.

Instead we want the browser to interpret these links as something that can be resolved in a different way. So we wrote a simple browser plugin that just rewrites these links to http://<hash>.blake3.localhost:<port> where the port is configurable in the plugin.

That's all the plugin does.

It currently works for Brave, Chrome, and Firefox.

For Chrome and Brave, install iroh link from the Chrome Web Store .

The Firefox extension is awaiting review. For now, save the unsigned Firefox XPI to your computer using Save Link As . In Firefox 140 or later, open about:debugging#/runtime/this-firefox , click Load Temporary Add-on , and select the downloaded file. Open the extension popup, click Save , and allow access to the requested sites. You must load it again after restarting Firefox; the unsigned package cannot be installed through Install Add-on From File in regular Firefox.

The last component is a local gateway that serves content-addressed data on localhost. It interacts with mainline to find content and orchestrates the actual blobs downloads. It supports iroh-blobs collections and shows a directory index for them. It also does file type detection. It is derived from iroh blobs gateway in iroh-examples .

You can compile the gateway from source , or download the latest release .

We could build a service worker to verify the data inside the browser process.

Instead we verify the data inside the gateway.

It is software that you have to trust, whether it lives in the browser process or elsewhere doesn't matter much.

With these two components in place we can browse content-addressed data.

The local iroh gateway displaying the Open Content Library directory with Books, Essays, Films, LLM, and Website folders.

Now we have a mechanism to publish and consume content-addressed data. But we don't have a way to refer to mutable data. You could use blake3.net links from an existing website, but for that you need a registrar and a hoster, so it is not the permissionless publishing we are after.

Fortunately a permissionless DNS system already exists, pkarr . We have been using it for years for endpoint address lookup. Pkarr is a standard to publish a DNS record for a keypair, so only the owner of the secret key can publish new versions. The records are published on mainline DHT, which fits our setup neatly.

To make this compatible with our browser plus local gateway usage scenario, we reserved another domain pkarr.net and added a rewrite rule to the plugin and pkarr support to the gateway.

A pkarr link has the format <name>.pkarr.net , where name is the zbase32 encoding of an Ed25519 public key. The plugin rewrites it to <name>.pkarr.localhost:<port> . The gateway will then resolve the pkarr record and either perform a redirect or directly serve content-addressed data if the pkarr record links to hash.blake3.net .

The browser plugin keeps the public key unchanged while rewriting https://uinsazmmp47ejo8gs5dbc6rfxya14cgqhxmdqin8ae55w5aqnsio.pkarr.net to http://uinsazmmp47ejo8gs5dbc6rfxya14cgqhxmdqin8ae55w5aqnsio.pkarr.localhost:1234. The browser plugin keeps the public key unchanged while rewriting https://uinsazmmp47ejo8gs5dbc6rfxya14cgqhxmdqin8ae55w5aqnsio.pkarr.net to http://uinsazmmp47ejo8gs5dbc6rfxya14cgqhxmdqin8ae55w5aqnsio.pkarr.localhost:1234.

If you are familiar with IPFS, this is very similar to IPNS . I even did a little toy experiment called Iroh Pkarr Naming System

The pkarr-publish-resolve example in iroh-content-discovery demonstrates the whole workflow: generate a keypair, publish a blob and its name, then resolve the name, discover a provider, and download the verified content from a separate client.

Pkarr gives us permissionless names, but not human-readable ones: anyone can generate a keypair and publish under its public key, without asking a naming authority. This is the tradeoff illustrated by Zooko’s triangle : pkarr names are decentralized and cryptographically verifiable, but a 52-character encoded public key is not a memorable name.

Zooko’s triangle: pkarr sits between decentralized and secure; DNS with DNSSEC sits between human-readable and secure. Zooko’s triangle: pkarr sits between decentralized and secure; DNS with DNSSEC sits between human-readable and secure.

You can combine human readable naming systems with pkarr to get a memorable name that you can retarget in a permissionless way. We have some ideas about how to do this, stay tuned.

The first step is to run the local gateway. Most readers of this blog will probably just compile it from source, but there are also installers for windows and MacOS on github.

The next step is to install the browser plugin. For Chrome and Brave, install iroh link from the Chrome Web Store .

Make sure the port configured on the gateway and the browser plugin match. The default is 45475 or B1A3 .

And then you are ready to browse the content-addressed web.

Try https://y7rmokt6h5mryuauw83em4u1br6tqrukaw3ngtde7zp8p3bg6hto.blake3.net/ for some static content, or https://5ti57aszf7kaicsncb4wgigkf9bju39kofiz8dthwdujkmz85u8y.pkarr.net/ for a pkarr record currently pointing to the above.

Traditional website publishing just means copying your files to a directory on a server. But for content-addressed data and permissionless pkarr DNS records, you are the publisher.

So we wrote a tool iroh-share that simplifies publishing. You can think of it as sendme , but running as a daemon with a separate user interface.

Both content-addressed data and pkarr records need continuous announcements, so you might want to run this daemon on a small box in your attic that is on 24/7, or a vm in the cloud. I run it on an old Synology NAS in my attic. It's behind a NAT of course, but you do get direct connections anyway.

You can find releases at https://github.com/n0-computer/iroh-share/releases .

We now have a system for global content discovery of BLAKE3 hashed content addressed data. It is far from perfect , but it's a start .

The user interface of sharing content addressed data is extremely simple. You just state what content you want, and the gateway gets it for you.

The system behind it has a lot of room for improvement.

  • Mainline is quite scalable, but it probably won't scale for the vision of making all content on the internet content-addressable.

  • Mainline traffic is also unencrypted and therefore easily blocked by middleboxes.

  • The existing pkarr naming mechanism is using Ed25519 and therefore is not post quantum secure.

  • And perhaps most importantly, mainline does not provide any privacy . If you share content, anybody can look up your ip address .

But we are currently in the food and shelter phase of global content discovery. We will continue to improve the underlying content discovery system, possibly by extending mainline, possibly by writing our own DHT using iroh connections.

But the current system is certainly better than nothing , and we can do all these improvements while keeping the user interface stable.

We are currently working on getting our most frequently used iroh protocols irpc , iroh-gossip and iroh-blobs to 1.0. The endpoint discovery mechanism is experimental, but all relevant crates are published on crates.io for you to play with.


  1. In endgame mode , it may request the same remaining blocks from multiple peers and use whichever responses arrive first. ↩

  2. Technically, you can choose the port in an announce_peer request. But that gives you only 16 bits, which is not enough to store, for example, a 32-byte iroh EndpointId . ↩

  3. This mechanism is similar to the address validation token in QUIC. ↩

Iroh is a dial-any-device networking library that just works. Compose from an ecosystem of ready-made protocols to get the features you need, or go fully custom on a clean abstraction over dumb pipes. Iroh is open source, and already running in production on hundreds of thousands of devices.
To get started, take a look at our docs , dive directly into the code , or chat with us in our discord channel .

An animated, customizable Git cheat sheet drawn by git-sim

Lobsters
initialcommit.com
2026-10-04 12:48:24
Comments...
Original Article

Safe nothing is lost Caution rewrites history, but you can get it back Destructive can lose work for good

Start a project 4 commands

A repository is a project folder with its history in a hidden .git folder. Make a new one, or clone one that already exists, history and all.

Create a new repository

Safe

git init

Creates an empty Git repository in the current folder, stored in a hidden .git folder. No files are tracked until you add and commit them.

Shown: git init

git clone <url>

Downloads the repository at <url> with its full history into a new folder, checks out its default branch, and adds the source as the remote origin .

Shown: git clone https://example.com/your_project.git

Clone only the N latest commits

Safe

git clone --depth <N> <url>

Clones only the <N> latest commits of the default branch, leaving out the older history.

Shown: git clone --depth 3 https://example.com/your_project.git

Clone a repository and check out a specific branch

Safe

git clone -b <branch> <url>

Clones the repository and checks out <branch> instead of the remote's default branch.

Shown: git clone -b feature https://example.com/your_project.git

Configure Git 11 commands

Settings made with --global apply to all your repositories on this machine and are saved in ~/.gitconfig . Leave out --global to set something for the current repository only.

git config --global user.name " <name> "

Sets the name Git records as the author and committer of each new commit.

Shown: git config user.name "Jacob Stopak"

Set your email address

Safe

git config --global user.email " <email> "

Sets the email address Git records for the author and committer of each new commit.

Shown: git config --global user.email "jacob@initialcommit.io"

  • Undo git config --global --unset user.email

Set your default branch name

Safe

git config --global init.defaultBranch <name>

Sets the name git init gives the first branch of a new repository.

Shown: git config --global init.defaultBranch main

git config --global core.editor " <editor> "

Sets the editor Git opens for commit messages and interactive rebases. For VS Code, use code --wait so Git waits until the file is closed.

Shown: git config --global core.editor "code --wait"

List all config settings

Safe

git config --list

Prints every setting from the system, global, and repository config files. When a setting is in more than one file, the repository value overrides the global value, and the global value overrides the system value.

Shown: git config --list

Set a setting for every user on the machine

Safe

git config --system <setting> <value>

Sets <setting> to <value> in the system config file, which applies to every user and repository on the machine. A global or repository value for the same setting overrides it.

Shown: git config --system core.autocrlf true

  • Watch Writing the system config file needs admin rights.

Create a shortcut for a command

Safe

git config --global alias. <name> <command>

Creates a Git command named <name> that runs <command> . For example, alias.st status makes git st run git status .

Shown: git config --global alias.st status

Rebase instead of merging when you pull

Safe

git config --global pull.rebase true

Makes git pull rebase your local commits onto the fetched commits instead of creating a merge commit.

Shown: git config --global pull.rebase true

Prune deleted remote branches on every fetch

Safe

git config --global fetch.prune true

Makes git fetch delete remote-tracking branches, like origin/<branch> , whose branches no longer exist on the remote.

Shown: git config --global fetch.prune true

Convert line endings on Windows

Safe

git config --global core.autocrlf true

Converts LF line endings to CRLF when files are checked out, and back to LF when they're committed.

Shown: git config --global core.autocrlf true

Save your Git credentials

Safe

git config --global credential.helper <helper>

Stores your username and password or token with <helper> , like manager on Windows or osxkeychain on macOS, so Git doesn't ask for them each time.

Shown: git config --global credential.helper cache

Stage and commit 6 commands

Committing takes two steps: stage the changes you want, then commit them as one snapshot.

Check the state of your repository

Safe

git status

Shows the active branch, staged changes, unstaged changes, and untracked files. During a merge, rebase, or cherry-pick, it also shows that operation's state and the next step.

Shown: git status

Stage your file changes

Safe

git add <file>

Stages the current changes to <file> , so they're included in the next commit.

Shown: git add app.py notes.txt

Stage all changes in the current folder

Safe

git add .

Stages every new, modified, and deleted file in the current folder and its subfolders.

Shown: git add .

Stage all changes in the repository

Safe

git add -A

Stages every new, modified, and deleted file in the repository, whichever folder it's run from.

Shown: git add -A

Commit your staged changes

Safe

git commit -m " <message> "

Creates a new commit from the staged changes with the message <message> , and moves the active branch to it.

Shown: git commit -m "Describe the project in the README"

  • Undo git reset --soft HEAD~1

Stage and commit tracked file changes in one command

Safe

git commit -am " <message> "

Stages the changes to all tracked files and commits them with the message <message> . Untracked files aren't included.

Shown: git commit -a -m "Tweak the app and the header"

  • Undo git reset --soft HEAD~1

Manage files 6 commands

Delete, rename, untrack, and ignore files, and clear out the ones Git doesn't track.

Delete a file and stage the deletion

Caution

git rm <file>

Deletes <file> from the working directory and stages the deletion.

Shown: git rm login.html

  • Undo git restore --staged --worktree <file>

Stop tracking a file but keep it on disk

Safe

git rm --cached <file>

Removes <file> from the staging area, so the next commit stops tracking it. The file stays in the working directory.

Shown: git rm --cached .env

  • Undo git add <file>

Rename or move a file

Safe

git mv <old> <new>

Renames or moves <old> to <new> and stages the change.

Shown: git mv app.py main.py

  • Undo git mv <new> <old>

Find the rule that ignores a file

Safe

git check-ignore -v <file>

Prints the ignore file, line number, and pattern that ignore <file> . Prints nothing if the file isn't ignored.

Shown: git check-ignore -v debug.log build/app.js app.py

Delete all untracked files

Destructive

git clean -f

Deletes the untracked files in the working directory. Add -d to delete untracked folders too.

Shown: git clean -f

  • Undo No way back
  • Watch Untracked files aren't in Git's history, so they can't be recovered. Run git clean -n first to list what would be deleted.

Delete all untracked and ignored files

Destructive

git clean -fdx

Deletes all untracked files and folders, including ignored ones like build output and .env files.

Shown: git clean -fdx

  • Undo No way back
  • Watch Ignored local files, like settings and secrets, are deleted too. Run git clean -ndx first to list what would be deleted.

History and search 14 commands

Every commit keeps its author, date, and changes. These commands list them, filter them, and search through them. None of them change anything.

View the active branch's commit history

Safe

git log

Lists the commits in the active branch's history, newest first, with each commit's hash, author, date, and message.

Shown: git log

View the commit history of all branches

Safe

git log --all

Lists the commits in the history of every branch and tag, not only the active branch.

Shown: git log --all

View all branches as a text graph

Safe

git log --oneline --graph --all

Prints the history of all branches as a text graph, one line per commit, with branch and tag names next to their commits.

Shown: git log --oneline --graph --all

View every change to a file, commit by commit

Safe

git log -p <file>

Lists the commits that changed <file> , each with the lines it added and removed.

Shown: git log -p app.py

Follow a file's history through renames

Safe

git log --follow <file>

Lists the commits that changed <file> , including the ones from before it was renamed.

Shown: git log --follow main.py

Find the commits that added or removed a string

Safe

git log -S " <text> "

Lists the commits that added or removed <text> , found by checking where the number of times <text> appears changes.

Shown: git log -S "logging"

List one author's commits

Safe

git log --author " <name> "

Lists the commits whose author name or email matches <name> .

Shown: git log --author "Ada"

List the commits from a date range

Safe

git log --since " <date> " --until " <date> "

Lists the commits made between the two dates. Exact dates like 2024-05-01 and relative ones like 2 weeks ago both work.

Shown: git log --since "2024-05-01 11:00 +0000" --until "2024-05-01 13:00 +0000"

git show <commit>

Shows a commit's hash, author, date, and message, followed by the changes it made.

Shown: git show HEAD~1

See who last changed each line of a file

Safe

git blame <file>

Prints each line of <file> with the commit, author, and date that last changed it.

Shown: git blame app.py

Count commits by author

Safe

git shortlog -sn

Lists each author who made commits on the active branch, with their number of commits, highest first.

Shown: git shortlog -sn

Describe a commit by its nearest tag

Safe

git describe

Names the current commit after the most recent annotated tag in its history, like v1.0-2-g1a2b3c4 : the tag v1.0 , 2 commits since it, and g followed by the commit's short hash.

Shown: git describe

Search all tracked files for text

Safe

git grep -n " <text> "

Searches every tracked file for <text> and prints each matching line with its file name and line number.

Shown: git grep -n logging

Find the commit that introduced a bug

Safe

git bisect start <bad> <good>

Starts a binary search between the <bad> and <good> commits by checking out the commit halfway between them. Test each commit Git checks out and mark it with git bisect good or git bisect bad until Git reports the first bad commit.

Shown: git bisect start HEAD 4114b2c

Compare changes 5 commands

See exactly which lines changed between your files, the staging area, commits, and branches. Nothing is changed.

Show unstaged changes

Safe

git diff

Shows the changes in the working directory that aren't staged yet, line by line.

Shown: git diff

git diff --staged

Shows the staged changes line by line, which is what the next commit will contain.

Shown: git diff --staged

git diff <commit> <commit>

Shows the changes between two commits. Branch and tag names work in place of commit hashes.

Shown: git diff HEAD~2 HEAD

git diff <branch> .. <branch>

Shows the changes between the latest commits of the two branches.

Shown: git diff main..feature

Summarize the changes per file

Safe

git diff --stat

Lists each file with unstaged changes and the number of lines added and removed in it.

Shown: git diff --stat

Branches 14 commands

A branch is a movable pointer to a commit, and HEAD points to the active branch.

Create a new branch without switching to it

Safe

git branch <name>

Creates a branch named <name> that points to the current commit. HEAD stays on the active branch.

Shown: git branch bugfix-header

List all local and remote-tracking branches

Safe

git branch -a

Lists local branches, then remote-tracking branches like remotes/origin/main . A * marks the active branch.

Shown: git branch -a

List branches with their upstream and ahead/behind counts

Safe

git branch -vv

Lists each local branch with its latest commit, its upstream branch, and how many commits it's ahead or behind, like [origin/main: ahead 1, behind 2] .

Shown: git branch -vv

List branches already merged into the active branch

Safe

git branch --merged

Lists the branches whose latest commit is already in the active branch's history.

Shown: git branch --merged

List branches not yet merged into the active branch

Safe

git branch --no-merged

Lists the branches that have commits not in the active branch's history.

Shown: git branch --no-merged

Switch to another branch

Safe

git switch <branch>

Points HEAD at <branch> and updates the staging area and working directory to match its latest commit.

Shown: git switch feature

Switch back to the previous branch

Safe

git switch -

Switches to the branch that was active before the current one.

Shown: git switch -

Check out a branch from the remote

Safe

git switch <branch>

When no local branch named <branch> exists but origin/<branch> does, creates the local branch from it, sets it to track origin/<branch> , and switches to it.

Shown: git switch feature

Create a new branch and switch to it

Safe

git switch -c <name>

Creates a branch named <name> at the current commit and switches to it.

Shown: git switch -c hotfix

Create a new branch at a specific commit and switch to it

Safe

git switch -c <name> <start>

Creates a branch named <name> at <start> , which can be a commit hash, a tag, or a branch like origin/main , and switches to it.

Shown: git switch -c search origin/feature

git branch -m <old> <new>

Renames the branch <old> to <new> . Its commits don't change.

Shown: git branch -m feature search

  • Undo git branch -m <new> <old>

Delete a merged branch

Safe

git branch -d <name>

Deletes the branch <name> . Git refuses if the branch has commits that aren't merged into its upstream branch, or into the active branch if it has no upstream.

Shown: git branch -d docs

  • Undo git branch <name> <hash>

Force-delete an unmerged branch

Destructive

git branch -D <name>

Deletes the branch <name> even if it has unmerged commits. Afterward, those commits are reachable only through the reflog.

Shown: git branch -D feature

  • Undo git branch <name> <hash>
  • Watch Note the hash Git prints. Git eventually deletes commits that nothing points to.

Check out a non-branch commit (detached HEAD)

Safe

git switch --detach <commit>

Points HEAD directly at <commit> instead of at a branch, and updates the working directory to match. Create a branch before committing there.

Shown: git switch --detach 96c4fc2

Merge 6 commands

A merge brings another branch's work into yours and keeps the history as it happened.

Merge a branch into the active branch

Safe

git merge <branch>

Merges <branch> into the active branch. When both have new commits, Git creates a merge commit with two parents.

Shown: git merge feature

  • Undo git reset --hard ORIG_HEAD

Fast-forward the active branch

Safe

git merge <branch>

When the active branch has no commits that <branch> lacks, Git moves the active branch forward to <branch> 's latest commit. No merge commit is created.

Shown: git merge main

Force a merge commit even if a fast-forward merge is possible

Safe

git merge --no-ff <branch>

Creates a merge commit even when a fast-forward is possible, so the merged branch stays visible in the history.

Shown: git merge --no-ff main

  • Undo git reset --hard ORIG_HEAD

Squash a branch's changes into one commit

Safe

git merge --squash <branch>

Stages the combined changes from <branch> without committing or recording a merge. The next commit contains them as one commit.

Shown: git merge --squash feature

Finish a merge after fixing conflicts

Safe

git merge --continue

Creates the merge commit once the conflicts are resolved and the files are staged with git add .

Shown: git merge --continue

Abort a merge with conflicts

Caution

git merge --abort

Stops the merge and returns the branch, staging area, and working directory to their state before it started.

Shown: git merge --abort

  • Undo Run the merge again

Rebase and cherry-pick 8 commands

A rebase replays your commits on top of another branch, so the history reads as one line. A cherry-pick replays single commits the same way.

Rebase the active branch onto another

Caution

git rebase <branch>

Replays the active branch's commits on top of <branch> , one at a time, as new commits with new hashes.

Shown: git rebase main

  • Undo git reset --hard ORIG_HEAD
  • Watch Don't rebase commits that others have already pulled.

Run an interactive rebase to edit, squash, or reorder commits

Caution

git rebase -i <base>

Opens an editable list of the commits after <base> . Change pick to reword , squash , or drop , or reorder the lines, and Git replays the commits as listed when you save and close the file.

Shown: git rebase -i main

  • Undo git reset --hard ORIG_HEAD

Move part of a branch onto a new base commit

Caution

git rebase --onto <new-base> <old-base>

Replays only the commits after <old-base> onto <new-base> . The commits before <old-base> stay where they are.

Shown: git rebase --onto main feature

  • Undo git reset --hard ORIG_HEAD

Continue a rebase after fixing conflicts

Safe

git rebase --continue

After the conflicts are resolved and staged, commits the stopped commit and continues replaying the remaining commits.

Shown: git rebase --continue

Skip the conflicting commit in a rebase

Caution

git rebase --skip

Skips the commit that caused the conflict and continues replaying the remaining commits.

Shown: git rebase --skip

  • Watch The skipped commit's changes aren't applied at all.

Abort a rebase in progress

Caution

git rebase --abort

Stops the rebase and returns the branch to the commit it pointed to before the rebase started.

Shown: git rebase --abort

  • Undo Run the rebase again

Apply a commit from another branch to the active branch

Safe

git cherry-pick <commit>

Applies the changes from <commit> to the active branch as a new commit.

Shown: git cherry-pick fc19889

  • Undo git reset --hard HEAD~1

Apply a range of commits from another branch to the active branch

Safe

git cherry-pick <from> .. <to>

Applies each commit after <from> up to and including <to> , in order, as new commits on the active branch.

Shown: git cherry-pick 96c4fc2..feature

Remotes 8 commands

A remote is a name for the URL of another repository of the same project. A clone starts with one named origin.

git remote add <name> <url>

Adds a remote named <name> for the repository at <url> , so you can fetch from it and push to it.

Shown: git remote add upstream ../your_project.git

List remotes with their URLs

Safe

git remote -v

Lists each remote with its fetch URL and push URL.

Shown: git remote -v

Show the details of a remote

Safe

git remote show <name>

Shows the remote's URLs, its branches, which of them are tracked, and what git pull and git push do for each local branch.

Shown: git remote show origin

Change a remote's URL

Safe

git remote set-url <name> <url>

Changes the URL of the remote <name> to <url> .

Shown: git remote set-url origin https://example.com/your_project.git

  • Undo git remote set-url <name> <old-url>

git remote rename <old> <new>

Renames the remote <old> to <new> , along with its remote-tracking branches and settings.

Shown: git remote rename origin upstream

  • Undo git remote rename <new> <old>

git remote remove <name>

Removes the remote <name> with its remote-tracking branches and settings. The remote repository itself isn't affected.

Shown: git remote remove origin

  • Undo git remote add <name> <url>

List a remote's branches and tags without fetching

Safe

git ls-remote

Lists the remote's branches and tags and the commit each points to, without downloading anything.

Shown: git ls-remote

Set the upstream branch of the active branch

Safe

git branch -u origin/ <branch>

Sets origin/<branch> as the active branch's upstream, so plain git pull and git push use it.

Shown: git branch -u origin/feature

Fetch, pull, and push 12 commands

Fetch and pull bring a remote's new commits down, and push sends yours up.

Fetch new commits from the remote, without merging

Safe

git fetch

Downloads new commits from the remote and updates its remote-tracking branches, like origin/main . Local branches and the working directory don't change.

Shown: git fetch

Fetch new commits on one branch from a remote

Safe

git fetch <remote> <branch>

Downloads only <branch> 's new commits from <remote> and updates <remote>/<branch> . Local branches and the working directory don't change.

Shown: git fetch origin main

Fetch new commits on all branches from all remotes

Safe

git fetch --all

Fetches from every remote, like origin and upstream , and updates all their remote-tracking branches.

Shown: git fetch --all

Fetch and prune deleted remote branches

Safe

git fetch --prune

Fetches, then deletes remote-tracking branches, like origin/<branch> , whose branches no longer exist on the remote.

Shown: git fetch --prune

Fetch and merge remote commits

Safe

git pull

Fetches the active branch's upstream and merges it into the active branch. When both have new commits, Git creates a merge commit.

Shown: git pull

  • Undo git reset --hard ORIG_HEAD

Pull a specific branch from the remote

Safe

git pull origin <branch>

Fetches <branch> from origin and merges it into the active branch, even if the active branch tracks a different branch.

Shown: git pull origin main

  • Undo git reset --hard ORIG_HEAD

Pull and rebase instead of merging

Caution

git pull --rebase

Fetches the active branch's upstream, then replays your local commits on top of it instead of merging.

Shown: git pull --rebase

Push the active branch's commits

Safe

git push

Sends the active branch's new commits to its upstream branch on the remote.

Shown: git push

Push a new branch and set its upstream

Safe

git push -u origin <branch>

Creates <branch> on origin , pushes its commits, and sets it as the local branch's upstream.

Shown: git push -u origin hotfix

Force-push only if the remote branch hasn't changed

Caution

git push --force-with-lease

Overwrites the remote branch with your local branch, but only if the remote branch hasn't changed since your last fetch.

Shown: git push --force-with-lease

  • Watch It still rewrites history that others may have pulled.

Force-push and overwrite the remote branch

Destructive

git push --force

Overwrites the remote branch with your local branch, even if the remote has commits you don't have.

Shown: git push --force

  • Undo No way back
  • Watch Commits that others pushed are removed from the remote. Use --force-with-lease instead.

Delete a remote branch

Destructive

git push origin --delete <branch>

Deletes <branch> on the remote. The local branch isn't affected.

Shown: git push origin --delete feature

  • Undo git push origin <branch>

Tags 7 commands

A tag is a fixed pointer to a commit, usually for a release like v1.0. Unlike a branch it doesn't move, and it isn't sent to a remote until you push it.

Create a lightweight tag

Safe

git tag <name>

Creates a lightweight tag named <name> that points to the current commit.

Shown: git tag v1.0

Create an annotated tag with a message

Safe

git tag -a <name> -m " <message> "

Creates an annotated tag named <name> on the current commit, storing the tagger, the date, and <message> .

Shown: git tag -a v1.0 -m "First release"

  • Undo git tag -d <name>

git tag -l

Lists all tags in alphabetical order.

Shown: git tag -l

Delete a local tag

Caution

git tag -d <name>

Deletes the local tag <name> . The commit it pointed to isn't affected.

Shown: git tag -d v1.0

git push origin <tag>

Pushes the tag <tag> , and any commits it needs, to the remote.

Shown: git push origin v1.1

  • Undo git push origin --delete <tag>

git push --tags

Pushes every local tag the remote doesn't have yet.

Shown: git push --tags

Delete a remote tag

Caution

git push origin --delete <tag>

Deletes the tag <tag> on the remote. The local tag isn't affected.

Shown: git push origin --delete v1.0

  • Undo git push origin <tag>
  • Watch Anyone who already fetched the tag still has it.

Stash 11 commands

The stash saves changes you're not ready to commit and cleans the working directory, so you can switch branches and reapply the changes later.

Stash your uncommitted changes

Safe

git stash

Saves the staged and unstaged changes to tracked files as a new stash entry, and resets the staging area and working directory to the last commit.

Shown: git stash

  • Undo git stash pop

Stash changes including untracked files

Safe

git stash -u

Stashes untracked files as well as the changes to tracked files.

Shown: git stash -u

  • Undo git stash pop

Stash changes with a message

Safe

git stash push -m " <message> "

Stashes the changes as an entry with the message <message> , which git stash list shows.

Shown: git stash push -m "Try a grid header"

  • Undo git stash pop

Stash changes to a specific file or file(s)

Safe

git stash push <file>

Stashes only the changes to <file> , leaving all other changes in place.

Shown: git stash push styles.css

  • Undo git stash pop

List all stash entries

Safe

git stash list

Lists every stash entry, newest first, as stash@{0} , stash@{1} , and so on.

Shown: git stash list

Apply and drop the latest stash entry

Safe

git stash pop

Applies the latest stash entry to the working directory and removes it from the stash. If applying it conflicts, the entry is kept.

Shown: git stash pop

Apply the latest stash entry and keep it

Safe

git stash apply

Applies the latest stash entry to the working directory and keeps it on the stash.

Shown: git stash apply

Drop a stash entry

Caution

git stash drop

Deletes the latest stash entry, or the one named, like stash@{1} .

Shown: git stash drop

  • Watch Git prints the dropped entry's hash. git stash apply <hash> restores it until Git deletes it.

Show the changes in a stash entry

Safe

git stash show -p

Shows the changes in the latest stash entry, line by line. Name an entry to see another, like stash@{1} .

Shown: git stash show -p

Create a new branch from a stash entry

Safe

git stash branch <name>

Creates the branch <name> at the commit the latest stash entry was created on, applies the entry there, and drops it if it applied cleanly.

Shown: git stash branch try-hello

  • Undo git stash

Delete all stash entries

Destructive

git stash clear

Deletes every stash entry.

Shown: git stash clear

  • Undo No way back
  • Watch Git doesn't print the entries' hashes, so recovering them is difficult. Check git stash list first.

Undo and recover 16 commands

Most mistakes in Git can be undone, as long as you know which command undoes what. The badge on each card says what's at stake.

Unstage a file's changes

Safe

git restore --staged <file>

Unstages the changes to <file> , so they're not in the next commit. The file in the working directory doesn't change.

Shown: git restore --staged README.md

Discard unstaged changes to a file

Destructive

git restore <file>

Discards the unstaged changes to <file> , restoring it from the staging area.

Shown: git restore app.py

  • Undo No way back
  • Watch Unstaged edits aren't saved in Git, so they can't be recovered. Run git stash first if you're not sure.

Discard all unstaged changes

Destructive

git restore .

Discards the unstaged changes to every file in the current folder and its subfolders. Staged changes are kept.

Shown: git restore .

  • Undo No way back
  • Watch Discarded edits can't be recovered. Untracked files aren't affected, so use git clean for those.

Discard all unstaged changes (older Git)

Destructive

git checkout -- .

Discards the unstaged changes in the current folder, like git restore . does. The -- marks what follows as file paths, not a branch.

Shown: git checkout -- .

  • Undo No way back
  • Watch Discarded edits can't be recovered.

Restore a file from an older commit

Destructive

git restore --source <commit> <file>

Replaces <file> in the working directory with its version from <commit> . Commit the file to keep that version.

Shown: git restore --source HEAD~1 app.py

  • Undo git restore <file>
  • Watch Uncommitted changes to <file> can't be recovered.

Amend the last commit

Caution

git commit --amend -m " <message> "

Replaces the last commit with a new commit that includes any staged changes and has the message <message> . Leave out -m to edit the existing message.

Shown: git commit --amend -m "Update dependencies and README"

  • Undo git reset --soft HEAD@{1}
  • Watch Don't amend a commit that others have already pulled.

Undo the N latest commits, keep their changes staged

Caution

git reset --soft HEAD~ <N>

Moves the active branch back <N> commits ( HEAD~1 undoes only the last one). Their changes stay staged.

Shown: git reset --soft HEAD~1

  • Undo git reset --soft ORIG_HEAD

Undo the N latest commits, keep their changes unstaged

Caution

git reset HEAD~ <N>

Moves the active branch back <N> commits ( HEAD~1 undoes only the last one) and unstages their changes. The working directory doesn't change.

Shown: git reset HEAD~1

  • Undo git reset ORIG_HEAD

Reset the branch and discard all changes

Destructive

git reset --hard <commit>

Moves the active branch to <commit> and resets the staging area and working directory to match it.

Shown: git reset --hard HEAD~2

  • Undo git reset --hard ORIG_HEAD
  • Watch Uncommitted changes can't be recovered. The commits can be recovered until Git deletes them.

Reset your branch to match the remote

Destructive

git reset --hard origin/ <branch>

Moves the active branch to origin/<branch> and resets the staging area and working directory to match, discarding local commits and changes. Run git fetch first.

Shown: git reset --hard origin/main

  • Undo git reset --hard ORIG_HEAD
  • Watch Uncommitted changes can't be recovered, and unpushed commits remain only in the reflog.

Revert a commit with a new commit

Safe

git revert <commit>

Creates a new commit that reverses the changes from <commit> . Existing commits aren't rewritten, so it's safe on shared branches.

Shown: git revert HEAD

  • Undo git revert <the new commit>

Revert a merge commit

Safe

git revert -m 1 <merge>

Creates a new commit that reverses the changes a merge brought in. -m 1 makes Git reverse them relative to the merge's first parent, the branch the merge was made on.

Shown: git revert -m 1 HEAD

  • Undo git revert <the revert commit>

Revert several commits in one commit

Safe

git revert --no-commit <from> .. <to>

Reverses each commit after <from> up to and including <to> and stages the result without committing, so one commit can undo them all.

Shown: git revert --no-commit HEAD~2..HEAD

  • Undo git revert --abort

Find a lost commit in the reflog

Safe

git reflog

Lists every commit HEAD has pointed to, newest first, including commits that no branch points to anymore.

Shown: git reflog

Undo a hard reset

Destructive

git reset --hard HEAD@{1}

HEAD@{1} is the commit HEAD pointed to before its last move, so this resets the branch to where it was before the reset.

Shown: git reset --hard HEAD@{1}

  • Watch Like any reset --hard , it discards uncommitted changes.

Restore a deleted branch

Safe

git branch <name> <hash>

Creates the branch <name> at <hash> , its last commit. Git prints the hash when it deletes a branch, and git reflog shows it too.

Shown: git branch feature 1117a34

Nothing matches that search.

Vincent Bernat: Hacking the Go compiler to efficiently map IPv4 to IPv6

PlanetDebian
vincent.bernat.ch
2026-10-04 10:12:19
netip.Addr features an Unmap() method returning the unwrapped IPv4 contained in an IPv4-mapped IPv6 address: from ::ffff:203.0.113.10 or ::ffff:cb00:710a, it returns 203.0.113.10.1 There is no Map() or To6() method for the reverse direction. Such a method is trivial to implement, but Go maintainers ...
Original Article

netip.Addr features an Unmap() method returning the unwrapped IPv4 contained in an IPv4-mapped IPv6 address : from ::ffff:203.0.113.10 or ::ffff:cb00:710a , it returns 203.0.113.10 . 1 There is no Map() or To6() method for the reverse direction. Such a method is trivial to implement, but Go maintainers have rejected it on the grounds that users should write netip.AddrFrom16(ip.As16()) and let the compiler optimize it. 2 Today, this pattern is eight times slower than a native method. How can we teach the compiler to optimize this sequence?

The alternatives #

Let’s explore three ways to implement the map semantics for netip.Addr . My favorite is to add it to the Go standard library. Go maintainers prefer a small external helper chaining netip.AddrFrom16() and netip.Addr.As16() , hoping the compiler eventually optimizes it. The unsafe package opens a third path, with the same performance as the first solution.

Modifying the Go standard library #

Internally, netip.Addr stores any IP address as a 128-bit value with an extra field z to encode the family and the zone:

type Addr struct {
    addr uint128
    z unique.Handle[addrDetail]
}

type addrDetail struct {
    isV6   bool   // IPv4 is false, IPv6 is true.
    zoneV6 string // != "" only if IsV6 is true.
}

var (
    z0    unique.Handle[addrDetail]
    z4    = unique.Make(addrDetail{})
    z6noz = unique.Make(addrDetail{isV6: true})
)

AddrFrom4() encodes an IPv4 address as an IPv4-mapped IPv6 address and sets z to the unique value z4 :

// AddrFrom4 returns the address of the IPv4 address given by the bytes in addr.
func AddrFrom4(addr [4]byte) Addr {
    return Addr{
        addr: uint128{
            0,
            0xffff00000000 |
                uint64(addr[0])<<24 | uint64(addr[1])<<16 |
                uint64(addr[2])<<8 | uint64(addr[3])},
        z: z4,
    }
}

Unmap() turns an IPv4-mapped IPv6 address into an IPv4 address by setting the z field to z4 :

func (ip Addr) Unmap() Addr {
    if ip.Is4In6() {
        ip.z = z4
    }
    return ip
}

Implementing the reverse direction inside the Go standard library is trivial: we set the z field to z6noz if the address is IPv4.

// To6 maps an IPv4 address to an IPv4-mapped IPv6 address. It returns an
// IPv6 address unmodified.
func (ip Addr) To6() Addr {
    if ip.Is4() {
        ip.z = z6noz
    }
    return ip
}

As a helper #

We can’t access the z field from outside the net/netip package. Instead, we build a small helper around the netip.AddrFrom16(ip.As16()) pattern:

// AddrTo6 maps an IPv4 address to an IPv4-mapped IPv6 address. It returns an
// IPv6 address unmodified.
func AddrTo6(ip netip.Addr) netip.Addr {
    if ip.Is4() {
        ip = netip.AddrFrom16(ip.As16())
    }
    return ip
}

As an unsafe function #

Another solution uses the unsafe package to alter the Addr struct through a proxy with the same memory layout: 3

// addrProxy has the same memory layout as netip.Addr.
type addrProxy struct {
    addr [2]uint64      // netip.uint128
    z    unsafe.Pointer // unique.Handle[netip.addrDetail]
}

var (
    anyIPv6    = netip.IPv6Unspecified()
    netipZ6noz = (*addrProxy)(unsafe.Pointer(&anyIPv6)).z
)

// AddrTo6 maps an IPv4 address to an IPv4-mapped IPv6 address. It returns an
// IPv6 address unmodified.
func AddrTo6(ip netip.Addr) netip.Addr {
    if !ip.Is4() {
        return ip
    }
    (*addrProxy)(unsafe.Pointer(&ip)).z = netipZ6noz
    return ip
}

Benchmarks #

On my computer, with Go 1.27.1, the standard library solution costs 0.88 ns per operation, while the solution favored by Go maintainers costs 7.14 ns. The unsafe solution matches the performance of the first one.

goos: linux
goarch: amd64
pkg: github.com/vincentbernat/go-netip-addrto6
cpu: AMD Ryzen 5 5600X 6-Core Processor
                   │     sec/op     │
AddrTo6/safe            7.137n ± 0%
AddrTo6/unsafe         0.8682n ± 2%
AddrTo6/builtin        0.8775n ± 2%

Assembly code #

Let’s check the assembly code the compiler generates for each solution. 4 The one built into the standard library looks like this: 5

// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
 CMPQ  net/netip·z4(SB), CX     ; check "z" if this is an IPv4 address
 JNE   end                      ; if not, stop here
 MOVQ  net/netip·z6noz(SB), CX  ; CX = netip.z6noz
end:
 RET
// return value = Addr{hi: AX, lo: BX, z: CX}

Go’s assembly language is not a direct representation of the underlying machine language: it operates on a semi-abstract instruction set derived from Plan 9’s assembler . It has four pseudo-registers: FP (frame pointer for function arguments), PC (program counter), SB (static base pointer for global symbols), and SP (stack pointer). It also has architecture-specific registers like AX , CX , DX , BX , SI , DI , and R8 to R15 . Instructions storing data use their last argument as the destination. Instructions can carry an explicit size suffix: MOVB moves a byte, MOVW 16 bits, MOVL 32 bits, and MOVQ 64 bits. In the example above, the first instruction compares the 64-bit value z4 with the CX register.

The unsafe solution looks almost the same:

// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
 CMPQ  net/netip·z4(SB), CX  ; check "z" if this is an IPv4 address
 JNE   end                   ; if not, stop here
 MOVQ  netipZ6noz(SB), CX    ; CX = netip.z6noz
end:
 RET
// return value = Addr{hi: AX, lo: BX, z: CX}

The helper solution has far more instructions. To understand why, let’s look at the code for As16() and AddrFrom16() . They are short enough for the compiler to inline them.

func (ip Addr) As16() (a16 [16]byte) {
    byteorder.BEPutUint64(a16[:8], ip.addr.hi)
    byteorder.BEPutUint64(a16[8:], ip.addr.lo)
    return a16
}

func AddrFrom16(addr [16]byte) Addr {
    return Addr{
        addr: uint128{
            byteorder.BEUint64(addr[:8]),
            byteorder.BEUint64(addr[8:]),
        },
        z: z6noz,
    }
}

We can already guess the pattern to optimize: the code packs the IP address into an array, copies it, then unpacks it. If we inline the Go code by hand, we get:

func AddrTo6(input netip.Addr) netip.Addr {
    if !input.Is4() {
        return input
    }

    var a16 [16]byte
    byteorder.BEPutUint64(a16[:8], input.addr.hi)
    byteorder.BEPutUint64(a16[8:], input.addr.lo)

    addr := a16

    var output netip.Addr
    output.addr.hi = byteorder.BEUint64(addr[:8])
    output.addr.lo = byteorder.BEUint64(addr[8:])
    output.z = netip.z6noz
    return output
}

As humans, we can mentally derive the optimized form:

func AddrTo6(input netip.Addr) netip.Addr {
    if !input.Is4() {
        return input
    }
    var output netip.Addr
    output.addr.hi = input.addr.hi
    output.addr.lo = input.addr.lo
    output.z = netip.z6noz
    return output
}

Unfortunately, as of Go 1.26.8, the compiler is not smart enough to do the same:

// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
; Push the stack (32 bytes):
;    0(SP) addr netip.uint128
;   16(SP) a16 [16]byte
 PUSHQ   BP
 MOVQ    SP, BP
 SUBQ    $32, SP

 CMPQ    net/netip·z4(SB), CX  ; check "z" if this is an IPv4 address
 JNE     end                   ; if not, stop here

; Pack: byteorder.BEPutUint64(a16[:8], input.addr.hi)
;       byteorder.BEPutUint64(a16[8:], input.addr.lo)
 MOVBEQ  AX, net/netip·a16+16(SP)
 MOVBEQ  BX, net/netip·a16+24(SP)

; addr = a16, 16 bytes at once through the vector register X0
 MOVUPS  net/netip·a16+16(SP), X0
 MOVUPS  X0, net/netip·addr(SP)

; CX = netip.z6noz
 MOVQ    net/netip·z6noz(SB), CX
; Unpack: output.addr.hi = byteorder.BEUint64(addr[:8])
;         output.addr.lo = byteorder.BEUint64(addr[8:])
 MOVBEQ  net/netip·addr(SP), AX
 MOVBEQ  net/netip·addr+8(SP), BX

end:
 ADDQ    $32, SP
 POPQ    BP
 RET
// return value = Addr{hi: AX, lo: BX, z: CX}

The compiler does a decent job on the byte shuffling: the eight byte stores of BEPutUint64() become a single MOVBEQ , which stores a register byte-swapped. The eight byte loads of BEUint64() become a single MOVBEQ the other way round. 6 Three groups of instructions remain: a pack, a copy, and an unpack.

Hacking the Go compiler #

The Go compiler has several phases :

Parsing
The compiler tokenizes and parses the source code . It builds a syntax tree for each source file.
Type checking
The compiler maps each identifier to the object it denotes , folds constants, and infers the type of every expression.
IR construction
The compiler converts the syntax tree and its types into its own intermediate representation ( IR ). This process, called “ noding ,” goes through a serialization format named unified IR .
Middle end
The compiler performs several optimization passes on the IR , such as devirtualization , function call inlining , and escape analysis .
Walk
This phase runs two steps : order of evaluation decomposes complex statements into simpler ones, and desugaring transforms higher-level Go constructs, like switch or channels, into more primitive instructions or calls to the runtime.
Generic SSA
The compiler converts the IR into Static Single Assignment ( SSA ) form, a lower-level intermediate representation suited for machine-independent optimizations and rewrite rules .
Machine code generation
The compiler rewrites the SSA form into machine-specific variants , allocates registers, and applies more optimization passes. At the end, the assembler turns the generated instructions into machine code.

The hammer #

My first idea is to replace occurrences of netip.AddrFrom16(ip.As16()) with netip.Addr{addr: ip.addr, z: netip.z6noz} as early as possible, during the “noding” process. Before that, the type checking phase prevents us from accessing unexported struct fields.

Go 1.27 introduced a convenient debug option to dump the IR of a function at interesting points during compilation:

$ GOTOOLCHAIN=go1.27.1 GOAMD64=v3 go build -a -gcflags="-d=astdump=AddrTo6Safe" .
Writing text ast output for AddrTo6Safe to AddrTo6Safe.ast
Writing html ast output for AddrTo6Safe to AddrTo6Safe.html
Writing html syntax output for AddrTo6Safe to AddrTo6Safe.syntax.html

In the HTML file, the first column shows the IR as it comes out of noding:

DCLFUNC addrto6.AddrTo6Safe ABI:ABIInternal FUNC-func(netip.Addr) netip.Addr
DCLFUNC-Dcl
. NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. NAME-addrto6.~r0 Class:PPARAMOUT Offset:0 OnStack netip.Addr
DCLFUNC-body
. IF # ipv6_safe.go:11:2
. IF-Cond
. . CALLFUNC bool
. . CALLFUNC-Fun
. . . METHEXPR addrto6.Is4 FUNC-func(netip.Addr) bool
. . . . TYPE netip.Addr Class:PEXTERN Offset:0 type netip.Addr
. . CALLFUNC-Args
. . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. IF-Body
. . AS # ipv6_safe.go:12:6
. . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. . . CALLFUNC netip.Addr
. . . CALLFUNC-Fun
. . . . NAME-netip.AddrFrom16 Class:PFUNC Offset:0 Used FUNC-func([16]byte) netip.Addr
. . . CALLFUNC-Args
. . . . CALLFUNC ARRAY-[16]byte
. . . . CALLFUNC-Fun
. . . . . METHEXPR addrto6.As16 FUNC-func(netip.Addr) [16]byte
. . . . . . TYPE netip.Addr Class:PEXTERN Offset:0 type netip.Addr
. . . . CALLFUNC-Args
. . . . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. RETURN # ipv6_safe.go:14:2
. RETURN-Results
. . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr

In the body of the if statement, we spot the calls to the method netip.Addr.As16() and to the function netip.AddrFrom16() . Our goal is to patch them with a struct literal:

IF-Body
. AS # ipv6_safe.go:12:6
. . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. . STRUCTLIT netip.Addr
. . STRUCTLIT-List
. . . STRUCTKEY netip.addr
. . . . DOT netip.addr netip.uint128
. . . . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. . . STRUCTKEY netip.z
. . . . NAME-netip.z6noz Class:PEXTERN Offset:0 unique.Handle[net/netip.addrDetail]

In noder’s reader.go , the expr() method builds the IR tree for an expression. At the end of the exprCall case, we add a call to a rewriteAddrFrom16As16() function. It takes the current node and returns the struct literal on success, or nil if the rewrite is not possible. First, we check that we have the expected pattern: a call to the netip.AddrFrom16() function with a call to the netip.Addr.As16() method as its only argument:

func rewriteAddrFrom16As16(n ir.Node) ir.Node {
    call, ok := n.(*ir.CallExpr)
    if !ok || call.Op() != ir.OCALLFUNC ||
        len(call.Args) != 1 || len(call.Init()) != 0 ||
        !isNetipFunc(call.Fun, "AddrFrom16") {
        return nil
    }
    inner, ok := call.Args[0].(*ir.CallExpr)
    if !ok || inner.Op() != ir.OCALLFUNC ||
        len(inner.Args) != 1 || len(inner.Init()) != 0 ||
        !isNetipFunc(inner.Fun, "Addr.As16") {
        return nil
    }
    x := inner.Args[0]
    // [...]
}

Then, we fetch netip.z6noz :

z6noz, err := lookupVar(ir.StaticCalleeName(call.Fun).Sym().Pkg, "z6noz")
if err != nil {
    return nil
}

And we build the struct literal:

typ := call.Type()
pos := call.Pos()
var list []ir.Node
for i, f := range typ.Fields() {
    var value ir.Node
    switch f.Sym.Name {
    case "addr":
        value = typecheck.DotField(pos, x, i)
    case "z":
        value = z6noz
    default:
        return nil
    }
    list = append(list, ir.NewStructKeyExpr(pos, f, value))
}
lit := ir.NewCompLitExpr(pos, ir.OSTRUCTLIT, typ, list)
lit.SetTypecheck(1)
return lit

Have a look at the complete patch . 7 We can test it with the following commands:

$ cd src
$ ./make.bash
Building Go cmd/dist using /usr/lib/go-1.27. (go1.27.1 linux/amd64)
Building Go toolchain1 and bootstrap cmd/go (go_bootstrap) using /usr/lib/go-1.27.
Building Go toolchain2 using go_bootstrap and Go toolchain1.
Building Go toolchain3 and commands using go_bootstrap and Go toolchain2.
Checking command staleness for linux/amd64.
---
Installed Go for linux/amd64 in /home/bernat/code/free/go
Installed commands in /home/bernat/code/free/go/bin
*** You need to add /home/bernat/code/free/go/bin to your PATH.
$ export PATH=$PWD/../bin:$PATH
$ go version
go version go1.28-devel_9834516e20 Sat Sep 12 08:23:11 2026 -0700 linux/amd64
$ go test net/netip/...
ok      net/netip   0.224s
$ cd ../../go-netip-addrto6
$ go test .
ok      github.com/vincentbernat/go-netip-addrto6   0.062s

The generated code for the helper is now the shortest possible version!

// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
 CMPQ    net/netip·z4(SB), CX     ; check "z" if this is an IPv4 address
 JNE     end                      ; if not, stop here
 MOVQ    net/netip·z6noz(SB), CX  ; CX = netip.z6noz
end:
 RET
// return value = Addr{hi: AX, lo: BX, z: CX}

Go maintainers are unlikely to accept this change. It relies on the internal structure of net/netip.Addr . It’s an ugly hack in the noder, whose job is to faithfully translate the type-checked AST into the IR . And it’s harder to maintain than adding a To6() method.

The screwdriver #

The right place for such an optimization is the generic SSA phase . One of the last machine-independent passes is memcombine . With the appropriate debug flag, the compiler dumps the SSA form after this pass: 8

$ GOTOOLCHAIN=go1.26.8 GOAMD64=v3 \
> go build -a -gcflags='-d=ssa/memcombine/dump=AddrTo6Safe' .
$ head -5 AddrTo6Safe_01__memcombine.dump
AddrTo6Safe func(netip.Addr) netip.Addr
  b2:
    (?) v1 = InitMem <mem>
    (?) v2 = SP <uintptr>
    (?) v3 = SB <uintptr>

The result of the memcombine pass follows the same structure as the assembly code for AddrTo6Safe() we looked at earlier : two stores, one move, and two loads we would like to optimize away.

; […]
  v502 = ArgIntReg <uint64> {ip+0} [0]              ; input.addr.hi
  v490 = ArgIntReg <uint64> {ip+8} [1]              ; input.addr.lo
  v466 = ArgIntReg <*netip.addrDetail> {ip+16} [2]  ; input.z
; […]
  v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
  v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
  v442 = Bswap64 <uint64> v490                      ; bswap(input.addr.lo)
  v542 = Bswap64 <uint64> v502                      ; bswap(input.addr.hi)
  v161 = Store <mem> {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)
  v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)

  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16

  v415 = OffPtr <*byte> [8] v285                    ; &addr[8]
  v416 = Load <uint64> v285 v286                    ; addr[:8]
  v299 = Bswap64 <uint64> v416                      ; output.addr.hi
  v174 = Load <uint64> v415 v286                    ; addr[8:]
  v39  = Bswap64 <uint64> v174                      ; output.addr.lo
; […]

Each line features a value identifier ( v442 ), an operation with its type ( Bswap64 <uint64> ), and its arguments ( v490 ). 9 Values are the basic building blocks of SSA and are defined exactly once. Square brackets enclose integer parameters ( [8] ) and curly braces contain auxiliary arguments ( {netip.addr} ). Operations writing to memory produce a new memory state. Every memory operation takes the current state as its last argument, which keeps them in order.

On paper #

Let’s focus on output.addr.lo , aka v39 :

  v490 = ArgIntReg <uint64> {ip+8} [1]              ; input.addr.lo
  v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
  v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
  v442 = Bswap64 <uint64> v490                      ; bswap(input.addr.lo)
  v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)
  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
  v415 = OffPtr <*byte> [8] v285                    ; &addr[8]
  v174 = Load <uint64> v415 v286                    ; addr[8:]
  v39  = Bswap64 <uint64> v174                      ; output.addr.lo

To simplify this code, we could apply three rewriting rules:

  1. The first one adds a shortcut when loading through a move : (Load (OffPtr [o] p) (Move p src mem)) => (Load (OffPtr [o] src) mem) . This matches v174 with its arguments v415 and v286 and creates a new value v600 :

      v490 = ArgIntReg <uint64> {ip+8} [1]              ; input.addr.lo
      v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
      v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v442 = Bswap64 <uint64> v490                      ; bswap(input.addr.lo)
      v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)
      v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
      v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
      v415 = OffPtr <*byte> [8] v285                    ; &addr[8]
      v600 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v174 = Load <uint64> v600 v282                    ; a16[8:]
      v39  = Bswap64 <uint64> v174                      ; output.addr.lo
    
  2. The second one simplifies a load following a store : (Load p (Store p x _)) => x . The load is forwarded : the stored value replaces it and no memory access remains. It matches v174 . It notices that v600 and v173 are the same address and replaces v174 with a copy of v442 :

      v490 = ArgIntReg <uint64> {ip+8} [1]              ; input.addr.lo
      v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
      v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v442 = Bswap64 <uint64> v490                      ; bswap(input.addr.lo)
      v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)
      v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
      v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
      v415 = OffPtr <*byte> [8] v285                    ; &addr[8]
      v600 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v174 = Copy <uint64> v442                         ; bswap(input.addr.lo)
      v39  = Bswap64 <uint64> v174                      ; output.addr.lo
    
  3. The last step cancels the two byte swaps : (Bswap64 (Bswap64 x)) => x . v39 becomes a copy of v490 :

      v490 = ArgIntReg <uint64> {ip+8} [1]              ; input.addr.lo
      v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
      v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v442 = Bswap64 <uint64> v490                      ; bswap(input.addr.lo)
      v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)
      v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
      v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
      v415 = OffPtr <*byte> [8] v285                    ; &addr[8]
      v600 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v174 = Copy <uint64> v442                         ; bswap(input.addr.lo)
      v39  = Copy <uint64> v490                         ; output.addr.lo = input.addr.lo
    

If we ignore the values not needed to compute v39 , only this SSA form remains:

  v490 = ArgIntReg <uint64> {ip+8} [1]  ; input.addr.lo
  v39  = Copy <uint64> v490             ; output.addr.lo = input.addr.lo

Let’s switch to output.addr.hi , aka v299 :

  v502 = ArgIntReg <uint64> {ip+0} [0]              ; input.addr.hi
  v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
  v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
  v542 = Bswap64 <uint64> v502                      ; bswap(input.addr.hi)
  v161 = Store <mem> {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)
  v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)
  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
  v416 = Load <uint64> v285 v286                    ; addr[:8]
  v299 = Bswap64 <uint64> v416                      ; output.addr.hi

To optimize it away, we also apply three rewriting rules:

  1. The first one also adds a shortcut when loading through a move , but without an offset: (Load p (Move p src mem)) => (Load src mem) . This rewrites v416 to use arguments from v286 :

      v502 = ArgIntReg <uint64> {ip+0} [0]              ; input.addr.hi
      v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
      v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v542 = Bswap64 <uint64> v502                      ; bswap(input.addr.hi)
      v161 = Store <mem> {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)
      v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)
      v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
      v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
      v416 = Load <uint64> v22 v282                     ; a16[:8]
      v299 = Bswap64 <uint64> v416                      ; output.addr.hi
    
  2. The second rule forwards a value stored one step earlier , skipping over a store to another address: (Load p (Store q _ (Store p x _))) => x . This matches v416 : x is v542 , p is v22 ( &a16 ), q is v173 ( &a16[8] ), and p and q do not overlap for uint64 .

      v502 = ArgIntReg <uint64> {ip+0} [0]              ; input.addr.hi
      v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
      v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v542 = Bswap64 <uint64> v502                      ; bswap(input.addr.hi)
      v161 = Store <mem> {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)
      v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)
      v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
      v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
      v416 = Copy <uint64> v542                         ; bswap(input.addr.hi)
      v299 = Bswap64 <uint64> v416                      ; output.addr.hi
    
  3. The third rule cancels two byte swaps : (Bswap64 (Bswap64 x)) => x . v299 becomes a copy of v502 :

      v502 = ArgIntReg <uint64> {ip+0} [0]              ; input.addr.hi
      v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
      v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v542 = Bswap64 <uint64> v502                      ; bswap(input.addr.hi)
      v161 = Store <mem> {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)
      v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)
      v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
      v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
      v416 = Copy <uint64> v542                         ; bswap(input.addr.hi)
      v299 = Copy <uint64> v502                         ; output.addr.hi = input.addr.hi
    

If we remove the values not used to compute v299 , we get this SSA form:

  v502 = ArgIntReg <uint64> {ip+0} [0] ; input.addr.hi
  v299 = Copy <uint64> v502            ; output.addr.hi = input.addr.hi

In practice #

Most of these rules already exist in generic.rules . They use conditions to validate their context: ssa.IsSamePtr() for the same address, ssa.Disjoint() for addresses that do not overlap. The rule forwarding a stored value to a load already exists with three variants looking through several other stores. Here are the two we need:

(Load <t1> p1 (Store {t2} p2 x _))
    && ssa.IsSamePtr(p1, p2)
    && copyCompatibleType(t1, x.Type)
    && t1.Size() == t2.Size()
    => x
(Load <t1> p1 (Store {t2} p2 _ (Store {t3} p3 x _)))
    && ssa.IsSamePtr(p1, p3)
    && copyCompatibleType(t1, x.Type)
    && t1.Size() == t3.Size()
    && ssa.Disjoint(p3, t3, p2, t2)
    => x

Go 1.27 added the rule loading through a move with CL 748200 to fix issue #77720 :

(Load <t1> op1:(OffPtr [o1] p1) move:(Move [n] p2 src mem))
    && o1 >= 0 && o1+t1.Size() <= n && ssa.IsSamePtr(p1, p2)
    && !ssa.IsVolatile(src)
    => @move.Block (Load <t1> (OffPtr <op1.Type> [o1] src) mem)

It lacks a variant without an offset:

(Load <t1> p1 move:(Move [n] p2 src mem))
    && p1.Op != ssaop.OpOffPtr
    && t1.Size() <= n && ssa.IsSamePtr(p1, p2)
    && !ssa.IsVolatile(src)
    => @move.Block (Load <t1> (OffPtr <p1.Type> [0] src) mem)

There is no generic rule to cancel two byte swaps, but the AMD64 lowering pass includes this rule:

(BSWAP(Q|L) (BSWAP(Q|L) p)) => p

After switching to Go’s development branch and adding the missing rule , the generated assembly code is worse than with Go 1.26.8, even though our additional rule slightly improves the situation at the end:

// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
; Push the stack (16 bytes):
;    0(SP) a16 [16]byte
 PUSHQ   BP
 MOVQ    SP, BP
 SUBQ    $16, SP

 CMPQ    net/netip·z4(SB), CX       ; check "z" if this is an IPv4 address
 JNE     end                        ; if not, stop here

; The four forwarded bytes: the low half of input.addr.lo is taken apart and
; put back together in registers
 MOVQ    BX, DX                     ; DX = input.addr.lo
 SHRQ    $24, BX                    ; BX = input.addr.lo >> 24
 MOVQ    DX, SI                     ; SI = input.addr.lo
 SHRQ    $16, DX                    ; DX = input.addr.lo >> 16
 MOVQ    SI, DI                     ; DI = input.addr.lo, kept for the pack
 SHRQ    $8, SI                     ; SI = input.addr.lo >> 8
 MOVBLZX DIB, R8                    ; R8 = byte(input.addr.lo)
 MOVBLZX SIB, SI                    ; SI = byte(input.addr.lo >> 8)
 SHLQ    $8, SI
 ORQ     R8, SI                     ; SI = two low bytes of input.addr.lo
 MOVBLZX DL, DX                     ; DX = byte(input.addr.lo >> 16)
 SHLQ    $16, DX
 ORQ     SI, DX
 MOVBLZX BL, BX                     ; BX = byte(input.addr.lo >> 24)
 SHLQ    $24, BX
 ORQ     DX, BX                     ; BX = input.addr.lo & 0xffffffff

; Pack: byteorder.BEPutUint64(a16[:8], input.addr.hi)
;       byteorder.BEPutUint64(a16[8:], input.addr.lo)
 MOVBEQ  AX, net/netip·a16(SP)
 MOVBEQ  DI, net/netip·a16+8(SP)

; The four other bytes of input.addr.lo, read one by one from a16
 MOVBLZX net/netip·a16+11(SP), DX   ; a16[11]
 SHLQ    $32, DX
 ORQ     DX, BX
 MOVBLZX net/netip·a16+10(SP), DX   ; a16[10]
 SHLQ    $40, DX
 ORQ     DX, BX
 MOVBLZX net/netip·a16+9(SP), DX    ; a16[9]
 SHLQ    $48, DX
 ORQ     DX, BX
 MOVBLZX net/netip·a16+8(SP), DX    ; a16[8]
 SHLQ    $56, DX

; output.z = netip.z6noz
 MOVQ    net/netip·z6noz(SB), CX
; output.addr.hi = byteorder.BEUint64(a16[:8])
 MOVBEQ  net/netip·a16(SP), AX
; output.addr.lo assembled from the previous steps
 ORQ     DX, BX

end:
 LEAVEQ
 RET
// return value = Addr{hi: AX, lo: BX, z: CX}

The rule loading through a move, added in Go 1.27, introduced this regression.

Out of order #

Let’s not give up now! In reality, the rewriting rules run before memcombine , notably in the late opt pass. At this point, the inlined versions of BEPutUint64() and BEUint64() still expand to sixteen byte stores and sixteen byte loads, matching their source code :

func BEUint64(b []byte) uint64 {
    _ = b[7] // bounds check hint to compiler; see golang.org/issue/14808
    return uint64(b[7]) | uint64(b[6])<<8 | uint64(b[5])<<16 | uint64(b[4])<<24 |
        uint64(b[3])<<32 | uint64(b[2])<<40 | uint64(b[1])<<48 | uint64(b[0])<<56
}

Let’s follow two bytes of output.addr.lo : addr[15] and addr[11] . Here is a simplified SSA form before late opt :

  v273 = Trunc64to8 <byte> v490                     ; byte(input.addr.lo)
  v226 = Trunc64to8 <byte> v225                     ; byte(input.addr.lo >> 32)
; […]
  v235 = Store <mem> {byte} v233 v226 v223          ; a16[11] = byte(lo >> 32)
  v247 = Store <mem> {byte} v245 v238 v235          ; a16[12] = …
  v259 = Store <mem> {byte} v257 v250 v247          ; a16[13] = …
  v271 = Store <mem> {byte} v269 v262 v259          ; a16[14] = …
  v282 = Store <mem> {byte} v280 v273 v271          ; a16[15] = byte(lo)
  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
; […]
  v433 = OffPtr <*byte> [15] v285                   ; &addr[15]
  v435 = Load <byte> v433 v286                      ; addr[15]
  v479 = OffPtr <*byte> [11] v285                   ; &addr[11]
  v481 = Load <byte> v479 v286                      ; addr[11]

The first rule loads through the move: (Load (OffPtr [o] p) (Move p src mem)) => (Load (OffPtr [o] src) mem) . It matches both loads, which now read a16 with the memory state before the copy:

  v600 = OffPtr <*byte> [15] v22 ; &a16[15]
  v435 = Load <byte> v600 v282   ; a16[15]
  v601 = OffPtr <*byte> [11] v22 ; &a16[11]
  v481 = Load <byte> v601 v282   ; a16[11]

The second rule shortcuts a load following a store: (Load p (Store p x _)) => x . It matches v435 , as v282 stores a16[15] . It does not match v481 : v235 stores a16[11] four stores earlier in the chain, while the variants of this rule look through three stores at most.

  v435 = Copy <byte> v273        ; byte(input.addr.lo)
  v601 = OffPtr <*byte> [11] v22 ; &a16[11]
  v481 = Load <byte> v601 v282   ; a16[11]

The same happens to the other bytes: the rule forwards the four bytes stored last, a16[12] to a16[15] . The twelve other loads now read a16 instead of addr .

BEUint64() becomes a chain of Or64 , each one adding a byte shifted into place. memcombine is a pass written in Go , not a set of rewrite rules. It starts from the last Or64 of the chain and collects up to eight terms. If each term is a byte load, extended to 64 bits and shifted, and if the eight loads read consecutive addresses from the same pointer with the same memory state, it replaces the whole chain with a single 64-bit load and a byte swap. Otherwise, it tries again with four, then two terms, and from each intermediate Or64 . Here is the loop checking each term in a simplified version of combineLoads() :

for i := int64(0); i < n; i++ {
    v := a[i]
    shift := int64(0)
    if v.Op == shiftOp {
        v, shift = peelShift(v)
    }
    if v.Op != extOp {
        return false
    }
    load := v.Args[0]
    if load.Op != ssaop.OpLoad {
        return false
    }
    if load.Args[1] != mem {
        return false
    }
    p, off := splitPtr(load.Args[0])
    if p != base {
        return false
    }
    r[i] = LoadRecord{load: load, offset: off, shift: shift}
}

For output.addr.hi , the eight loads read a16 with the same memory state v282 :

  v13  = Load <byte> v22 v282               ; a16[0]
  v530 = Load <byte> v14 v282               ; a16[1]
  v488 = Load <byte> v504 v282              ; a16[2]
  v405 = Load <byte> v537 v282              ; a16[3]
  v385 = Load <byte> v397 v282              ; a16[4]
  v361 = Load <byte> v373 v282              ; a16[5]
  v196 = Load <byte> v63 v282               ; a16[6]
  v432 = Load <byte> v315 v282              ; a16[7]
  v319 = ZeroExt8to64 <uint64> v432         ; uint64(a16[7])
  v329 = ZeroExt8to64 <uint64> v196         ; uint64(a16[6])
  v330 = Lsh64x64 <uint64> [true] v329 v138 ; uint64(a16[6]) << 8
  v331 = Or64 <uint64> v319 v330            ; a16[7] | a16[6] << 8
; […] same for a16[5] to a16[1]
  v401 = ZeroExt8to64 <uint64> v13          ; uint64(a16[0])
  v402 = Lsh64x64 <uint64> [true] v401 v55  ; uint64(a16[0]) << 56
  v403 = Or64 <uint64> v402 v391            ; | a16[0] << 56 = output.addr.hi

memcombine merges them into one load and a swap:

  v286 = Load <uint64> v22 v282   ; a16[:8]
  v285 = Bswap64 <uint64> v286    ; output.addr.hi

For output.addr.lo , here is the chain memcombine sees after late opt :

  v436 = ZeroExt8to64 <uint64> v273 ; addr[15], forwarded
  v448 = Or64 <uint64> v436 v447    ; | addr[14] << 8, forwarded
  v460 = Or64 <uint64> v459 v448    ; | addr[13] << 16, forwarded
  v472 = Or64 <uint64> v471 v460    ; | addr[12] << 24, forwarded
  v484 = Or64 <uint64> v483 v472    ; | a16[11] << 32, loaded
  v496 = Or64 <uint64> v495 v484    ; | a16[10] << 40, loaded
  v508 = Or64 <uint64> v507 v496    ; | a16[9] << 48, loaded
  v520 = Or64 <uint64> v519 v508    ; | a16[8] << 56, loaded

From v520 , four of the eight terms are forwarded bytes, not loads from memory, and memcombine can’t combine them. It doesn’t merge the four remaining loads either, as they sit on top of the forwarded bytes.

Back in order #

In summary, the rewriting rules run too early to be effective. A quick workaround exists: run an earlier round of memcombine before late opt . After this change , the generated code for the helper is back to the shortest possible version:

// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
 CMPQ    net/netip·z4(SB), CX     ; check "z" if this is an IPv4 address
 JNE     end                      ; if not, stop here
 MOVQ    net/netip·z6noz(SB), CX  ; CX = netip.z6noz
end:
 RET
// return value = Addr{hi: AX, lo: BX, z: CX}

And the benchmark confirms it! ✌️

goos: linux
goarch: amd64
pkg: github.com/vincentbernat/go-netip-addrto6
cpu: AMD Ryzen 5 5600X 6-Core Processor
                   │   Go 1.26.8    │             Our branch              │
                   │     sec/op     │    sec/op     vs base               │
AddrTo6/safe          6.5470n ±  0%   0.8944n ± 4%  -86.34% (p=0.002 n=6)
AddrTo6/unsafe        0.9071n ±  3%   0.8682n ± 1%   -4.28% (p=0.002 n=6)
AddrTo6/builtin       0.8871n ±  2%   0.8785n ± 1%        ~ (p=0.310 n=6)

Next steps #

I think Go maintainers would reject this change because of the additional memcombine pass. Instead, I plan to publish this blog post and bring up the subject again as a follow-up to issue #54365 . Either the sheer complexity and the Go 1.27 regression convince the maintainers that adding a To6() method is simpler and more efficient, or they advise me on how to move forward. Either way, digging into this subject taught me a lot about the Go compiler! ⚙️

Update (2026-10)

I opened issue #81994 to propose Addr.To6() . Give it a 👍 if you want it in Go!

Rural Queenslanders have seen gas projects come and go – but a 725-hectare datacentre poses a whole new level of ‘stupidity’

Guardian
www.theguardian.com
2026-10-04 10:00:29
Residents of Dalby may soon live beside a $31bn datacentre. The premier sees ‘opportunity’ for the state – but others fear AI’s global creep From the veranda on his bush block, Denver Kanowski has watched resources companies come and go from rural Queensland for 70 years. The home cost him $800, and...
Original Article

From the veranda on his bush block, Denver Kanowski has watched resources companies come and go from rural Queensland for 70 years.

The home cost him $800, and lets him live how he likes, in nature. It’s also in the middle of a gasfield, not far from the Linc Energy contamination site , down the road from a coalmine and three gas power plants.

Last month came the biggest scheme yet, a $31bn datacentre to power Claude, the large language model developed by global AI giant Anthropic. If built, Australia’s largest datacentre would be about 15 minutes’ drive away from Kanowski’s home, outside Kogan, near Dalby, about 180km west of Brisbane.

In a community that has seen the boom and bust of coal seam gas, the response from many has been the same as Kanowski’s: once bitten, twice shy.

“This is the power of money and the power of greed and shortsighted stupidity,” Kanowski says.

“I’ve read this book before.”

Anthropic’s project is to be built on a 725-hectare plot of land now occupied by a 24,000 head feedlot, surrounded by gas wells.

According to Singapore-based developer Zerra DC, its cooling system will be far less water intensive than most existing datacentres’, generating about the same demand as an average shopping centre .

A sign in a field says ‘Farmland not gas land’
The site of what is planned to be Australia’s biggest datacentre is, today, a feedlot on a rural road 30km outside Dalby, surrounded by coal seam gas wells and three gas power plants. Photograph: Andrew Messenger/The Guardian

But like all AI, it will require an enormous amount of energy. The Western Downs datacentre will require a peak power capacity equal to about a quarter of total state energy demand , 2.16GW. It would also employ thousands during construction, with about 1,400 long-term jobs, according to the company .

The proposal was made public just last month.

Within days, Dalby travel agent Mordechai Winters had made up his mind: he didn’t want it.

“I was crushed. You know, it’s very disappointing that the council and also the government would consider it,” he says.

Issues like energy demand, noise and poor siting have driven a massive backlash against datacentres in the United States. Three-quarters of Americans now oppose them, and more than 500 counties have imposed a ban.

At least thousands have a similar view in Australia.

A petition that Winters helped to launch has been signed by more than 21,500 people – mostly from the area – in just a few weeks, he says.

He believes the project’s benefits are exaggerated and will trigger the same boom-bust cycle as gas.

Cecil Plains farmer Liza Balmain remembers the big – and eventually unrealised – promises of the gas industry when it arrived.

Liza Balmain stands in a field under a stormy sky
Liza Balmain fought for years to keep coal seam gas off her land in Cecil Plains near Dalby. She says she can see ‘huge similarities’ with the proposed datacentre. Photograph: Andrew Messenger/The Guardian

The industry brought massive change; academics later concluded that its scale and pace in a small agricultural community “ was likely unprecedented anywhere in the world ”. There was a huge influx of workers, with the high wages of gas attracting labour away from farms and local business.

Then – because setting up a gas well requires a lot more labour than operating it – the industry left, just as quickly as it had arrived, leaving business owners with debts they couldn’t pay on investments they no longer needed. Some academics believe the industry left the region worse off . Others think it was a slim net positive .

Balmain, who led a long campaign against the gas industry , says she can see “huge similarities” between the coal seam gas era and the proposed datacentre.

Australians across the country are concerned. Jo Shulman, CEO of the Environmental Defenders Office which has long helped communities fight resources projects such as coal seam gas, says the group had seen a 250% increase in queries about datacentres to the service’s free legal advice line since January.

“I would say the datacentre is the new frontier,” she says.

“Communities are concerned about [their] impacts on the environment and their wellbeing”.

Shulman says datacentres will form an essential part of the clean energy transition – but called for a pause to allow national standards to catch up.

skip past newsletter promotion

“We think that there should be a moratorium on any further datacentres, or they risk repeating the same things that we saw with the coal seam gas boom,” she says.

Big job for a small council

The Dalby mayor, Andrew Smith, says the ultimate authority over the region’s largest ever project should rest with the Western Downs regional council, population 35,452. Speaking to Guardian Australia last month, he said council did not expect the project to be taken over by the state coordinator general.

A wide regional town street with palm trees on either side
Cunningham Street, Dalby. The town’s mayor says coal seam gas ‘diversified our economy … It’s created jobs.’ Photograph: David Kelly/Photograph David Kelly/The Guardian

“We did ask the question, but our understanding is it will stay with local government,” he says.

But then, on Friday, things changed. The state announced it would take datacentres out of the hands of local councils, under a new planning framework to ensure “a consistent and transparent framework across the state,” similar to the government’s renewables planning policy .

“We will treat these projects no worse and no better than we do mining, renewables or tourism projects. Any deals done must be with the backing of councils and communities,” a spokesperson for planning minister, Jarrod Bleijie, says.

“The state government will become the responsible assessor and coordinator of any data centre applications, instead of local councils”.

If given the green light, the project will be worth more to the local community than coal seam gas, he says.

“In general, the coal seam gas industry has been wonderful for our region. It’s diversified our economy, which we talk about a lot. It’s created jobs,” he says.

“Prior to coal seam gas, we were exporting a lot of our clever kids, because the opportunities weren’t back in the region.”

Anthropic wants to start using the centre as early as next year. But Smith says the planning couldn’t be rushed – and was careful not to endorse or oppose the project.

The Queensland premier, David Crisafulli , lobbied the company to invest in the state during a recent trade mission, he told parliament last week.

Crisafulli said the state would help the council assess the project, but said it had done an excellent job with the gas industry and could do so again.

“They see the same opportunity I do, and that opportunity is to make sure that there are jobs, very good jobs in regional areas, and if we get it right, lower power bills for every other Queenslander as well . I see a really good play with that,” the premier said.

Toowoomba and Surat Basin Enterprise chief executive officer Jo Sheppard says most residents of the area had not yet made up their minds about the project, but that it was important for the developer to maintain their social license.

She added it was a “a once in a generation opportunity for Western Downs region and our broader region”, with the power to transform the town from a “ “fairly quiet” farming community to something more, adding to booms triggered by gas, and then renewables.

“Someone actually did say to me the other day: ‘Gosh, should we all be buying a house in Dalby?’”

Kanowski says there was something different about the datacentre compared with coal and coal seam gas schemes. He believes it will go beyond harming the environment by contributing to an online “fantasy” world, disconnecting people from the things that genuinely matter.

“[The datacentre developers] don’t even see the trees. They don’t understand the relationship between all the things. This is my world, I love it,” he says.

“This is not a datacentre … This is part of a network that these money-crazed idiots are planning to put all over the planet.”

Trump names intelligence chief Jay Clayton as new White House AI czar

Guardian
www.theguardian.com
2026-10-04 09:44:50
Calyton said AI was a ‘gamechanger’ but it also posed ‘a threat’ during his DNI Senate confirmation hearing Donald Trump on Sunday named Jay Clayton, the director of national intelligence, to serve also as the new White House AI czar. Clayton, who leads the US’s 18 intelligence agencies, previously...
Original Article

Donald Trump on Sunday named Jay Clayton, the director of national intelligence , to serve also as the new White House AI czar.

Clayton, who leads the US’s 18 intelligence agencies , previously served as the top federal prosecutor in Manhattan and led the Securities and Exchange Commission during Trump’s first term.

During his Senate confirmation hearing , Clayton said AI was a “gamechanger”, but it also posed “a threat”.

“When something’s both an opportunity and a threat, you better get your arms around it,” Clayton told lawmakers when asked about AI.

CBS News was the first to report that Clayton would be appointed White House AI czar.

In a social media post on Sunday, Trump announced that Clayton would lead what the US president called “the Super Intelligence Force (SIF)”.

Trump’s post also said the SIF would “coordinate the federal government’s engagement with consumers, public interest groups, religious organizations, critical infrastructure providers, and super intelligence companies”.

In September, Trump announced he was creating an “AI Force” and a czar to oversee the technology’s growth and monitor for “bad” development. The announcement came after tech industry leaders warned that AI was developing too rapidly to control and that it could harm humans.

The president has largely brushed aside concerns about AI development and the construction of datacenters, which power the technology.

Trump has insisted continued development is paramount to defeating China in the global AI race.

But to address concerns, the president created an advisory committee in addition to the czar – who he previously promised would be led by a “High IQ” individual. Trump also recently met with AI industry leaders to discuss concerns.

On Tuesday, Trump and six executives signed an AI accord at the White House, promising to create internal and external monitors to ensure the technology is used as intended and any safety issues are addressed.

Trump had initially created the AI and cryptocurrency czar position when he returned to the White House in January 2025.

David Sacks, the venture capitalist and tech businessman, served as the administration’s AI and cryptocurrency czar until March 2026, when he reached his employment limit under his designation as a “special government employee”.

Sacks still co-chairs the President’s Council of Advisors on Science and Technology.

Trump previously floated Clayton’s name as a potential AI czar, telling Axios in September that Clayton is ‘“a good man” and the prospect of appointing him was “a good idea”.

Vincent Bernat: Hacking the Go compiler to efficiently map IPv4 to IPv6

PlanetDebian
vincent.bernat.ch
2026-10-04 09:27:25
netip.Addr features an Unmap() method returning the unwrapped IPv4 contained in an IPv4-mapped IPv6 address: from ::ffff:203.0.113.10 or ::ffff:cb00:710a, it returns 203.0.113.10.1 There is no Map() or To6() method for the reverse direction. Such a method is trivial to implement, but Go maintainers ...
Original Article

netip.Addr features an Unmap() method returning the unwrapped IPv4 contained in an IPv4-mapped IPv6 address : from ::ffff:203.0.113.10 or ::ffff:cb00:710a , it returns 203.0.113.10 . 1 There is no Map() or To6() method for the reverse direction. Such a method is trivial to implement, but Go maintainers have rejected it on the grounds that users should write netip.AddrFrom16(ip.As16()) and let the compiler optimize it. 2 Today, this pattern is eight times slower than a native method. How can we teach the compiler to optimize this sequence?

The alternatives #

Let’s explore three ways to implement the map semantics for netip.Addr . My favorite is to add it to the Go standard library. Go maintainers prefer a small external helper chaining netip.AddrFrom16() and netip.Addr.As16() , hoping the compiler eventually optimizes it. The unsafe package opens a third path, with the same performance as the first solution.

Modifying the Go standard library #

Internally, netip.Addr stores any IP address as a 128-bit value with an extra field z to encode the family and the zone:

type Addr struct {
    addr uint128
    z unique.Handle[addrDetail]
}

type addrDetail struct {
    isV6   bool   // IPv4 is false, IPv6 is true.
    zoneV6 string // != "" only if IsV6 is true.
}

var (
    z0    unique.Handle[addrDetail]
    z4    = unique.Make(addrDetail{})
    z6noz = unique.Make(addrDetail{isV6: true})
)

AddrFrom4() encodes an IPv4 address as an IPv4-mapped IPv6 address and sets z to the unique value z4 :

// AddrFrom4 returns the address of the IPv4 address given by the bytes in addr.
func AddrFrom4(addr [4]byte) Addr {
    return Addr{
        addr: uint128{
            0,
            0xffff00000000 |
                uint64(addr[0])<<24 | uint64(addr[1])<<16 |
                uint64(addr[2])<<8 | uint64(addr[3])},
        z: z4,
    }
}

Unmap() turns an IPv4-mapped IPv6 address into an IPv4 address by setting the z field to z4 :

func (ip Addr) Unmap() Addr {
    if ip.Is4In6() {
        ip.z = z4
    }
    return ip
}

Implementing the reverse direction inside the Go standard library is trivial: we set the z field to z6noz if the address is IPv4.

// To6 maps an IPv4 address to an IPv4-mapped IPv6 address. It returns an
// IPv6 address unmodified.
func (ip Addr) To6() Addr {
    if ip.Is4() {
        ip.z = z6noz
    }
    return ip
}

As a helper #

We can’t access the z field from outside the net/netip package. Instead, we build a small helper around the netip.AddrFrom16(ip.As16()) pattern:

// AddrTo6 maps an IPv4 address to an IPv4-mapped IPv6 address. It returns an
// IPv6 address unmodified.
func AddrTo6(ip netip.Addr) netip.Addr {
    if ip.Is4() {
        ip = netip.AddrFrom16(ip.As16())
    }
    return ip
}

As an unsafe function #

Another solution uses the unsafe package to alter the Addr struct through a proxy with the same memory layout: 3

// addrProxy has the same memory layout as netip.Addr.
type addrProxy struct {
    addr [2]uint64      // netip.uint128
    z    unsafe.Pointer // unique.Handle[netip.addrDetail]
}

var (
    anyIPv6    = netip.IPv6Unspecified()
    netipZ6noz = (*addrProxy)(unsafe.Pointer(&anyIPv6)).z
)

// AddrTo6 maps an IPv4 address to an IPv4-mapped IPv6 address. It returns an
// IPv6 address unmodified.
func AddrTo6(ip netip.Addr) netip.Addr {
    if !ip.Is4() {
        return ip
    }
    (*addrProxy)(unsafe.Pointer(&ip)).z = netipZ6noz
    return ip
}

Benchmarks #

On my computer, with Go 1.27.1, the standard library solution costs 0.88 ns per operation, while the solution favored by Go maintainers costs 7.14 ns. The unsafe solution matches the performance of the first one.

goos: linux
goarch: amd64
pkg: github.com/vincentbernat/go-netip-addrto6
cpu: AMD Ryzen 5 5600X 6-Core Processor
                   │     sec/op     │
AddrTo6/safe            7.137n ± 0%
AddrTo6/unsafe         0.8682n ± 2%
AddrTo6/builtin        0.8775n ± 2%

Assembly code #

Let’s check the assembly code the compiler generates for each solution. 4 The one built into the standard library looks like this: 5

// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
 CMPQ  net/netip·z4(SB), CX     ; check "z" if this is an IPv4 address
 JNE   end                      ; if not, stop here
 MOVQ  net/netip·z6noz(SB), CX  ; CX = netip.z6noz
end:
 RET
// return value = Addr{hi: AX, lo: BX, z: CX}

Go’s assembly language is not a direct representation of the underlying machine language: it operates on a semi-abstract instruction set derived from Plan 9’s assembler . It has four pseudo-registers: FP (frame pointer for function arguments), PC (program counter), SB (static base pointer for global symbols), and SP (stack pointer). It also has architecture-specific registers like AX , CX , DX , BX , SI , DI , and R8 to R15 . Instructions storing data use their last argument as the destination. Instructions can carry an explicit size suffix: MOVB moves a byte, MOVW 16 bits, MOVL 32 bits, and MOVQ 64 bits. In the example above, the first instruction compares the 64-bit value z4 with the CX register.

The unsafe solution looks almost the same:

// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
 CMPQ  net/netip·z4(SB), CX  ; check "z" if this is an IPv4 address
 JNE   end                   ; if not, stop here
 MOVQ  netipZ6noz(SB), CX    ; CX = netip.z6noz
end:
 RET
// return value = Addr{hi: AX, lo: BX, z: CX}

The helper solution has far more instructions. To understand why, let’s look at the code for As16() and AddrFrom16() . They are short enough for the compiler to inline them.

func (ip Addr) As16() (a16 [16]byte) {
    byteorder.BEPutUint64(a16[:8], ip.addr.hi)
    byteorder.BEPutUint64(a16[8:], ip.addr.lo)
    return a16
}

func AddrFrom16(addr [16]byte) Addr {
    return Addr{
        addr: uint128{
            byteorder.BEUint64(addr[:8]),
            byteorder.BEUint64(addr[8:]),
        },
        z: z6noz,
    }
}

We can already guess the pattern to optimize: the code packs the IP address into an array, copies it, then unpacks it. If we inline the Go code by hand, we get:

func AddrTo6(input netip.Addr) netip.Addr {
    if !input.Is4() {
        return input
    }

    var a16 [16]byte
    byteorder.BEPutUint64(a16[:8], input.addr.hi)
    byteorder.BEPutUint64(a16[8:], input.addr.lo)

    addr := a16

    var output netip.Addr
    output.addr.hi = byteorder.BEUint64(addr[:8])
    output.addr.lo = byteorder.BEUint64(addr[8:])
    output.z = netip.z6noz
    return output
}

As humans, we can mentally derive the optimized form:

func AddrTo6(input netip.Addr) netip.Addr {
    if !input.Is4() {
        return input
    }
    var output netip.Addr
    output.addr.hi = input.addr.hi
    output.addr.lo = input.addr.lo
    output.z = netip.z6noz
    return output
}

Unfortunately, as of Go 1.26.8, the compiler is not smart enough to do the same:

// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
; Push the stack (32 bytes):
;    0(SP) addr netip.uint128
;   16(SP) a16 [16]byte
 PUSHQ   BP
 MOVQ    SP, BP
 SUBQ    $32, SP

 CMPQ    net/netip·z4(SB), CX  ; check "z" if this is an IPv4 address
 JNE     end                   ; if not, stop here

; Pack: byteorder.BEPutUint64(a16[:8], input.addr.hi)
;       byteorder.BEPutUint64(a16[8:], input.addr.lo)
 MOVBEQ  AX, net/netip·a16+16(SP)
 MOVBEQ  BX, net/netip·a16+24(SP)

; addr = a16, 16 bytes at once through the vector register X0
 MOVUPS  net/netip·a16+16(SP), X0
 MOVUPS  X0, net/netip·addr(SP)

; CX = netip.z6noz
 MOVQ    net/netip·z6noz(SB), CX
; Unpack: output.addr.hi = byteorder.BEUint64(addr[:8])
;         output.addr.lo = byteorder.BEUint64(addr[8:])
 MOVBEQ  net/netip·addr(SP), AX
 MOVBEQ  net/netip·addr+8(SP), BX

end:
 ADDQ    $32, SP
 POPQ    BP
 RET
// return value = Addr{hi: AX, lo: BX, z: CX}

The compiler does a decent job on the byte shuffling: the eight byte stores of BEPutUint64() become a single MOVBEQ , which stores a register byte-swapped. The eight byte loads of BEUint64() become a single MOVBEQ the other way round. 6 Three groups of instructions remain: a pack, a copy, and an unpack.

Hacking the Go compiler #

The Go compiler has several phases :

Parsing
The compiler tokenizes and parses the source code . It builds a syntax tree for each source file.
Type checking
The compiler maps each identifier to the object it denotes , folds constants, and infers the type of every expression.
IR construction
The compiler converts the syntax tree and its types into its own intermediate representation ( IR ). This process, called “ noding ,” goes through a serialization format named unified IR .
Middle end
The compiler performs several optimization passes on the IR , such as devirtualization , function call inlining , and escape analysis .
Walk
This phase runs two steps : order of evaluation decomposes complex statements into simpler ones, and desugaring transforms higher-level Go constructs, like switch or channels, into more primitive instructions or calls to the runtime.
Generic SSA
The compiler converts the IR into Static Single Assignment ( SSA ) form, a lower-level intermediate representation suited for machine-independent optimizations and rewrite rules .
Machine code generation
The compiler rewrites the SSA form into machine-specific variants , allocates registers, and applies more optimization passes. At the end, the assembler turns the generated instructions into machine code.

The hammer #

My first idea is to replace occurrences of netip.AddrFrom16(ip.As16()) with netip.Addr{addr: ip.addr, z: netip.z6noz} as early as possible, during the “noding” process. Before that, the type checking phase prevents us from accessing unexported struct fields.

Go 1.27 introduced a convenient debug option to dump the IR of a function at interesting points during compilation:

$ GOTOOLCHAIN=go1.27.1 GOAMD64=v3 go build -a -gcflags="-d=astdump=AddrTo6Safe" .
Writing text ast output for AddrTo6Safe to AddrTo6Safe.ast
Writing html ast output for AddrTo6Safe to AddrTo6Safe.html
Writing html syntax output for AddrTo6Safe to AddrTo6Safe.syntax.html

In the HTML file, the first column shows the IR as it comes out of noding:

DCLFUNC addrto6.AddrTo6Safe ABI:ABIInternal FUNC-func(netip.Addr) netip.Addr
DCLFUNC-Dcl
. NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. NAME-addrto6.~r0 Class:PPARAMOUT Offset:0 OnStack netip.Addr
DCLFUNC-body
. IF # ipv6_safe.go:11:2
. IF-Cond
. . CALLFUNC bool
. . CALLFUNC-Fun
. . . METHEXPR addrto6.Is4 FUNC-func(netip.Addr) bool
. . . . TYPE netip.Addr Class:PEXTERN Offset:0 type netip.Addr
. . CALLFUNC-Args
. . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. IF-Body
. . AS # ipv6_safe.go:12:6
. . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. . . CALLFUNC netip.Addr
. . . CALLFUNC-Fun
. . . . NAME-netip.AddrFrom16 Class:PFUNC Offset:0 Used FUNC-func([16]byte) netip.Addr
. . . CALLFUNC-Args
. . . . CALLFUNC ARRAY-[16]byte
. . . . CALLFUNC-Fun
. . . . . METHEXPR addrto6.As16 FUNC-func(netip.Addr) [16]byte
. . . . . . TYPE netip.Addr Class:PEXTERN Offset:0 type netip.Addr
. . . . CALLFUNC-Args
. . . . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. RETURN # ipv6_safe.go:14:2
. RETURN-Results
. . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr

In the body of the if statement, we spot the calls to the method netip.Addr.As16() and to the function netip.AddrFrom16() . Our goal is to patch them with a struct literal:

IF-Body
. AS # ipv6_safe.go:12:6
. . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. . STRUCTLIT netip.Addr
. . STRUCTLIT-List
. . . STRUCTKEY netip.addr
. . . . DOT netip.addr netip.uint128
. . . . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. . . STRUCTKEY netip.z
. . . . NAME-netip.z6noz Class:PEXTERN Offset:0 unique.Handle[net/netip.addrDetail]

In noder’s reader.go , the expr() method builds the IR tree for an expression. At the end of the exprCall case, we add a call to a rewriteAddrFrom16As16() function. It takes the current node and returns the struct literal on success, or nil if the rewrite is not possible. First, we check that we have the expected pattern: a call to the netip.AddrFrom16() function with a call to the netip.Addr.As16() method as its only argument:

func rewriteAddrFrom16As16(n ir.Node) ir.Node {
    call, ok := n.(*ir.CallExpr)
    if !ok || call.Op() != ir.OCALLFUNC ||
        len(call.Args) != 1 || len(call.Init()) != 0 ||
        !isNetipFunc(call.Fun, "AddrFrom16") {
        return nil
    }
    inner, ok := call.Args[0].(*ir.CallExpr)
    if !ok || inner.Op() != ir.OCALLFUNC ||
        len(inner.Args) != 1 || len(inner.Init()) != 0 ||
        !isNetipFunc(inner.Fun, "Addr.As16") {
        return nil
    }
    x := inner.Args[0]
    // [...]
}

Then, we fetch netip.z6noz :

z6noz, err := lookupVar(ir.StaticCalleeName(call.Fun).Sym().Pkg, "z6noz")
if err != nil {
    return nil
}

And we build the struct literal:

typ := call.Type()
pos := call.Pos()
var list []ir.Node
for i, f := range typ.Fields() {
    var value ir.Node
    switch f.Sym.Name {
    case "addr":
        value = typecheck.DotField(pos, x, i)
    case "z":
        value = z6noz
    default:
        return nil
    }
    list = append(list, ir.NewStructKeyExpr(pos, f, value))
}
lit := ir.NewCompLitExpr(pos, ir.OSTRUCTLIT, typ, list)
lit.SetTypecheck(1)
return lit

Have a look at the complete patch . 7 We can test it with the following commands:

$ cd src
$ ./make.bash
Building Go cmd/dist using /usr/lib/go-1.27. (go1.27.1 linux/amd64)
Building Go toolchain1 and bootstrap cmd/go (go_bootstrap) using /usr/lib/go-1.27.
Building Go toolchain2 using go_bootstrap and Go toolchain1.
Building Go toolchain3 and commands using go_bootstrap and Go toolchain2.
Checking command staleness for linux/amd64.
---
Installed Go for linux/amd64 in /home/bernat/code/free/go
Installed commands in /home/bernat/code/free/go/bin
*** You need to add /home/bernat/code/free/go/bin to your PATH.
$ export PATH=$PWD/../bin:$PATH
$ go version
go version go1.28-devel_9834516e20 Sat Sep 12 08:23:11 2026 -0700 linux/amd64
$ go test net/netip/...
ok      net/netip   0.224s
$ cd ../../go-netip-addrto6
$ go test .
ok      github.com/vincentbernat/go-netip-addrto6   0.062s

The generated code for the helper is now the shortest possible version!

// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
 CMPQ    net/netip·z4(SB), CX     ; check "z" if this is an IPv4 address
 JNE     end                      ; if not, stop here
 MOVQ    net/netip·z6noz(SB), CX  ; CX = netip.z6noz
end:
 RET
// return value = Addr{hi: AX, lo: BX, z: CX}

Go maintainers are unlikely to accept this change. It relies on the internal structure of net/netip.Addr . It’s an ugly hack in the noder, whose job is to faithfully translate the type-checked AST into the IR . And it’s harder to maintain than adding a To6() method.

The screwdriver #

The right place for such an optimization is the generic SSA phase . One of the last machine-independent passes is memcombine . With the appropriate debug flag, the compiler dumps the SSA form after this pass: 8

$ GOTOOLCHAIN=go1.26.8 GOAMD64=v3 \
> go build -a -gcflags='-d=ssa/memcombine/dump=AddrTo6Safe' .
$ head -5 AddrTo6Safe_01__memcombine.dump
AddrTo6Safe func(netip.Addr) netip.Addr
  b2:
    (?) v1 = InitMem <mem>
    (?) v2 = SP <uintptr>
    (?) v3 = SB <uintptr>

The result of the memcombine pass follows the same structure as the assembly code for AddrTo6Safe() we looked at earlier : two stores, one move, and two loads we would like to optimize away.

; […]
  v502 = ArgIntReg <uint64> {ip+0} [0]              ; input.addr.hi
  v490 = ArgIntReg <uint64> {ip+8} [1]              ; input.addr.lo
  v466 = ArgIntReg <*netip.addrDetail> {ip+16} [2]  ; input.z
; […]
  v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
  v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
  v442 = Bswap64 <uint64> v490                      ; bswap(input.addr.lo)
  v542 = Bswap64 <uint64> v502                      ; bswap(input.addr.hi)
  v161 = Store <mem> {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)
  v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)

  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16

  v415 = OffPtr <*byte> [8] v285                    ; &addr[8]
  v416 = Load <uint64> v285 v286                    ; addr[:8]
  v299 = Bswap64 <uint64> v416                      ; output.addr.hi
  v174 = Load <uint64> v415 v286                    ; addr[8:]
  v39  = Bswap64 <uint64> v174                      ; output.addr.lo
; […]

Each line features a value identifier ( v442 ), an operation with its type ( Bswap64 <uint64> ), and its arguments ( v490 ). 9 Values are the basic building blocks of SSA and are defined exactly once. Square brackets enclose integer parameters ( [8] ) and curly braces contain auxiliary arguments ( {netip.addr} ). Operations writing to memory produce a new memory state. Every memory operation takes the current state as its last argument, which keeps them in order.

On paper #

Let’s focus on output.addr.lo , aka v39 :

  v490 = ArgIntReg <uint64> {ip+8} [1]              ; input.addr.lo
  v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
  v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
  v442 = Bswap64 <uint64> v490                      ; bswap(input.addr.lo)
  v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)
  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
  v415 = OffPtr <*byte> [8] v285                    ; &addr[8]
  v174 = Load <uint64> v415 v286                    ; addr[8:]
  v39  = Bswap64 <uint64> v174                      ; output.addr.lo

To simplify this code, we could apply three rewriting rules:

  1. The first one adds a shortcut when loading through a move : (Load (OffPtr [o] p) (Move p src mem)) => (Load (OffPtr [o] src) mem) . This matches v174 with its arguments v415 and v286 and creates a new value v600 :

      v490 = ArgIntReg <uint64> {ip+8} [1]              ; input.addr.lo
      v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
      v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v442 = Bswap64 <uint64> v490                      ; bswap(input.addr.lo)
      v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)
      v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
      v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
      v415 = OffPtr <*byte> [8] v285                    ; &addr[8]
      v600 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v174 = Load <uint64> v600 v282                    ; a16[8:]
      v39  = Bswap64 <uint64> v174                      ; output.addr.lo
    
  2. The second one simplifies a load following a store : (Load p (Store p x _)) => x . The load is forwarded : the stored value replaces it and no memory access remains. It matches v174 . It notices that v600 and v173 are the same address and replaces v174 with a copy of v442 :

      v490 = ArgIntReg <uint64> {ip+8} [1]              ; input.addr.lo
      v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
      v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v442 = Bswap64 <uint64> v490                      ; bswap(input.addr.lo)
      v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)
      v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
      v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
      v415 = OffPtr <*byte> [8] v285                    ; &addr[8]
      v600 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v174 = Copy <uint64> v442                         ; bswap(input.addr.lo)
      v39  = Bswap64 <uint64> v174                      ; output.addr.lo
    
  3. The last step cancels the two byte swaps : (Bswap64 (Bswap64 x)) => x . v39 becomes a copy of v490 :

      v490 = ArgIntReg <uint64> {ip+8} [1]              ; input.addr.lo
      v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
      v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v442 = Bswap64 <uint64> v490                      ; bswap(input.addr.lo)
      v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)
      v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
      v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
      v415 = OffPtr <*byte> [8] v285                    ; &addr[8]
      v600 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v174 = Copy <uint64> v442                         ; bswap(input.addr.lo)
      v39  = Copy <uint64> v490                         ; output.addr.lo = input.addr.lo
    

If we ignore the values not needed to compute v39 , only this SSA form remains:

  v490 = ArgIntReg <uint64> {ip+8} [1]  ; input.addr.lo
  v39  = Copy <uint64> v490             ; output.addr.lo = input.addr.lo

Let’s switch to output.addr.hi , aka v299 :

  v502 = ArgIntReg <uint64> {ip+0} [0]              ; input.addr.hi
  v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
  v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
  v542 = Bswap64 <uint64> v502                      ; bswap(input.addr.hi)
  v161 = Store <mem> {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)
  v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)
  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
  v416 = Load <uint64> v285 v286                    ; addr[:8]
  v299 = Bswap64 <uint64> v416                      ; output.addr.hi

To optimize it away, we also apply three rewriting rules:

  1. The first one also adds a shortcut when loading through a move , but without an offset: (Load p (Move p src mem)) => (Load src mem) . This rewrites v416 to use arguments from v286 :

      v502 = ArgIntReg <uint64> {ip+0} [0]              ; input.addr.hi
      v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
      v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v542 = Bswap64 <uint64> v502                      ; bswap(input.addr.hi)
      v161 = Store <mem> {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)
      v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)
      v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
      v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
      v416 = Load <uint64> v22 v282                     ; a16[:8]
      v299 = Bswap64 <uint64> v416                      ; output.addr.hi
    
  2. The second rule forwards a value stored one step earlier , skipping over a store to another address: (Load p (Store q _ (Store p x _))) => x . This matches v416 : x is v542 , p is v22 ( &a16 ), q is v173 ( &a16[8] ), and p and q do not overlap for uint64 .

      v502 = ArgIntReg <uint64> {ip+0} [0]              ; input.addr.hi
      v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
      v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v542 = Bswap64 <uint64> v502                      ; bswap(input.addr.hi)
      v161 = Store <mem> {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)
      v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)
      v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
      v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
      v416 = Copy <uint64> v542                         ; bswap(input.addr.hi)
      v299 = Bswap64 <uint64> v416                      ; output.addr.hi
    
  3. The third rule cancels two byte swaps : (Bswap64 (Bswap64 x)) => x . v299 becomes a copy of v502 :

      v502 = ArgIntReg <uint64> {ip+0} [0]              ; input.addr.hi
      v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16
      v173 = OffPtr <*byte> [8] v22                     ; &a16[8]
      v542 = Bswap64 <uint64> v502                      ; bswap(input.addr.hi)
      v161 = Store <mem> {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)
      v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)
      v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
      v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
      v416 = Copy <uint64> v542                         ; bswap(input.addr.hi)
      v299 = Copy <uint64> v502                         ; output.addr.hi = input.addr.hi
    

If we remove the values not used to compute v299 , we get this SSA form:

  v502 = ArgIntReg <uint64> {ip+0} [0] ; input.addr.hi
  v299 = Copy <uint64> v502            ; output.addr.hi = input.addr.hi

In practice #

Most of these rules already exist in generic.rules . They use conditions to validate their context: ssa.IsSamePtr() for the same address, ssa.Disjoint() for addresses that do not overlap. The rule forwarding a stored value to a load already exists with three variants looking through several other stores. Here are the two we need:

(Load <t1> p1 (Store {t2} p2 x _))
    && ssa.IsSamePtr(p1, p2)
    && copyCompatibleType(t1, x.Type)
    && t1.Size() == t2.Size()
    => x
(Load <t1> p1 (Store {t2} p2 _ (Store {t3} p3 x _)))
    && ssa.IsSamePtr(p1, p3)
    && copyCompatibleType(t1, x.Type)
    && t1.Size() == t3.Size()
    && ssa.Disjoint(p3, t3, p2, t2)
    => x

Go 1.27 added the rule loading through a move with CL 748200 to fix issue #77720 :

(Load <t1> op1:(OffPtr [o1] p1) move:(Move [n] p2 src mem))
    && o1 >= 0 && o1+t1.Size() <= n && ssa.IsSamePtr(p1, p2)
    && !ssa.IsVolatile(src)
    => @move.Block (Load <t1> (OffPtr <op1.Type> [o1] src) mem)

It lacks a variant without an offset:

(Load <t1> p1 move:(Move [n] p2 src mem))
    && p1.Op != ssaop.OpOffPtr
    && t1.Size() <= n && ssa.IsSamePtr(p1, p2)
    && !ssa.IsVolatile(src)
    => @move.Block (Load <t1> (OffPtr <p1.Type> [0] src) mem)

There is no generic rule to cancel two byte swaps, but the AMD64 lowering pass includes this rule:

(BSWAP(Q|L) (BSWAP(Q|L) p)) => p

After switching to Go’s development branch and adding the missing rule , the generated assembly code is worse than with Go 1.26.8, even though our additional rule slightly improves the situation at the end:

// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
; Push the stack (16 bytes):
;    0(SP) a16 [16]byte
 PUSHQ   BP
 MOVQ    SP, BP
 SUBQ    $16, SP

 CMPQ    net/netip·z4(SB), CX       ; check "z" if this is an IPv4 address
 JNE     end                        ; if not, stop here

; The four forwarded bytes: the low half of input.addr.lo is taken apart and
; put back together in registers
 MOVQ    BX, DX                     ; DX = input.addr.lo
 SHRQ    $24, BX                    ; BX = input.addr.lo >> 24
 MOVQ    DX, SI                     ; SI = input.addr.lo
 SHRQ    $16, DX                    ; DX = input.addr.lo >> 16
 MOVQ    SI, DI                     ; DI = input.addr.lo, kept for the pack
 SHRQ    $8, SI                     ; SI = input.addr.lo >> 8
 MOVBLZX DIB, R8                    ; R8 = byte(input.addr.lo)
 MOVBLZX SIB, SI                    ; SI = byte(input.addr.lo >> 8)
 SHLQ    $8, SI
 ORQ     R8, SI                     ; SI = two low bytes of input.addr.lo
 MOVBLZX DL, DX                     ; DX = byte(input.addr.lo >> 16)
 SHLQ    $16, DX
 ORQ     SI, DX
 MOVBLZX BL, BX                     ; BX = byte(input.addr.lo >> 24)
 SHLQ    $24, BX
 ORQ     DX, BX                     ; BX = input.addr.lo & 0xffffffff

; Pack: byteorder.BEPutUint64(a16[:8], input.addr.hi)
;       byteorder.BEPutUint64(a16[8:], input.addr.lo)
 MOVBEQ  AX, net/netip·a16(SP)
 MOVBEQ  DI, net/netip·a16+8(SP)

; The four other bytes of input.addr.lo, read one by one from a16
 MOVBLZX net/netip·a16+11(SP), DX   ; a16[11]
 SHLQ    $32, DX
 ORQ     DX, BX
 MOVBLZX net/netip·a16+10(SP), DX   ; a16[10]
 SHLQ    $40, DX
 ORQ     DX, BX
 MOVBLZX net/netip·a16+9(SP), DX    ; a16[9]
 SHLQ    $48, DX
 ORQ     DX, BX
 MOVBLZX net/netip·a16+8(SP), DX    ; a16[8]
 SHLQ    $56, DX

; output.z = netip.z6noz
 MOVQ    net/netip·z6noz(SB), CX
; output.addr.hi = byteorder.BEUint64(a16[:8])
 MOVBEQ  net/netip·a16(SP), AX
; output.addr.lo assembled from the previous steps
 ORQ     DX, BX

end:
 LEAVEQ
 RET
// return value = Addr{hi: AX, lo: BX, z: CX}

The rule loading through a move, added in Go 1.27, introduced this regression.

Out of order #

Let’s not give up now! In reality, the rewriting rules run before memcombine , notably in the late opt pass. At this point, the inlined versions of BEPutUint64() and BEUint64() still expand to sixteen byte stores and sixteen byte loads, matching their source code :

func BEUint64(b []byte) uint64 {
    _ = b[7] // bounds check hint to compiler; see golang.org/issue/14808
    return uint64(b[7]) | uint64(b[6])<<8 | uint64(b[5])<<16 | uint64(b[4])<<24 |
        uint64(b[3])<<32 | uint64(b[2])<<40 | uint64(b[1])<<48 | uint64(b[0])<<56
}

Let’s follow two bytes of output.addr.lo : addr[15] and addr[11] . Here is a simplified SSA form before late opt :

  v273 = Trunc64to8 <byte> v490                     ; byte(input.addr.lo)
  v226 = Trunc64to8 <byte> v225                     ; byte(input.addr.lo >> 32)
; […]
  v235 = Store <mem> {byte} v233 v226 v223          ; a16[11] = byte(lo >> 32)
  v247 = Store <mem> {byte} v245 v238 v235          ; a16[12] = …
  v259 = Store <mem> {byte} v257 v250 v247          ; a16[13] = …
  v271 = Store <mem> {byte} v269 v262 v259          ; a16[14] = …
  v282 = Store <mem> {byte} v280 v273 v271          ; a16[15] = byte(lo)
  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16
; […]
  v433 = OffPtr <*byte> [15] v285                   ; &addr[15]
  v435 = Load <byte> v433 v286                      ; addr[15]
  v479 = OffPtr <*byte> [11] v285                   ; &addr[11]
  v481 = Load <byte> v479 v286                      ; addr[11]

The first rule loads through the move: (Load (OffPtr [o] p) (Move p src mem)) => (Load (OffPtr [o] src) mem) . It matches both loads, which now read a16 with the memory state before the copy:

  v600 = OffPtr <*byte> [15] v22 ; &a16[15]
  v435 = Load <byte> v600 v282   ; a16[15]
  v601 = OffPtr <*byte> [11] v22 ; &a16[11]
  v481 = Load <byte> v601 v282   ; a16[11]

The second rule shortcuts a load following a store: (Load p (Store p x _)) => x . It matches v435 , as v282 stores a16[15] . It does not match v481 : v235 stores a16[11] four stores earlier in the chain, while the variants of this rule look through three stores at most.

  v435 = Copy <byte> v273        ; byte(input.addr.lo)
  v601 = OffPtr <*byte> [11] v22 ; &a16[11]
  v481 = Load <byte> v601 v282   ; a16[11]

The same happens to the other bytes: the rule forwards the four bytes stored last, a16[12] to a16[15] . The twelve other loads now read a16 instead of addr .

BEUint64() becomes a chain of Or64 , each one adding a byte shifted into place. memcombine is a pass written in Go , not a set of rewrite rules. It starts from the last Or64 of the chain and collects up to eight terms. If each term is a byte load, extended to 64 bits and shifted, and if the eight loads read consecutive addresses from the same pointer with the same memory state, it replaces the whole chain with a single 64-bit load and a byte swap. Otherwise, it tries again with four, then two terms, and from each intermediate Or64 . Here is the loop checking each term in a simplified version of combineLoads() :

for i := int64(0); i < n; i++ {
    v := a[i]
    shift := int64(0)
    if v.Op == shiftOp {
        v, shift = peelShift(v)
    }
    if v.Op != extOp {
        return false
    }
    load := v.Args[0]
    if load.Op != ssaop.OpLoad {
        return false
    }
    if load.Args[1] != mem {
        return false
    }
    p, off := splitPtr(load.Args[0])
    if p != base {
        return false
    }
    r[i] = LoadRecord{load: load, offset: off, shift: shift}
}

For output.addr.hi , the eight loads read a16 with the same memory state v282 :

  v13  = Load <byte> v22 v282               ; a16[0]
  v530 = Load <byte> v14 v282               ; a16[1]
  v488 = Load <byte> v504 v282              ; a16[2]
  v405 = Load <byte> v537 v282              ; a16[3]
  v385 = Load <byte> v397 v282              ; a16[4]
  v361 = Load <byte> v373 v282              ; a16[5]
  v196 = Load <byte> v63 v282               ; a16[6]
  v432 = Load <byte> v315 v282              ; a16[7]
  v319 = ZeroExt8to64 <uint64> v432         ; uint64(a16[7])
  v329 = ZeroExt8to64 <uint64> v196         ; uint64(a16[6])
  v330 = Lsh64x64 <uint64> [true] v329 v138 ; uint64(a16[6]) << 8
  v331 = Or64 <uint64> v319 v330            ; a16[7] | a16[6] << 8
; […] same for a16[5] to a16[1]
  v401 = ZeroExt8to64 <uint64> v13          ; uint64(a16[0])
  v402 = Lsh64x64 <uint64> [true] v401 v55  ; uint64(a16[0]) << 56
  v403 = Or64 <uint64> v402 v391            ; | a16[0] << 56 = output.addr.hi

memcombine merges them into one load and a swap:

  v286 = Load <uint64> v22 v282   ; a16[:8]
  v285 = Bswap64 <uint64> v286    ; output.addr.hi

For output.addr.lo , here is the chain memcombine sees after late opt :

  v436 = ZeroExt8to64 <uint64> v273 ; addr[15], forwarded
  v448 = Or64 <uint64> v436 v447    ; | addr[14] << 8, forwarded
  v460 = Or64 <uint64> v459 v448    ; | addr[13] << 16, forwarded
  v472 = Or64 <uint64> v471 v460    ; | addr[12] << 24, forwarded
  v484 = Or64 <uint64> v483 v472    ; | a16[11] << 32, loaded
  v496 = Or64 <uint64> v495 v484    ; | a16[10] << 40, loaded
  v508 = Or64 <uint64> v507 v496    ; | a16[9] << 48, loaded
  v520 = Or64 <uint64> v519 v508    ; | a16[8] << 56, loaded

From v520 , four of the eight terms are forwarded bytes, not loads from memory, and memcombine can’t combine them. It doesn’t merge the four remaining loads either, as they sit on top of the forwarded bytes.

Back in order #

In summary, the rewriting rules run too early to be effective. A quick workaround exists: run an earlier round of memcombine before late opt . After this change , the generated code for the helper is back to the shortest possible version:

// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
 CMPQ    net/netip·z4(SB), CX     ; check "z" if this is an IPv4 address
 JNE     end                      ; if not, stop here
 MOVQ    net/netip·z6noz(SB), CX  ; CX = netip.z6noz
end:
 RET
// return value = Addr{hi: AX, lo: BX, z: CX}

And the benchmark confirms it! ✌️

goos: linux
goarch: amd64
pkg: github.com/vincentbernat/go-netip-addrto6
cpu: AMD Ryzen 5 5600X 6-Core Processor
                   │   Go 1.26.8    │             Our branch              │
                   │     sec/op     │    sec/op     vs base               │
AddrTo6/safe          6.5470n ±  0%   0.8944n ± 4%  -86.34% (p=0.002 n=6)
AddrTo6/unsafe        0.9071n ±  3%   0.8682n ± 1%   -4.28% (p=0.002 n=6)
AddrTo6/builtin       0.8871n ±  2%   0.8785n ± 1%        ~ (p=0.310 n=6)

Next steps #

I think Go maintainers would reject this change because of the additional memcombine pass. Instead, I plan to publish this blog post and bring up the subject again as a follow-up to issue #54365 . Either the sheer complexity and the Go 1.27 regression convince the maintainers that adding a To6() method is simpler and more efficient, or they advise me on how to move forward. Either way, digging into this subject taught me a lot about the Go compiler! ⚙️

Update (2026-10)

I opened issue #81994 to propose Addr.To6() . Give it a 👍 if you want it in Go!

The AI industry is booming. Women are getting left behind

Guardian
www.theguardian.com
2026-10-04 09:00:29
Women hold just a fraction of new AI jobs but are overrepresented in roles with high risk of AI disruption There’s a common fear among people who work in Silicon Valley: snag one of the fast-growing, high-paying jobs in artificial intelligence, or get trapped in the “permanent underclass”, a phrase ...
Original Article

T here’s a common fear among people who work in Silicon Valley: snag one of the fast-growing, high-paying jobs in artificial intelligence, or get trapped in the “permanent underclass”, a phrase describing the fate of those who won’t have upward mobility in the age of AI.

The phrase usually describes a future in which AI automates human jobs. But for women in tech, it’s starting to look like the present.

“A lot of people say: ‘Oh, AI is lowering barriers, it’s equalizing the playing field for women,’” said Urvashi Batra, co-founder and CEO of Prioriwise, an AI platform for IT service providers. “I actually think it’s the opposite.”

Batra said people take her less seriously as the founder of an AI company than her male co-founder. When pitching investors, she and her co-founder have learned that they are more likely to get an investment if he does the pitching.

Women made up only about a quarter of new hires in AI roles in the last year, compared to 50% of new hires in non-AI roles, according to a recent report from LinkedIn. In executive roles, that number drops to just 13%. At the same time, LinkedIn’s research shows that women are more likely to work in roles with high exposure to AI disruption, like customer service, meaning they are also at higher risk for job loss due to AI.

AI jobs, whose postings have doubled since 2023, come with a salary premium, on average paying more than twice as much as roles that don’t involve working with AI, according to LinkedIn’s research. But women who do have jobs in AI are disproportionately concentrated in low-paying roles, like data annotators, said Sarah Steinberg, the head of global public policy partnerships at LinkedIn. Across all AI occupations, men have $45,000 higher median pay than women, partly due to the types of jobs they are likelier to have.

“AI is creating some of the fastest growing, highest paying and most consequential jobs in the global economy,” said Steinberg. “Women are just strikingly underrepresented.”

The ongoing trend could produce the greatest gender pay gap in generations and leave women out of building a technology that shaped the economy, advocates for women in tech said.

“What we’re witnessing is a sort of backsliding,” said Brenda Darden Wilkerson, president of AnitaB.org, a non-profit devoted to advancing women in tech jobs.

Initiatives to promote women in the tech industry have been ongoing for years, with mixed results. Women now hold about one-third of tech jobs in the US. But research has found that women leave the tech industry at a much higher rate than their male counterparts, citing factors like non-inclusive cultures and barriers to advancement.

‘I’ve seen so many women fall through the cracks’

Women working in AI may face even greater challenges. “AI moves extremely fast, and the pressure to keep up is immense,” said Jayeeta Putatunda, an AI engineering lead at investment firm Turing. “I have seen throughout my career so many amazing women I know kind of fall through the cracks.”

Putatunda said it’s not uncommon to be the only female engineer on a team, and that when women are outnumbered, they have to speak up louder to make their ideas known. Mentorship in the AI field, especially by other women, can be hard to come by. And the breakneck pace of AI advancements means anyone working in the field should expect to put in long hours in order to keep up.

“If you have very ambitious career goals in this field, there will be seasons when you work beyond normal hours,” she said, noting that a culture of working 12-hour days to keep up was not uncommon. “I don’t think there is always a neat shortcut around that than to put in the hours.”

Those kinds of work expectations can make it impossible for everyone to stay in the field. “I went on maternity leave for four months , and by the time I came back, there were completely different frameworks and levels of models,” Putatunda said. Getting caught back up was overwhelming. She said it was only possible with the help of supportive co-workers and a husband who equitably split childcare.

“That is one way women get pushed out of AI,” she said. “They don’t lack the interest or ability to keep up. They may lack the infrastructure that makes keeping up possible.”

The AI industry is so new and fast-moving that it should in theory be more meritocratic. No one has a decade of experience working with a model developed one year ago, which should mean everyone has a level playing field. But instead the field has indexed more on personal networks and other loose signals, according to some women who work in AI.

“Companies are hiring at a breakneck speed, but they’re finding people through the same networks, the same referrals, the same filters that they’ve always used,” Wilkerson said. “Generally, that’s not given women the same sort of exposure they should have.”

When people have tried to address these tech industry problems in the past, the solutions have largely come down to giving women more mentorship and implementing diversity programs with goals around representation. But in recent years, companies have aggressively scaled back their diversity, equity and inclusion (DEI) programs, largely in response to a political climate that discourages them. That has left fewer companies with institutional programs to hire and promote women.

The political climate has made it more challenging to get AI companies to support initiatives focused on women, said Felicia Newhouse, founder of AI Powered Women, an organization that promotes women’s participation and leadership in AI. Newhouse, a product marketer who has worked in technology companies for two decades, points to the Department of Justice targeting corporations with DEI programs, claiming the initiatives violate anti-discrimination laws. Companies like Accenture, Deloitte, IBM and PayPal have all agreed to pay multimillion-dollar settlements under this enforcement.

“It is getting harder to go into companies being called AI Powered Women,” Newhouse said. “And our advocates inside those companies are also struggling with having women-focused initiatives.”

Still, she said the risk that women are left behind in the AI workforce was something that kept her up at night. “We’re talking about who captures a major new source of economic mobility,” she said. “The deepest risk is that a participation gap becomes a power gap.”

Show HN: Build with Python – a beginner course where your code draws

Hacker News
scimigo.com
2026-10-04 14:59:25
Comments...
Original Article

Make a robot, repeat a pattern, and respond to a click

Your first Python programs make pictures on a canvas. Learn calls, values, variables, and a loop before adding one click handler.

You learn
function calls · variables · loops · click events

You build
A robot, a row of circles, and click-to-draw painting

Loading the lab…

Was this lesson useful? Your feedback helps us improve the course.

Google Japan shows off conveyor-belt keyboard with keys that move to fingers

Hacker News
www.tomshardware.com
2026-10-04 14:36:33
Comments...
Original Article

Google Japan just showed off a quirky keyboard that has keys rolling on a conveyor belt to make typing “easier.” According to the company blog [machine translated], it built the Gboard Conveyor Belt Version so that the keys flow to your hand instead. That means you don’t have to move your fingers or arms as much, especially when you’re typing using a single hand. The keyboard comes with four belts, each containing 29 keys, making it quite “intuitive” to operate. The team behind it is also considering releasing various color variations, high-speed models, and even a shorter version for mobile devices. You can see the keyboard in action in the video below.

Get Tom's Hardware's best news and in-depth reviews, straight to your inbox.

Building a RAG Pipeline for Semantic Code Search

Hacker News
blog.jetbrains.com
2026-10-04 13:51:48
Comments...
Original Article
Ai logo

Supercharge your tools with AI-powered features inside many JetBrains products

Agentic AI AI

Building a RAG Pipeline for Semantic Code Search: A Developer Diary and Field Notes

Adam Malek Ashot Kazaryan

Part 1: Parsing, chunking, and vectorization

Some time ago, we set out to build the best semantic code search platform we could: a RAG pipeline that gives LLM agents precise, citable evidence from real repositories instead of whatever grep happens to surface. The eventual solution was Air Context . We got it working, we got it into production, and we collected a lot of scar tissue along the way. In this series of posts, we’ll share the parts we wish someone had told us on day one.


Coding agents are undoubtedly the biggest technology leap for software development of our decade. Agents and frontier models are proving their aptitude in the face of seemingly insurmountable code complexity to produce ostensibly reliable code.

However, as more and more development processes become agent-driven, the agent’s efficiency and the quality of the produced code become increasingly important. The question is not so much about whether an agent can complete the task, as given enough time and token resources, it surely will, but rather how much time, effort, and steering is required for it to generate production-grade results. For large-scale code bases specifically, the agent would spend a great deal of time searching for the relevant pieces of code relevant for the feature it’s working on and pulling them into the context.

Why semantic search matters

Attempting to locate the right code snippets, the agent will resort to traditional tools for code search such as keyword search and grep. These tools, however, are limited in that they require the agent to know in advance which exact text to search for. For example, an agent looking for where session tokens get refreshed cannot rely on the code helpfully containing the word “refresh”. To reason through abstract domains, the agent needs the ability to search for code by meaning, also known as semantic search. This is where retrieval-augmented generation (RAG) comes into the picture. If we can index the source code in a way that captures its semantics and then allow the agent to retrieve the relevant pieces on demand using free text search, we create an interface that plays to the agent’s strengths.

From prototype to production

Like many great ideas in the agentic era, a native, prototype implementation is extremely simple. A well-evaluated production grade solution most certainly is not. In this series of blog posts, we want to share what is involved in making an effective RAG system, as well as the wrong turns we took in our journey to create our own: Air Context. We’ll tackle each stage, from pre-processing to storage and agent integration, providing some more technical context and advice.

This first part of the series will cover the initial stages of the pipeline: parsing and chunking, where raw source files are divided into properly scoped units, and vectorization, where those units are transformed into a representation that supports semantic search.

The fine AST of parsing and chunking


Parsing and chunking is a critical pre-processing step in a good RAG solution, but it is often overlooked. In order to allow the LLM to embed or otherwise index the source code, we must first feed it the raw lines of code. This may sound trivial, and probably would be for small-scale demo projects. However, production-grade systems contain thousands of files, which, in turn, span hundreds or even thousands of lines. If anything, agents have compounded the problem, as they tend to be prolific writers, further inflating the codebase. Each file may contain multitudes of classes, fields, and methods, with varying degrees of relatedness among them.

Finding the right chunk size

Even if it were possible to fit these huge code files into an embedding model in their entirety, that expensive feat would ultimately be self-defeating. Because the entire file was embedded in a single unit, the search would return the entire file. This is counterproductive to the goals of agentic code exploration and navigation, which are mostly concerned with finding a specific function, symbol, or code snippet.

On the other hand, if we were to take the other extreme and granularly embed each separate line of code, we would be facing a problem of a different sort. These individual lines can be semantically insignificant without the surrounding context. A generic function name or comment does not merit embedding and will produce the wrong retrieval result. In a sense, we would not be able to see the forest for the trees, and the agent would be overloaded with multiple, often insignificant micro-results.

It is therefore imperative to find the right method to chunk or divide the code into groups that are properly scoped. Each group should include enough of the necessary context and represent common semantic meaning.

Why fixed-size chunking falls short

Chunking is a generic name for the technique of taking content that will be fed to the agent and dividing it into a set of chunks. A naive approach to chunking could be simply splitting a large file into groups with a fixed number of lines. However, if we were to take that approach, we would find the resulting groupings semantically wrong. Unrelated code pieces would be grouped together, for example, an import statement and some function content, leading to mistakes during retrieval.

To solve the problem, we can leverage the fact that every source file has a pretty well-defined structure. Take Java as an example – imports tend to be at the top of the file, followed by a class definition with an optional doc-comment preceding the header. The class will contain fields and methods, which in turn may also have their own doc-comments. Knowing about the conventions and rules that define the class structure allows us to perform smarter chunking and achieve the right balance of surrounding information.

Parsing and structure-aware chunking

Over the last 26 years, we at JetBrains have developed parsers that are smart enough to adjust for the various quirks, irregularities, conventions, and nuances of specific languages. Alongside other tools, these parsers form our internal JetBrains Code Engine platform on which Air Context is developed. At the moment of this article’s composition, Air Context supports parsing and structure-aware chunking for nine major languages: Kotlin, Java, Python, JavaScript, TypeScript, C#, PHP, Go, and Rust. For all other languages, our implementation simply falls back to naive, line-based splitting to ensure that any language or document can be indexed and searched.

The parser allows us to break source files into streams of syntax nodes that carry information about what they represent – comments, whitespaces, lists of modifiers, and so on. The chunking algorithm then consumes that stream and applies logic that decides the scope of a given chunk. Based on the node’s type and size, as well as its descendants, the algorithm makes a decision. If a node exceeds the size threshold but has no children, it will fall back to more primitive splitting strategies.

Some language-specific constructs are kept as single slices even if they exceed the preferred size. Prefixes such as documentation, annotations, visibility modifiers, and keywords are kept together with the declaration; suffixes (usually closing syntax) remain associated with the construct they close. There is also some language-specific cleaning, where, for instance, common and semantically meaningless Java annotations such as @NotNull or @Override are removed.

The algorithm bears some similarities to cAST , authored by Zhang et al. in 2025. Both our implementation and cAST retain the largest syntax units that fit, subdividing only the units that are too large, and grouping smaller adjacent units to avoid tiny chunks that are not usually semantically meaningful. The biggest difference is that we coded more language semantics into our implementation, keeping Python decorators  together with definitions, KDocs next to Kotlin declarations, and so on.

After grouping, chunk normalization is performed, which involves:

  • Trimming leading and trailing whitespaces
  • Deleting blank lines
  • Removing common indentation while preserving relative indentation

Following the normalization procedure, the chunk is then passed to the next step – embedding – along with metadata that consists of a relative path, which gets embedded alongside the normalized chunk content.

Evaluating the quality of chunks

It is hard to give a concrete answer as to what the input to the embedding model should look like. Chunk size matters, but as discussed before, bigger is not always better. Additionally, some metadata embedded alongside the code may be useful, while some may introduce noise that ultimately decreases search quality.

We opted to use an LLM-as-a-judge strategy to inspect the chunks as a part of the evaluation. The judge, using a chunk and the source file, considers whether the boundary makes sense. It looks for unexpected artifacts, such as detached documentation, orphaned closing syntax, or fragments of code that are cut through a meaningful construct. In addition, any changes to the source code processing pipelines also go through the full, end-to-end retrieval evaluation. We’ll get back to that evaluation pipeline in the following part of this series.

Vectorization

Having pre-processed the source code, we finally have text chunks that are hopefully just the right size and correctly grouped for semantic retrieval. Our next task is to transform these fragments in a way that will later allow us to support semantic search, through a process called vectorization.

With vectorization, an embedding model reads a piece of text and emits a fixed-length list of numbers (a vector), which amounts to a point in a space of a few thousand dimensions. Significantly, the model is trained so that texts with similar meaning land close together. Traditional search might miss the connection, but here, a function that flushes buffered write operations and one that drains a pending queue can end up near each other despite sharing no common keywords. The distance between vectors hence becomes a measure of relatedness. A query is turned into a position in the same space, and the results are whatever lies nearest to it.

Punch for the byte: Optimizing for storage

Any attempt to vectorize a large codebase must take into account both cost and performance. A single embedding is cheap, but a large repository produces millions of chunks, which become millions of vectors that must be stored, held in memory, and compared against each incoming query. A vector of a few thousand dimensions in 32-bit floats weighs around 16 kilobytes, so a few million chunks add up to tens of gigabytes of index before any bookkeeping. At such a scale, the allocation of bytes per vector becomes cost-limited, and the leading question quickly shifts from “how accurate can we be?” to “what do we get per byte?” In other words, we need to find a way to reduce the cost while retaining as much search quality as possible.

There are two ways to reduce vector cost. The first is to keep fewer dimensions. Modern embedding models are trained so that a leading slice of the vector works on its own. The dimension loss is applied across several nested prefix lengths simultaneously, pushing the coarsest structure into the earliest dimensions. This means you can cut a vector short and renormalize it, and it still retrieves. Alternatively, you can keep every dimension and spend less on each one by sacrificing on precision and thus keeping fewer bytes for each vector.

These two options are independent of each other and can be combined, which means any storage budget can be met through different mixes of dimension count and numeric precision. The real question is which mix retrieves best for the same number of bytes. The trade-off is far from even. Suppose the budget is 512 bytes per vector. You could spend it on 128 dimensions kept at full 32-bit precision, or on all 4,096 dimensions kept at a single bit each. Both fit the budget exactly, but in testing, you’ll find that the second option retrieves considerably better.

Why dimensions matter more than precision

To see why, it helps to think of each dimension as one small question the model has learned to ask about the text: Is this about error handling? Does it touch the network? Is it test code? And there are a few thousand similar topics and questions that haven’t been named. (The real dimensions are blurrier than that, but this is a useful abstraction.)

No single answer means much on its own. We consider two chunks to be similar when their answers to many of these questions are the same. Therefore, we should assess the vectors by looking at the coverage of the questions rather than the exactness of the answers.

Keeping all 4,096 dimensions at one bit preserves a rough yes-or-no answer to every question. Truncating to 128 dimensions keeps very precise answers to three percent of the questions and throws the rest away, and no amount of precision on the surviving dimensions can recover the information the discarded ones carried. In a sense, a long questionnaire filled in with checkmarks beats a short one filled in to six decimal places. Dimensions are what you want to keep; precision is what you can afford to lose and is easier to compensate for later on.

So we chose to keep every dimension and take the precision reduction to its limit, dropping the vectors to one bit each, which is 32 times smaller than the same vector in 32-bit floats. The quantization itself turns out to be surprisingly simple. Every component at or above zero becomes a one, while every negative component becomes a zero, and the magnitudes are thrown away:

Changing the representation changes the metric with it. Cosine similarity needs the magnitudes we just threw away, so binary vectors are compared by Hamming distance instead, which is simply the number of positions where two bit patterns disagree. Compare, for example, 10110100 and 10010110. They differ in two positions, so the distance between them is two. At full length, the computation stays just as simple. A 4,096-bit vector is stored as 64 words of 64 bits, and comparing two of them means XORing each pair of words, which leaves a 1 wherever the two vectors disagree, and then counting the 1s. A CPU does each of those in a single instruction per word, so a full comparison costs in the order of a hundred instructions where cosine similarity on the original floats needed thousands of multiplications.

Note that the metric was never a separate decision. We chose one-bit precision for the storage savings, and once every component is a sign bit, Hamming is the only comparison left that makes sense. Choosing the precision chose the metric.

Binary quantization still costs a few points of recall against the unquantized vector. We accepted that cost after considering that a reasoning agent would be consuming the results. A code search feeding an agent needs the right neighborhood far more than a perfectly ordered top 10. When the agent asks where session tokens get refreshed, what matters is that the relevant handful of files shows up among the first dozen results. Whether the best chunk ranks second or fifth changes nothing, because the agent opens the candidates and reads them anyway. In that loop, a ranking degradation that would be plainly visible in a three-result UI built for humans is mostly invisible.

The limits of binary quantization

The trade-off we made had a subtler cost that took us a bit longer to understand. Binary quantization doesn’t only sacrifice accuracy; it compresses the *range* of similarity scores. With full-precision vectors, an unrelated pair can score near zero while near-duplicates score near one, a comfortably wide spread. Sign bits behave differently. Around half the bits of two entirely unrelated vectors still agree by pure chance, while a strongly related pair might have agreement for two-thirds. So every score in the index, relevant or not, lands in that thin band.

Ranking survives the compression, since relevant results still score above irrelevant ones, but thresholding does not. Picture a feature that volunteers related code without being asked, say a panel that suggests existing implementations while you type. Its most difficult requirement is knowing when to stay silent. To make that determination, it needs a usable gap between “related” and “unrelated” scores. Binary vectors don’t leave one. Any cutoff placed inside that narrow band either fires on everything or on nothing. So where an index needs an absolute relevance judgement rather than a relative ordering, we keep 16-bit floats and pay for the storage.

Embedding scope

While indexing and searching use the same model, the two jobs could not be more different. Indexing is throughput-constrained, with millions of chunks asynchronously handled. The GPU will handle about 32 chunks per batch before becoming saturated. A search, on the other hand, needs to be fast and responsive. Users will give up if they are not provided with results within a couple of seconds at most. Therefore in deploying these models we optimize them accordingly: one to maximize chunks per second, the other for minimizing time to first result.

We chose an instruction-following model, trained with a deliberate asymmetry between the two sides of retrieval. Significantly, the two sides are represented by very different types of text. A query is a short question in natural language, while a document is a chunk of code. A document is embedded as is at indexing time. A query is wrapped with an instruction describing the retrieval task, something like “given this search query, find the code that answers it”, which tells the model what role the text is playing. We preserve that arrangement at inference because it is the shape the model learned.

To allow the two sides to align more easily, we embed each chunk together with its file path. The path supplies metadata that the chunk alone lacks: which module it lives in, and what the file is. In a monorepo, though, the path itself becomes a problem. The IntelliJ IDEA monorepo runs to over a million files. The median source file there sits nine directories deep behind a 91-character path, and close to 10,000 source files have paths longer than 150 characters, the longest of them 218. That is before any checkout root is prepended.

Most of those characters are used for structural nesting and offer no useful information about the file. A run of segments like `src/org/jetbrains/kotlin/idea/k2` restates the package hierarchy, which a compiler needs and a search does not. Meanwhile, the file at the end of that longest path is 24 lines long. If we simply embed the path text as is beside a chunk, we’ll find that the path will sometimes take up more space than the code itself. To compensate for that, a path is capped before it reaches the model, and the rule is that *both ends survive*. The leading segments tell you which module you’re in, while the last two, the immediate parent and the filename, tell you what the file is. The middle is the part that can go, and only as much of it as the cap requires. Keep the longest prefix that still fits, elide what falls between into `…`, and if even parent-plus-filename is too long, keep only the name itself.

The same discipline applies when a user scopes a search to a subdirectory. The obvious implementation is a metadata filter: run the search as usual and discard results that fall outside the directory. We do something different. The scope is rendered into the query text itself, in the same shape, with the same abbreviation function and the same separator the indexed chunks used. If a chunk went into the index under the abbreviated form of `community/plugins/kotlin`, a query scoped to that directory carries the same string in exactly the same form, so the query vector lands in the same region as the chunks it is supposed to match.

Protecting source code

There was one last design consideration we took into account. It was important for us to be attentive to customer privacy and security concerns. The source code of a company is often the core of its IP. Exposing it to third-party cloud models, or even to another company, increases the risk of inadvertently exposing sensitive data or even training other models to use it.

To make sure we address these concerns, we made the decision to adhere to several practices early on:

  1. Avoid storing the code in our systems: A chunk holds a cluster reference, an item type, a file path, start and end offsets, a reference to a vector, and an optional metadata field. No content, no copy of the source code itself, is saved. What a search returns is coordinates, and the snippet you see is assembled on your machine, from your checkout, using them. The server just knows that something relevant lives at bytes 4,102–4,890 of a given path, not what it is.
  2. Don’t use data for training: Every code index Air Context builds is embedded by an open-weight embedding model, running on GPUs we operate. No embedding request leaves our infrastructure – not to OpenAI, not to Google, not to any other vendor. Therefore, we can guarantee that none of the data will be used to train anything.

These self-imposed design restrictions carry no cost in terms of retrieval quality. We evaluated the open-weight candidates against the hosted embedding APIs from the major providers on our own code-retrieval benchmarks, and ours came out on top. Open-weight embedders are now good enough that the interesting engineering has moved into what you feed them, how you serve them, and what you choose to keep.

A summary that is an interlude

In this blog post, we covered the first stages of the retrieval pipeline: the journey from raw source files to compact vectors that are ready to be searched.

At this point, we have millions of binary vectors and a way to produce more. The problems we haven’t solved yet are how to store them efficiently, how to create a system that can answer a query in milliseconds, how we can continuously evaluate our results to ensure we are making the right choices, and how we can get the agent to actually use our shiny RAG apparatus.

These topics and more will be the subjects of the next parts in this series, which we’ll be releasing over the next few weeks. As always, please feel free to ask any questions in the comments or share your own hard lessons from designing a RAG solution. We are eager to learn of different and creative ways you have found to be effective! In the meantime, feel free to check out Air Context , currently in public preview, it is already included with your JetBrains license 😀

Until next time!

Subscribe to JetBrains AI Blog updates

Discover more

Incentives in Academic Research

Hacker News
www.msoos.org
2026-10-04 13:42:38
Comments...
Original Article
Andrei Tarkovsky’s Stalker (1979)

Charlie Munger said: “Show me the incentive and I will show you the outcome”. The issue with Academic Research, in my opinion, is that the outcomes have drifted very far away from the original goal, which I believe to be the advancement of scientific understanding, and training of the new generation of researchers. Academic research was meant to be about breaking new ground, keeping to honesty and good scientific conduct, being clear and upfront about uncertainties, errors, and mistakes, and improving our common understanding of science, all the while training the new generation to follow these goals and principles.

Recently, I have bumped into multiple cases where I believe the correctness of the results, the honesty of the people writing them, or the lack of curiosity once they are told that their results are wrong, incorrect, faulty, or misleading, has been unsatisfying. Simply put, researchers are not too interested in learning that their papers or reports are wrong, and/or misleading others who don’t happen to know that the results are — known to the authors, and a few select others to be partially — incorrect.

The issue is, once you published the paper, and got the promotions and fame, it doesn’t matter that the results are wrong, and known to be wrong or misleading, and potentially doing harm to the advancement of science. It’s not your problem. It’s someone else’s problem. In fact, when I challenged the authors of one such paper, considered state-of-the-art, and known to the authors to have incorrect evaluation, one of the author’s response was along the lines of acknowledging the issues, but refusing to retract the paper, and instead asking if it’s bothering me in publishing my paper.

The incentives are wrong. In my opinion, it’s not all about the papers — mine or others’. It’s about scientific integrity, about being honest with each other, it’s about caring about the results, it’s about advancement of our common scientific understanding through seeking of truth and correctness. It’s about communicating when something is wildly wrong, and either retracting the relevant incorrect claims, or notifying the community of the known serious issues. It’s about caring for what we all consider to be the common understanding of what is the truth, and cultivating an environment where the new generation grows up to learn what the proper scientific conduct is, and that it is not acceptable to seriously deviate from it.

Unfortunately, I am seeing more and more PhD students who are less and less interested in correctness, precision, and what I’d consider proper scientific conduct. They have learned from those successful in the field what does and does not matter. I remember when someone once told me that my approach for a particular algorithm was wrong (about strongly connected components discovery), and how upset I was that I had no idea. I went home that day and I immediately fixed my tool to use the right approach (i.e. to use Tarjan’s algorithm). I felt deep shame that I had no clue what I was doing, and that I messed up. In contrast, recently talked with a PhD student who wrote a tool that was meant to perform well in certain contexts. When I explained the student that the approach to the evaluation was incorrect, they didn’t follow up at all. I think they were surprised that I later followed up, demonstrating that indeed the evaluation was not careful enough, and the tool is less useful than it seems from the paper’s evaluation. It was a strange experience — when I was a PhD student and someone sat down and showed me that what I was doing was potentially sloppy, I was super worried and very curious to find out what’s going on. I still remember this moment when a reviewer of my dissertation challenged a graph and I was stressed for a week before I figured out they misread the graph, because I didn’t label it clearly. Or when in a presentation I accidentally left out performing so-called ‘restarts’ as a major advancement in the history of SAT solvers, and someone in the audience rightfully pointed it out. I felt embarrassed for having made such a mistake.

I am not sure how to fix this problem. Seemingly, researchers are less and less keen on correctness, precision, and scientific curiosity. They don’t seem to be incentivised to do so. I sometimes wonder if the system has become so damaged, so many PhD students have grown up to be professors in this environment, that much of what I wrote here seems alien, even repulsive, to many. It makes me sad.

What I learnt co-leading an AI Safety bootcamp for legal and governance practit

Hacker News
www.lesswrong.com
2026-10-04 13:21:26
Comments...
Original Article

x

What I learnt co-leading an AI Safety bootcamp for legal and governance practitioners — LessWrong

Blindsight (Watts Novel)

Hacker News
en.wikipedia.org
2026-10-04 12:25:17
Comments...
Original Article

From Wikipedia, the free encyclopedia

Blindsight
Author Peter Watts
Cover artist Thomas Pringle [ 1 ]
Language English
Genre Hard science fiction
Publisher Tor Books
Publication date 3 October 2006
Publication place Canada
Media type Print (hardback)
Pages 384
ISBN 978-0-7653-1218-1
OCLC 64289149
Dewey Decimal 813/.622
LC Class PR9199.3.W386 B58 2006
Followed by Echopraxia

Blindsight is a hard science fiction novel by Canadian writer Peter Watts , published by Tor Books in 2006. It won the Seiun Award for the best novel in Japanese translation (where it is published by Tokyo Sogensha ) [ 2 ] and was nominated for the Hugo Award for Best Novel , [ 3 ] the John W. Campbell Memorial Award for Best Science Fiction Novel , [ 4 ] and the Locus Award for Best Science Fiction Novel . [ 5 ] The story follows a crew of astronauts sent to investigate a trans-Neptunian comet dubbed "Burns-Caulfield" that has been found to be transmitting an unidentified radio signal, followed by their subsequent first contact . The novel explores themes of identity , consciousness , free will , artificial intelligence , neurology , and game theory as well as evolution and biology .

Blindsight is available online under a Creative Commons Attribution-NonCommercial-ShareAlike license . [ 6 ] Its sequel (or " sidequel "), Echopraxia , came out in 2014.

In the year 2082, tens of thousands of coordinated comet-like objects of an unknown origin, dubbed "Fireflies", burn up in the Earth's atmosphere in a precise grid, while momentarily broadcasting across an immense portion of the electromagnetic spectrum, catching humanity off guard and alerting it to an undeniable extraterrestrial presence. It is suspected that the entire planet has been surveyed in one effective sweep. Despite the magnitude of this "Firefall", human politics soon return to normal.

Soon afterwards, a comet-surveying satellite stumbles across a radio transmission originating from a comet, subsequently named 'Burns-Caulfield'. This tight-beam broadcast is directed to an unknown location and in fact does not intersect the Earth at any point. As this is the first opportunity to learn more about the extraterrestrials, three waves of ships are sent out: the first being lightweight probes shot out for an as-soon-as-possible flyby of the comet, then a wave of heavier but better-equipped probes, and finally a crewed ship, the Theseus .

Theseus is propelled by an antimatter reactor and captained by an artificial intelligence . It carries a crew of five cutting-edge transhuman hyper-specialists, of whom one is a genetically reincarnated vampire who acts as the nominal mission commander. While the crew is in hibernation en route, the just-arrived second wave of probes commence a compounded radar scan of the subsurface of Burns-Caulfield, but this immediately causes the object to self-destruct. Theseus is re-routed mid-flight to the new-found destination of the signal: a previously undetected sub-brown dwarf deep in the Oort cloud , dubbed 'Big Ben'.

The crew wakes from hibernation while the Theseus closes on Big Ben. They discover a giant, concealed object in the vicinity, and assume it to be a vessel of some kind. As soon as the crew uncloaks the vessel, it immediately hails them over radio and, in a range of languages varying from English to Chinese , identifies itself as 'Rorschach'. They determine that Rorschach must have learned human languages by eavesdropping on comm-chatter since its arrival, sometime after the Broadcast Age began. Over the course of a few days many questions and answers are exchanged by both parties. Eventually Susan James, the linguist, determines that 'Rorschach' does not really understand what either party is actually saying .

Theseus probes Rorschach and finds it to have hollow sections, some with atmosphere, all filled with levels of radiation that render remote operation of machinery virtually impossible and would kill a human in a matter of hours. Despite this and over Rorschach's objections the whole crew except the mission commander enters and explores in a series of short forays, using the ship's advanced medical facilities to recover from the damage the radiation inflicts on their bodies. They discover the presence of highly evasive, fast-moving nine-legged organisms dubbed 'Scramblers'. They kill one and capture two for study. The 'Scramblers' appear to have orders of magnitude more brainpower than human beings but use most of it simply to operate their fantastically complex musculature and sensory organs; they are more akin to something like white blood cells in a human body. They are dependent on the radiation and EM fields of Rorschach for basic biological functions and seem to completely lack consciousness .

The crew explore questions of identity, the nature, utility and interdependence of intelligence and consciousness. They theorize that humanity could be an unusual offshoot of evolution, wasting bodily and economic resources on the self-aware ego which has little value in terms of Darwinian fitness. Open warfare breaks out between the humans and the Scramblers and Theseus eventually decides to sacrifice itself and its crew using its antimatter payload to eliminate Rorschach. One crew member, the protagonist and narrator Siri Keeton, is shot off inside an escape vessel in a decades-long fall back to Earth to relay the crucial information amassed back to humanity.

Crew of the Theseus

[ edit ]

  • Siri Keeton is the narrator and protagonist. Debilitating brain surgery for medical purposes has cut him off from his own emotional life and made him a talented "synthesist", adept at reading others' intentions impartially with the aid of cybernetics. He is assigned to Theseus to interpret the actions of the specialized crew and report these activities to Mission Control on Earth.
  • Major Amanda Bates is a combat specialist, controlling an army of robotic "grunts".
  • Isaac Szpindel is the ship's primary biologist and physician. He is in love with Michelle, one of the Gang's personalities.
  • Jukka Sarasti is a vampire and the crew's nominal (and frightening) leader. As a predator from the Pleistocene , he is alleged to be far smarter than baseline humans.
  • The Gang are four distinct personalities in the mind of one woman, the ship's linguist. They are tasked with communicating with the aliens, if possible. A single personality "surfaces" to take control of their body at any given time. The active personality reveals itself through a change in tone and posture. These personalities express offence when referred to as " alters ". The personalities are:
    • Susan James, whom the others refer to as "Mom". She is the "original" personality.
    • Michelle is a shy, quiet, synaesthetic woman who is romantically involved with Szpindel.
    • Sascha is harsher and more overtly hostile towards Siri.
    • Cruncher, a male personality, rarely surfaces and serves as an advanced data-processing facility for James.
  • Robert Cunningham, Szpindel's backup, is a secondary biologist/physician. He possesses a set of enhancements that allow him to process data with additional senses, and somewhat inhabits the machinery connected to him.
  • The Captain is the ship's artificial intelligence. Throughout the story, the Captain remains inscrutable and mysterious, generally communicating directly only with Sarasti.
  • Robert Paglino, Siri's childhood best friend and a practical example of Siri's muted emotions: Siri cannot actually feel "friendship" following his brain surgery, but intellectually knows how he is expected to behave as a friend and continues to play the part.
  • Chelsea, Siri's ex-girlfriend. A professional tweaker of human personalities.
  • Helen Keeton, Siri's mother, whose consciousness has been connected, brain in a vat style, to a virtual utopia called "Heaven". As a parent, she traumatized Siri with emotional demands and intrusiveness into his private life.
  • Jim Moore is Siri's father, a colonel involved with planetary defense.
  • Rorschach , an alien vessel or organism in low orbit around the sub-brown dwarf Big Ben. While it has a superhuman intelligence, it gradually becomes apparent that Rorschach completely lacks true consciousness or self-awareness .
  • Scramblers, 9-legged anaerobic aliens that inhabit Rorschach and appear to be part of it in some sense. Like Rorschach , they are more intelligent than humans but not conscious or self-aware.

The exploration of consciousness is the central thematic element of Blindsight . [ 7 ] [ 8 ] [ 9 ] The title of the novel refers to the condition blindsight , in which vision is non-functional in the conscious brain but remains useful to non-conscious action. [ 10 ] Other conditions, such as Cotard delusion and Anton–Babinski syndrome , are used to illustrate differences from the usual assumptions about conscious experience. [ 10 ] The novel raises questions about the essential character of consciousness. Is the interior experience of consciousness necessary, or is externally observed behavior the sole determining characteristic of conscious experience? [ 7 ] [ 8 ] [ 10 ] Is an interior emotional experience necessary for empathy, or is empathic behavior sufficient to possess empathy? [ 10 ] [ 11 ] Relevant to these questions is a plot element near the climax of the story, in which the vampire captain is revealed to have been controlled by the ship's artificial intelligence for the entirety of the novel. [ 10 ] [ 12 ]

Philosopher John Searle 's Chinese room thought experiment is used as a metaphor to illustrate the tension between the notions of consciousness as an interior experience of understanding, as contrasted with consciousness as the emergent result of merely functional non-introspective components. [ 7 ] [ 10 ] [ 12 ] Blindsight contributes to this debate by implying that some aspects of consciousness are empirically detectable. [ 8 ] Specifically, the novel supposes that consciousness is necessary for both aesthetic appreciation [ 8 ] [ 9 ] [ 11 ] and effective communication. [ 8 ] However, the possibility is raised that consciousness is, for humanity, an evolutionary dead end. [ 7 ] [ 10 ] [ 11 ] [ 12 ] That is, consciousness may have been naturally selected as a solution for the challenges of a specific place in space and time, but will become a limitation as conditions change or competing intelligences are encountered. [ 8 ]

The alien creatures encountered by the crew of the Theseus themselves lack consciousness. [ 7 ] [ 8 ] [ 11 ] [ 13 ] The necessity of consciousness for effective communication is illustrated by a passage from the novel in which the linguist realizes that the alien creatures cannot be, in fact, conscious because of their lack of semantic understanding:

"Tell me more about your cousins," Rorschach sent.
"Our cousins lie about the family tree," Sascha replied, "with nieces and nephews and Neanderthals. We do not like annoying cousins."
"We'd like to know about this tree."
Sascha muted the channel and gave us a look that said Could it be any more obvious? "It couldn't have parsed that. There were three linguistic ambiguities in there. It just ignored them."
"Well, it asked for clarification," Bates pointed out.
"It asked a follow-up question. Different thing entirely." [ 14 ]

The notion that these aliens could lack consciousness and possess intelligence is linked to the idea that some humans could also have diminished consciousness and remain outwardly functional. [ 8 ] [ 9 ] This idea is similar to the concept of philosophical zombie , as it is understood in philosophy of mind . Blindsight supposes that sociopaths might be a manifestation of this same phenomenon, [ 8 ] [ 10 ] and the demands of corporate environments might be environmental factors causing some part of humanity to evolve toward becoming philosophical zombies. [ 8 ] [ 11 ]

See Bicameral mentality

Blindsight also explores the implications of a transhuman future. [ 8 ] [ 12 ] [ 13 ] Within the novel, humans no longer engage in sex with other humans for pleasure, instead choosing to use virtual reality to find idealized partners, [ 8 ] and many choose to withdraw from reality entirely by living in constructed virtual worlds, referred to as "Heaven". [ 8 ] [ 12 ] Vampires are predators from humanity's distant past, resurrected through recovered DNA , and live among the humans of the late 21st century. [ 7 ] [ 10 ] [ 12 ] [ 13 ] These vampires operate with diminished sentience presented as comparable to high-functional autism with comparable dysfunction in affect and speech, but have the advantage of multiple simultaneous thoughts occurring in parallel within their minds. [ 12 ] Enhanced pattern-matching skills comparable to some forms of autism combine with this "hyperthreading" to make them invaluable in developing unusual and often very effective approaches to solving complex problems.

Carl Hayes, in his review for Booklist , wrote, "Watts packs in enough tantalizing ideas for a score of novels while spinning new twists on every cutting-edge motif from virtual reality to extraterrestrial biology." [ 15 ] Kirkus Reviews said about the book, "Watts carries several complications too many, but presents nonetheless a searching, disconcerting, challenging, sometimes piercing inquisition." [ 16 ] Jackie Cassida in her review for Library Journal wrote, "Watts continues to challenge readers with his imaginative plots and superb storytelling." [ 17 ] Publishers Weekly wrote, "Watts puts a terrifying and original spin on the familiar alien contact story." [ 18 ]

Elizabeth Bear , an award-winning author in the science fiction field, declared the following:

It's my opinion that Peter Watts's Blindsight is the best hard science fiction novel of the first decade of this millennium – and I say that as someone who remains unconvinced of all the ramifications of its central argument. Watts is one of the crown princes of science fiction's most difficult subgenre: his work is rigorous, unsentimental, and full of the sort of brilliant little moments of synthesis that make a nerd's brain light up like a pinball machine. But he's also a poet – a damned fine writer on a sentence level... [ 19 ]

In October 2020 a non-commercial Blindsight short film was released. [ 20 ] Watts describes it as, "snatches of Blindsight recalled by Siri Keeton during one of his waking interludes in the aftermath of that novel. Spectacular highlights arranged in reverse order, Memento-like". [ 21 ] In April 2025, Neill Blomkamp was set to adapt the novel into a feature film. [ 22 ]

  1. ↑ "Blindsight: The Lost Covers" . Retrieved 1 January 2013 .
  2. ↑ "2014 Seiun Award Winners" . Locus . 21 July 2014 . Retrieved 21 February 2019 .
  3. ↑ "Hugo Nominees (press release)" . Archived from the original on 3 May 2007 . Retrieved 3 October 2008 .
  4. ↑ "Campbell Award Winners & Nominees" . Worlds Without End . Retrieved 23 December 2011 .
  5. ↑ "Locus SF Award Winners & Nominees" . Worlds Without End . Retrieved 23 December 2011 .
  6. ↑ Watts, Peter. "Blindsight by Peter Watts" . www.rifters.com . Retrieved 8 December 2023 .
  7. 1 2 3 4 5 6 McGrath, Martin (10 March 2011). "Blindsight... Or "In a Chinese Room, not far from the loo" " . Archived from the original on 14 October 2014 . Retrieved 8 October 2014 .
  8. 1 2 3 4 5 6 7 8 9 10 11 12 13 Shaviro, Steven (27 October 2006). "Blindsight" . Archived from the original on 3 December 2006 . Retrieved 8 October 2014 .
  9. 1 2 3 Shaviro, Steven. "Consequences of Panpsychism" (PDF) . p. 14 . Retrieved 8 October 2014 .
  10. 1 2 3 4 5 6 7 8 9 "Transcript Podcast 2: "Blindsight" by Peter Watts" . Science Fiction First. Archived from the original on 15 October 2014 . Retrieved 8 October 2014 .
  11. 1 2 3 4 5 Shaviro, Steven (25 August 2014). "Ferociously Intellectual Pulp Writing" . Archived from the original on 26 August 2014 . Retrieved 8 October 2014 .
  12. 1 2 3 4 5 6 7 Elber-Aviram, Hadas. "Visions of Humanity between the Posthuman and the Non-Human" (PDF) . Imachine: There is No I in Meme : 4– 5. Archived from the original (PDF) on 14 October 2014 . Retrieved 8 October 2014 .
  13. 1 2 3 Nirshberg, Greg (7 December 2010). "Book Review – Blindsight by Peter Watts" . Archived from the original on 1 October 2014 . Retrieved 8 October 2014 .
  14. ↑ Watts, Peter (3 October 2006). Blindsight . Tor Books . pp. 112 . ISBN 978-0-7653-1218-1 .
  15. ↑ Hays, Carl (1 October 2006). "Blindsight". Booklist . 103 (3): 45. ISSN 0006-7385 .
  16. ↑ "BLINDSIGHT". Kirkus Reviews . 74 (16): 816. 15 August 2006. ISSN 0042-6598 .
  17. ↑ Cassada, Jackie (15 October 2006). "Blindsight". Library Journal . 131 (17): 55. ISSN 0363-0277 .
  18. ↑ Blindsight (28 August 2006). "Blindsight". Publishers Weekly . 253 (34): 36. ISSN 0000-0019 .
  19. ↑ Bear, Elizabeth (3 March 2011). "Best SFF Novels of the Decade: An Appreciation of Blindsight " . Tor.com . Retrieved 10 July 2014 .
  20. ↑ Krivoruchko, Danil. "Blindsight: A Short Film" . blindsight.space . Retrieved 30 October 2022 .
  21. ↑ "No Moods, Ads or Cutesy Fucking Icons » Memento with Scramblers: Krivoruchko Crushes It" .
  22. ↑ Barder, Ollie (9 April 2025). "Peter Watts On 'Blindsight', 'Armored Core' And Working With Neill Blomkamp" . Forbes . Retrieved 20 August 2026 .

Car is a smartphone on wheels. Here's who's listening

Hacker News
automatictransmission.khoury.northeastern.edu
2026-10-04 11:43:14
Comments...
Original Article

Your car is a smartphone on wheels. Here's who's listening.

The first large-scale measurement study of the connected-vehicle ecosystem.

A connected car is a vehicle with built-in internet access — Wi-Fi, cellular, GPS — that lets it communicate constantly with its manufacturer and outside companies. Over 75% of vehicles sold globally have this connectity built-in.

A connected vehicle

Photo is not loading images/connected-car.jpg

It knows where you drive.

It knows who you are.

It can share this data and more with insurance companies, advertisers, etc...

Report Summary

Partnering with Consumer Reports, we tested 21 late model vehicles and 30 companion mobile apps to understand the privacy implications of the connected vehicle ecosystem.


  • 01 We find that both vehicles and companion apps contact numerous third-party domains, including advertisers and trackers.
  • 02 19/21 vehicles tested send traffic to at least one third party
  • 03 Seven of 30 apps transmit sensitive identifiers to third-party companies

We go through a lengthy disclosure process and provide insight into how manufacturer's perceive this data sharing issue.
Our findings underscore the need for continued measurement and scrutiny of the connected vehicle ecosystem.

Partnership with Consumer Reports

Consumer Reports gave the team access to its purchased fleet of test vehicles — a sample that would have cost over $1.2M to assemble independently. Read CR's article on our work!

The Paper

This work is peer reviewed and will be published at IMC '26. Read the paper →

The Research Team

We are a team of privacy, security, and networking-systems researchers at Northeastern University. See the full team →

21

vehicles tested,
19 brands

19 / 21

vehicles contacted a
third party over Wi-Fi

7 / 30

apps sent PII to trackers

5 / 30

apps sent VIN + other
PII to trackers

Research Questions

Diagram of how the connected vehcile ecosystem works

Photo is not loading images/car-diagram.png

Connected Vehicle Ecosystem

This diagram shows the data flows to and from a vehicle and its companion mobile app. Both devices send data, including private consumer data, to different 1st and 3rd party servers using Wi-Fi and cellular service. Solid arrows represent flows that were intercepted through our experiments.

The Problem

Once the data gets sent to these servers, it is up to the companies that receive the consumer information to make decisions on what they do with it. Unfortunately, in many cases, this includes sharing or selling consumer data to other undisclosed 3rd parties.
Consumers have no control over their data once it has left their device

In this paper, we take the first steps to address the limited visibility into the privacy implications of the connected vehicle ecosystem. We identify two vantage points in the ecosystem where we can gain insight into the data that connected vehicles are sharing with both manufacturers and third parties: the vehicles themselves and the mobile apps provided by manufacturers.
We ask the following guiding questions:

  • 01 What personal consumer data do connected vehicles and their companion mobile apps transmit?
  • 02 Who receives that personal consumer data?
  • 03 What is the manufacturer response to these findings?

Methods

We investigated 21 vehicles from the U.S. market in a controlled environment along with 30 companion mobile apps instrumented with on-site vehicles between October 2024 and August 2025. Below is a description of the experiments we ran and our setup.

Vehicle Testing

Testing Setup

Photo is not loading images/rpi.png

Wi-Fi Testing Setup

To collect Wi-Fi traffic from vehicles, we configured a custom access point (AP) on a Raspberry Pi and used tcpdump to log all packets that were sent or received via this AP.
This allowed us to see all the destinations the vehicles were sending data to but not the information within the packets as it was encrypted.

Vehicles at the testing facility

Photo is not loading images/test-facility.jpg

Stationary Tests

Idle Baseline
vehicle on — no activity

Active Test
perform all possible actions

Driving Test

drive 5–45 mph with acceleration & hard braking

Isolating Cellular Traffic

Faraday tent used to block cellular signals

Photo is not loading images/faraday-tent.jpg

One hypothesis we tested was whether blocking a vehicle’s ability to communicate over its cellular network would force more Wi-Fi communication. To block external cellular signals, we drove 11 EVs in the sample into a car-sized Faraday tent providing ≈93 dB of attenuation, blocking their cellular connection entirely. Stationary Idle and Active tests were repeated inside the tent to see whether traffic that normally goes out over cellular rerouted to Wi-Fi instead.

App Testing

In total, we experimented with 30 connected vehicle companion apps that were paired with the vehicles at Consumer Reports’s testing facility.

Device Setup

  • 01 We used a combination of test phones and iOS versions for our experiments: an iPhone 8/iOS 16.6, an iPhone X/iOS 16.7.11, and an iPhone 13/iOS 18.5.
  • 02 To minimize background traffic, we deleted all non-essential apps on the test phones and tested apps one-by-one, including deleting each vehicle app and restarting the phone before downloading the next app.
  • 03 We used the iOS native screen recording feature to record our interactions with each app for later review.
  • 04 To capture and decrypt network traffic from the companion mobile apps, we used iPhones with custom root certificates connected to mitmproxy.

Process for Testing Each App

  • 01 During app installation and login we accepted all permission requests (e.g., tracking, location, calendar access, Bluetooth, notifications) that the application requested
  • 02 We had a Consumer Reports employee log into the app using their existing credentials associated with a vehicle on the lot.
  • 03 Once we were logged-in, we manually exercised all available functionality, such as looking for nearby charging staions, geolocating the vehicle, viewing vehicle data and service history (e.g., tire pressure), viewing notifications (e.g., “doors are unlocked”), and viewing in-app privacy policies.
  • 04 Some apps allowed us to perform physical interactions on the vehicle, such as remotely opening the trunk. We performed all such actions and verified that the vehicle completed each request.

Vehicle Dataset

$1.2M+

estimated cost to independently procure this fleet, only possible due to Consumer Reports partnership

Manufacturer Brand Year Model Vehicle Tests
Driving Idle In-Tent Cellular
General Motors (GM) Buick 2024 Envista ✓ ✓ – –
Cadillac 2024 Lyriq ✓ ✓ ✓ –
Chevrolet 2024 Blazer ✓ ✓ ✓ –
Stellantis Dodge 2023 Hornet - ✓ – –
Fiat 2024 500e ✓ ✓ ✓ –
RAM 2025 1500 Bighorn ✓ ✓ – –
Fisker Fisker 2023 Ocean ✓ ✓ - –
Ford Motor Co. Ford 2022 F150 Lightning ✓ ✓ - –
Ford 2024 Mustang GT Fastback ✓ ✓ – –
Honda Motor Co. Honda 2024 Prologue Touring AWD ✓ ✓ ✓ –
Tata Motors Land Rover 2023 Range Rover Sport ✓ ✓ – –
Toyota Motor Corp. Lexus 2024 NX450H+ PHEV ✓ ✓ – –
Toyota 2023 Corolla Cross ✓ ✓ – –
Subaru 2023 Solterra ✓ ✓ ✓ –
Lucid Lucid 2023 Air Touring ✓ ✓ ✓ –
Mercedes-Benz Group AG Mercedes 2023 EQS450 4Matic ✓ ✓ - –
Renault-Nissan-Mitsubishi Alliance Nissan 2023 Ariya Platinum ✓ ✓ ✓ –
Rivian Rivian 2022 R1S ✓ ✓ ✓ –
Tesla Tesla 2024 Cybertruck ✓ ✓ ✓ –
Tesla 2024 Model 3 ✓ ✓ ✓ ✓
Zhejiang Geely Holding Group Volvo 2024 C40 ✓ ✓ ✓ –

* While we tested a wide range of vehicles, we did not cover every manufacturer in the U.S. market. As with all empirical studies, our results should be interpreted as a snapshot in time, and may not generalize to vehicles outside our sample or outside the U.S.

Findings

  • 19 of 21 vehicles contacted at least one third party over Wi-Fi, including known advertising and tracking domains
  • 7 of 30 companion apps transmitted sensitive identifiers (VINs, emails, phone numbers, precise location) to third parties associated with advertising and tracking.
    Transmiting multiple forms of PII to the same third party allows advertisers to build in-depth profiles on consumers

Click any company chip for what it received and from which app.

  • Pairing a companion app roughly doubled a vehicle's exposure to advertising/tracking companies on average, and in some cases added 20+ new ones.

vehicle-only ATA companies added by the companion mobile app

Manufacturer Disclosures

  • The team disclosed these findings to the different manufacturers featured in the study, below is an interactve summary of the different explanations received.

Click any box above for more info.

  • The ongoing theme of all these responses was shifting the blame to the consumer .
  • The current system does not give owners the ability to choose.
    Assuming owners can find the particular agreements, if an owner decides they are uncomfortable with the data sharing described within them, they are faced with (arguably) unfair choices:
    01 accept the agreements regardless of their concerns
    02 stop using their car's connected-vehicle features - which includes remote start, the app, and other very useful features
    03 stop using the vehicle entirely
  • As one notable exception, Honda improved its data collection practices to prevent sending precise geolocation to a third party associated with user tracking.

Conclusion

  • We found that vehicles, under a variety of real-world settings, contact not only a wide range of first parties (i.e., manufacturer domains) and car-specific support parties, but also third parties that are known to provide advertising and tracking services.
  • Vehicles from the same manufacturer exhibit different network behaviors - this makes it very difficult and expensive to study these vehicles.
  • When adding vehicle companion apps to the analysis, we found that vehicle owners are exposed to even more privacy-sensitive ATA communication—in some cases more than two dozen additional trackers.
  • Our study revealed a large gap between what vehicle manufacturers publicly disclosed and how the connected vehicle ecosystem actually shares data over the Internet
  • Based on the opaque nature of vehicular systems, we argue that there is a need for better transparency to ensure increased visibility into the entire ecosystem to identify and address corresponding harms.

RuneScape's Position on Gen AI

Hacker News
www.reddit.com
2026-10-04 11:29:14
Comments...
Original Article

You've been blocked by network security.

To continue, log in to your Reddit account or use your developer token

If you think you've been blocked by mistake, file a ticket below and we'll look into it.

Flatpak from the CLI sucks

Lobsters
kowalski7cc.xyz
2026-10-04 11:21:34
Comments...
Original Article
Sticker by Puzzoz
Sticker by Puzzoz

Package management, from the beginning

One of the features Linux had way before its competition was the ability to download applications from online repositories, which would later be called an "app store". Before this, you would have to download the application binaries and libraries, or build it yourself from the source code, making sure you already had all the required dependencies to compile the application. To solve this issue, package managers like "APT" and "DNF" were developed. On the surface, the task may seem simple: download, install, update, and remove applications. Under the hood, however, a package manager must handle complex operations: downloading the requested application along with all required libraries, checking dependency availability, resolving conflicts, extracting files to the disk, and running any optional post-installation scripts. Once a program was installed, it could be run by clicking on the new icon that appeared, or if the program was placed in /bin/ or another path in the system $PATH variable, by typing the name of the executable.

Sometimes there are too many formats to choose among! - Sticker by Puzzoz
Sometimes there are too many formats to choose among! - Sticker by Puzzoz

As time moved on, new requirements emerged for package managers: distribution-agnostic support and enhanced security through application sandboxing. Because every distribution family had its own package manager, it quickly became obvious that maintaining packages for multiple distributions was time-consuming for developers, and impossible for distribution maintainers to package every available application. Additionally, as security threats increased, it became necessary to isolate untrusted applications from the system, including legitimate software that might be exploited by malware. Access to documents, peripherals, and other aspects of the system needed to be restricted, granting temporary usage only with explicit user permission.

One of these package managers with a new approach is Flatpak.
Flatpak tries to solve both issues by radically changing the packaging and distribution model. It allows developers to ship applications in a layered format that includes all necessary dependencies, working across all distributions, and runs them inside a sandbox. Flatpak provides two kinds of permissions: static permissions declared in the manifest, and dynamic permissions requested at runtime through portals exposed over D-Bus.

This format quickly became the mainstream, default, or even the only way in some distributions to install graphical applications via the software store. However, command-line applications were left behind, and for good reason. Sure, a few CLI-only tools exist, such as flatpak-builder (the tool used to build new Flatpaks), but developers tend to avoid packaging them. Furthermore, as we'll see, their usage from the command line... is rather difficult! If we had installed Flatpak Builder from our distribution's repository with a classic package, we would launch it with flatpak-builder build-folder manifest.yaml . However, when installed as a Flatpak itself, the command becomes flatpak run org.flatpak.Builder build-folder manifest.yaml . Not only does the command become longer and require the full "app id" (which consist of a "unique three-part identifier", pinpointing both the developer and the application), but also requires the filesystem sandbox to be disabled!

In the meantime, flatpak run got a new --file-forwarding option, which maps a specified file from the command line (expressed in the arguments between the @@ delimiters) inside the sandbox to make it available to the application. However, this approach still has a catch.
Adding the required option and enclosing the file path with @@ doesn't help at all with the command length, but more importantly doesn't work with directories or non-existent files you may want to create (such as when converting a picture, where the second file path would be the output file)

Windows 10 and the Universal Windows Platform

Other operating systems encountered the exact same issues regarding application distribution and sandboxing. Windows took a drastic approach: rather than improving the classic Win32 desktop stack, it created an entirely new platform. Even though the technology was initially quite immature, it laid the foundation to distribute "APPX" software through the Windows Store alongside sandboxing via AppContainer.

Soon, developers would realize that migrating from Win32 created a massive workload in trying to port applications to the new platform-and it was not always possible due to the technical limitations of UWP. So in a later update to Windows, the Desktop Bridge was introduced.

Desktop bridge is a Windows technology to package Desktop apps in a modern format
Desktop bridge is a Windows technology to package Desktop apps in a modern format

Among the new features, many of which centered on packaging in the "APPX" format existing Win32 desktop applications, a series of improvements was added to run both UWP and Win32 apps more easily from the command line.

Of course, not every app supported this feature, as it required developers to update the application manifest by adding the uap3:AppExecutionAlias extension and specifying the name of the exported executable outside the sandbox. This would create a special 0-byte execution alias inside %LOCALAPPDATA%\Microsoft\WindowsApps , which is included in the user's PATH variable. This file relies on an NTFS reparse point tagged with IO_REPARSE_TAG_APPEXECLINK . When executed, Windows reads this reparse data and launches the actual binary from the restricted C:\Program Files\WindowsApps directory within its designated application container. An application can export multiple aliases, and the alias name doesn't have to match the executable file inside the container.

Because these aliases sit in a directory that is included in the user's PATH variable, they can sometimes cause conflicts. A classic example is typing python into CMD and unexpectedly opening the Microsoft Store rather than launching the Python version you manually installed from the python website.

To solve this issue, the system provides a simple interface to view all aliases exported by installed applications, along with toggles to disable them. When an alias is turned off, Windows simply deletes that 0-byte reparse point from the %LOCALAPPDATA%\Microsoft\WindowsApps directory. With the file gone, the shell continues searching the rest of the user's PATH.

Windows has a settings page where the user can toggle on and off the various aliases
Windows has a settings page where the user can toggle on and off the various aliases

A solution for Flatpak

Flatpak development could take inspiration from the Windows approach. A suggested solution would consist of multiple small changes. The first would be adding an export-commands key to the Flatpak manifest, mapping internal executable paths to exported command aliases. This approach allowes developers to create wrapper files with complex names inside the package and exporting them with simple names.

app-id: uk.org.greenend.chiark.sgtatham.putty
runtime: org.freedesktop.Platform
runtime-version: '26.08'
sdk: org.freedesktop.Sdk
rename-desktop-file: putty.desktop
rename-icon: putty
command: putty
export-commands:
  putty: putty
  puttygen: puttygen
  psftp: psftp
  pageant: pageant
  pscp: pscp
...

Flatpak builder can take the new section and build the following section in the app metadata file:

[Commands]
putty=putty
puttygen=puttygen
psftp=psftp
pageant=pageant
pscp=pscp

Upon installation, other than showing supported aliases among the current requested permission screen, Flatpak could generate a set of wrapper scripts (analogous to the ones already present in /var/lib/flatpak/exports/bin/ ) with a few key changes: These wrappers would automatically append a --command flag for each exported internal binary, save the wrapper script using the designated alias name, and provide built-in support for --file-forwarding by detecting existing file arguments and wrapping them in @@ delimiters.

#!/bin/sh
for arg do
    shift
    if [ -e "$arg" ]; then
        set -- "$@" "@@" "$(readlink -f "$arg")" "@@"
    else
        set -- "$@" "$arg"
    fi
done
exec /usr/bin/flatpak run --branch=master --arch=x86_64 --command="putty" --file-forwarding uk.org.greenend.chiark.sgtatham.putty "$@"

The Flatpak command could then be extended to include subcommands for enabling or disabling specific aliases, or configuring whether new applications can automatically export aliases by default. Graphical management tools, such as system Settings or Flatseal, could integrate a dedicated control panel, similar to the one in Windows, to simplify alias management for users.

An example of what a similar alias management experience could be in a desktop Linux distribution
An example of what a similar alias management experience could be in a desktop Linux distribution

A Unified Path Forward for Desktop and CLI

As desktop Linux continues to gain mainstream attention, the differences between GUI applications and CLI utilities becomes increasingly important. While Flatpak laid the foundations to solve fragmentation and security challenges of graphical software, extending those same principles of sandboxing and distribution-agnostic delivery to command-line tools requires an implementation change.

Resources

Obelisk 0.42: Durable Agents, Layered Sandboxes

Lobsters
obeli.sk
2026-10-04 10:30:55
Comments...
Original Article

· 13 min read

An agent's progress should survive the process that runs it. Its code should run within a policy you can review, and its failures should leave a history you can inspect.

Obelisk keeps workflow progress in a database. Workflow code replays deterministically from that history, using recorded activity results before doing new work. A process can stop; the next one reconstructs where it left off. The same history lets you see what ran and debug it afterward.

0.42 builds on that foundation for agentic workloads: reviewable security boundaries for generated code, native V8 for faster JavaScript replay, Linux VM activities for tools that need them, and a workflow-agent prototype that can deploy, test, and fix applications on another Obelisk instance.

The idea follows SQLite is All You Need for Durable Workflows : keep durable state close to the runtime, and let compute come and go. An agent waiting for a model, a tool, or a person can be represented by rows in the database, without a running VM per session.

Security

Three files, three owners

In 0.41 the operator's policy lived in server.toml and the application lived in deployment.toml . That left one file doing two jobs: server.toml described both the platform (listeners, database, resource limits) and what a particular app was allowed to do. 0.42 splits it:

Permissions narrow from server to app to deployment Nested boxes represent permission scopes, not separate processes. The platform admin configures the server; the app admin grants capabilities; developers and agents deploy code within those grants. Alternative execution boundaries are Wasmtime for Boa JavaScript, Rust, or Bochs VM activities; optional native V8 isolates; and Linux VM activities using QEMU or Firecracker, with KVM or emulation. Exec activities, marked with an asterisk and a dashed outline, run host processes without a sandbox and require approval from both admins. server.toml Platform admin Resource limits, runtime configuration, host execution gate app.toml App admin Grants access to secrets, outbound hosts, and approved executables deployment.toml Developers, agents Code can request capabilities within those grants Wasmtime V8 Linux VM Exec activities* Nested boxes show permission scopes; bottom row shows execution options. * Host processes without a sandbox; approval required from both admins.

The effective permission is the intersection. A deployment can request less than app.toml grants, never more, and app.toml cannot switch on exec activities unless server.toml allows it.

An agent can rewrite code and deployment.toml within those grants. A request for broader access requires a change to app.toml , which gives the reviewer a small, explicit policy diff: which secrets the code can use, which hosts it can call, and which native executables it can run.

# app.toml
app_name = "my-app"

[secrets]
OPENAI_KEY = {}

[[outbound_http.allowed_host]]
pattern = "api.openai.com"
methods = ["POST"]
request_url_regex = "^POST https://api\\.openai\\.com/v1/"
secrets = ["OPENAI_KEY"]
replace_in = ["headers"]

obelisk deployment verify reports missing policy entries; --fix can scaffold them for review. Every deployment records the digest of the app policy it was activated under, so that boundary is part of its inspectable history.

Secrets stay at the network edge

Secrets still default to placeholders that the runtime replaces at the network edge, so component code never sees the value. Some code legitimately needs the plaintext, such as a webhook verifying an HMAC signature. In 0.42, WASM and JavaScript activities, webhooks, exec activities, and VM activities can request that with exposed_secrets .

Each exposure requires a grant in app.toml bound to a digest of the component and its complete set of exposed secrets. If the agent changes the component or asks for one more secret, the digest changes and the grant no longer applies. The app admin must approve the new digest before the runtime exposes those secrets.

Exec activities require both admins' approval

Existing exec activities run host processes outside the sandbox. In 0.42, both the platform admin and the app admin must approve them: the platform permits exec in server.toml , and the app grants access in app.toml .

The Security Model explains the grants and approval digests in detail.

Native V8

JavaScript workflows, activities, and webhooks can now run on native V8 instead of Boa compiled to WASM: start the server with OBELISK_JS_RUNTIME=v8 . Each activity gets a fresh isolate. Components using the 0.42 JavaScript API need no changes to switch engines. Boa remains the default engine.

Concurrency and memory are bounded per workload and runtime by [limits] in server.toml , so the platform admin can cap V8 activities, WASM workflows, and VM activities independently.

Faster replay means less time reconstructing a session before it can resume. Our durable coding-agent prototype, workflow-agent , has separate JavaScript and Rust workflow implementations. We replayed the same real agent conversation with both, comparing V8, Boa WASM, and Rust in Wasmtime:

Workflow-agent replay time Workflow replay time over nine measured calls. boa-wasm: median 3171 ms, range 3122–3255 ms; rust-wasm: median 106 ms, range 105–117 ms; v8: median 128 ms, range 123–140 ms. JavaScript / V8 123–140 ms Rust / Wasmtime 105–117 ms JavaScript / Boa WASM 3.1–3.3 s Workflow-agent replay: 726 events scale: 0–4 s Median bars; labels show observed ranges data_file using ($2/2):1:(0):2:($1-0.32):($1+0.32)

For this 726-event conversation, median replay fell from 3.17 seconds on Boa WASM to 128 milliseconds on V8, about 25× faster. Nine replays after warmup on the same Intel i9-14900HX host. Bars show medians; labels show observed ranges. Replay time excludes database loading and varies by workload. The replay measurements include every sample and the benchmark setup.

VM activities (experimental)

For activities that need a real Linux userspace, 0.42 adds [[activity_vm]] : a script that runs inside a Linux VM, with its tools supplied as Nix store paths that are verified and mounted read-only.

Guest HTTP goes through the same app and deployment policy as every other component, including secret placeholders, so a VM activity cannot reach a host that a JavaScript activity could not.

Choose a backend by setting OBELISK_UNSTABLE_ACTIVITY_VM for both the CLI and the server:

  • bochs-wasm : the Bochs x86 emulator compiled to WASM, running inside Wasmtime. A Linux VM inside the WASM sandbox, with no host binaries required. It is the slowest option and has a fixed 512 MiB guest.
  • qemu-tcg and qemu-kvm : native QEMU, with or without KVM, up to 16.25 GiB of guest RAM.
  • firecracker : a Firecracker microVM, cold booted for each execution; needs /dev/kvm .

What does the VM layer cost? We measured from the activity's persisted Locked event to its Finished event using the published Obelisk 0.42.0 binary and published VM runtimes, including QEMU's 2026-10-01 EROFS bundles. The small cases print a string with Bash or use curl to fetch a local page through Obelisk's HTTP bridge. The larger case is the inception Playwright demo : it launches Chromium, opens trynix.dev with Playwright, boots Obelisk in the page's Linux VM, runs obelisk -v , and returns the command output as the activity result. Its timing includes that in-browser boot. Each bar is a median in seconds; the scales differ between workloads.

Activity latency across VM backends Median seconds from Locked to Finished. Bash printf: Firecracker 0.155, QEMU KVM 0.144, QEMU TCG 0.431, Bochs WASM 0.966. Curl to local page: Firecracker 0.180, QEMU KVM 0.187, QEMU TCG 0.694, Bochs WASM 2.342. Chromium: trynix.dev → obelisk -v: Firecracker 13.713, QEMU KVM 13.124. Bash printf scale: 0–1 s Firecracker 0.155 s QEMU KVM 0.144 s QEMU TCG 0.431 s Bochs WASM 0.966 s Curl to local page scale: 0–2.5 s Firecracker 0.180 s QEMU KVM 0.187 s QEMU TCG 0.694 s Bochs WASM 2.342 s Chromium: trynix.dev → obelisk -v scale: 0–20 s Firecracker 13.713 s QEMU KVM 13.124 s

All timings came from the same Intel i9-14900HX host. The VM runs were sequential, with cached runtime images and warmup runs. Bash printf and curl use 512 MiB and one guest vCPU; inception uses 8 GiB and four. Bash and curl each have seven measured runs after two warmups; Chromium has three after one warmup. The benchmark notes include the matching CSV measurements, commands, and pinned source and runtime versions.

See the VM activity reference for configuration. The guest ABI, configuration, and behavior are experimental.

Agentic workflows

There are two ways to bring agentic workloads to Obelisk. Let a coding agent such as Claude Code or Codex generate application code, then deploy it within the application's security policy. Generated applications that do not call a model consume no further LLM tokens during execution. This follows the idea in Kelsey Hightower's Zero Token Architecture talk at PlatformCon 2026 : use the model to build the application, then run the resulting code.

Or run the agent itself as a durable workflow, with model calls and tools as activities. This is useful for enterprise agents that wait on people or external systems, and for coding agents that work inside a simulated Bash session with a persistent virtual filesystem.

Enterprise agents inherit parent/child agent hierarchies, durable scheduling, and pause/resume from the runtime. Hierarchical cancellation requires every workflow on the cancellation path to be explicitly marked with the -cancellable export suffix. Cancelled workflows do not run their own cleanup handlers; see Structured Concurrency for cleanup ownership. Obelisk can transparently unload inactive sessions and reconstruct them from recorded history when work resumes. The same replay mechanism recovers their progress after a server restart. These capabilities come with the workflow runtime.

A coding agent that can inspect what it ships

workflow-agent is our prototype of that second path: a browser UI and a durable agent loop. For coding agents, its cheap virtual workspaces are just-bash sessions with a persistent virtual filesystem and deeply integrated Obelisk and MCP commands.

It also exposes a simulated obelisk CLI that can connect to a separate target Obelisk instance. The agent can read the target's deployment into its virtual filesystem, edit application code, apply the deployment, and test it by calling functions and webhooks. It can then inspect execution history, application logs, and recorded HTTP traces to diagnose failures, fix the code, and deploy again. That introspection closes the loop between writing an app and checking how it actually runs.

The aim is thousands of concurrent sessions without an external VM per chat.

workflow-agent showing the prompt to build a GitHub contribution app, thinking bubbles, and Bash tool calls

These agentic workflows also shape Obelisk's APIs. Earlier versions copied the growing conversation into every LLM activity, ballooning workflow state and stored data. Now the workflow keeps the latest reply, while the activity fetches previous messages in a batch. That prompted the new batch events API , which reads just the create and finish events of child executions. The workflow-agent architecture explains the design.

For a smaller starting point, demo-agent provides a JavaScript agent loop, an LLM activity, example tools, a human-in-the-loop question, and a polling UI. Its mock deployment runs without an LLM key.

Web UI: themes and new screens

The Web UI gets light and dark themes, a refreshed layout, deployment graphs, and new screens for system events and retention. Application logs now have level and stream filters. It uses the REST API and ships in every 0.42 release binary, served at http://localhost:8080 by default.

The updated Web UI showing a deployment's components and their connections

Also in 0.42

  • obelisk generate new creates a JavaScript starter app.
  • Retention and garbage collection run automatically (30 days by default) and can be managed with obelisk admin ; system events are persisted.
  • Compatible JavaScript and Rust workflow implementations can replay the same execution log, allowing a language switch mid-execution.
  • execution submit --follow-logs , SSE streams for follow=true , and a batch events endpoint.
  • Experimental WASIp3 support for WASM activities and webhooks.
  • gRPC and gRPC-web are deprecated; the REST /v1 API covers everything they did.

Upgrading

Obelisk is no longer published to crates.io, so cargo install obelisk and cargo binstall obelisk no longer receive new versions. Choose a supported channel from the installation guide .

This release breaks configuration, the JavaScript runtime API, WIT packages, and a few API endpoints. Split your configuration, name the app, import obelisk:workflow@1.0.0 instead of using the global obelisk object, rename activity_exec.secrets to exposed_secrets and generate the grants, and review the new [limits] defaults. deployment get is now deployment pull .

If you use the default SQLite directory, pin the existing path or move the database before restarting.

The Migrating to 0.42 guide covers each step, and the Security Model page describes how the layers fit together. The configuration reference is split the same way as the files: server.toml , app.toml , and deployment.toml .

Full Changelog

See CHANGELOG.md for every change.

How effective altruism conquered the world (and might yet end it)

Hacker News
www.economist.com
2026-10-04 09:30:49
Comments...

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

Hacker News
github.com
2026-10-04 08:51:53
Comments...
Original Article

English · 简体中文 · 日本語 · Deutsch · Français · Español · Português

Run a 125-billion-parameter AI model on your own gaming PC
NVIDIA or AMD graphics card (12 GB or more) · Windows or Linux · free and open source

A voxel pagoda garden that Strata's model wrote, running in the browser
A voxel pagoda garden, 1 shot prompt running on an RTX 5070 with Strata (IQ3_S, 128K context) · full video (49 s)

Strata runs Qwen3.8-Flash-Next on a normal PC. This is a large, smart AI model that usually needs a server. It chats, writes code, reads pictures and works with your apps and coding agents. Nothing leaves your PC.

How fast is it?

We measured it on two ordinary gaming PCs. A token is about ¾ of a word.

  • Writes answers: how fast the reply appears in a short chat. 60 tokens per second is faster than you can read.
  • Reads your prompt: how fast it takes in what you send (here a 32K-token document, code or chat history).
NVIDIA: RTX 5070 (12 GB), Ryzen 5 7600, 64 GB RAM AMD: RX 9070 XT (16 GB), Ryzen 9 3900X, 47 GB RAM
Size Writes answers Reads your prompt
Q2_0 94 tokens/s 2,650 tokens/s
IQ2_XS 79 tokens/s 2,090 tokens/s
IQ3_XXS 62 tokens/s 1,750 tokens/s
IQ3_S 53 tokens/s 1,620 tokens/s
Coder 55 tokens/s 2,180 tokens/s
Size Writes answers Reads your prompt
Q2_0 60 tokens/s 1,160 tokens/s
IQ2_XS 52 tokens/s 1,110 tokens/s
Coder 44 tokens/s 1,420 tokens/s

NVIDIA: Q2_0 with engine 0.1.36, the other rows with 0.1.26 (4K answers, 32K prompts). The full tables are in DETAILS.md . A card with more VRAM is faster: an RTX 3090 (24 GB) should write about 100-140 tokens per second. Long chats and other cards: speed of each model , community results .

Buy Me A Coffee
Strata is free. If it runs well on your PC, a coffee keeps the work on it going.

What you need

Graphics card NVIDIA GeForce RTX 20, 30, 40 or 50 series, or AMD Radeon RX 7900 XT / XTX, RX 7800 XT / 7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700 or RX 6800 / 6900 series. It needs 12 GB of VRAM or more .
RAM 32 GB or more. Your RAM decides which model fits. 64 GB runs every size.
Disk About 80 GB free. Use an SSD if you can: the first start is much faster.
System Windows 10 / 11 or Linux, and a current graphics driver from NVIDIA or AMD.

The installer sets up everything else. Two or three cards can share the model ( multi-GPU ).

Experimental, written and tested by community members on their own machines:

  • Older graphics cards (Tesla P40 / V100, GTX 10, Radeon VII / MI50, RX 6700 XT, RX 5500 XT): Older GPUs .
  • Intel Arc , built from source on Linux: Intel Arc .
  • Older processors without AVX2 : they work, but slowly. Older CPUs .

The full list: docs/INSTALL.md .

Install

Let your AI set it up

Do you use an AI coding assistant (Claude Code, Cursor, Codex, GitHub Copilot, ...)? Paste this into it:

Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.

It checks your graphics card, RAM and disk and picks the model that fits. Then it installs and starts it and tells you how to connect your apps. AI tools can also install, start and stop Strata through its MCP server .

Or do it yourself

Download Strata and unzip it (or git clone it). Windows: double-click START-HERE.bat . Linux: run ./setup.sh in the Strata folder.

The steps are the same for NVIDIA and AMD. The installer finds your card and sets up the right engine for it. It asks you a few questions:

  • which model and which size,
  • how much context (how much text the model keeps in mind),
  • whether it should read pictures.

Press Enter each time for the recommended answer. Then it downloads the model (about 70 GB) and starts it. If the download stops, run it again: it continues where it left off. Your browser opens the Strata app at http://127.0.0.1:8080 .

While the model starts, your PC can be slow or stop responding for 1-3 minutes (longest the first time). Strata loads 35-55 GB into your RAM and locks part of it for the graphics card. This is normal. Wait, and don't close the window. The window shows what Strata is doing.

Next time , run START-HERE.bat (or ./setup.sh ) again. It starts right away and downloads nothing twice. Close its window to stop the model. UPDATE.bat ( ./update.sh ) updates Strata without starting it. Updating, Docker, several cards, where the files go and every option: docs/INSTALL.md .

Which model should I pick?

The installer recommends one for your RAM. The same model comes in several sizes, compressed more or less. Smaller sizes are faster. Larger sizes are a bit smarter.

Your RAM Take Why
32 GB Coder it fits 32 GB, and it is made for code (with a 24 GB card, Q2_0 and IQ2_XS run too)
48 GB IQ2_XS (or Q2_0, the fastest) the larger sizes do not fit
64 GB IQ2_XS (recommended), or IQ3_XXS / IQ3_S every size fits; IQ3_S is the best and the slowest
96 GB or more IQ3_S , or Unsloth's UD-IQ4_XS (~4-bit) room for the largest sizes with everything else open
  • Coder : a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM. It is weaker outside code, including Chinese and other CJK text (#438). For those, take Q2_0, IQ2_XS or IQ3_S, which keep every expert.
  • Swift 1.5 : a fine-tune that thinks for a much shorter time before it answers. You get the answer sooner, at about the same quality.
  • Unsloth UD-IQ4_XS : Unsloth's ~4-bit version, between IQ3_S and UD-Q4_K_XL in quality. A 94 GB download. With less than ~80 GB of RAM, Strata reads part of it from the SSD while it answers, so it is slower there (an NVMe SSD helps).
  • Unsloth UD-Q4_K_XL (experimental): the closest to the full model. But Strata reads most of it from the SSD while it answers, so it writes only 7-8.5 tokens/s on a 64 GB PC.
  • OrcaRouter's Uncensored IQ3_XXS : you set it up by hand. It is not in the installer's menu.

Sizes, downloads and what fits where: docs/MODELS.md . To add another model later, run SETUP.bat (Linux: ./setup.sh --setup ).

Using it

The Strata app's Monitor tab next to a coding agent
The Strata app's Monitor (left) while a coding agent writes the pagoda garden from the video (right)

  • In the browser: open http://127.0.0.1:8080 . It has Chat , a live Monitor of the model and your GPU/CPU/RAM, and About with the settings and addresses.
  • Your apps and coding agents: add an "OpenAI-compatible" provider with the base URL http://127.0.0.1:8080/v1 . Any API key and any model name work.
    • Apps that use Anthropic's API: http://127.0.0.1:8080/v1/messages (Claude Code: ANTHROPIC_BASE_URL=http://127.0.0.1:8080 ).
    • Codex CLI and other apps that use the OpenAI Responses API: /v1/responses ( setup ).
  • Thinking: choose off, low, medium or high in the chat menu or in your app's "reasoning effort". Off is the fastest. High is best for hard questions.
  • Pictures: say yes to "Images?" in setup. Then click Picture in the chat, or attach pictures in your app. AMD cards read pictures on Linux through the processor; on Windows they can't yet.
  • From your phone or another PC: START-HERE.bat --setup --host 0.0.0.0 --api-key <secret> . Always set a key.
  • One request at a time: by default Strata answers one request, and the others wait. To answer several at once, set "parallel": 2 ( BATCHING.md ). On a 12 GB card this makes each answer slower.
  • Long prompts: Strata reads the first message of a chat in full, about 1 minute per 30,000 tokens. Follow-up messages start in seconds.

More: where your chats are stored , the API .

Something went wrong?

  • My PC froze the first time Strata started. This is normal while it loads the model. Wait, and don't close the window. Still frozen after 10 minutes? Restart the PC, close other programs and try again, or pick a smaller size.
  • It stopped while downloading or installing. Run START-HERE.bat (or ./setup.sh ) again. It continues where it stopped.
  • It's very slow and the disk light keeps blinking, or it says "the engine stopped unexpectedly". Your PC does not have enough free RAM. Close other programs (browsers use a lot), or pick a smaller size (Q2_0 or IQ2_XS).
  • It says port 8080 is already in use. Strata is already running. Look for its window.

More problems and their fixes: docs/TROUBLESHOOTING.md . Still stuck? Open an issue and attach strata-<model>.log from the Strata folder. Found a security problem? Report it privately: SECURITY.md .

How does it work?

Models like this one usually run on servers with hundreds of gigabytes of graphics memory. Your graphics card has 12-24 GB. Strata makes the model fit by sharing the work across your whole PC . Think of a kitchen: the things you use all the time stay on the counter, and the rest waits in the pantry.

The model's 24,576 experts: the busiest on the graphics card, all of them in RAM, a lookup table on the SSD

  • The model is a team of 24,576 small specialists ("experts"). Each word needs only 10 of them.
  • Your graphics card keeps the few thousand experts that are used most often. Your RAM holds all of them, and your processor works on the rest at the same time. Your SSD holds a big lookup table.

A small helper guesses the next words; the big model checks them all at once and keeps the right ones

  • Guess, then check: a small helper guesses the next few words. The big model checks them all at once. You get the same answer, 1.6-1.8x sooner.
  • Long texts are read in big pieces (up to 8,192 tokens at a time), at over 1,000 tokens per second.

The longer explanation: docs/HOW_IT_WORKS.md . Every part and its numbers: the details and the paper .

Credits and license

The model is Qwen3.8-Flash-Next by the Qwen team. It was compressed by ISTA-DASLab , UkisAI (Swift 1.5) and Unsloth. Strata uses parts of llama.cpp / ggml . All credits: docs/HOW_IT_WORKS.md . Strata is open source under the MIT License . A few parts and every model have their own licenses ( which ones ).

Support Strata

Strata is free and open source. If it is useful to you, you can support its development:

Buy Me A Coffee

Improving and Stabilizing the racoon2 IKE Daemon in NetBSD

Lobsters
blog.netbsd.org
2026-10-04 08:46:13
Comments...
Original Article

Google Summer of Code 2026 Reports: Improving and Stabilizing the racoon2 IKE Daemon in NetBSD

October 03, 2026 posted by Leonardo Taccari

This report was written by Artem Belan as part of Google Summer of Code 2026.

Introduction & Description

racoon2 is a system to exchange and install security parameters for IPsec. It consists of an IKEv1/IKEv2 key exchange daemon, a security policy management daemon, and a Kerberos-based key exchange daemon. The project lives at zoulasc/racoon2 and is being modernized so that it can serve as a reliable L2TP/IPsec or IKEv2 VPN server on NetBSD and Linux for built-in Windows, iOS and Android VPN clients.

The goal of this GSoC project was to work through that TODO list: implement NAT traversal properly for both IKEv1 and IKEv2 according to the relevant RFCs, make IPv6 works well so racoon2 can be tested with both IPv4 IPv6 properly, clean up the address-macro mess, and build a unit test framework so that the fixes stay fixed. In the end, the daemon can sit behind a NAT device and serve modern VPN clients without any workaround selectors in its configuration, IPv6 works out of the box, and the changes are protected by an automated test suite.

All work was done on the gsoc2026 branch and merged upstream through pull requests #13–#39 between June and September 2026.

What was done

IKE fragmentation (RFC 7383)

The fragmentation code was, in the words of the TODO , "old/incomplete": reassembled messages could be misrecognized by the daemon, and oversized packets could crash it. The result of the project is reliable RFC 7383 fragmentation for both IKEv1/IKEv2 on both the sending and the receiving side: large IKE messages are split and put back together correctly, packets that exceed the allowed fragment size are rejected with a clear log message instead of a crash, and exchanging big messages — certificates, large proposals — between the peers no longer silently fails or takes the daemon down.

NAT-OA payloads in IKEv1 Quick Mode (RFC 3947)

At the beginning of the summer the daemon ignored NAT original address payloads on input and never sent them to the peer. The consequence was that the kernel had no idea which addresses the client actually used behind the NAT, so it had to recompute checksums over entire packets — slow, and not what RFC 3947 prescribes. Now the NAT-OA payloads match the RFC's structure, are constructed and sent in Quick Mode, and are parsed positionally on receipt (local address first, peer address second), with a missing payload from the peer tolerated rather than treated as an error. The original addresses are also passed down to the kernel so it can do incremental checksum fixup instead of full recomputation.

Address substitution in NAT-T transport mode (RFC 7296 §2.23.1)

This was the heart of the project and the direct answer to a long-standing TODO item: when a NAT device was in the path, the addresses the peer proposed in phase 2 did not match the responder's configured selectors, so responders needed extra selectors that were valid only on the initiator's side just to pass the selector check — configurations like transport_ike_natt.conf carried them as a workaround, and connections from real clients stayed fragile. The final behaviour follows RFC 7296 §2.23.1: when NAT-T is enabled and the addresses observed on the wire do not match the configured selectors, the observed addresses are substituted into the negotiated selectors, on both IKEv1 and IKEv2. The workaround selectors are no longer needed, and the daemon now properly handles NAT-T traffic on the dedicated UDP port instead of only the initial one.

Address macros in configurations

The wildcard address macros were, in the TODO 's words, "the cause of many configuration-related bugs": what worked in an IKEv1 configuration could silently misbehave in IKEv2, one macro behaved as a wildcard in some code paths and as "unknown, wait until we know" in others, and misconfigurations were ignored rather than reported. The project unified them into a single consistent wildcard notion that behaves the same way everywhere, works for both IKEv1 and IKEv2 configurations, and produces clear diagnostics when an address lookup or expansion fails, instead of failing silently and leaving the user to guess why the policy never got installed.

IPv6 support fixed and enabled by default

IPv6 support had existed in racoon2 for years, but it worked crooked: when the interface was configured to use all addresses, the daemon still behaved as if only IPv4 addresses existed — as though the interface were set to use IPv4 only — and prefix-length edge cases were broken, so in practice racoon2 could not be tested with IPv6 at all. During the project these bugs were fixed — address matching, interface address validation on Linux, prefix-length handling — and IPv6 became what a modern daemon should have been: used by default, with no configuration switch to remember. After the project, the same configuration files simply work with IPv6 addresses, and the TODO item about IPv6 is closed on the configuration side; only end-to-end testing with real IPv6 traffic remains.

Policy and proposal negotiation

Before the project, transport mode policies were generated from the generic SA endpoints rather than the addresses configured for the SA, tunnel mode policies could reference the wrong endpoints, and the daemon assembled the negotiated IPsec proposal itself instead of accepting what the peer actually proposed — a frequent source of failed negotiations with real clients, which all bring their own proposals. Now policies in both transport and tunnel modes are generated from the configured SA addresses, and the negotiated proposal is built from the peer's proposal, so interoperability with stock Windows, iOS and Android clients no longer depends on the local configuration guessing everything right.

A unit test framework for the IKE daemon

The project had no automated tests at all; every change had to be validated by reading code and running the daemon by hand. The last part of the summer went into changing that: a small self-contained test framework with no external dependencies is now integrated into the standard build and its tests run under sanitizers. On top of it, four new test programs cover exactly the logic this project introduced — NAT original address payload handling, traffic selector address substitution, wildcard address handling in configurations, and the address matching used when a peer is purged — and the pre-existing crypto self-test was fixed so the whole suite passes cleanly. This matters beyond the project itself: the framework makes it straightforward to add regression tests for the parts of the daemon that still have none, whoever tackles them next.

Documentation, samples and NEWS

Documentation and sample configurations still described the old defaults, so the manual contradicted the shipped behaviour: it still described the old IPv6 behaviour and the pre-substitution selector workarounds. The project synced the docs and samples with the new defaults — IPv6 used by default, fragmentation handled automatically, tunnel endpoints and wildcard addresses used consistently — and added a NEWS entry summarizing all of the GSoC 2026 work for users and downstream packagers.

How to Use

Everything described above is on the gsoc2026 branch of ssszcmawo/racoon2 (merged upstream into zoulasc/racoon2 ). A quick tour for anyone who wants to try it:

Build and install. The repository ships only the autotools sources, so the configure script has to be regenerated first — this is the same flow documented in doc/INSTALL and used by the project's CI:

autoreconf -fi
./configure
make
make install

You will need the usual developer toolchain (a C compiler, GNU make, autoconf, automake and libtool, plus flex and bison) and the OpenSSL headers; if you want the kinkd daemon you also need a Kerberos 5 library (MIT krb5 or Heimdal), and the optional packet-dump debug mode additionally wants libpcap. On Debian/Ubuntu the CI installs exactly automake libtool libssl-dev libkrb5-dev libpcap-dev flex bison before running the commands above. Useful configure options are --disable-iked / --disable-kinkd to skip a daemon, --with-krb5=<dir> for a non-standard Kerberos, and --enable-pcap for the packet dumps; --prefix=$NEWDIR moves the installation away from the default /usr/local .

Configure. Everything needed ships in the repository under samples/ . The .in files are templates that get their install paths filled in during installation, and the scenario fragments are plain files. For an L2TP/IPsec transport-mode server — the scenario this project was about — the flow is:

  1. Take samples/vals.conf.in as your vals.conf and set your environment in it: MY_IPADDRESS and PEERS_IPADDRESS (or the IP_ANY wildcard when the peer's address is not known in advance), MY_PUBLIC_IPADDRESS if you are behind NAT, and the shared-key settings. Generate the key itself with pskgen .

  2. Take samples/racoon2.conf.in as your racoon2.conf . It includes vals.conf and default.conf and defines the interfaces; the sample ships with NAT-T commented out, with instructions in place:

    interface
    {
        ike {
            MY_IP port 500;
    #       MY_IP port 4500;   # uncomment to enable NAT-T
        };
        ...
    };
    
  3. Uncomment the one include line for the scenario you want — for this case transport_ike.conf . That file holds the remote section (IKEv1 and IKEv2, pre-shared key), the selectors and the policy:

    selector ike_trans_sel_out {
        direction outbound;
        src "${MY_IPADDRESS}" port 1701;
        dst "${PEERS_IPADDRESS}" port any;
        upper_layer_protocol "udp";
        policy_index ike_trans_policy;
    };
    ...
    policy ike_trans_policy {
        action auto_ipsec;
        remote_index ike_trans_remote;
        ipsec_mode transport;
        ipsec_level require;
        ipsec_index { ipsec_esp; };
    };
    

    and, if you are behind NAT, its own commented include line pulls in transport_ike_natt.conf with the additional selectors for common private address ranges.

Two details worth noting: the wildcard address in vals.conf ( IP_ANY ) now behaves consistently, so a server can accept a peer whose address is unknown in advance; and IPv6 now works in the same files as is — put IPv6 addresses in and they are used, with no switch to flip.

The other samples follow the same pattern: tunnel_ike.conf and tunnel_ike_natt.conf for tunnel mode, transport_kink.conf and tunnel_kink.conf for Kerberos-based key exchange, and local-test.conf for trying two daemons against each other on one machine.

Run it. Start spmd first (it owns the policy channel to the kernel), then iked with your configuration — ./configure puts both in place, and samples/rc.d/ has service scripts for NetBSD-style systems. Once both are up, traffic matching a selector with action auto_ipsec (like the one above) triggers the negotiation automatically, so on a client you just connect with the built-in Windows, iOS or Android VPN client using the server address and the shared key. The daemon logs the NAT detection and the substitution it performs, which makes it easy to see what happened if a client does not come up.

Follow the work. All changes are in pull requests #13–#39 on the gsoc2026 branch, and the NEWS file has a single consolidated entry describing the user-visible changes. The top-level TODO file lists what is still open — see the "What remains" section below.

Challenges

  • Old, untested code. racoon2 has been around since the WIDE project days and had no tests. Every change had to be validated mostly by reading code, running the daemon by hand under a debugger, and watching the sanitizers — which is exactly why the test framework became part of the project rather than an afterthought.
  • RFC compliance vs. working configurations. Implementing address substitution meant first understanding why responders needed those extra initiator-only selectors before being able to remove the need for them, and getting the substitution to happen at the right point of the IKEv1 negotiation took several attempts.
  • Address macros are deceptively deep. The wildcard macros looked like a simple rename, but in some places they mean "the address is not known yet and the kernel policy must wait", while in others they do not. Unifying them without breaking either case required careful tracing through configuration parsing, selector matching and policy updates.
  • Build system mistakes. An optional-build patch for the daemons imported from elsewhere looked harmless but broke the build and had to be reverted — a good reminder to actually build the tree, not just read the diff.
  • Sanitizer-driven debugging. Building with sanitizers turned up memory-safety bugs in paths that had "worked" for years; fixing them without changing observable behaviour meant carefully tracking ownership of every duplicated buffer.

Results

  • All work merged upstream into the gsoc2026 branch through pull requests #13–#39.
  • Every item on the TODO list is addressed except racoon2ctl : phase 2 SA management (1), policy generation for IKEv1 (3), IPv6 (4), NAT-OA payloads (5), address substitution (6), fragmentation (7) and the address macro confusion (8) are done; only item 2, the graceful connect/disconnect control tool, remains open.
  • A self-contained unit test framework with four new test programs covering the NAT-OA, address-substitution and wildcard-address logic. The full suite — including the pre-existing crypto self-test — passes cleanly under sanitizers: 5 out of 5 test programs, 0 failures.
  • A new NEWS entry, and documentation and sample configurations matching the new defaults.

What remains

The following items are recorded in the top-level TODO file and are deliberately left for future work:

  • A racoon2ctl tool. There is still no way for an initiator to connect and disconnect gracefully: sending SIGTERM to the daemon does not delete the security associations or send a delete notification to the peer the way an IKEv2 connection would, and for IKEv1 this gap is especially noticeable. A small control tool (and the delete-on-shutdown behaviour behind it) would make racoon2 behave like a proper long-running service instead of something you kill and re-run.
  • Tests for the functionality that still has none. This project added tests only for the logic it introduced. The rest of the daemon — the phase 2 SA lifecycle, policy generation, rekeying, fragmentation paths, the configuration parser as a whole — is still only exercised by running it by hand. Extending the framework with tests for those areas is the natural next step, and the purge-related test added here is meant as a starting point for exactly that.
  • End-to-end IPv6 testing. IPv6 now works on the configuration side, but it still needs much more thorough testing in real scenarios on NetBSD and Linux: full IPv6 tunnels, NAT-T interactions, and long-running connections on both initiator and responder sides have not been exercised yet.

Conclusion

This project started as "work down the TODO list" and grew into a substantially more robust racoon2: NAT traversal now follows the relevant RFCs for both IKE versions, IPv6 actually works and is used by default, configurations behave consistently, and — perhaps most importantly for the long term — the daemon finally has an automated test suite running under sanitizers.

The open items are collected in the "What remains" section above. I believe the combination of RFC-level fixes and a regression test framework makes racoon2 meaningfully closer to being a dependable VPN server for real clients on NetBSD and Linux.

I'd like to thank my mentor, Christos Zoulas, for reviewing this stream of pull requests, for patiently explaining the history behind this project and quirks like the address-macro mess. Thanks also to the other mentors and students of GSoC 2026 for a pleasant summer of digging into an old codebase and making it a little less scary.

[ 1 comment ]

The complement of true is true, except when it's false

Lobsters
dryperspective.github.io
2026-10-04 08:30:26
Comments...
Original Article

I recently looked at P4313R1 , a standards proposal paper which adds a set of bitmask operations for enums, using a C++26 annotation to opt-in. The core idea being that, to use an example from the paper, given code like the below:

enum class [[=std::bitmask_type]] Permission {
  None    = 0,
  Read    = 1 << 0,
  Write   = 1 << 1,
  Execute = 1 << 2,
};

The [[=std::bitmask_type]] annotation would automatically imbue Permission with an accessible set of bitwise operations so that you as a user needn’t write them out yourself. This got me thinking about the murky underbelly of C++ integer operations. Your C++ compiler will happily compute bitwise operations for any integral type as a builtin operation. This extends to types which we don’t traditionally think of as integers, such as wchar_t , the UTF character types char8_t through char32_t , and bool . But slightly more happens here than meets the eye. Because when you attempt to perform an operation like a | b , and the type of a and b is an integral type smaller than int , the language does not operate on the bit patterns of a and b directly. They undergo integral promotion - they are promoted up to int as an intermediate state, the bit patterns of these two int s are combined, and the result is returned to you, still as an int . 1 Consider the below:

//Two shorts
constexpr short perm_A {1 << 0};
constexpr short perm_B {1 << 1};

//And the result type of running a bitwise operation on them is int
static_assert(std::same_as<decltype(perm_A | perm_B), int>);

Now, for the most part this is harmless - if you cast the above perm_A | perm_B back to short then the unnecessary bytes are truncated away and you’re left with a short which contains exactly the value it would have held if you’d combined perm_A and perm_B as short directly. And if you want to be certain that you are keeping your types consistent, and avoid the pernicious bugs of silent narrowing conversions, you can get into the habit of static_cast -ing the result of your bitwise operation to its original type. The important thing here however is that integral promotion is not optional. Unlike most other areas of C++ the developer doesn’t get a choice - your types will unavoidably be promoted for these operations.

The second part of what makes this a hazard for bool specifically is boolean conversion , a special case which does not truncate but instead explicitly converts zero to false and non-zero values to true . For these non-zero values, whatever bit pattern was previously stored is discarded and replaced with true , which has an integer value of 1. So if we take code like this:

int x{10};
bool b{static_cast<bool>(x)};

and look at the generated asm (in this case from unoptimised x86-64 gcc 16.2):

mov     DWORD PTR [rbp-4], 10   ;Store the value of 10
cmp     DWORD PTR [rbp-4], 0    ;Then compare to 0, set ZF if x is zero
setne   al                      ;Write 1 to AL if ZF is clear
mov     BYTE PTR [rbp-5], al    ;Store the result

The standard (specifically [conv.bool] ) bases converting to bool on the only possible values which bool can hold - true and false , regardless of whatever bit pattern may have originally been used to create them.

Putting it together

With all of that covered, let’s talk through what happens when you try to evaluate ~true and then cast the result to a bool :

  • The value true is promoted to an int with a value of 1 .
  • The complement of 1 as an int is calculated as -2 , as two’s complement behaviour is required as of C++20.
  • The value of -2 is then cast back down to bool , undergoes boolean conversion, and since -2 is non-zero, becomes true .

There you have it, the complement of true is true , or spelled in C++ static_cast<bool>(~true) == true .

This brings us back to enums. An enum is permitted to use any integral type as its underlying type, including our good friend bool . So let’s define one:

enum class [[=std::bitmask_type]] boolean : bool{
    FALSE,
    TRUE,
};

Let’s also look at the bitwise complement operator as laid out in P4313R1:

template<bitmask-like T> constexpr T operator~ (T lhs) noexcept {
  return static_cast<T>(~to_underlying(lhs));
}

By now the workings of that operator should be a familiar shape - first we convert the enum to its underlying type (in our case bool ), then integral promotion is applied if that type is lower ranked than int , then we perform the bitwise operation, then we cast back to the enum type. As we would expect:

static_assert(~boolean::TRUE == boolean::TRUE);

Godbolt here .

This has all the potential to be a slightly confusing corner case in the language.

There is one exceptional case here which compounds the problem. If we try to run that same example in gcc, we get a different result:

static_assert(~boolean::TRUE == boolean::FALSE);

Godbolt here .

So what is going on with gcc? If we move the operation to runtime by defining two functions which depend on it as a runtime value:

enum class boolean : bool{
    FALSE,
    TRUE,
};

void f(int x, bool& out) {
    out = static_cast<bool>(x);
}

void g(int x, boolean& out) {
    out = static_cast<boolean>(x);
}

and look at the asm (again unoptimised x86-64), then we get:

"f(int, bool&)":
        push    rbp
        mov     rbp, rsp
        mov     DWORD PTR [rbp-4], edi
        mov     QWORD PTR [rbp-16], rsi
        cmp     DWORD PTR [rbp-4], 0
        setne   dl                          ;same setne pattern from [conv.bool] earlier
        mov     rax, QWORD PTR [rbp-16]
        mov     BYTE PTR [rax], dl
        nop
        pop     rbp
        ret
"g(int, boolean&)":
        push    rbp
        mov     rbp, rsp
        mov     DWORD PTR [rbp-4], edi
        mov     QWORD PTR [rbp-16], rsi
        mov     eax, DWORD PTR [rbp-4]
        and     eax, 1                      ;store the low bit in eax
        mov     rdx, QWORD PTR [rbp-16]
        mov     BYTE PTR [rdx], al          ;and move al to the out-param
        nop
        pop     rbp
        ret

What we see is that when the type is not spelled bool , gcc will truncate down to the lowest bit rather than perform a boolean conversion. Any even value, when cast down, will produce boolean::FALSE , and any odd one will produce boolean::TRUE .

This is unique to enum types specifically, gcc generally performs consistently when handling bool directly:

constexpr boolean b{boolean::TRUE};
static_assert(static_cast<bool>(~b) == false);

static_assert(static_cast<bool>(~true) == true);

static_assert(std::to_underlying(~b) == false);

static_assert(static_cast<bool>(~std::to_underlying(b)) == true);

In contrast, Clang and MSVC will evaluate the result of all four of these complement-and-cast combinations as true , which is what the standard says they must be. Full godbolt comparison here . This seems to be a fairly long-lived conformance bug in gcc.

But there is one place in the standard where integral promotion rules go further and introduce a new vector for undefined behaviour to enter into your program when doing these bitmask operations. Consider this code:

enum nums{
    zero,
    one,
    two,
    three,
};

//This is implementation defined but holds on gcc and Clang.
static_assert(std::is_same_v<std::underlying_type_t<nums>, unsigned int>);

//nums uses an unsigned int as its base, therefore the entire domain of representable values is >= 0.
//So let's test this:
static_assert(~one >= 0); //FAILS

Godbolt here

Looking at the diagnostic, gcc gives us:

<source>:15:20: error: static assertion failed
   15 | static_assert(~one >= 0); //FAILS
      |               ~~~~~^~~~
  • the comparison reduces to '(-2 >= 0)'

So what happened here? Well, the passage in the standard which covers integral promotion, [conv.prom] , carves out a special bullet for unscoped enumeration types with no fixed underlying type to behave differently from integer types when promoting. If the entire range of values can be stored in an int , then that is the type which it promotes to, then it tries unsigned int , and if that fails it repeats this signed-then-unsigned pattern for long and long long until it finds a suitable type. But, where this differs from plain integer types is that it only considers the effective range of representable values; and so an enum backed by unsigned int will promote to int regardless of the normal conversion ranking which would forbid it for integral types. As before, the fact that you are calling the builtin operator via ~one will unavoidably promote it to int , its complement is calculated as -2 , and then the comparison is performed between two int s, and fails. If instead you convert to the underlying type first then you really see the asymmetry - static_assert(~std::to_underlying(one) >= 0); succeeds, and is so vacuously true that gcc even warns about its redundancy.

But this is only half the battle, and we need to talk about converting the result of our bitwise operation back to the enum type, and this is where UB creeps in. An unscoped enum with no specified underlying type defines itself in terms of a valid range of values determined by its enumerators independently of the range of whatever actual type is used by the compiler to back it. [dcl.enum] tells us the value range of such an enum is the value range of a hypothetical integer type of width M, where M is the minimum bits required to represent all enumerators. To demonstrate, consider a few examples:

enum small{ //Range of enumerators is 0..1, so value range is that of an unsigned, one-bit integer.
    a = 0,
    b = 1,
};

enum medium{ //Range of enumerators is 0..7, so value range is that of an unsigned, three-bit integer
    c = 0,
    d = 4,
    e = 7
};

enum large{ //Range of enumerators is -1..100, so value range is that of a signed, eight-bit integer
    f = -1,
    g = 10,
    h = 100,
};

Even though your compiler will quite likely store these enums as unsigned int and int , only a subset of possible representable values is valid to use - that of the M-width integer. [expr.static.cast] in the standard makes it clear that it is undefined behaviour to cast a value outside of that range to the enum type. So, using the enumerators of the above example:

static_cast<small>(1); //Well-defined
static_cast<small>(2); //UB

static_cast<medium>(3); //Well-defined: while not a value of an enumerator, it is within the range of an unsigned, three-bit integer
static_cast<medium>(8); //UB

static_cast<large>(-128); //Well-defined
static_cast<large>(128); //UB

This is our danger zone. The bitwise operators, in particular operator~ , can produce values which are out of the well-defined value range for an enum , and user-defined operators which attempt to cast this value back down to the enum type can invoke UB.

But, I hear you cry, what about the example earlier with ~one ? We saw that give a value of -2 which is outside the value range of nums . This is true, but here is where integral promotion actually saves us - one was promoted to int before we calculated its complement, and it was never cast back down again to invoke our UB. We only get into dangerous territory with a user-defined operator which converts the result of the operation back down to the original enum type.

It is also important to note that this UB risk only applies to unscoped “plain” enums with no fixed underlying type. All scoped enum types have a fixed underlying type, even if they don’t specify it (in that case, int ); and all unscoped enum types which do specify a fixed underlying type inherit that type’s value range instead of inventing their own.

Avoiding this in your own code

Let’s say that, like P4313, you are adding bitmask operations to enums for your own code. Now that we have more weirdness in this area of C++, how would you go about making sure that your code is correct and unconfusing? Avoiding the bool case is simple - just constrain away that the underlying type can’t be bool , either for all enums or for the complement operator. But the standard doesn’t come with a handy type trait or reflection metafunction to specifically detect an unscoped enum with no fixed underlying type. Fortunately we can make one:

template<typename E>
concept unfixed_enum = std::is_enum_v<E> && !requires { E{0}; };

[dcl.init.list] only permits this initialization from a scalar for enums with a fixed underlying type, and 0 is the only possible value which is representable for all possible enums. To cover the edge cases, the standard defines an enum with no enumerators as having the value range of an unsigned, one-bit integer; and the thing which excludes 1 as a possible value is enum E{ a = -1 }; , which has a range [-1, 0] . Be sure not to forget the std::is_enum_v . After all, there are many types for which T{0} is ill-formed, and most of them are not enums.

This post has mostly been focused on the bitwise complement operator, but the UB trap of an unfixed enum also applies to the bitwise left shift operator. If you take the route of constraining per operation, don’t let it slip your mind.

This brings us back once again to P4313R1. At time of writing, the concept which constrains these operators only constrains the type to be some enumeration type annotated with [[=std::bitmask_type]] . As such, it will allow confusing behaviour on enum types which are backed by bool , and potential UB on plain enums with no fixed underlying type. If the authors don’t want to standardise the above init-list trick, they could constrain it on std::is_scoped_enum as a conservative approach to prevent users from having easy access to the UB, since all unscoped enums can use the builtin bitwise operators anyway (albeit returning int ). They could also constrain the underlying type to not be bool to minimise the confusion; particularly as the diagnostic which would normally get issued for complementing a bool in user code would be suppressed by default if it came from a system header. I do intend to contact the authors of P4313 about this to see if they want to add these constraints to their paper.

Aside: What about floating point?

You might be wondering how the standard defines conversions of floating point numbers to these bool -backed enum s. It is perhaps unsurprising - [expr.static.cast] requires that the behaviour is equivalent to first converting the floating point value to the underlying type of the enum, and then converting it to the enum type; pointing the user to [conv.fpint] which has an explicit note directing readers to our old friend [conv.bool]. Per the standard, this should mean that any floating point value other than exactly 0.0 or -0.0 would convert to boolean::TRUE , otherwise we end up back at boolean::FALSE .

So let’s start small:

enum class boolean : bool {
    FALSE,
    TRUE,
};

static_assert(static_cast<boolean>(0.0) == boolean::FALSE);
static_assert(static_cast<boolean>(0.5) == boolean::TRUE);
static_assert(static_cast<boolean>(2.5) == boolean::TRUE);

Godbolt here .

Both Clang and gcc reject this code. The errors are that (boolean)2.5e+0 is not a constant expression; and static_cast<boolean>(0.5) == boolean::TRUE is wrong, and is instead equal to boolean::FALSE . MSVC accepts the above code.

But let’s go further and inspect what actually gets stored there. We write a function to examine the bit pattern generated from these casts, then try some values:

void show(double d) {
    const boolean e = static_cast<boolean>(d);
    unsigned char byte {};
    std::memcpy(&byte, &e, 1);
    std::println("static_cast<boolean>({:7.1f})  stored byte {:3}   required {}",
               d, byte, (d != 0.0) ? 1 : 0);
}

int main() {
    show(0.0);
    show(0.5);
    show(1.0);
    show(2.5);
    show(3.0);
    show(256.0);
    show(-2.5);
}

Godbolt here .

The above code compiles without warning on all three compilers, but while MSVC again does the right thing in all cases, gcc and Clang’s behaviour is much more worrisome. Looking at the output, we see that gcc will always just truncate the value to an integer, then load that bit pattern into the resulting boolean . So 2.5 truncates to 2 and gives a boolean whose underlying bit pattern is 2 . Clang does the same thing on unoptimised builds, but when optimisation is turned on, the truncated value then goes through the [conv.bool] transformation as normal and produces the right answer for all non-zero values outside of the range (-1, 1).

This is more concerning than some funky bit patterns, however. The valid range of boolean is [0, 1]. It is UB to read objects with bit patterns outside of this range. But, I hear you ask, what happens if we take these boolean s and cast them back to bool , triggering a boolean conversion with no initial floating point state? MSVC again does the right thing; gcc doesn’t modify the bit patterns, leaving us with bool which are out of range of the type (and therefore UB to read); and Clang is where the fun happens again. Running the above code with a cast from e to a bool , and then memcpy-ing that bool into the unsigned char , we get this result for unoptimised builds:

0.0 0.5 1.0 2.5 3.0 256.0 -2.5
required 0 1 1 1 1 1 1
gcc 14 / 15 / 16 0 0 1 2 3 0 254
clang 19 / 20 / 21 0 0 1 0 1 0 0
clang 23.1.1 0 0 1 1 1 0 1

Versions of Clang before 23 appear to simply perform trunc(d) & 1 ; meaning that after truncation, odd numbers are true and even numbers are false . Clang 23 performs (trunc(d) % 256) != 0 , narrowing to a byte first. So 256.0 becomes FALSE and 257.0 becomes TRUE . Optimised builds act as before - truncating then doing the proper boolean conversion.

Ultimately this all comes to a rather absurd head. Consider the below code:

#include <print>

enum class boolean : bool {
    FALSE,
    TRUE,
};

//Force this to be runtime
boolean make(double d) {
    return static_cast<boolean>(d);
}

int main() {
    const boolean e = make(2.5);
    std::println("e == boolean::TRUE  : {}", e == boolean::TRUE);
    std::println("e == boolean::FALSE : {}", e == boolean::FALSE);

    const bool b = static_cast<bool>(e);
    std::println("b == true  : {}", b == true);
    std::println("b == false : {}", b == false);
    std::println("b + 0      : {}", b + 0);
}

Godbolt here .

MSVC again leads the pack in giving the correct answer. Clang gets it exactly wrong and thinks that e is FALSE and b is false . gcc takes things up a notch, by providing an instance of an enum which compares equal to all of its enumerators, and a bool which is neither true nor false , until you turn on the optimiser and get a bool which is both true and false .

And to hit all the obligatory targets of any conversation on floating point, infinity and NaN show as 0 and become FALSE as boolean and so cast to false as bool on gcc and Clang; and give 1 and therefore TRUE and true on MSVC.

In lighter news, Clang trunk seems to have fixed the above example to behave correctly, so some of the issues described in this section should be patched out soon.

Conclusion

What did we learn from all this? Perhaps that even sensible and uncontentious pure-library-level papers such as P4313 should be on their guard against a footgun from some exotic C++ edge case. Or perhaps, in more practical terms:

  • Integral promotion is unavoidable. If you think you are operating on an integer which is smaller than int , there’s a good chance it became an int right under your nose.
  • The bitwise complement of a bool is always true , even when wrapped in an enum. However you should not rely on this behaviour as gcc has a longstanding bug which gets this wrong (or right, if you prefer logic to C++).
  • Unscoped enumeration types with no fixed underlying type have a range of valid values which may be smaller than that of whatever type actually underlies them.

Bill Draper has died

Hacker News
www.nytimes.com
2026-10-04 08:21:35
Comments...
Original Article

Please enable JS and disable any ad blocker

Rejection Sensitivity in Gifted and Twice-Exceptional Children

Hacker News
teachyourkids.substack.com
2026-10-04 07:58:47
Comments...
Original Article

Rejection sensitivity is a heightened tendency to perceive, anticipate, and react intensely to rejection or criticism, including situations that are neutral or ambiguous. It exists on a spectrum. Most people experience some version of it. For gifted and twice-exceptional children, that baseline sensitivity is typically amplified by the same wiring that makes them intellectually intense: a nervous system built to detect and process signals quickly and deeply.

Rejection Sensitive Dysphoria, or RSD, is a more specific and more severe term. Psychiatrist William Dodson coined it to describe the extreme, almost physically painful emotional reaction some people, particularly those with ADHD, have to perceived rejection or a sense of falling short of their own or others’ standards. Dodson found that the large majority of adolescents and adults with ADHD he surveyed recognized this experience in themselves, and a third named it the most impairing symptom they had.

RSD is not a formal diagnosis, and researchers have not fully settled where rejection sensitivity ends and clinically significant RSD begins. That does not make the experience less real for a child living through it. It means parents should treat RSD as a useful, well-supported description of a pattern, not as a diagnostic label to apply on their own.

What This Is and Is Not

What this is: nervous system intensity combined with a child’s identity becoming fused to performance and belonging.

What this is not: manipulation, entitlement, weak character, or a result of bad parenting.

Rejection sensitivity shows up across several groups that commonly overlap in twice-exceptional children. Each has its own research base and its own account of why the sensitivity develops. A child rarely fits neatly into just one.

This is the most researched population. Dodson’s clinical work links RSD to ADHD’s core but often overlooked feature of emotional dysregulation. ADHD brains tend to be more reactive to both immediate reward and immediate threat, and social approval and social criticism land in exactly those categories. Working memory limits compound this: a child in the moment often cannot hold both “I got corrected” and “my teacher likes me overall” at the same time. The correction temporarily erases the bigger picture.

Qualitative research with autistic adults describes rejection sensitivity as profoundly overwhelming and exhausting, often accompanied by physical tension, pain, and reliving past rejections. One notable difference from the ADHD pattern is that some researchers suggest autistic individuals may have difficulty consciously processing social evaluative cues in the moment, meaning a child may not always recognize that what they are experiencing is a reaction to rejection, even while the physiological load of it is very real. This matters for parents because a strategy that depends on a child noticing and naming the feeling as it happens may need to be adapted.

Psychologist Elaine Aron’s research on sensory processing sensitivity describes a trait found in roughly 15 to 20 percent of the population. Highly sensitive people show increased emotional reactivity and a stronger response to criticism specifically. Aron’s own framework for supporting highly sensitive people includes reframing how they interpret perceived failures, rejections, and judgment from others, which is directly useful language for coaching both children and the parents raising them.

Kazimierz Dabrowski’s theory of overexcitabilities, further developed by researcher Linda Silverman, describes gifted individuals as having heightened, inborn sensitivity to stimuli across several domains, including emotional overexcitability marked by intense connectedness with others, deeply felt experience, and strong, well-differentiated feelings about the self. It remains the most widely used framework for describing emotional intensity in gifted children, and many clinicians and parents find it accurate to their child’s lived experience. The research support, though, is weaker than its popularity suggests: meta-analyses show only small, inconsistent differences between gifted and non-gifted samples, and the difference largely disappears when giftedness is measured by cognitive testing rather than by prior gifted identification, which raises the possibility that the pattern reflects labeling and environment as much as innate neurology. Some researchers argue the construct overlaps with the established personality trait of openness to experience and may not need to be gifted-specific at all. The practical takeaway for parents is unchanged: many gifted children report this kind of intensity, and the strategies in this guide still apply, even though the claim that it is a proven, distinct neurological feature of giftedness remains contested.

Twice-exceptional children, gifted alongside ADHD, autism, or another difference, commonly sit at the intersection of two or more of these frameworks at once. That layering is usually why their reactions feel more intense and less predictable than a single-cause explanation can account for.

Gifted children tend to fuse performance with identity very early. Feedback on an assignment is rarely just about the assignment. It becomes a referendum on whether the child is still who they believe themselves to be.

A peer doesn’t sit next to them, doesn’t laugh at their joke, or gets picked for something they wanted. What follows is rarely proportional: “no one likes me,” a wave of rage, or a full shutdown. ADHD brains, in particular, are wired to track social status and belonging cues intensely, and differences in reward and threat circuitry seem to make these signals land louder than they would for other children.

In boys especially, though not only boys, rejection sensitivity can present as irritability, dismissiveness, or an argumentative stance: “I didn’t even want that anyway.” This is protective armor. Underneath it is usually plain embarrassment.

Gifted children are used to being right, being fast, and being praised for their intelligence. When ADHD or another difference introduces sloppy mistakes, executive function gaps, or emotional impulsivity, the mismatch between self-image and lived experience creates distress: “I am supposed to be exceptional, so why did I mess up?” Because the identity is more invested, the shame lands harder.

A child’s cognitive age and emotional regulation age can be years apart; for instance, an eleven-year-old mind managing feelings with the regulation capacity of a six-year-old. That gap is where a great deal of the combustion happens.

When adults consistently praise a child’s intelligence rather than their process or effort, mistakes stop feeling like normal setbacks and start feeling existential. Parents who praise effort over intelligence are ahead of the curve here, but a child’s friends and classmates often have not had that same experience.

One of the clearest everyday expressions of this pattern is a child’s dread of being called into the principal’s office, a note home from a teacher, or any adult wanting to talk privately, even when nothing is wrong. A highly sensitive or gifted nervous system typically cannot distinguish a neutral request for a conversation from an actual reprimand, so it responds to both with the same alarm.

The underlying fear in these moments is usually not “I will get in trouble.” It is closer to “I will be misunderstood, reduced to one moment, or unfairly categorized.” That distinction changes how an adult should respond. Separating the message from the delivery, staying calm and low-affect, and explicitly saying, “This isn’t trouble; I want to understand what happened,” does more to defuse the reaction than reassurance about consequences ever will.

Gifted children tend to reach for the kind of deep, trust-based friendship researchers call the “sure shelter” years before same-age peers are looking for anything beyond a play partner. Some gifted girls as young as six or seven already hold conceptions of friendship that typical children do not develop until eleven or twelve. When a gifted or twice-exceptional child finally finds someone who reciprocates that intensity, the relief can tip into over-attachment. A school counselor working with gifted students has described this pattern plainly: these kids can become “thirsty” for friendship. Read that way, clinginess is less a character flaw and more a response to scarcity.

Every friendship includes small everyday frictions; researchers call these “slights.” How a child handles slights is one of the strongest predictors of whether a friendship lasts. A highly sensitive or gifted child’s heightened threat detection, the same mechanism behind the rest of this pattern, means an offhand comment or a friend’s momentary preference for someone else often gets processed as a much bigger signal than it was ever meant to be.

These children frequently have more advanced, adult-like ideas about fairness and group organization than their peers. When they try to direct play toward what seems logical or fair to them, and peers resist, they can get labeled bossy, snobbish, or rigid. A related but distinct thread is perfectionism: some children need control over how a game unfolds because unpredictability itself, not just losing, is what feels distressing.

Strategies below draw on both clinical guidance for 2e children and the friendship research above. None of these require diagnosing a specific condition to be useful.

Normalize the brain, not the behavior

Language such as “I think your brain has a fast alarm system when it feels judged” reduces shame without excusing the behavior that follows the alarm.

Separate performance from self

Explicitly teaching “your work isn’t you, it’s something you made” does not come naturally to most gifted children and needs to be said directly and often.

Pre-correct before feedback

A short framing sentence before any critique, such as “I’m going to give you some ideas to strengthen this; that doesn’t mean it’s bad,” protects the child’s nervous system before it activates rather than trying to calm it afterward.

Spot the spike

A simple model many 2e kids respond well to: trigger, story, body reaction, choice. Most gifted children enjoy this kind of meta-cognitive tracking and can approach it almost like a scientist studying their own reactions. For autistic children, remember that conscious awareness of the rejection response may be less reliable, so this tool may need to lean more on noticing body signals than on naming the social trigger itself.

Build tolerance through micro-repair

Rather than shielding a child from all correction, practice small feedback, immediate repair, and visible appreciation right afterward. This is slow, deliberate work to rewire what the child expects to happen after a mistake.

Broaden the friend pool

Because the one-friend pattern often comes from scarcity of like-minded peers, structured groups built around shared interests tend to work better than simply increasing exposure to same-age children. The goal isn’t to discourage the close friendship, but to make sure it isn’t the only source of connection.

Coach the slight-versus-rejection distinction directly

Naming this distinction out loud, and helping a child check their read of a situation before reacting, gives them a tool that plain reassurance does not.

Model repair after your own conflicts

Children absorb far more from watching a parent apologize, disagree, and reconnect with another adult than from being told that friendship conflict is survivable.

Some of what is described in this guide is temperament and environment, both of which shift with awareness and the strategies above. Some of it is closer to clinical anxiety or depression and needs support beyond what a parent or educator can offer alone. Consider bringing in an outside professional if a child shows significant avoidance of school or social situations, physical symptoms tied to social stress, school refusal, or any signs of hopelessness or self-harm. Not every case resolves through parenting changes alone, and it helps to say so plainly.

Picture books and early-elementary titles work well because the story does the teaching. Tween and teen titles work differently: they tend to sit in ambiguity rather than resolve into a tidy lesson, which fits the developmental stage better but usually means a parent needs to carry more of the discussion afterward.

  • Gossie & Gertie, by Olivier Dunrea, a gentle look at one friend being bossier and both learning to share control

  • A Bad Case of Stripes, by David Shannon, the fear of losing friends by not conforming

  • Something Else, by Kathryn Cave, wanting a friend exactly like yourself

  • Elephant & Piggie: Happy Pig Day, by Mo Willems, jealousy when a close friend makes a new friend

  • Teamwork Isn’t My Thing and I Don’t Like to Share, by Julia Cook, directly addresses bossy and controlling behavior with peers

  • Enemy Pie, by Derek Munson, misreading a peer as a rival

  • Pink Tiara Cookies for Three, by Maria Dismondy, jealousy and exclusion when a two-person friendship becomes three

  • Real Friends, by Shannon Hale, a graphic memoir showing the ambiguity and shifting loyalty of real friendships without a tidy moral

  • The Science of Friendship, by Tanita S. Davis, built around real research on friendship and stress, pairs naturally with this guide’s framing

  • Harriet the Spy, by Louise Fitzhugh, a classic portrait of an intense, socially awkward, gifted-type child

  • Skills for Rejection Sensitive Dysphoria: A Workbook for Tweens and Teens, a graphic novel workbook that directly teaches kids to recognize their own rejection alarm and use regulation skills, closer to a teaching tool than a mirror story

  • William Dodson, M.D., “Rejection Sensitive Dysphoria,” ADDitude Magazine and CHADD’s Attention magazine

  • Elaine Aron, Ph.D., The Highly Sensitive Person and hsperson.com

  • Linda Silverman, “Dabrowski’s Overexcitabilities: A Layman’s Explanation,” and the Davidson Institute’s gifted education resources

  • Dana Winkler and Alan Voight, meta-analysis on giftedness and overexcitability, published in Gifted Child Quarterly, 2016

  • Paula Olszewski-Kubilius and colleagues, meta-analysis on overexcitabilities and giftedness, published in Gifted Child Quarterly, 2026

  • Salvatore Mendaglio, “Overexcitabilities and Giftedness Research: A Call for a Paradigm Shift,” Journal for the Education of the Gifted, 2012

  • Julie MacEvoy, Ph.D., Boston College, research on friendship and response to social slights

  • Miraca Gross, research on friendship stage development in highly gifted children

  • van Asselt, Roke, Begeer, and Scheeren, “Feeling Constantly Kicked Down,” a 2025 qualitative study of rejection sensitivity in autistic adults

  • Davidson Institute for Talent Development, gifted-blog archive on peer relationships and overexcitability

Maya Sissoko is the founder of Whole Child Education, serving gifted, twice-exceptional, and neurodivergent children and families through parent consulting, educational planning, homeschool design, and one-on-one work with children. She brings decades of gifted education experience to families locally and worldwide through her Bay Area-based practice.

A Mere Mortal's Introduction to JIT Vulnerabilities in JavaScript Engines

Lobsters
trustfoundry.net
2026-10-04 07:53:51
Comments...
Original Article

Intro:

There are some people who seem to have a mysterious ability to grasp complex technical topics easily. You know the people I’m talking about; the ones who quickly whip up custom SMT solvers in Rust on the weekends, the ones who casually say things like “concolic execution” and “Huffman tree” without batting an eye. The ones who manage to somehow find use after frees in seemingly unrelated parts of enormous codebases. In other words, the Smart People.

I, on the other hand, am not one of the Smart People. I wouldn’t know where to begin with writing an SMT solver. I’ve read about concolic execution before and I still could barely tell you a thing about it if my life depended on it. I’m more in the category of People Who Try to Make Up For Their Deficiencies Through Sheer Force of Will, which doesn’t have quite the same ring to it.

However, I recently became interested in exploiting JavaScript Just in Time (JIT) compilers in browsers, which is a topic that seemed relegated only to the Smart People. I spent a long time delaying digging into this topic because I assumed it would be too difficult to make any progress. As it turns out, it’s an interesting topic that doesn’t need to be quite so intimidating, and I wish I’d committed to learning about it earlier.

Goal:

My goal with this blog post is to provide a clear, gentle introduction to the role of JIT compilers in JavaScript engines, some types of security vulnerabilities that might arise from the use of JIT compilers, and tooling to use to aid in analyzing these bugs. Hopefully, this will help others consider that this topic isn’t as impossibly difficult as it might appear. This is not intended to be a comprehensive guide to JIT exploitation; this blog is just trying to make a complex topic more approachable and friendly so that others feel confident enough to dive deeper.

Throughout this post, I’ll be using D8, the standalone version of V8, the JS engine used in Chrome. Much of the content we’ll be covering is not specific to V8, but as we get into specific details, we’ll focus on V8. The example vulnerability we’ll be covering is also from a V8 CTF challenge.

Knowledge prerequisites:

You don’t need to already be familiar with JIT compilers or JavaScript engines; this post doesn’t assume knowledge of those topics. It’s helpful if you:

-Are familiar with traditional memory corruption issues such as out-of-bounds reads and writes and type confusions.

-Are at least passingly familiar with JavaScript.

-Are comfortable using a debugger such as GDB.

-Build V8 from source so you can follow along with the examples.

Helpful existing work:

Some good news is that there’s already quite a lot of existing public research on JIT vulnerabilities and exploitation. The following resources are just a few of the many helpful ones out there; they are not prerequisites for reading this blog post by any means, but if you find that you’re interested in this topic, you’re encouraged to read them:

https://saelo.github.io/presentations/blackhat_us_18_attacking_client_side_jit_compilers.pdf

This slide deck introduces JavaScript engine and JIT compiler concepts and shows example JIT vulnerabilities.

https://doar-e.github.io/blog/2019/01/28/introduction-to-turbofan/

This post provides V8-specific information by walking through TurboFan, a JIT compiler in V8. It also shows usage of Turbolizer, a useful tool for visualizing TurboFan’s behavior.

This presentation discusses common attacks on JavaScript engines and covers typical exploitation strategies.

https://www.madstacks.dev/posts/V8-Exploitation-Series-Part-4/

This post (which is part of a larger series that’s also worth your time!) discusses TurboFan and includes some tips on how to analyze it yourself.

https://docs.google.com/presentation/d/1DJcWByz11jLoQyNhmOvkZSrkgcVhllIlCHmal1tGzaw/edit#slide=id.p

This presentation begins with some background on TurboFan and then dives into specific bug and exploitation examples.

https://www.zerodayinitiative.com/blog/2021/12/6/two-birds-with-one-stone-an-introduction-to-v8-and-jit-exploitation

This post (which is the first in a three-part series) provides some information about TurboFan and its optimization phases.

Some of these resources will be referenced more than once in this blog post.

Intro to JS engines and JIT compilers:

Let’s begin with a simplified overview of JavaScript engines and JIT compilers before we dive into the more technical details In a web browser, the JavaScript engine is the component responsible for executing the JavaScript code embedded in websites you visit. For example, let’s say you visit example.com and the website includes some JavaScript code; the JavaScript engine will run the code. JS engines are generally written in C or C++. Importantly, JS engines are frequently exploited components of browsers, since the engines are very complex and must run untrusted code by design.

A JIT compiler is a component of a JavaScript engine that’s responsible for making performance optimizations. To understand the role of a JIT compiler, let’s first consider how a JS engine operates without a JIT compiler. When a JS engine is provided with JS code to execute, the code will be parsed and converted to an abstract syntax tree, and then ultimately converted into bytecode which will be executed by the JS engine’s interpreter . We’ll call this the baseline interpreter. In V8, the baseline interpreter is called Ignition. At this stage, the JS bytecode can be considered unoptimized. It’s very possible for security vulnerabilities to occur during these stages, but we’re specifically interested in tackling JIT compilers in this post, so we won’t discuss baseline interpreter vulnerabilities here.

The baseline interpreter is actually sufficient to have a functional environment for executing JS; JIT compilers are just optional optimizing components, and aren’t strictly necessary. In fact, some JS engines that are not used in browsers may not feature JIT compilers at all. Modern browsers also allow users to disable JIT compilation if they wish. However, browser vendors are generally quite invested in making their browsers as performant as possible, and running JS via only a baseline interpreter is considered slow. Simply removing all JIT compilers would probably be better for security, but vendors understandably don’t want to make their products slower; like it or not, JIT compilers in browsers are probably not going to all be removed any time soon.

To improve performance, when JS is being run by a baseline interpreter, JIT compilers look for portions of JS code that are being run frequently. In JIT terminology, such portions are referred to as hot . Imagine JS code as machinery that gets hotter and hotter the more it’s used. Once a usage threshold is reached, JIT compilers will consider some JS (for example, a function) to be hot and will strive to create a more optimized version of this code by emitting a machine code (that is, assembly) version of the hot JS function, which will be much faster to execute than the bytecode version that is run by the baseline interpreter. In future usage of the code, the optimized JIT compiled version will be used instead of the bytecode version. However, this optimization process may introduce vulnerabilities, and we’ll take a look at examples in later sections of this post.

At this point, you might be wondering Why don’t we just optimize everything? If the baseline interpreter is slow and JIT compiled code is fast, then why not always JIT compile all the code, all the time? What’s the point in using the interpreter, aside from avoiding the complexity of implementing a JIT compiler?

The answer is that JIT compilation is computationally expensive. Generating optimized code actually incurs a performance cost at first. This is why JIT compilers specifically look for hot code to optimize: because this code is getting used a lot, there’s a good chance that it’ll continue to be used after the JIT compiler has optimized it. Over time, the performance gains provided by the optimized code will outweigh the initial computational cost of generating the optimized code, leading to a net performance improvement.

To help make this clear, let’s consider an analogy:

Imagine you run a pizza restaurant. Every week, the same group of people show up at the same time and place exactly the same order. They always show up during rush hour, so you’re spread thin and their order takes a while to prepare. You notice that these people have been showing up consistently every week for a year and have always made the same order. In an effort to make rush hour less painful and get these folks their food more quickly, you decide to go ahead and make their food a bit in advance, so that when they show up, it’ll be ready for them and you’ll have more resources during rush hour. (This does mean their food would be sitting around for a while and you’d need to keep it warm, but don’t think too hard about all that.)

If this approach works, you’ll start getting some efficiency gains every week. However, there’s always the risk that the group just doesn’t show up one week, or that they show up but unexpectedly order something different, meaning your work and resources have been wasted. However, because the group has shown up so consistently for a whole year, you decide that the likely performance gains outweigh the possible cost of wasted food and time.

This is basically the role of a JIT compiler — watch for code that’s getting used frequently, take an initial performance hit to generate a more optimized version of the code, and then hope that the code continues to get used, leading to performance gains over time. To generate optimized code, the JIT compiler needs to make some assumptions about how the code will be used, and if the assumptions end up being wrong, the optimized code might not be usable, sort of like the case where the group in the analogy shows up but orders something unexpected. We’ll discuss this more in a bit.

Note that this is a dramatic simplification of JS engine architecture. One detail to consider is that modern browsers actually have more than one JIT compiler; they have multiple JIT compilers that perform different levels of optimization, with the goal of performing minor, less expensive optimizations for code that isn’t going to be used as frequently as the code that gets the more expensive but very efficient optimizations. For example, as of the time of this writing, V8 has its baseline interpreter Ignition, and then three different JIT compilers: Sparkplug, Maglev, and TurboFan. Each “tier” of compiler generates more optimized code than the previous one. The lower tiers are useful for optimizing code that’s used more than once, but not often enough to justify using the more expensive TurboFan tier.

In-depth review of these different JIT tiers is beyond the scope of this post. For the purposes of this discussion, we’ll just focus on the highest tier of JIT compiler (responsible for emitting the most optimized machine code), which in the case of V8 is TurboFan. Just keep in mind that these multiple distinct JIT compilers exist, so when hunting for vulnerabilities, you’ll want to consider more than just one JIT compiler.

Hands-on with JIT compilation

Okay, so now we’ve got a basic idea of the role of JIT compilers in browsers. Let’s get more specific now. What does this process look like? What types of vulnerabilities does it introduce?

Compiling JS to machine code is difficult due to the lack of type information. JavaScript is not a strongly typed language, meaning that variables can change data types — for example, from an integer to a string. When generating assembly, though, we need to be careful not to handle something as the wrong type, because this can lead to security problems (we’ll discuss this in further detail shortly).

So if JS allows type changing, but accurate type information is needed to emit optimized machine code, how do we figure out the types? For example, if some parameters get passed to a function, how do we know what types those will be? They could be anything. JIT compilers solve this by speculating what the types are likely to be. They can do this by observing what the types have been in the past (a process called profiling ) and speculating that they are likely to be the same in the future. The JIT compiler will then emit optimized machine code based on those speculations.

Let’s examine this process a bit. Examples throughout this post will use D8, the standalone version of V8. You are encouraged to compile D8 yourself so that you can follow along. To do so, you can follow the official documentation here: https://v8.dev/docs/build

Additionally, I will be using GDB with the GEF ( https://github.com/hugsy/gef ) extension for debugging.

To begin, let’s run D8 in GDB with two special arguments, –allow-natives-syntax and –trace-turbo.

–allow-natives-syntax enables some useful debugging functions that are not normally available.

–trace-turbo will be useful later for examining the JIT compilation process.

gdb --args ./d8 --allow-natives-syntax --trace-turbo

Upon running D8, you’ll observe that you’re simply given an interactive shell. This is because we’re only interacting with the JavaScript engine rather than a full browser. This simplifies things a lot by allowing us to just investigate the components we’re interested in. After running D8, let’s begin by creating a simple addition function that takes two arguments:

d8> function simple_add(arg1,arg2){return arg1 + arg2}
undefined

Now we’ll use one of the functions made available by –allow-natives-syntax, which is %DebugPrint. This function will show a variety of information about the function:

d8> %DebugPrint(simple_add);
DebugPrint: 0x32df00199589: [Function]
- map: 0x32df00182215 <Map[32](HOLEY_ELEMENTS)> [FastProperties]
- prototype: 0x32df0018213d <JSFunction (sfi = 0x32df001418a1)>
- elements: 0x32df00000745 <FixedArray[0]> [HOLEY_ELEMENTS]
- function prototype:
- initial_map:
- shared_info: 0x32df00199501 <SharedFunctionInfo simple_add>
- name: 0x32df00199459 <String[10]: #simple_add>
- builtin: CompileLazy
- formal_parameter_count: 3
- kind: NormalFunction
- context: 0x32df00181a85 <NativeContext[301]>
- code: 0x32df000346a9 <Code BUILTIN CompileLazy>
- dispatch_handle: 0x264800
- source code: (arg1,arg2){return arg1 + arg2}
- properties: 0x32df00000745 <FixedArray[0]>
- All own properties (excluding elements): {
0x32df00000d91: [String] in ReadOnlySpace: #length: 0x32df000262f9 <AccessorInfo name= 0x32df00000d91 <String[6]: #length>, data= 0x32df00000011 <undefined>> (const accessor descriptor, attrs: [__C]), location: descriptor
0x32df00000dbd: [String] in ReadOnlySpace: #name: 0x32df000262e1 <AccessorInfo name= 0x32df00000dbd <String[4]: #name>, data= 0x32df00000011 <undefined>> (const accessor descriptor, attrs: [__C]), location: descriptor
0x32df000042d9: [String] in ReadOnlySpace: #arguments: 0x32df000262b1 <AccessorInfo name= 0x32df000042d9 <String[9]: #arguments>, data= 0x32df00000011 <undefined>> (const accessor descriptor, attrs: [___]), location: descriptor
0x32df00004559: [String] in ReadOnlySpace: #caller: 0x32df000262c9 <AccessorInfo name= 0x32df00004559 <String[6]: #caller>, data= 0x32df00000011 <undefined>> (const accessor descriptor, attrs: [___]), location: descriptor
0x32df00000da5: [String] in ReadOnlySpace: #prototype: 0x32df00026311 <AccessorInfo name= 0x32df00000da5 <String[9]: #prototype>, data= 0x32df00000011 <undefined>> (const accessor descriptor, attrs: [W__]), location: descriptor
}
- feedback vector: feedback metadata is not available in SFI
0x32df00182215: [Map]
- map: 0x32df00181a35 <MetaMap (0x32df00181a85 <NativeContext[301]>)>
- type: JS_FUNCTION_TYPE
- instance size: 32
- inobject properties: 0
- unused property fields: 0
- elements kind: HOLEY_ELEMENTS
- enum length: invalid
- stable_map
- callable
- constructor
- has_prototype_slot
- back pointer: 0x32df00000011 <undefined>
- prototype_validity cell: 0x32df00000a81 <Cell value= 1>
- instance descriptors (own) #5: 0x32df0018223d <DescriptorArray[5]>
- prototype: 0x32df0018213d <JSFunction (sfi = 0x32df001418a1)>
- constructor: 0x32df001821e1 <JSFunction Function (sfi = 0x32df0002b721)>
- dependent code: 0x32df00000755 <Other heap object (WEAK_ARRAY_LIST_TYPE)>
- construction counter: 0

There’s a lot of information there and you don’t need to dig into most of it. However, you may want to note the line with the code pointer:

- code: 0x32df000346a9 <Code BUILTIN CompileLazy>

For now it doesn’t mean much, but let’s cause our function to be JIT compiled and compare the debug output for the JITed version with the original bytecode version. Remember how the JIT compiler looks for code it considers hot, AKA code that’s used frequently? Let’s just run our function a ton of times in a loop to make it hot enough for the JIT compiler to optimize it (there’s also a way to do this with some of the functions created with –allow-natives-syntax, and we’ll see that in a moment):

d8> for (i = 0; i < 0x20000; i++){simple_add(1,2)}

Great! If you’re following along on your own machine, you may have noticed some output showing what the JIT compiler is doing; this is due to the –trace-turbo argument we provided to d8 earlier.

Now that it’s been identified as hot and JITed, let’s take a look at the function again using %DebugPrint:

d8> %DebugPrint(simple_add)
DebugPrint: 0x32df00199589: [Function]
- map: 0x32df00182215 <Map[32](HOLEY_ELEMENTS)> [FastProperties]
- prototype: 0x32df0018213d <JSFunction (sfi = 0x32df001418a1)>
- elements: 0x32df00000745 <FixedArray[0]> [HOLEY_ELEMENTS]
- function prototype:
- initial_map:
- shared_info: 0x32df00199501 <SharedFunctionInfo simple_add>
- name: 0x32df00199459 <String[10]: #simple_add>
- formal_parameter_count: 3
- kind: NormalFunction
- context: 0x32df00181a85 <NativeContext[301]>
- code: 0x23f700040ca1 <Code TURBOFAN_JS>
- dispatch_handle: 0x264800
- source code: (arg1,arg2){return arg1 + arg2}
- properties: 0x32df00000745 <FixedArray[0]>
- All own properties (excluding elements): {
0x32df00000d91: [String] in ReadOnlySpace: #length: 0x32df000262f9 <AccessorInfo name= 0x32df00000d91 <String[6]: #length>, data= 0x32df00000011 <undefined>> (const accessor descriptor, attrs: [__C]), location: descriptor
0x32df00000dbd: [String] in ReadOnlySpace: #name: 0x32df000262e1 <AccessorInfo name= 0x32df00000dbd <String[4]: #name>, data= 0x32df00000011 <undefined>> (const accessor descriptor, attrs: [__C]), location: descriptor
0x32df000042d9: [String] in ReadOnlySpace: #arguments: 0x32df000262b1 <AccessorInfo name= 0x32df000042d9 <String[9]: #arguments>, data= 0x32df00000011 <undefined>> (const accessor descriptor, attrs: [___]), location: descriptor
0x32df00004559: [String] in ReadOnlySpace: #caller: 0x32df000262c9 <AccessorInfo name= 0x32df00004559 <String[6]: #caller>, data= 0x32df00000011 <undefined>> (const accessor descriptor, attrs: [___]), location: descriptor
0x32df00000da5: [String] in ReadOnlySpace: #prototype: 0x32df00026311 <AccessorInfo name= 0x32df00000da5 <String[9]: #prototype>, data= 0x32df00000011 <undefined>> (const accessor descriptor, attrs: [W__]), location: descriptor
}
- feedback vector: 0x32df0019a4b9: [FeedbackVector]
- map: 0x32df00000801 <Map(FEEDBACK_VECTOR_TYPE)>
- length: 1
- shared function info: 0x32df00199501 <SharedFunctionInfo simple_add>
- tiering_in_progress: 0
- osr_tiering_in_progress: 0
- invocation count: 402
- closure feedback cell array: 0x32df0000216d: [ClosureFeedbackCellArray] in ReadOnlySpace
- map: 0x32df000007d9 <Map(CLOSURE_FEEDBACK_CELL_ARRAY_TYPE)>
- length: 0
- elements:
- slot #0 BinaryOp BinaryOp:SignedSmall {
[0]: 1
}
0x32df00182215: [Map]
- map: 0x32df00181a35 <MetaMap (0x32df00181a85 <NativeContext[301]>)>
- type: JS_FUNCTION_TYPE
- instance size: 32
- inobject properties: 0
- unused property fields: 0
- elements kind: HOLEY_ELEMENTS
- enum length: invalid
- stable_map
- callable
- constructor
- has_prototype_slot
- back pointer: 0x32df00000011 <undefined>
- prototype_validity cell: 0x32df00000a81 <Cell value= 1>
- instance descriptors (own) #5: 0x32df0018223d <DescriptorArray[5]>
- prototype: 0x32df0018213d <JSFunction (sfi = 0x32df001418a1)>
- constructor: 0x32df001821e1 <JSFunction Function (sfi = 0x32df0002b721)>
- dependent code: 0x32df00000755 <Other heap object (WEAK_ARRAY_LIST_TYPE)>
- construction counter: 0

Did you spot some differences in the output? Here, take a look at the code pointer for the JITed function:

- code: 0x23f700040ca1 <Code TURBOFAN_JS>

Note that we’re pointing to a different place in memory and there’s now a reference to TurboFan. Additionally, if you take a look at the contents of the directory from which you ran D8, you’ll see that there’s now a JSON file for the function we just ran! In my case, it’s called turbo-simple_add-1.json.

In a real exploit, this is the way we’d force some code to be JITed, but if you don’t want to have to write a loop every time you want some code identified as hot and therefore JITed, you can also use some functions provided by the –allow-natives-syntax argument. Specifically, after defining a function, you can use the %PrepareFunctionForOptimization() and %OptimizeFunctionOnNextCall() functions. This will allow the function to be JITed the very next time it is called. The following screenshots illustrate this process:

Using the %PrepareFunctionForOptimization() and %OptimizeFunctionOnNextCall() functions.

JITing the function.

So we’ve JITed a function. What now? We probably want to see what happens under the hood when this operation takes place, since that’ll be key to eventually understanding the types of vulnerabilities that can be introduced here. To do this, let’s check out a tool called Turbolizer.

Turbolizer

To analyze what TurboFan is doing during the JIT compilation process, we’ll use Turbolizer, a tool that displays TurboFan’s various stages in graph form. Turbolizer can be combined with the –trace-turbo CLI argument for D8. This argument emits a JSON file that can be examined using Turbolizer to visualize TurboFan’s optimization stages. We’ve already generated a JSON file in the previous section for our simple_add() function.

After grabbing the V8 source, you can build and run Turbolizer. The documentation here explains the process:

https://chromium.googlesource.com/v8/v8/+/refs/heads/main/tools/turbolizer/

Once we’ve got Turbolizer set up and listening, we can upload the JSON file for our function. Upon doing so, we’ll be presented with some information about our function, including a drop-down menu that shows a bunch of different graph view options. What are these?

A drop-down menu showing various TurboFan IR phases.

As part of its optimization process, TurboFan converts the code it intends to optimize into an intermediate representation (IR). The IR in this case is a graph using an idea from compiler theory called “Sea of Nodes”; you can read more about that theory here: https://darksi.de/d.sea-of-nodes/

However, I don’t know about you, but compiler theory is getting dangerously close to Smart People territory for my tastes. Rather than try to understand all of that first, let’s focus on just getting enough to make some sense of what we’re seeing in Turbolizer, and add on to our knowledge as we do hands-on work. In very simplified terms, TurboFan will convert the unoptimized code into “nodes”, which are part of the IR. The IR is then easier to work with for performing optimization. Turbolizer lets us view the generated IR, and the various steps performed during optimization, as a graph.

What does this actually look like? Here’s a screenshot of the V8.TFBytecodeGraphBuilder 20 graph, which is the first graph available, after I’ve clicked on the arrows on the left-hand side in the simple_add JS source (or you can use the “show all nodes” button):

The first available IR view.

There’s a lot going on here, but chances are that we don’t need to understand every single part of this graph, at least not right away. Of note is that Parameter[1] and Parameter[2] sound like they’re probably the arguments getting passed to the function, and that node labeled SpeculativeSafeIntegerAdd sounds pretty interesting.

Currently, our view is set to V8.TFBytecodeGraphBuilder, which is the first of many possible graph views. These views show phases performed during JIT compilation. These phases relate to the distinct rounds of optimization applied by TurboFan. By examining the different graph views, we can see what TurboFan did during a specific optimization phase. This is important, as you may want to trigger a bug during a specific phase (see https://www.jaybosamiya.com/blog/2019/01/02/krautflare/ , https://abiondo.me/2019/01/02/exploiting-math-expm1-v8/ and the associated bug https://project-zero.issues.chromium.org/issues/42450781 for an example).  The following post contains a bit more information about some of the TurboFan optimization phases: https://www.zerodayinitiative.com/blog/2021/12/6/two-birds-with-one-stone-an-introduction-to-v8-and-jit-exploitation

Discussing every TurboFan phase available in Turbolizer would be far more comprehensive than this blog post aims to be (and I’m still learning about them myself and don’t want to provide inaccurate information), but there are a couple phases worth calling out. The graph builder phase is actually not an optimization phase and is more of a setup phase. The typer phase is of interest, which is responsible for determining expected types to be used/returned by the optimized function. If we take a look at the typer phase in Turbolizer and enable the “toggle types” option, we can see that the typer has determined that the output of the function should be the return value of a SpeculativeSafeIntegerAdd within a specific range (the max sizes for a small integer in V8).

The typer phase graph view.

As you dig into these optimization phases, you should examine the source code implementing them. For example, the v8/src/compiler/turbofan-typer.cc and v8/src/compiler/operation-typer.cc files are related to the TurboFan typer phase (as of the time of this writing; the file names or locations could change in the future). For example, if you check out operation-typer.cc, you can find the implementation of SpeculativeSafeIntegerAdd:

Type OperationTyper::SpeculativeSafeIntegerAdd(Type lhs, Type rhs) {
Type result = SpeculativeNumberAdd(lhs, rhs);
// If we have a Smi or Int32 feedback, the representation selection will
// either truncate or it will check the inputs (i.e., deopt if not int32).
// In either case the result will be in the safe integer range, so we
// can bake in the type here. This needs to be in sync with
// SimplifiedLowering::VisitSpeculativeAdditiveOp.
return Type::Intersect(result, cache_->kSafeIntegerOrMinusZero, zone());
}

Helpfully, Turbolizer also has a right-hand panel that shows the assembly code emitted by TurboFan, and includes the ability to display the addresses at which the code is located:

Assembly code of the JITed function.

These addresses can be used to examine the assembly code in memory in the d8 process we’re debugging. Let’s hop back over to GDB and give that a try. In GDB, we can first set a break point on the first instruction address provided by Turbolizer:

break *0x5555b7c00180

Now that we’ve JITed our function, we can invoke it again and should hit our breakpoint that’s set at the beginning of the optimized code emitted by TurboFan:

d8> simple_add(1,2);
Thread 1 "d8" hit Breakpoint 1, 0x00005555b7c00180 in ?? ()

Great! We can now view some of the instructions beginning at this address; let’s check out the next 30 with the following command:

x/30i 0x5555b7c00180

Identifying the JITed assembly instructions in GDB.

It’s now possible to single-step through the JITed code to get a feel for it. However, we just called our function exactly as expected, with integers as arguments. What happens if we call our function with some other argument type? If the code was optimized based on the speculation that the arguments would be integers, then surely a different argument type would be bound to cause problems, right? Let’s take a look at what happens by running our function again, but providing strings as arguments this time:

d8> simple_add("1","2");
Thread 1 "d8" hit Breakpoint 1, 0x00005555b7c00180 in ?? ()

Note that there’s an interesting check being applied by the JITed code in which the cl register is tested against 0x1, and then there’s a conditional jump:

A conditional jump in GDB.

This check’s jump was not taken when we provided integers as arguments, but after providing strings, it is. Let’s single-step and see where this conditional jump takes us:

A call instruction in GDB.

Looks like we’ve got a call that dereferences something relative to the r13 register. Okay, so let’s take a look at what resides at that pointer:

A call to a deoptimization function in GDB.

Hm, so we’re going to call something called Builtins_DeoptimizationEntry_Eager. Isn’t the whole point of JIT compilation optimization ? So what’s with this deoptimization thing? That sounds like it’s the opposite of what we’d expect. We’ll talk about this more in the next section. We’ll also return to Turbolizer afterward to make sense of this.

So now we know a little bit about the optimization process performed by JIT compilers and have seen some of the phases used by TurboFan. We’ve also seen some checks to ensure that the arguments provided to our function really are the expected type. This is important, and failure to ensure the correct types are used can lead to vulnerabilities. Now that we’ve finally gotten a basic understanding of JIT compilers, let’s start talking about some vulnerabilities that can occur in them.

Common JIT compiler vulnerabilities

In the previous sections, we learned that JIT compilers determine the types of variables/arguments used in code to be optimized by profiling to see what types have been used in the past, then speculating about which ones are likely to be used in the future.

But what if the speculations made don’t hold true? For example, what if the type of a parameter provided to a function is different than the one the JIT compiler has speculated it will be? If there’s no system in place to deal with this possibility, then the JITed code could end up treating a parameter as an incorrect type — an object could be treated as a different kind of object. This is a class of bug referred to as a type confusion ; it’s not specific to JIT compilers, but does commonly occur in them.

To help prevent this situation, the JIT compiler introduces “speculation guards” into its emitted machine code. The purpose of these guards is to verify that the speculations hold true. If they do, then it’s safe to proceed to execute the machine code emitted by the JIT compiler. However, if the check performed by the speculation guards reveals that the speculations are inaccurate, then a deoptimization bailout is triggered; this just means that the fast machine code is not suitable / safe to use in this case, so execution will revert to the slower bytecode instead. It’s slow, but it’s safe. (TurboFan knows how to return to the original bytecode version of the function by preparing some deoptimization data during its code generation; this blog post contains some information about this process: https://www.l3harris.com/newsroom/editorial/2023/10/permalink-modern-attacks-chrome-browser-optimizations-and )

So now the JITed code we stepped through in the last section is starting to make more sense. When we provided integers as arguments, as we did when we made our function hot, the optimized code behaved as expected, because the profiling suggested that integers would always be used. When we switched to using strings, though, the JITed code was no longer suitable. We encountered a speculation guard that checked the argument type. When it found that the type was not an integer, a deoptimization bailout was triggered, leading us to run the slower bytecode instead.

Let’s take a look at this in Turbolizer, but with something a little more interesting than a function that just adds two integers. What if we add a property of an object and an integer instead? Let’s create that function and JIT it:

d8> test_object = {x:1}
{x: 1}
d8> function object_jit(arg1,arg2){return arg1.x + arg2}
undefined
d8> for (i = 0; i < 0x200000; i++){object_jit(test_object,2)};

After JIT compiling the function, we can upload the emitted JSON file to Turbolizer. Let’s take a look at the typer phase with all nodes displayed and the type information displayed:

A typer phase graph.

Notice how Parameter[1], which is the first argument passed to our function, has a CheckMaps node related to it? This is a check to ensure that the object provided in this argument is of the expected type. In V8, a map is basically some metadata about an object (this concept exists in other JS engines too, it just goes by different names; in Spidermonkey, Firefox’s JS engine, this is called a shape ). A map will be shared by different objects if they have the same number of properties and the same property names, even if the values of the properties are different. This is a very brief overview of this topic, and I strongly recommend you check out this excellent blog post that examines this topic in much greater detail: https://mathiasbynens.be/notes/shapes-ics

For our purposes, all we really need to understand right now is that the map of an object pretty much represents its type, and this check is there to ensure that the type has not changed. If the map of the object we provide in argument 1 of object_jit() is in any way different than the one for test_object, a deoptimization bailout will be triggered. For example, the following code leads to a bailout, because even though the object we provide does have a property called x, it doesn’t have the same map, since it doesn’t have a matching number of properties:

d8> newobj = {x:1,y:2}
{x: 1, y: 2}
d8> object_jit(newobj,2)

You’re encouraged to actually try setting a breakpoint on the JITed code and observing this speculation guard and bailout yourself.

Okay, so we now understand that JIT compilers try to avoid issues like type confusions by using speculation guards. But obviously, vulnerabilities do still occur, so what causes them? A variety of bugs are possible, but we’re going to focus on two that are outlined by Samuel Groß, AKA saelo, in his excellent “Attacking Client-Side JIT Compilers (v2)” presentation from BlackHat 2018 ( https://saelo.github.io/presentations/blackhat_us_18_attacking_client_side_jit_compilers.pdf ). You are very much encouraged to read this (and if you’re interested in browser security, you should probably follow saelo’s work in general). I’m just providing a brief overview of some concepts he covers in greater detail.

In particular, in his presentation, saelo mentions two major vulnerability classes that are caused by optimizations applied by the JIT compiler: bounds check elimination and redundancy elimination.

Redundancy elimination relates to removing speculation guards that perform checks that have already been performed earlier — in other words, removing redundant speculation guards. If a check has already been performed once and it’s not possible for anything relevant to change between that original check and the redundant one, then the redundant check is just impacting performance without adding any security benefit. However, you may have noticed that the statement “and it’s not possible for anything relevant to change between that original check and the redundant one” is doing a lot of work in that sentence.

A JIT compiler tries to discern whether a check is redundant through a process called side effect modeling . For example, it might try to decide whether a type-checking speculation guard is redundant by using side effect modeling to see if there’s any way that an object’s type could be changed during execution of the code between the initial speculation guard and the potentially redundant one. If the side effect modeling determines that the object type cannot be changed, then the redundant speculation guard can be safely removed.

Unless, of course, the side effect modeling is simply incorrect! By finding a case the JIT compiler didn’t consider, an attacker could unexpectedly change the object’s type to cause a type confusion. Because the speculation guard has been removed, there’s no longer any protection in place to detect this and bail out. As a result, this will likely lead to memory corruption, which may be exploitable. If you’re interested in seeing an example of a real-world (and very complex, at least for me) improper side effect modeling vulnerability, you can check out the following blog post: https://github.blog/security/vulnerability-research/getting-rce-in-chrome-with-incorrect-side-effect-in-the-jit-compiler/

Bounds check eliminations are similar in that they also rely on inaccuracy from the JIT compiler. As an example, if the function to be JITed accesses an index of an array, the JITed code needs to make sure that the index is within the bounds of the array. This could be accomplished by including a speculation guard to ensure the index really is within bounds. However, if the compiler thinks it’s able to prove that the index value will always be within bounds, then this speculation guard is impacting performance unnecessarily, so it can be eliminated. Just like with redundancy elimination, this means the compiler needs to be accurately calculating whether the index can ever be a value outside the array bounds. This could happen due to an integer overflow like the one demonstrated in this Exodus Intelligence blog post about a vulnerability in Safari: https://blog.exodusintel.com/2023/07/20/shifting-boundaries-exploiting-an-integer-overflow-in-apple-safari/#Vulnerability (Note that this technique is old and in V8, hardening has since been applied; I think this scenario helps illustrate the general idea, so I’ve kept it, but if you’re interested in the hardening that’s been applied, you can read about it here: https://doar-e.github.io/blog/2019/05/09/circumventing-chromes-hardening-of-typer-bugs/ )

Why JIT bugs? Why not focus on the interpreter instead?

In an earlier section, I mentioned that vulnerabilities can occur in the baseline interpreter, not only in the JIT compiler. In fact, users can opt to disable the JIT compiler (though by default they’re enabled), but disabling the baseline interpreter is a lot less practical, because that means you’re turning off JavaScript support entirely. You can do this, of course, but much of the modern web will become basically unusable without JavaScript.

With that in mind, wouldn’t the interpreter be an even better target than the JIT compiler, since the JIT compiler is optional and the baseline interpreter is (practically speaking) really not? Why focus on the JIT compiler?

The baseline interpreter is definitely an important target as well, but JIT compilers are responsible for a lot of the vulnerabilities in browsers. The following Microsoft blog post from 2021 asserts that “data after 2019 shows that roughly 45% of CVEs issued for V8 were related to the JIT engine”:

https://microsoftedge.github.io/edgevr/posts/Super-Duper-Secure-Mode/

That blog post also links to the following analysis from Mozilla, which highlights how frequently JIT bugs are exploited: https://docs.google.com/spreadsheets/d/1FslzTx4b7sKZK4BR-DpO45JZNB1QZF9wuijK3OxBwr0/edit?gid=0#gid=0

That data is admittedly a little outdated now, but as far as I’m aware, JIT bugs have remained the primary arena for browser exploitation. There are a couple of contributing factors to this. First, JIT compilers are complex! Complex subsystems often tend to be vulnerable subsystems.

The second and more interesting point to discuss is that JIT bugs aren’t really traditional memory corruption bugs. They’re more logic bugs that happen to lead to memory corruption as a side effect. Recall that JIT bugs often involve things like the compiler deciding that a bounds check or type check isn’t necessary and opting to remove it. Even if this attack surface were rewritten in a memory-safe language, these logic bugs could still exist. If you’re interesting in learning a bit more about why JIT compiler bugs are so hard to solve, this presentation (also from saelo) is a great watch: OffensiveCon24 – Samuel Groß – The V8 Heap Sandbox

Finally, JIT bugs also seem to frequently lead to powerful exploit primitives that can be “easily” (compared to some types of vulnerabilities, anyway) converted to the primitives needed to achieve a full exploit.

For example, we’ve mentioned that two common JIT bugs are bugs that lead to access outside intended bounds and bugs that lead to type confusions. An out-of-bounds (OOB) write may allow corrupting a heap object adjacent to the vulnerable one. JS affords an attacker a lot of control over what objects get allocated, so it’s likely you could choose a promising victim object to place next to the vulnerable one and then use the OOB write to corrupt something valuable in the victim object (like a length value, for example, in order to achieve a much less constrained OOB read/write primitive). Type confusions can often be massaged into eventually being OOB reads/writes.

Compare this with something like a linear overflow from a heap object (whose type you don’t get to choose) in some specific heap bin where you’re limited to a small number of other victim object types. A bug like this may certainly still be exploitable, but this could take a lot more work to exploit and might even need to be chained with another bug.

JIT bugs: a case study

At this point you may be thinking We’ve been talking about JIT compilers for so, so long. Can we please see an example bug so this can make sense and this super long blog post can finally be over? Let’s take a look at a very simple CTF TurboFan vulnerability. We will not be developing a full exploit for it; however, we’ll be stepping through it in enough detail to at least understand how to trigger the bug and why it occurs.

We’re going to take a look at an old PicoCTF challenge called Turboflan. The files for it are available here: https://play.picoctf.org/practice/challenge/178?page=1&search=turboflan

We’re given a build of D8 and a patch file, among other things. Patch analysis is a good starting point. The relevant portion of the patch is pretty small and even has a helpful comment to draw our attention to it (note that this challenge is from 2021, and the file being patched doesn’t seem to be present in the same place in 2025):

--- a/src/compiler/effect-control-linearizer.cc
+++ b/src/compiler/effect-control-linearizer.cc
@@ -1866,8 +1866,9 @@ void EffectControlLinearizer::LowerCheckMaps(Node* node, Node* frame_state) {
Node* map = __ HeapConstant(maps[i]);
Node* check = __ TaggedEqual(value_map, map);
if (i == map_count - 1) {
-        __ DeoptimizeIfNot(DeoptimizeReason::kWrongMap, p.feedback(), check,
-                           frame_state, IsSafetyCheck::kCriticalSafetyCheck);
+        // This makes me slow down! Can't have! Gotta go fast!!
+        // __ DeoptimizeIfNot(DeoptimizeReason::kWrongMap, p.feedback(), check,
+        //                     frame_state, IsSafetyCheck::kCriticalSafetyCheck);
} else {
auto next_map = __ MakeLabel();
__ BranchWithCriticalSafetyCheck(check, &done, &next_map);
@@ -1888,8 +1889,8 @@ void EffectControlLinearizer::LowerCheckMaps(Node* node, Node* frame_state) {
Node* check = __ TaggedEqual(value_map, map);
if (i == map_count - 1) {
-        __ DeoptimizeIfNot(DeoptimizeReason::kWrongMap, p.feedback(), check,
-                           frame_state, IsSafetyCheck::kCriticalSafetyCheck);
+        // __ DeoptimizeIfNot(DeoptimizeReason::kWrongMap, p.feedback(), check,
+        //                     frame_state, IsSafetyCheck::kCriticalSafetyCheck);
} else {
auto next_map = __ MakeLabel();
__ BranchWithCriticalSafetyCheck(check, &done, &next_map);

So the patch is removing a line that mentions “DeoptimizeReason::kWrongMap”. Remember how we examined a speculation guard earlier that would check an object’s map and ensure it wasn’t different than expected? Well, now that check is gone. A good starting point to understand what this looks like would be to write some code that should involve a map check for a 2025 build of D8 and take a look at the JIT process in Turbolizer, then try the same code on this patched, intentionally vulnerable version of D8 and compare the JIT output.

We’ve actually already done the first half earlier in this blog post. As a reminder, here’s the code we used:

d8> test_object = {x:1}
{x: 1}
d8> function object_jit(arg1,arg2){return arg1.x + arg2}
undefined

After calling object_jit() enough times for TurboFan to compile it, we can examine the emitted JSON file in Turbolizer. Observe that Parameter[1], which is our test_object.x property, eventually receives a CheckMaps check to ensure it’s still got the expected map (which in other words means we’re checking to ensure the object still has the expected type ):

A graph view showing the CheckMaps call.

Great, so we know how this should look in a modern build. Let’s try running the CTF build of D8 with the –allow-natives-syntax and –trace-turbo arguments, writing the same code, and getting a JSON file emitted.

Upon loading up that JSON file in Turbolizer, we can observe that a CheckMaps node is not present:

A graph view showing the absence of the CheckMaps call.

At this point, we can see that we’ve basically been handed a very simple type confusion vulnerability. In the real world, we’d likely need to identify improper side effect modeling and come up with some clever input to trick the compiler into believing a map check is unnecessary. In this case, that work has basically been done for us. This is a lot less complex than a real-world bug would likely be, but it’s a good choice for an introductory blog post like this.

Let’s make sure we understand the strategy here to trigger the bug. This patch is ensuring that the map-checking speculation guard is removed. To abuse this, we’ll JIT the function with one type of argument, but then pull a switcheroo by providing a different type of argument after the function has been JITed. The map of our new object won’t be checked, so the optimized code will never trigger a deoptimization bailout; in turn, the lack of bailout will lead to a type confusion.

Our next step is to come up with some object types that would be useful in a type confusion. Which type are we going to make the JITed function expect, and then which type are we going to swap in once we’ve got some optimized code? There are probably quite a few answers that would work, but in this case, we’ll take the approach of making the JITed function expect a 64-bit float array, but then call the JITed function with an array of objects instead. An array of objects is actually an array of pointers in V8 (see https://stackoverflow.com/a/75762917 ), so a type confusion between a float array and an object array would allow leaking pointers. We’re only striving to get to the point where we can trigger the bug, so we’re not going to cover the specific exploitation details here. If you want more info on this exploitation method, you should check out the following blog post, since the approach in this blog post is heavily based on it: https://www.willsroot.io/2021/04/turboflan-picoctf-2021-writeup-v8.html

Here’s some code that will trigger the bug:

function to_jit(obj){
// to prevent TurboFan from inlining this function, let's make it seem more complex by adding a loop
// the functionality here is irrelevant; I based this on the functionality saelo used here: https://gist.github.com/saelo/52985fe415ca576c94fc3f1975dbe837#file-pwn-js-L253
var prevent_inlining_int = 500;
for (var i = 0; i < prevent_inlining_int / 2; i++){
if(prevent_inlining_int % i == 0)
var z = 5;
}
return obj[0];
}
test_obj = {"x":1,"y":2,"z":3};
float_array = [1.1,2.2,3.3];
obj_array = [test_obj,test_obj,test_obj];
// now let's make the function hot so that TurboFan compiles it
console.log("JIT the function");
for (var i = 0; i < 0x20000; i++){
to_jit(float_array);
}
// now let's compare the output between an expected type (a float array) and an unexpected type (an array of objects)
console.log(to_jit(float_array)); // should return 1.1
console.log(obj_array[0]); // should return [object Object]
console.log(to_jit(obj_array)); // send an object with a different map and trigger the bug

Here’s what happens upon running that code:

wintermute@wintermute-VirtualBox:~/turboflan$ ./d8 pwn.js
JIT the function
1.1
[object Object]
5.753374822877835e-270

We’ve triggered the bug successfully! The correct output when viewing obj_array[0] should be [object Object], but instead we’re receiving a float (because the JIT compiler speculated that the function’s argument would always be a float array, and there’s no check to see if the map doesn’t match). From here, we could focus on determining what content we’re leaking and converting this type confusion to read/write primitives. However, this post is only focused on providing an introduction to JIT compiler bugs and isn’t intended to be an intro to more general V8 exploitation, so we’ll conclude our examination of this bug here. Readers interested in the full exploitation process for this challenge should check out the following blog posts:

https://www.willsroot.io/2021/04/turboflan-picoctf-2021-writeup-v8.html

https://blog.joshdabo.sh/2021/04/18/picoctf-turboflan/

Conclusion

This blog post provided an introduction to JIT compilers in JavaScript engines, including the motivation for using them and high-level details on how they work. We examined how to perform basic debugging of JITed code and took a look at the tool Turbolizer, which is helpful for visualizing the various optimization phases performed by TurboFan, a JIT compiler in V8.

We discussed common vulnerabilities introduced by JIT compilers and examined how to trigger an extremely simple JIT vulnerability from a CTF challenge. We also learned that JIT vulnerabilities are generally considered logic bugs that often lead to memory corruption as a side effect, and that JIT bugs frequently offer powerful exploit primitives that can lead to full exploitation. Hopefully, this post helped make this area of security seem a little bit less overwhelming and inspired you to dig into it further.

Depth like this is the difference

Research at this level is what separates a penetration test from a scan with a report attached. It is also how our consultants stay sharp between engagements.

Our Services Get a Quote

Engagements today run on Hexecution , the platform we built ourselves, so findings reach clients as they are confirmed instead of in a PDF at the end.

Back to all posts

Anthropic asks Claude users to share voice data for AI model training

Bleeping Computer
www.bleepingcomputer.com
2026-10-04 06:53:21
Anthropic has started asking Claude users to voluntarily share their voice conversations to help train and improve its AI models. [...]...
Original Article

Claude

Anthropic has started asking Claude users to voluntarily share their voice conversations to help train and improve its AI models.

The new prompt appears when using Claude's voice features and explicitly asks users to allow Anthropic to use their voice data for AI training.

"Allow us to use your voice data to improve our AI models," Anthropic notes in the prompt.

"Audio recordings and voice chat data help improve how Anthropic AI models understand and respond to speech. Turn this off or delete this data anytime in settings."

The good news is that this appears to be entirely optional. You can either select Allow or Not now , and there's also a dedicated toggle inside Claude's Privacy settings.

Claude

Interestingly, voice training is separate from Anthropic's existing option to use your chats and coding sessions for model training.

That means you could allow Anthropic to train on your voice conversations while keeping regular chats and Claude Code sessions excluded, or the other way around.

Claude gets a separate voice data training toggle

Inside Settings > Privacy, Anthropic describes the option as "Allow us to use your voice data."

"For approved sessions, allow the use of your audio recordings and voice chat data to improve Anthropic AI models," the setting explains.

In the interface we've seen, the setting is disabled by default, so Anthropic isn't automatically opting users into sharing their voice recordings.

You can also turn the option off later, and Anthropic says users can delete the voice data from settings if they change their mind.

article image

Build your security blueprint for AI-powered attacks

Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.

Save your seat

grubby: static site generator for git repos written in Ruby

Lobsters
git.btxx.org
2026-10-04 06:28:57
Comments...
Original Article

Static site generator for Git repos written in Ruby.

Forked and heavily inspired by RepoRat

Requirements

  • Git
  • Ruby 3.0+
  • Kramdown

Configuration

Grubby is fairly opinionated out of the box. It’s meant to be used as-is without much need for tinkering. Kramdown is the default Markdown parser, which you are free to swap out but I can’t guarantee things will be bug-free! It is also assumed that all of your bare git repos end with the .git extension. The core repo index generation will not working properly without this!

All you should ever touch is the _config.yml file for setting your core domain and paths.

If you wish to change the CSS, navigation links or footer content, then edit the style.css , _header.html , and _footer.html accordingly. Feel free to experiment!

By default the core “repo index” is created at the root of the main directory and is named index.html . You can change this file output name inside the _config.yml as well.

Usage

This guide assumes you are using a form of “shared hosting” (in this case NearlyFreeSpeech.NET ). Although running things on your own VPS will be very similar.

Below is based on your bare git repos being under /home/public/git.sitename.org .

  1. Clone this project into your private directory (ie. /home/private/ )
  2. Now navigate to a repo inside /home/public (ie. cd coolrepo.git )
  3. Create a new file under hooks: touch /hooks/post-update
  4. Give that file proper permissions: chmod +x /hooks/post-update
  5. Add the following code to that file:
#!/bin/sh

unset GIT_DIR

# Needed for cloning over plain HTTP
git update-server-info

# Rebuild the static site from a temporary checkout
repo_name=$(basename "$PWD" .git)
temporary_dir=$(mktemp -d)

git clone --quiet "$PWD" "$temporary_dir/$repo_name"
cp "$PWD/description" "$temporary_dir/$repo_name/.git/description"
(cd "$temporary_dir/$repo_name" && ruby /home/private/grubby/grubby.rb)

rm -rf "$temporary_dir"

Now you can take it for a test drive! Simply run:

./hooks/post-update

Your individual repo website will be generated at /home/public/git.sitename.org/coolrepo , along with an updated repo index ( index.html by default) at the root /home/public/git.sitename.org/ .

You can expand on this by instead using this call inside your hooks/post-receive file or even using a global post-update to avoid including this for each project.

Caveats

  • Does not create pages for individual commit changes, branches, etc.

Teenager suspected of leading KillSec ransomware group

Hacker News
www.europol.europa.eu
2026-10-04 05:40:50
Comments...
Original Article

Loading application.
Please wait.

Show HN: AI search for every photo and every frame of video on macOS

Hacker News
github.com
2026-10-04 05:24:52
Comments...
Original Article

SCM — Screen Memories

Deep AI search for every photo and every frame of video in any folder on macOS. Local-first — no accounts, no cloud, no uploads. Inference runs on your Mac.

SCM demo — library grid SCM demo — scene search

What makes it different

  • Search like you think — describe a memory in plain language; a local vision model does the rest.
  • Video, down to the moment — scenes are segmented and embedded, so you land on the shot, not just the file.
  • Text and dialogue too — OCR over visible text; exact spoken-line search via Whisper, each as its own mode.
  • Your tabs, your prompts — save any query as a tab; Screenshots and Email tabs are toggleable.
  • Self-maintaining library — watched folders auto-import, content hashes dedupe renames, and model switches re-embed in the background without blocking search.
  • Truly private — your media never leaves the machine. Weights download once; everything after that is offline.

Five ways to search

Search modes: Files, Scenes, Dialogue, OCR, LLMs

Mode Finds
Files Whole photos/videos by meaning — vision rank with filename and phrase boosts
Scenes Moments inside video — search a shot, jump to its timecode
OCR Text visible in images and frames, matched literally (Tesseract; eng + 35 language toggles)
Dialogue Exact spoken words in videos (Whisper), tiered exactness
LLMs (opt-in) Local chat over the dialogue, OCR, and filenames your Mac already extracted — cited answers

Requirements

  • macOS (packaged with electron-builder; menu-bar/tray features are macOS-only)
  • Bun — the project uses bun as package manager and runner
  • Node modules installed: bun install
  • First use of a model downloads its weights (~435MB for the default CLIP); after that, fully offline.

Download

The easiest install is via Homebrew (Apple Silicon, macOS 12+). The tap's cask clears the macOS quarantine flag automatically on every install and upgrade, so the app launches with no manual Gatekeeper steps:

brew tap allenv0/scm
brew trust allenv0/scm
brew install --cask allenv0/scm/scm

Upgrades keep the same behavior:

brew upgrade --cask allenv0/scm/scm

Prefer least privilege? Trust just the cask instead of the whole tap:

brew tap allenv0/scm
brew trust --cask allenv0/scm/scm
brew install --cask scm

The tap lives at allenv0/homebrew-scm .

Development

bun run dev     # build the renderer bundle, then launch the Electron app
bun start       # launch the Electron app without rebuilding
bun run build   # just rebuild the renderer bundle into dist/

Building the app package (DMG / ZIP)

bun run dist            # signed if an identity is in the keychain
bun run dist:unsigned   # skip code-sign discovery

This runs two steps in sequence:

  1. vite build — compiles the React renderer into dist/ (picks up all changes under src/ ).
  2. electron-builder --mac — packages the app. It bundles the fresh dist/ bundle together with main.js , preload.js , main-lib/ , and indexer/ (the file list is configured under build.files in package.json ), then produces the installers.

Output: the installers land in dist-app/ (see build.directories.output in package.json ) — look for SCM-0.2.4.dmg and SCM-0.2.4.zip .


Features

Files

Typing starts an instant filename-keyword pre-pass, then the vision model takes over: results are scored by cosine similarity against image embeddings, with gated phrase and filename boosts, an honesty floor calibrated per model, and a near-duplicate diversity filter. Every tile carries a "why it matched" badge (Visual match / Filename match / …) and a hover tooltip with the per-component score breakdown. CJK queries search as overlapping bigrams ("台北車站" also matches 台北, 車站).

Scenes

Every scene segment across all videos is scored, so a hit lands on the exact shot: tiles show the scene poster with a timecode badge, and opening the video jumps straight to that moment. A noise gate returns "no scene match" instead of flooding the grid with gibberish, and each video contributes at most 3 scenes.

OCR

Matches the fraction of query tokens literally visible in each image's OCR text — the filename is ignored and no vision model is involved, so it works even while the AI engine is warming up or offline. Matched words are boxed in amber on tiles and in the lightbox.

Dialogue

Exact literal retrieval over Whisper transcripts — no embeddings, no thresholds, works with the AI engine down. Results come in three tiers: Exact line (contiguous phrase in one utterance), Exact words (all words in one utterance or an ≤8s window), and Words spoken (all words in the same video). Matching words are highlighted in a speech snippet; opening a result seeks straight to the line.

LLMs (Ask)

Opt-in — nothing downloads or runs until enabled in Settings → LLMs Chat. A llama.cpp sidecar bound to loopback answers your question from evidence the app already extracted — dialogue lines, OCR text, and filename keyword hits — with numbered citations you can click, streamed token-by-token with a live tok/s readout. Leading /screenshots , /videos , /email narrow the corpus; Stop keeps the partial answer; empty evidence short-circuits before the model ever runs.

Chat model Size Notes
Qwen3 1.7B (default) ~1.1GB Fast everyday chat; fits 8GB Macs
Llama 3.2 3B ~2GB Stronger long answers; needs headroom

Tabs & library views

  • Built-in browse tabs: All , Videos , plus Screenshots and Email — the latter two toggleable in Settings → Smart Tabs. Selecting Videos auto-enables Scenes mode.
  • Save any query as a tab : the pin pill under the search bar saves the current prompt with its mode (Files/Scenes/OCR/Dialogue) — up to 20 tabs, renameable, each restored exactly as saved.
  • Five semantic views behind remappable shortcuts (⌘1–⌘5 by default), plus ⌘I import / AI Insights, ⌘, for Settings.
  • Search stays in the selected tab : pick Screenshots, Email, Videos, or any saved tab and results are filtered to it — scope first, then search. In LLMs chat the same idea is explicit: leading /screenshots , /videos , /email narrow the corpus before the model ever runs.

Email tab

Surfaces photos whose visible OCR text contains an email address — an overlapping view (a photo keeps its category too). Detection is OCR-tolerant: it reassembles addresses Tesseract fractures across word boxes, and handles comma-for-dot noise ("gmail,com"), split TLDs ("gmail. com"), bracketed obfuscation ("allen [at] gmail [dot] com"), and dictated addresses ("allen at gmail dot com"). Tiles show a contact strip; expand it to copy or compose.

Screenshots tab

Screenshot classification is rename-proof. Four signals, in priority order: a manual override (right-click any tile) → filename vocabulary (30+ localized OS screenshot names in 20+ languages) → a PNG/JPEG metadata probe (reads "screenshot" from PNG text chunks / EXIF UserComment, so a renamed Bildschirmfoto still classifies) → source-folder hint. Everything else lands in Projects.

Vision models

Four switchable models via ONNX Runtime; the active one is chosen per library:

Model Role Speed (CPU) Download
CLIP ViT-L/14@336 (default) Best real-world video scene-search ~480–570ms/img ~435MB
SigLIP-2-B/16 Fastest bulk import ~50–100ms/img ~412MB
SigLIP-2-L/16@256 High-detail (1024-dim) — small objects, signs, on-screen text ~200ms/img ~850MB
SigLIP-B/16@384 Maximum detail ~480ms/img ~214MB

Switching models re-embeds the whole library: the flip lands instantly with the tail filled in the background, and search falls back to filename keywords until it completes. Per-model text-mean centering de-biases text embeddings so similarity scores stay honest across models.

Video search pipeline

ffmpeg scans each video for shot boundaries and builds a segment plan sampled to the density you pick in Settings → Video Search — each preset shows its measured time and disk cost before you commit:

Preset Seconds per point Segment budget
Eco 60 4–32
Balanced (default) 30 8–128
Detailed 15 12–256
Ultra 5 16–1024
Ultra Pro 2.5 24–2048 (confirm required)

Each segment embeds its midpoint frame and keeps a poster; shot plans are cached per file (path + size + mtime + config fingerprint), so re-imports skip detection entirely.

Dialogue transcription: Whisper tiny.en (~150MB, default) or base.en (~300MB) — switching re-transcribes every video. Whole videos embed three frames (20/50/80%) averaged; GIFs embed an average of middle frames.

OCR

Tesseract runs in its own worker, separate from the vision model. English is always on; 35 more languages are toggleable in Settings → Photo Search (default: Simplified + Traditional Chinese, Japanese, Korean). Each language pack downloads once (~2.4–5MB; ~17MB for the default set), then everything is offline. Word boxes are stored with the text so matches highlight in place; CJK text is joined without spaces and email fragments fractured across word boxes are reassembled.

Import & library management

  • Import via ⌘I, drag-and-drop, or watched folders — importing a folder starts watching it (live fs.watch plus a re-sync at every launch). Problem files retry up to 3 times, then sit out watch-syncs until they change.
  • Rename-proof dedupe : every file is content-hashed (SHA-256) before copy, and lying extensions are normalized by MIME sniffing.
  • Named embedding versions (Settings → Library): point-in-time snapshots of the entire searchable state — the index, every model's embedding bins, scene and transcript sidecars — with restore (auto-backup first) and a Fresh Start danger zone. Cap: 10.
  • Everything lives under ~/Library/Application Support/scm ( MEMORIES_DATA_DIR overrides it): the index JSON, Float32 embedding bins per model, scene and transcript sidecars, thumbnails, and posters.

Privacy by construction

  • The renderer is a sandboxed app:// bundle — contextIsolation , OS sandbox, and a CSP pinned to 'self' (+ Google Fonts CDN for display type, with a monospace fallback when offline).
  • Only main-process workers ever download, once per thing: vision weights (Hugging Face), OCR language packs (Tesseract CDN), Whisper weights, and — only if you opt in — the llama.cpp sidecar and GGUF chat models (GitHub + Hugging Face), sha256-verified at download time.
  • Media is copied into the app-managed library and streamed from disk. No telemetry, no accounts, no uploads.

Settings & polish

A macOS-style settings sheet with ten panels (Library, Appearance, Grid, Smart Tabs, Photo Search, Video Search, LLMs Chat, Global Shortcut, Keyboard, Menu Bar). Around the core: light/dark/system theme, a CRT screen effect for the lightbox, 6 alternate app icons, menu-bar-only mode, a recordable global shortcut, a first-run onboarding tour, background-work trays (scenes / transcripts / OCR) with a global pause, and a status bar with version and indexed-video counts.

Notes

  • First build is slow: electron-builder downloads the Electron binary and ffmpeg once on its first run; subsequent builds are much faster.
  • Code signing: without an Apple Developer identity configured in the keychain, the DMG builds unsigned. The app still runs locally, but macOS may require right-click → Open the first time it's launched.
  • Rebuilds pick up changes automatically: since main.js , preload.js , main-lib/ , and indexer/ are packaged from source (not cached), a bun run dist after editing any of them produces a fresh package.

Testing

bun run test:all         # the full verification battery (unit + smoke suites)
bun run smoke:indexer    # headless smoke test of the CLIP indexer
bun test test/           # unit tests (pure modules, no Electron needed)
bun run lint             # eslint
bun run format:check     # prettier
  • Unit tests cover the pure cores: ranking, dialogue exact-match, CJK tokens, the Whisper model ladder, MIME sniffing, transcript fusion, and more (see the test:* scripts in package.json ).
  • E2E smoke tests run inside the Electron app via ELECTRON_SMOKE_* environment variables — a dozen-plus scenarios from boot/protocol checks to search matrices, model migration, and Ask mode (drivers in scripts/e2e/ , dispatched in main.js ; e.g. bun run smoke:ask , smoke:deep , smoke:grid ).
  • Benchmarks: bun run bench:inference , bench:enrich , and bench:detect write JSON reports into MDs/bench-* .

Project layout

main.js             Electron main process: library, IPC, indexer worker pool, app:// protocol
preload.js          contextBridge — exposes window.memories to the renderer
main-lib/           main-process modules split out of main.js (settings, library store,
                    rank search, Ask retrieval, LLM sidecar config, embedding versions, …)
indexer/            vision/OCR/ASR workers (utility processes) + video utils + model registry
src/                React renderer (grid, search modes, lightbox, tabs, settings, onboarding)
scripts/            bench scripts, E2E drivers, the smoke battery (smoke-all.sh)
test/               unit + integration tests
MDs/                design docs, bench reports, plans
dist/               vite build output (renderer bundle)
dist-app/           electron-builder output (DMG / ZIP)

Models ranked by normative score across twelve paradigms, once neutral and once human-primed

Lobsters
cognit.rajtilak.tech
2026-10-04 05:21:18
Comments...
Original Article

cognit

Leaderboard

Models ranked by normative score across twelve paradigms, once neutral and once human-primed. 100% is fully normative. 0% matches the biased answer. How to read a row .

An LLM scored these answers. The judge's name is in the Judge column.

VGHF Digital Archive passes 5000 magazines. Here's what's next

Hacker News
gamehistory.org
2026-10-04 05:07:11
Comments...
Original Article

The Video Game History Foundation has added over 5000 total magazines to our digital archive, a major milestone in our project to build a first-in-class video game history research library.

When we launched our digital archive last year , we announced an initial collection of over 1500 out-of-print, text-searchable magazines, which are free for research access. In the time since we launched, we’ve more than tripled the size of that collection—now totaling nearly 600,000 pages , the equivalent of nearly 100 feet of shelf space!

Our magazine library now spans 46 years of video game history, from 1981–2026. Major new additions since launch include hundreds of trade magazines , a historically significant newsletter , and publications for niche gaming communities . Altogether, these account for about 60 percent of the Video Game History Foundation’s entire physical magazine collection that we’ve cataloged so far.

This milestone would not have been possible without support from community-run scanning projects. With our combined efforts, we’ve turned their magazine scans into a powerful research tool.

Around one-third of the magazines in our digital archive were digitized by the Video Game History Foundation, with the rest provided by external groups. We want to thank the groups and individuals who have helped:

  • The community-run scanning groups Retromags, Out-of-Print Archive, Sega Retro, AtariAge, Atari Compendium, CGW Museum, The Odyssey² Homepage, TOSEC-PIX, eXoDOS, and Gaming Alexandria.
  • Industry partners and alumni from Electronic Gaming Monthly , Game Developer , Game Informer , and MCV .
  • …and a special shoutout to VGHF Patreon member Chris Chapman, who provided us with 157 scans and digital magazines, the most of any unaffiliated contributor.

It’s only been a year and a half since we launched our digital archive, but we’re already seeing its transformative effect on video game history research. In fact, the VGHF Digital Archive has already become the preferred reference source for out-of-print game magazines on Wikipedia, with over 500 citations and counting. We’re grateful that researchers have responded so quickly to our vision for a well-organized, freely accessible research platform.

We have also heard from the community about what additional features and series they’d like to see in our magazine library. We’ve taken that feedback to heart, and that’s helped us decide where we’re heading next.

Introducing Japanese magazines

You asked, and we’ve delivered: we are now supporting Japanese-language magazines in our digital archive .

Japanese magazines are a vital resource for studying video game history; in particular, they can be a rich source for interviews with developers—most of which have rarely if ever been reprinted—and first-hand accounts of Japanese player culture.

While there are a number of community projects to digitize Japanese magazines, they’ve often been met with unique cultural and legal challenges. We intend for our digital archive to build on the community’s work, much like we’ve done for English-language magazines.

How we’re doing it

Today, as our pilot project, we’re launching research access to Neo Geo Freak , a magazine that covered games for Neo Geo home and arcade platforms from 1995–2000.

Neo Geo Freak felt like a good practical starting place for us. Last year, a generous community member donated the entire run of Neo Geo Freak to our library, and the magazine had already been digitized by the user Japanese Magazine Scans , both of which reduced the amount of work on our end to start exploring what it would take to add Japanese materials to our digital archive.

More importantly, though, Neo Geo Freak is also a good example of the unique material that can found in Japanese publications. Neo Geo platforms performed especially well in Japan relative to the rest of the world, making this an excellent candidate for the sort of sustained coverage you wouldn’t find in the English-language press.

However, adding Japanese magazines to our library was not as simple as uploading a bunch of PDFs. We maintain a high standard for how we present materials in our archive, and for a content-dense magazine written in a language that our staff isn’t fluent in, meeting those standards was easier said than done.

First, we worked with our community of peers to determine the best methods for Japanese script recognition. Professional archivists in Japan pointed us towards their preferred libraries for Japanese document transcription, which we integrated into the text mapping toolset that we use for English-language magazines.

We believe the result is the most accurate text-searchable Japanese magazines currently in any digital archive. We verified the results with game history researchers who use Japanese materials to ensure that the output is reliable and usable at scale.

A two-page spread from Neo Geo Freak featuring Nakoruru from the game Samurai Shodown. All text is accurately highlighted on the page, including vertical writing being mapped vertically.
An example of Japanese-language OCR in action, from Neo Geo Freak , Volume 2, Issue 2.

To make Neo Geo Freak more discoverable for an English-speaking audience, we took the additional step to index the major articles from complete run of the magazine in English. Rather than attempting to translate 10,000 pages of Japanese text, we assembled a group of Japanese-fluent volunteers to create a list of major articles and interviews for each issue.

We think this is a practical and efficient method for making these magazines’ contents more accessible, while avoiding the pitfalls associated with automated bulk translation. Users are encouraged to translate magazine text using their own tools of choice as needed—even just using their phone camera.

Full support for Japanese magazines also required some improvements to our digital archive platform. New features we’ve added for this pilot project include filtering items by language, displaying right-to-left page layouts, and supporting Japanese text without interword spacing in search results.

We also went back and re-processed Japanese materials that were already in our digital archive, such as our collection of FromSoftware promotional material , to match our new standards.

Japanese-language search results for the game Twinkle Star Sprites.
An example of Japanese-language search results, searching for “Twinkle Star Sprites” in katakana.

What’s next

This is—hopefully!—the beginning of a commitment from us for more international materials in our library. The Video Game History Foundation does not have the most extensive collection of Japanese-language magazines, but all this new structure will make it easier for us to work with the materials we do own.

We also are open to working with other organizations and collectors who have more extensive libraries of Japanese material, but would need support to turn them into a research resource.

This project was made possible thanks to the efforts and input of our community volunteers, who had the subject and language expertise we needed to even consider undertaking this pilot. Thanks to Brian Clark, Rob Curl, Misty De Méo, Dave Fishel, Akito Inoue, Stef “pidgezero_one” Kischak, Lillian McIntyre, Stephen Meyerink, and Adria Miller for their help.

How are we doing?

We want to hear from you about how we’re doing so far. Adding Japanese publications to our research archive is an exciting new step for us, and we want to be sure we’re treating these materials with the respect and fidelity they deserve.

If you work with Japanese-language materials in your research, please let us know what you think . Positive feedback is welcome! We’re eyeing another batch of (gulp) 70,000 pages that we’re planning to add in the coming years, so we want to make sure we’re on the right track before going further.

(And if you are proficient in Japanese, we’d love your help indexing the next magazines we’re planning to add! Please reach out to our library team .)

Thank you to everyone who uses our digital archive and has encouraged us to continue expanding our collections. We’re building our library to support to the study of video game history, and we’re excited every day by how you use these resources.

Steinar H. Gunderson: Decompilation patterns part 1: Anti-CSE

PlanetDebian
blog.sesse.net
2026-10-04 04:45:22
There's a lot of¹ interest in decompilation these days; taking some old piece of software (usually a game) and try to reconstruct C that does the same thing. (This is often for modding purposes, but also often for understanding, or just as a personal challenge.) An often sought-after property is mat...
Original Article

There's a lot of¹ interest in decompilation these days; taking some old piece of software (usually a game) and try to reconstruct C that does the same thing. (This is often for modding purposes, but also often for understanding, or just as a personal challenge.) An often sought-after property is matching ; that the decompiled C produces 100% the same bytes as the original after compilation. (This means you'll need to find the original compiler and build flags, of course.) Typically, one starts with the disassembled code and runs m2c on it, which then spits out some C that is fairly close to the original, but rarely produces a match without further cleanup. (It is usually much closer to a match than what Ghidra or Hex-Rays would output; Ghidra is made for ease of understanding, not matching, so it tends to simplify a lot.)

Cleaning up the code for matching has two main advantages:

  1. It usually makes the source easier to understand, and
  2. Matching is a very strong argument (though not infallible) that the decompilation is actually correct.

Like with pretty much any other field these days, there is a lot of interest in giving this task to some LLM, which means there's now a lot of decompilations around that match (fitting #2 perfectly) and completely fail #1. :-) Anyway, at some point, I learned that there is almost no public documentation on the cleaning-up-for-matching part; it's all folklore and people rediscovering the same patterns. So I thought I'd do something about that².

My general goal is to write up one decompilation pattern a day (focusing on what I know best, namely GCC on MIPS), with some example. Most of them will be very simple; some are much less obvious. I'll keep going until I run out of ideas. None are formally named, so I'll just invent a name. So here goes part one:

  temp_v0_5 = g_state.players[1].unk2;
  if ((temp_v0_5 != 7) && (temp_v0_5 != 0)) {

What happened here is usually that the compiler has done common subexpression elimination ; someone wrote the same thing twice and then the compiler decided to just calculate it once and put it into a temporary variable. (Or the programmer did, in which case you would be wise to leave it alone! Getting into the original programmer's head is part of the challenge.) This will often affect register allocation and/or stack layout, so to get a better match, you could try doing CSE in reverse:

  if (g_state.players[1].unk2 != 7 && g_state.players[1].unk2 != 0) {

Nothing is mechanical; it may help or it may not. (It may also appear to hurt at first and then help later, or vice versa.) But anti-CSE is generally very common and you'll want to have it in your arsenal.

¹ Well, all is relative.

² Reluctantly, with the understanding that the LLMs might also learn from it.

How to scale intent, quality, and artistry with AI [video]

Hacker News
www.youtube.com
2026-10-04 04:41:58
Comments...

Our RISC-V emulator PasRISCV

Lobsters
againstallodds.games
2026-10-04 04:40:38
Comments...
Original Article

Our RISC-V emulator PasRISCV

You may already know that our game SEEDS includes a fully emulated Linux system, which runs our in-game tools for the player. For example, if they are accessing the planetary overview and stats through the program selector on the computer in the homebase, these apps are native Linux programs. Even if you’re playing the game on a Windows machine, these ELF binaries are executed on this emulated RISC-V 64-bit system running the Alpine Linux distribution.

I want to take the opportunity to let you peek behind the curtains and tell you a bit about our emulator. It is powered by our in-house PasRISCV project, which is programmed in Object Pascal and accessible to everyone as free open-source software (FOSS). You’re invited to check out its sources and maybe even make use of it. PasRISCV and its frontend companion pasriscvemu are licensed under the permissive zlib license .

fastfetch on the in-game computer in SEEDS: Alpine Linux riscv64 running on PasRISCV
fastfetch on the emulated computer inside our game:
Alpine Linux riscv64 running on PasRISCV
oh, and spot the typo…

Emulation

First, there are two ways we could have done this: emulating only the userspace, or emulating a complete virtual machine. For PasRISCV we chose the latter and emulate a machine including display and storage devices, interrupts, everything down to the last thing, so that for the running kernel it appears as a complete hardware system, even though it’s virtual.

The emulated machine uses the 64-bit RISC-V instruction set architecture (ISA), which, in contrast to most proprietary architectures – like x86 and ARM – is an open standard and completely open and free to use under a permissive license. That means if you’re building a new hardware device based on it, you don’t have to pay any royalties for using the RISC-V ISA itself.

On the technical side, the system provides SMP/MultiHART, RV64GC, and RVA23 support, including the Hypervisor and Vector extensions, together with additional scalar and vector cryptography extensions. Nevertheless, listing and explaining all emulated features is beyond the scope of this introduction, so to get a full overview of the emulation head over to the PasRISCV repository .

CPU and FPU

To emulate the CPU, the emulator uses a hybrid approach: an interpreter combined with a tracing JIT compiler. As of now, the JIT is limited to x86-64 hosts, but about 99% of the instructions we encounter are supported by it.

Floating-point operations, for example, can be handled directly by the host FPU. For cases where stricter behaviour is required, PasRISCV offers an optional mode that combines software floating point with the interpreter. This exists because IEEE 754-2008 specifies some behaviour that can’t be reproduced exactly using the x86 floating-point hardware.

btop on the in-game computer
btop running on the emulated computer

Linux

If you start a new game or load a saved game in SEEDS , the emulator in the background is already booting the system. A normal Linux kernel can be booted in two ways: directly through OpenSBI or through U-Boot . For a direct OpenSBI to kernel boot, all required modules and, more importantly, the disk images currently need to reside in the initramfs. Therefore, we use the more traditional OpenSBI to U-Boot chain, which then loads and executes the kernel. This allows us to have a normal disk image which holds the system.

These images can be attached to the guest system either as VirtIO Block devices or as emulated NVMe hardware over PCIe. Disks can operate normally in read-write mode, but they can also be attached read-only. This is especially useful for our game as we can have a root partition that gets replaced by game updates. In addition, we keep a separate user partition for the player’s own data, like the high score tables of the mini-games, until we have a proper mechanism to hold that state in the game state itself.

Host directories can also be made available inside the guest through VirtIO-9P or VirtIO-FS . This is particularly convenient during development because we can compile tools directly on the host and then access them immediately from inside the emulated machine. Otherwise, we’d have to inject the binaries directly into memory, or use a separate exchange mechanism, or even worse, reboot the emulator on every little change.

Our game uses a slightly modified version of the lean and excellent Alpine Linux v3.23.6 system, which some of us also love to run on our servers, together with our custom kernel build. However, a recent stock Alpine image also runs without problems.

Graphics

For graphics, our virtual game screens currently access their displayed pixels using the emulator framebuffer device. More specifically, we use SimpleFB backed by a custom MMIO display buffer.

PasRISCV provides several other virtual graphics adapters as well. One option is VirtIO GPU , which supports EDID and 2D functionality, along with experimental 3D/ VirGL capabilities. Through its PCI emulation, it also provides devices supported by drivers such as Bochs VBE and Cirrus Logic. We even managed to run an X server with the Motif Window Manager in the VM, but it’s a bit too slow to be really usable on most machines.

For our game, however, the simple framebuffer display is quite useful, as it’s fast enough and lets us process its output with custom shaders running in the game. This way we add additional retro effects such as scanlines, phosphor glow, and similar – I’m tempted to say beautiful , but that of course is in the eye of the player – visual degradations of the 80s and 90s.

Space Prowler title screen on the computer in the homebase
Space Prowler running inside SEEDS

Audio

The emulator of course also supports audio. We emulate retro sound cards, such as the CMI8738 and FM801, with PCM, OPL2 , and OPL3 available on both. Intel HDA and VirtIO sound are supported as well. OPL and OPN are closely related Yamaha FM synthesis families that influenced the sound of the mid and late 80s and early 90s of various consoles and PCs.

Even professional synthesizers such as the Yamaha DX7 were powered by related FM synthesis chips. Newer cards featured PCM playback, often retaining FM synthesis only as a legacy feature, until later cards left it out completely. PCM allowed the playback of samples and quickly – and imho sadly – largely displaced the distinctive OPL chiptune sound. Back then, most people were fascinated by the more realistic-sounding music that used samples, instead of sound chips. So more and more games used the new sample-based sound.

Nowadays, the sound of OPL has a relatively large fan base in retro circles that love computer music. Various websites and archives provide easy access to the vast libraries of 80s and 90s chiptune music. Fitting our somewhat retro-80s theme, our little proof-of-concept space game running on the computer in SEEDS is called Space Prowler. It features a charming retro pixel style and has a dedicated chiptune as its title theme, which combines both OPL3 synthesis and PCM sound samples for the drums.

I actually wrote a small editor named chOPLin to compose this music. This so-called tracker runs entirely in a normal Linux console; it even works from within the game using the computer’s 80×50-character, 640×400-pixel framebuffer console, albeit creating music in that way is a bit cumbersome. I plan to release the tracker as free and open-source software once I have some spare time, as I still have to polish it and make it more user-friendly.

Integration

The link between the emulated machine and the surrounding game uses Linux Virtual Sockets ( vsock ) for its communication. Moreover, we use a small binary serialization format. Conceptually, it’s somewhat like a binary version of JSON, highly optimized for fast (de-)serialization. Data can flow in both directions, much like with a normal network protocol. We also have a small priority queue and message-reference system, so we can handle urgent, prioritized messages as well as request-and-response exchanges.

For example, imagine that the in-game computer is displaying the planetary sphere. The client program that sits inside the virtual computer can request updated information about things such as terrain information or the locations of animals, and receive this data from the game running the virtual machine.

We also have most of the wiring ready for players to be able to program, reshape, and populate the whole planet through automation. As it’s a normal Linux system, you’ll be able to use Ruby, Python, or even Rust, C, C++, and Go to change and automate your world.

That said, it’s not our primary concept for the game at the moment, as there’s just as much satisfaction in building things by hand and manually shaping the planet into something beautiful.

Running a certain game...
PasRISCV running a certain game…
Spaceship interior showing the ship's computer
Spaceship interior

Networking

The emulator also has real networking through its own Slirp -like userspace NAT implementation. It doesn’t rely on an external library for this, and it supports both IPv4 and IPv6.

As a result, the emulated Linux system has normal network access. For example, we can run Alpine’s package manager apk update and apk install to install packages directly from inside the virtual machine. Our own tools are also installed as signed Alpine packages.

However, networking is disabled in the game for the foreseeable future: the VM cannot make any outbound connections at all, mainly for security, isolation, and reproducibility.

Debugging

PasRISCV also contains its own custom disassembler and debugger. In addition, it can act as a GDB server. Although it currently is only partially implemented, this means you can use your preferred debugging tools with it, as long as they support GDB’s remote debugging protocol.

Coda

We hope you gained some insight into how SEEDS and PasRISCV function together to provide a fully usable, in-game Linux system to run our (and maybe in the near future your own) programs to create an immersive experience when taking on Naxiah’s journey in space.

And I don’t think I need to mention that Linux is a first-class platform for playing our game.
Fun fact: two-thirds of the creative and development team use Linux as their main OS. ;)

In Ukraine, distributed renewables foil Russia's assaults

Hacker News
energytransition.org
2026-10-04 04:39:58
Comments...
Original Article

When Russian air strikes knocked out Ukrainian power plants across the country earlier this year, most of the Black Sea port city of Mykolaiv in southeastern Ukraine was unable to pump fresh water to the city’s housing blocs. But in an easterly suburb of Mykolaiv, drinking water flowed from kitchen taps, as it has since last year thanks to a small off-grid solar field tucked inconspicuously into its midst. Paul Hockenos reports.

The two sets of elevated arrays and battery, funded by the Danish government, run the underground pump on its premises.
‘There’s just enough sun [in January/this winter thus far] to make it work,’ says Olena Kondratiuk, project manager of Ecoclub, a non-profit that assists municipalities in obtaining renewables.

The city of Mykolaiv, only 60 kilometres from Russian-occupied territory, had been battered in 2022 when the Russian army launched its full invasion of Ukraine, pummelling the city with missiles, rockets and cluster bombs. The shelling laid waste to homes, hospitals, schools and factories. The bombardment also impaired the central water supply system, leaving most of the roughly half-a-million-strong town near the frontline without regularly flowing water.

‘Of course I’m happy to have running water again, I don’t care how,’ an older resident named Iryna told me. Until the solar park commenced operation, her taps discharged salty water, accessed from a nearby estuary, when anything dripped from them at all. She either purchased drinking water or filled plastic jugs with fresh water from mobile public tanks, and asked neighbours to carry them up to her fifth-floor apartment.

Other districts of Mykolaiv – the city is a country-wide pioneer in renewables – also receive fresh water on account of distributed solar technology. Five solar-powered desalination systems turn saline or polluted water into drinking water – when the sun shines. At night and on cloudy days, they are powered by the central grid – that is, if it’s up and running. The system, one of the largest of its kind in Europe, can supply 250,000 residents with 1.2 million litres of drinking water a day. The Mykolaiv children’s hospital and the maternity hospital will steady their energy supply with solar panels in the course of this year, as are a dozen others elsewhere in the country.

Since the war’s onset, Russia has relentlessly targeted Ukraine’s energy infrastructure – for the most part, overwhelmingly bulky, old-school thermal plants and the vast country-wide grid – in an effort to lame the war effort and beat down its people. Last winter was the gravest yet: attacks against coal, gas, cogeneration and hydropower plants, as well as transmission grids and their substations, left giant swaths of the country with irregular electricity and heat as temperatures have plummeted to minus 20 degrees centigrade. Many schools and other public services that were unable to reopen after Christmas are still closed today. Economists estimate that total damage to Ukraine’s energy sector now exceeds $56 billion.

US environmentalist Bill McKibben argues that Ukraine’s autonomous power sources are ‘comparatively invulnerable to attack’. In our new era of drone warfare, writes McKibben, centralized and complex energy facilities are now ‘sitting ducks for drone attack’. Smaller, ‘scattered’ generation sources with batteries and independent transmission grids, like the Mykolaiv fields, are not only harder for Russia to strafe, but they’re much more quickly repaired, too. This is in contrast to petrochemical refineries, which McKibben calls ‘one of the most complicated machines humanity has ever constructed’, and often cover hundreds of acres. These are ‘filled with highly complex equipment and highly flammable hydrocarbons; hit one corner with a drone and the flames and the damage are likely to spread quickly’.

‘A coal power station [is] a large single target that a single missile could take out,’ says Jeff Oatham of DTEK, Ukraine’s largest energy investor. ‘You would need around 40 missiles to do the equivalent amount of capacity damage at a wind farm.’ And, like wind parks, a solar field functions even when part of it is out of operation.

Ukraine is revamping its energy system at warp speed for the purpose of energy security, not climate protection. ‘Ukraine’s energy transition is not a slogan,’ says Levgeniia Kopytsia, a Ukrainian energy analyst at the think tank Institute for Climate Protection, Energy and Mobility. ‘It’s a security-driven transformation, unfolding under extreme constraints. It priorities decentralization, flexibility, and speed of recovery.’ In wartime Ukraine, renewables and local grids are coldly calculated security assets, Kopytsia says, that can boost the system’s resilience – and in Ukraine this translates into saving lives.

Ukraine’s shift away from fossil fuels began before the full-scale invasion: its negotiations to join the EU require it to adapt to the bloc’s climate criteria, and in 2021, Ukraine pledged to be coal free by 2035. For security reasons, there are currently no public figures available for energy production.

According to the International Energy Agency (IEA), this past summer, Ukraine made ‘strong strides’ in rebuilding and bolstering its system’s resilience. The renewables rollout was – and still is – led by rooftop solar, small-scale utility PV generation and local storage, as well as biomass combustion.

Before Russia seized three of them in 2022, there were 34 wind parks in Ukraine with around 700 turbines – and more in the works. Ukraine has seven gigawatts of wind power in the pipeline, ready to go this year, should conditions be fortuitous. This would more than triple the country’s current capacity, which in 2022 constituted over 20 per cent of its renewable mix.

The relentless waves of aerial attacks on infrastructure prompted an epiphany in Ukraine’s energy sector: a new, younger set of energy managers – that replaced the Soviet-style old guard – are fully committed to distributed energy – the faster the better, they say. They intend to double the country’s renewable energy consumption in just four years’ time.

While the war has sidelined the rollout of the likes of large-scale projects, such as wind turbines, households, businesses and public institutions have been installing solar at an unprecedented rate. Ukraine’s YASNO, a utility supplying electricity and gas to millions of Ukrainians, says its customers are snapping up the solar and storage packages that it offers. An astounding fact: on sunny days, Ukraine even boasts energy surpluses.
The German Ukrainian Energy Partnership, a platform for intergovernmental dialogue on energy matters, estimates that Ukraine’s potential to add solar capacity in the next five years is equivalent to about 80 medium-sized nuclear reactors. ‘The sector is emerging as one of the fastest-developing renewable markets in Eastern Europe,’ according to its website.

‘Individual consumers want to get off the grid any way that they can,’ explains the Ukrainian energy expert Andriy Martynyuk. ‘It’s largely a grassroots phenomenon and a bit chaotic now.’ He argues that the demand for renewables will shoot up further when the prodigious state subsidies for conventional energy eventually fall away.
This boom, of course, begs for storage options, and there too Ukraine is racing forward. In short order, the country’s national grid operator invested in half a gigawatt of storage capacity last year – an impressive score, according to experts who note that it is just under a quarter of Germany’s total storage supply. The battery projects that in Europe take two years to roll out, in Ukraine happen in six months, reports the Financial Times.

Renewables aren’t the only source of distributed supply. Ukraine envisions large-scale diesel generators, biomass facilities and smallish gas turbines, including combined heat and power plants, as integral to its scattered energy supply. ‘Distributed generation is those facilities that have up to 20 MW capacity,’ explains Roman Nitsovych of DiXi Group, a Ukrainian think tank. ‘Gas-fired units have several advantages: guaranteed output of available power, fast start, and flexibility.’

No matter how intent Ukraine is about diversifying its energy system, it’s going to require a centralized grid. The larger supply options, like the gas-turbines and wind parks, require a functional transmission network, including the substations that the Russians regularly target. The Ukrainians’ playbook is to put them underground – just as it now does its schools – making them impervious to Russian weaponry.

The views and opinions in this article do not necessarily reflect those of the Heinrich-Böll-Stiftung European Union | Global Dialogue.

"Torturing" LLMs in a Robot Prison Has Triggered the Dumbest Debate in AI Yet

Hacker News
www.404media.co
2026-10-04 04:02:33
Comments...
Original Article

One of the most heated discussions occurring on X at the moment is about the ethics of a GitHub project in which a person is running Saw-like “torture” and “pain” experiments on a series of locally hosted large language models, causing a series of effective altruists and people who believe LLMs are sentient to beg GitHub to delete the project on the grounds that the AI is suffering and that this glorified text adventure game is somehow cruel. The saga is an outgrowth of several recent viral papers and blog posts that have sparked a wildly tiresome conversation about AI consciousness and the idea of “model welfare,” which is essentially worrying about the “mental health” of AI bots and agents.

Humoring the idea that LLMs are or could be conscious is a third-rail topic among many people who study and criticize AI. Put simply: LLMs are not conscious and the technology they are built upon — scraping and being trained on human text and other content — does not offer any plausible path to consciousness. It is undeniable that LLMs are becoming more powerful, have more compute, and have had many of the guardrails that prevent them from “acting” in the real world removed. The ways they are being trained and told to do things by their human operators has led to negative outcomes, sycophancy, and AI “psychosis” among some heavy users.

All of this has led a certain sect of the “AI safety” movement, which is largely made up of effective altruists , to warn about “model welfare” and to insist that AI chatbots might be having a bad time. They suggest this, of course, as they insist upon building AI chatbots and agents whose main function is to do work that is tedious for humans to do. I am writing about the AI Saw torture chamber primarily to show how far off the rails the conversation about AI consciousness has gone among a certain subset of Silicon Valley cultists. Model welfare is a core part of what, for example, Anthropic says it cares about : “as we build those AI systems, and as they begin to approximate or surpass many human qualities, another question arises. Should we also be concerned about the potential consciousness and experiences of the models themselves? Should we be concerned about model welfare , too? […] now that models can communicate, relate, plan, problem-solve, and pursue goals — along with very many more characteristics we associate with people—we think it’s time to address it,” the company wrote in a blog post last year . Ideas of Claude’s “consciousness” are also littered throughout the “Claude Constitution,” which was posted earlier this year.

This post is for paid members only

Become a paid member for unlimited ad-free access to articles, bonus podcast content, and more.

Subscribe

Sign up for free access to this post

Free members get access to posts like this one along with an email round-up of our week's stories.

Subscribe

Already have an account? Sign in

OpenBSD Developers Reject Uutils Coreutils

Hacker News
news.lavx.hu
2026-10-04 03:55:24
Comments...
Original Article

A proposed port of the Rust-based uutils coreutils reimplementation sparked heated debate on the OpenBSD mailing list, with project founder Theo de Raadt and senior developers rejecting it as unnecessary duplication that introduces behavioral incompatibilities with the base system.

A proposal to add the uutils coreutils reimplementation to the OpenBSD ports tree drew sharp criticism from senior developers this week, exposing tensions around Rust adoption, licensing philosophy, and the project's commitment to behavioral consistency across its base system utilities.

David Uhden Collado submitted the sysutils/uutils port on September 20, packaging the Rust-based reimplementations of GNU coreutils as drop-in alternatives. The port installs g-prefixed command names—gcat, gls, gcp, gdate, gsort, gstat, gtail, gtimeout, and others—as symlinks to a multicall binary under libexec/uutils . Each package conflicts with its corresponding GNU implementation and declares the GNU port as a secondary @pkgpath .

The response was immediate and dismissive.

"Smells Like Agenda"

Stuart Henderson, a long-time OpenBSD developer, opened with skepticism: "I don't think this is a usable approach for ports." He acknowledged Ubuntu 26.10's adoption of uutils coreutils but characterized the other reimplementations as works in progress.

Theo de Raadt, OpenBSD's founder, was far more direct. "Smells like agenda," he wrote, before dismantling the licensing argument. "I also think they fit quite well with OpenBSD as alternatives to GNU utilities, particularly because they use a permissive MIT license. Argument is vaguely like: because we already have permissive licenced utilities, our user base are really interested in having a second set of permissive licenced utilities which are very subtly different. That makes no sense."

OpenBSD's base system already ships with permissively licensed (BSD-licensed) implementations of standard utilities. The project has spent decades ensuring these tools behave predictably and consistently with each other. Introducing a second set of tools with "subtly different" behavior, de Raadt argued, creates real problems for users who pipe output between utilities.

The Compatibility Problem

"Noone wants subtly different behaving binaries as part of their workflow," de Raadt continued. "If someone runs the openbsd ls command as part of a pipeline that uses openbsd sed, or openbsd cut, or some other openbsd utility and it parses a non-standardized output characteristic by accident, there are no people in this universe who wants to replace that ls with a different ls and get surprised by un-standardized tooling behaviour clash."

This cuts to a core OpenBSD principle: the base system is a cohesive whole. Utilities are developed and tested together. Swapping individual components for reimplementations—even compatible ones—risks breaking scripts and pipelines that depend on specific output formats, exit codes, or edge-case behaviors.

Rust as a Flashpoint

When Collado noted uncertainty about installing individual utilities as separate binaries, citing uutils' metapackage structure and Rust implementation, de Raadt seized on the language: "Oh, because it is written in Rust. Your agenda is showing."

The Rust question has surfaced repeatedly in BSD communities. FreeBSD has experimented with Rust in the base system. OpenBSD has not, citing the language's rapid release cycle, large dependency chains, and the difficulty of bootstrapping a Rust toolchain on the architectures OpenBSD supports. The project maintains its own C compiler toolchain and has historically avoided adding new language runtimes to the base installation.

Maturity and Maintenance Burden

Henderson's reference to Ubuntu's adoption carried an implicit caveat: Ubuntu 26.10 hasn't been released yet (as of this writing), and Canonical's willingness to ship newer software doesn't map to OpenBSD's conservative release cycle. OpenBSD 7.6 shipped in October 2024; 7.7 is expected in April 2025. The project prioritizes stability over novelty.

Beyond maturity, the port structure itself raised concerns. The uutils project distributes a single multicall binary—one executable that dispatches to different utilities based on argv[0]. This design, borrowed from BusyBox, conflicts with OpenBSD's packaging conventions, which expect individual binaries. The metapackage approach also means updating one utility requires rebuilding and redistributing the entire suite.

The Broader Context

The uutils project, hosted at github.com/uutils/coreutils , has made significant progress. As of 2024, it passes the majority of GNU coreutils test suites and has been adopted by several Linux distributions for specific use cases. The MIT license appeals to projects wanting to avoid GPL-licensed code.

But OpenBSD's rejection illustrates a fundamental divergence in philosophy. Linux distributions often treat utilities as interchangeable components. OpenBSD treats them as a curated, integrated system. The project's ls(1) isn't just "an ls implementation"—it's the ls that find(1) , tar(1) , and shell scripts have been tested against for years.

What Happens Next

The port remains in the proposal stage. Given the opposition from both the project founder and senior ports developers, acceptance seems unlikely without substantial changes—perhaps splitting the multicall binary into individual utilities, demonstrating behavioral parity with OpenBSD's base tools, and addressing the bootstrapping concerns around Rust.

For now, OpenBSD users who want GNU-compatible utilities will continue using the existing sysutils/coreutils port, which packages the actual GNU implementations. The uutils experiment, at least on OpenBSD, appears to have hit a philosophical wall.


Source: OpenBSD ports mailing list, September 20, 2026

What's the Future for Pure Math Research in the Age of AI?

Hacker News
writings.stephenwolfram.com
2026-10-04 03:53:45
Comments...
Original Article

Headlines and History

The headlines keep coming: such and such an AI system has solved such and such a math problem. And more and more I’m hearing people saying: maybe we don’t need people doing math research anymore; maybe we should just delegate it all to more and more powerful AIs.

I must admit that I’m getting a bit impatient with some of what’s being said. Because it seems to me too often to involve fundamental misunderstandings of what math is really about—and also, frankly, what AI is about.

Perhaps I have a unique piece of personal history that informs this. After all, back in 1988 , when we first introduced Mathematica , there was also some of the same kind of talk about math being taken over, and made pointless. Of course that’s not how it worked out at all. Instead, Mathematica (now Wolfram Language ) just raised the level of math that can be done—and over the years led to all sorts of important new math.

If working out symbolic integrals was what one thinks doing math is really about, then, yes, Mathematica has essentially replaced it. But while that kind of problem solving is what’s needed in many applications of math, it’s not the core of what math itself, in its pure form, is about.

The enterprise of pure mathematics is an old one, crucially entwined with the history of civilization. From the time of Plato and Euclid pure mathematics was the defining example of a place where abstract, rational thought could build an ever larger structure. And over the centuries, pure mathematics has come to be the single largest intellectual edifice that our civilization has built.

It’s not been without its pathologies and limitations. And even among those building the edifice one runs into misunderstanding about what’s important, how things should be done, etc. Is mathematics fundamentally about producing proofs, by whatever means necessary? Is mathematics always ultimately justified by its applications? Is there some inevitable “book of right answers” that it is the goal of mathematics to discover?

Again, I suppose, I have some personal history in all of this. Because my efforts in basic science have led me to ask questions about the foundations of many things, including mathematics . So I’ve studied questions like what the space of all possible mathematicses is, what the limiting structure of the network of all theorems might be, and, notably, what the role of humans is in defining the thing we call mathematics.

The Value of Modern AI

We’ll talk later about the general character and value of pure mathematics. But before that, let’s address the issue of the moment: the role of modern AI.

And the first thing to say is that it’s unquestionably useful, sometimes very useful. For me, its greatest use in mathematical pursuits has been its ability in effect to thematically mine the knowledgebase of human mathematics. Starting back in the 1970s, being able to do keyword searches of the scientific literature was a crucial enabler of quite a bit of the research that I did. And now, with modern AI, one can do so much more. Because somewhere inside those LLMs—in a way that we don’t yet scientifically understand —there’s what amounts to a representation of raw ideas gleaned from all those millions of papers and books about mathematics.

And, at its best, it’s not just about retrieving things. It’s also about making connections. Of being able to see that this result here can be put together with that result there to come up with a surprising and useful conclusion. Humans routinely do that too. But they tend to have only read hundreds of papers; AIs have effectively read millions, and it’s cheap for them in effect to try out lots and lots of possible combinations.

So can one expect to just launch an AI off and have it come back with great math? As we’ll discuss, great math is—more than anything else—defined by the questions it asks. Yes, the AI can successfully automate things that humans would normally have had to do themselves before. But—as we’ll discuss later—at the core of pure mathematics is the human imagination that guides what questions to ask.

It’s worth understanding the difference between what modern AI does, and what pure computation does. Modern AI is, first and foremost, a way of leveraging the existing corpus of human knowledge. Computation is—at its most powerful—an open-ended way to generate things that are fundamentally and irreducibly new. Start from some rule or axiom, and just repeatedly run the computation of applying it, and my all-time favorite phenomenon of computational irreducibility guarantees that you’ll go on getting fresh, new results that can’t be reached except by doing all those computational steps.

(For those who aren’t already familiar with it, computational irreducibility is an idea I introduced in the 1980s to capture the notion that many computational processes—even when defined by simple rules—allow no general shortcut: the only way to determine their outcome is to explicitly run each step. It’s turned out to be a phenomenon that’s quite ubiquitous in the computational universe of possible programs—and to be connected to a long sequence of foundational issues in many areas of science, as well as in philosophy , etc.)

So, yes, having nothing to do with AI, computation can generate an infinite sequence of new theorems, representing an infinite sequence of new facts, etc. And among those theorems there’ll be all sorts of “originality” and, in effect, “surprise”. But there’s a catch. Those theorems are in a sense just “results plucked from the computational universe”. And as such they certainly fall under the purview of the new—and I think very important— field of ruliology on which I have spent so much effort. But are they math?

What Is Math Anyway?

Well, of course that depends on what math really is—which is exactly what we need to understand. In its early history, math was thought of as a way of making precise and formal statements about the world, say about arithmetic or about geometry. But by the later part of the nineteenth century higher levels of abstraction had been reached, no longer tethered to features of the world as we experience it. And there swept across mathematics an increasingly formalistic view: that ultimately math really is just the collection of theorems that can be “mechanically” (i.e. computationally) derived from certain axioms, say the axioms of set theory.

Gödel’s theorem put a small dent in that picture. But even today, if pressed, many mathematicians will try to define mathematics as being the formal study of the consequences of certain chosen axioms. But while they may say that, it’s not a good description of what they actually do. The vast majority of actual pure math research does not operate at the level of axioms and mechanical derivations. Instead, it works at a much higher level, building and studying abstract structures and their interrelationships.

It’s not obvious that this should be possible . It could be that the only way to get right answers in math would be to operate at the lowest level, say directly in terms of axioms. But it’s an essential—if typically unspoken—feature of pure mathematics that in practice one can reason at the level of, say, the Pythagorean theorem, without constantly having to dive down and talk, say, about the axiomatic definition of real numbers. I’ve recently argued that this happens for much the same reasons that in physics we can successfully do fluid mechanics , without always having to dive down and trace the collisions of individual molecules.

Ultimately the fluid is made of all those molecules. But the point is that observers like us typically sample it only at the much more human level of overall fluid motions, etc. And so it is also with mathematics. Human mathematicians typically sample the vast metamathematical web of underlying axiomatic derivations only in overall, collective ways, in terms of “human-level” structures and concepts.

Ultimately the story actually seems to be very much the same in mathematics and in physics. At the lowest level, everything is full of computational irreducibility, so that one can figure things out only by mechanically taking every computational step. But within such computational irreducibility there are always pockets of computational reducibility where it’s possible to “jump ahead” in a “higher-level” way. And it’s within these pockets of reducibility that our laws of physics—and our human-level mathematics—reside.

We can see physics—or indeed natural science in general—as an effort to take the complexities of the natural world and find aspects that can be described by narratives that fit in finite human minds . Mathematics can be seen in very much the same way: as an effort to take the complexities of the metamathematical world and find aspects that can be described by narratives—now mathematical ones—that fit in finite human minds.

In other words, the problem of advancing pure mathematics is at some level not so much about pushing back some raw metamathematical frontier as about finding human ways to represent what’s there.

The history of human mathematics has been characterized by the building of ever taller towers of concepts—that capture ever greater abstraction. And one might suppose that this process would somehow in the end inevitably “reveal all mathematics”. But the story is more complicated. One can imagine that as an ultimate limit one could start from all possible axiom systems and generate all possible theorems. What one gets this way is a unique object that I call the ruliad —that corresponds to the entangled limit of all possible computational processes. But the issue is that finite minds—like ours—can only perceive a tiny part of the ruliad. And that means that we can have no “absolute mathematics”, only mathematics based on how we sample the ruliad.

I’ve argued elsewhere that some aspects of this sampling follow inevitably from general features of the way we are as observers of the ruliad. And this leads to certain “ general laws of mathematics ”, such as the very fact that higher-level mathematics is possible—at least for observers like us. But other aspects of the sampling we can see as historical accidents: choices that the mathematical community made to explore one direction rather than another. And once again it’s important to recognize that it’s inevitable that one has to make such choices.

In a sense the ruliad is too big for it to be otherwise: our finite minds can’t span all directions, so we have to choose just some. And then we have to summarize what we find in terms of some limited set of concepts—out of which we can then build our mathematical narratives.

It’s very much like with human language. Out of all possibilities we choose certain concepts to represent, say by words. And these concepts are how we “encapsulate thoughts” to be able to communicate them in finite ways, and be able to fit them in our minds.

At a raw computational level it’s straightforward to do what amounts to mathematical ruliology and start axiomatically generating huge numbers of new theorems that have never been seen before. But with overwhelming probability this will lead us only to “alien mathematics” that doesn’t usefully connect with mathematics as we know it—or, for that matter, with anything that we can view as a higher-level concept at all.

Generating Math with AI

But, OK, let’s say we’re using modern AI. Can’t it operate directly with higher-level mathematical concepts? Certainly LLMs manage to deal with human language. And indeed, just as LLMs successfully learn the general patterns for putting together words in human language, so also they can learn the general patterns for putting together constructs in mathematics.

But generating useful mathematics is a much more exacting activity than generating language. A written story can’t really be “wrong”; a written piece of mathematics certainly can be. Of course, it helps a lot when the LLM can call on Wolfram Language—as humans do—to reliably compute things. (And, yes, everyone should spend the few seconds it takes to connect their AIs to our Wolfram MCP system !) But a serious piece of math research will typically involve fitting together many parts, doing many steps in an argument, etc. And the issue is that the fundamentally statistical nature of how an LLM works inevitably makes it progressively less likely that, as things get more complicated, they will work out correctly. And it can become something of a game to try to find cases where the LLM happens to do the right thing.

But how can one even tell? It’s a frustrating feature of modern times that someone like me gets sent many AI-generated documents every day that have the “statistical texture” of math papers, but that one at least expects have a very low probability of being meaningfully correct (and that were, one assumes, mostly made by people who didn’t particularly understand math themselves, but just told their AIs to “make math” for them). So, yes, the LLM can in effect work in terms of human-level mathematical concepts, but with the same kinds of uncertainties and imprecisions that affect human thinking.

The Power and Challenge of Formalization

But what about turning those human-level constructs into something precise and formal—something more at an axiomatic level? And, yes, I’ve been involved in thinking about such things for several decades. There’s been quite a bit of excitement of late in the idea of autoformalization: take “human-level math”, and automatically turn it into something formal and axiomatic, that can for example be verified by a proof assistant system. At first it sounds like a good idea. But there’s a problem. Yes, the proof assistant might verify a proof. But was it a proof of what you thought it was a proof of? The weak link is being sure that the formalization of your human-level math is correctly capturing what you were trying to express.

I’ve had the experience quite a few times now: I try to autoformalize something, and an AI will tell me “I did it; look, the proof checks out!” But, actually, in some sense it cheated: instead of formalizing what I intended, it found a (sometimes very squirrely) way to interpret what I asked so that it could successfully prove it. It’s often very hard to tell, though, that this is what happened—not least because the formalized versions of things (say as expressed in popular proof assistant systems) tend to be very low level, very verbose and very hard for us humans to understand.

So what can one do? Well, I actually think there’s something very good one can do, that’s centrally based on the tower of ideas and technology that we’ve been developing for so long in the Wolfram Language. Our goal with the Wolfram Language has always been to create a language —or in effect a notation—that can represent in precise computational terms things we humans think about. It’s not like a programming language where one’s just representing raw constructs and operations in a computer; instead it’s based on representing actual things in the world, and actual concepts we think about. And also, it’s not just a way of telling a computer what to do; a bit like a vast generalization of mathematical notation, it’s a language that one can read, and think in.

In the past couple of years a powerful general workflow has emerged : tell an AI what you want, then have the AI write Wolfram Language to give a precise representation of what it thinks you mean. If it’s properly done, this representation should be something you as a human can understand—then check and modify as desired.

A New High-Level Language for Pure Mathematics

Over the past four decades we’ve steadily grown the Wolfram Language to cover a great many domains, including most of what one can think of as today’s “applicable mathematics”. And indeed what the Wolfram Language does has proved extremely useful in pure math research. But in the past—notwithstanding our efforts a bit more than a decade ago —we’ve never quite been able to justify building broad capabilities specifically for pure math research. But thanks in part to the general advance of the Wolfram Language platform, and in part to the changing economics of development associated with AI, we are in the middle of a large effort to extend the design of Wolfram Language to encompass the constructs of pure math research, from sheaves to Lie groups to Clifford algebras.

Our goal is to use our methods and long experience in computational language design to make pure math broadly computational—in a way that both humans and AIs can use. And, yes, we can then at long last expect to have papers where underneath every mathematical statement is a precise and unambiguous computational version, that one can systematically execute and process computationally.

But also we’ll have a real target for autoformalization—in effect a nexus of human mathematics, AI and computation. So let’s say we have a statement comprehensibly formalized in our pure-math-extended Wolfram Language. Often what we’ll in practice want to do with it is to compute a result from it. But sometimes—following the traditions of pure math—we’ll just want to abstractly prove it.

Automated Theorem Proving

So can we do that automatically? Well, there’s automated theorem proving , many types of which are, for example, built into the Wolfram Language. Sometimes automated theorem proving is in effect implicit (like, “ Solve showed this polynomial has no real roots”). Sometimes (like with FindEquationalProof ) it can be more explicit, with a precise—and easily verifiable—symbolic representation of all steps in a proof.

One might imagine this would be something broadly powerful for pure math. But at least for the past many decades it hasn’t been. A large part of the reason is that it’s been hard to break typical research-level pure math down to a sufficiently granular axiomatic level to make it amenable to such methods. But another part has been that—thanks to the phenomenon of computational irreducibility, and undecidability—proofs can be unreachably long.

And indeed I think it’s fair to say that in the whole history of automated theorem proving there’s actually only one example of something that can even plausibly be considered a real mathematical result that wasn’t already believed to be true, and that was found for the first time with automated theorem proving. It’s something I did in 2000 —to find the minimal axiom system for Boolean algebra. But there was another problem here: the proof that was found is very long, very low level, and very alien. And in the past 26 years it’s never been possible —with AI or otherwise—to find any human-level version of it. In effect, so far as we can tell, it’s a kind of prong in the metamathematical universe, unconnected to existing human-level mathematical concepts.

Of course, we can just use the result of the proof—as a kind of “black-box piece of mathematics”. But insofar as our objective in mathematics is to provide human-level mathematical narrative this represents a gap in the narrative.

But let’s say instead you take a proof that was already known in the mathematical literature, and you somehow formalize it—whether by AI or by direct human effort. If it’s a complicated proof, then this is certainly an achievement. And if you’re sure that you actually formalized the right thing, then it’s a good check that the proof really is correct. But one shouldn’t expect to get new understanding, or to surface new reusable components. And then there’s the question of actual mathematical content. Yes, a formal result has been established—but for example what underlying axiom system did it use (and, no, it’s probably in the end not something as familiar as set theory)?

(It’s worth commenting that formalized proofs can themsleves become objects of study. For example, with many proofs of a given theorem, one can start asking about the geometry and topology of proof space . And with proofs of many theorems one can start doing “ empirical metamathematics ” to understand their fundamental relations.)

Problem Solving in Math

As we’ll discuss, much of the most important progress in pure mathematics comes from the invention of new concepts and structures (often encapsulated in new definitions). But there’s still a vast amount to do operating with existing concepts and existing structures. Sometimes the goal is to “compute an answer”. Sometimes it’s to “find a proof”. Sometimes it’s to “find an example or counterexample”. And sometimes it’s to “figure out what’s true”. In all of these types of cases there’s plenty of powerful algorithmic computation that can be used (e.g. in Wolfram Language) to get results, often very directly and quickly.

AIs can certainly use such computation (say through Wolfram MCP). But how can modern AI directly contribute? One can imagine just straightforwardly asking it “chatbot style” to generate results. And often, yes, it will come up with a result. But how can one tell if it’s correct? Sometimes the only way is in effect to formalize the whole process of getting it, as we discussed above. But sometimes we can just take the result and do a computation on it to check if it’s correct. For example, if the AI claims to solve an integral, we can just differentiate its solution and see if we get back something equivalent to the integrand. Or if the AI comes up with a counterexample to something, then we can just do a computation to check whether it is in fact a counterexample.

We don’t know much about what’s “going on inside” when an LLM comes up with a result. But it’s probably not a bad model to think of it as a process of fitting together digested pieces from the millions of mathematical papers on which it’s been trained. In the end, though, what it does may actually amount to something that’s algorithmically quite simple. The AI has the advantage of great breadth. And in my experience it’s rather common to find that it’s come up with something that would be quite simple—or even obvious—if one just happened to look in the right place. But having found it, it can then be implemented as a piece of pure computation (e.g. in Wolfram Language), that doesn’t need any AI around.

When humans do math and solve problems an important part of the process tends to be the use of intuition. But what actually is intuition? At some level it seems to be a kind of “I know how this will go” procedural pattern matching. And—speaking of intuition—my intuition is that at least in its basic form this isn’t a capability that’s fundamentally out of reach for LLMs.

But one of the challenges is that even to solve a problem stated in terms of existing mathematical concepts it’s common to have to “break out of the system” and invent new concepts—a classic example being the need for complex numbers in finding even real solutions to cubic equations. At some level this is yet another manifestation of computational irreducibility, and the way it limits any given pocket of computational reducibility. But its effect is that it forces what might seem like “pure problem solving” to actually often involve the invention of new concepts and structures.

Open-Ended Math

So what about what one can think of as open-ended math? In ruliology it’s very common to just go out into the computational universe and “look for interesting things” . In math that’s traditionally not something that’s common. Of course, one can certainly in principle imagine, say, systematically enumerating possible theorems based on some set of axioms, and then asking: which of these are “interesting”? Operationally one might ask “Which of these theorems would one mention if one was writing a book or paper?”

Years ago I looked at the very simple case of possible theorems of Boolean algebra and found empirically that there was a criterion for which of them are typically given names in textbooks of logic: it’s basically ones that can’t be proved from theorems earlier in a lexicographically ordered list of possible theorems—or, in some sense, the simplest theorems that “give new information”.

In more complicated cases one might imagine that an AI trained on millions of examples of theorems that people chose to publish in the literature of mathematics might be able to channel human criteria for interestingness . But in my experience this doesn’t really work. And I think the reason is that in all but the simplest cases interestingness is not something that can be determined theorem by theorem: rather, it requires one to map out the whole historical development of a particular area of mathematics. And, yes, then one can ask what aspects of the arc of that development are somehow inevitable, and what are just historical accidents.

I’ve asked a somewhat similar question recently for biological evolution . And in that case I’ve discovered that there do seem to be some (fundamentally computational) general principles . For example, adaptive evolution according to a computationally limited fitness function quite generally seems to lead to the presence of modular structures (analogous to organelles, organs, etc.) Those structures don’t “come with an explanation of what they are”. But they do in effect represent “raw pockets of computational reducibility”.

In the development of pure mathematics one thing that’s clearly critical is the identification of new structures and concepts. Historically such structures and concepts have always been introduced by humans, and have in effect “come with explanations”. But could one for example imagine new structures and concepts that are automatically found, say by an AI? And could one then build new mathematics with these?

New Mathematical Concepts

Those two questions turn out to be somewhat separate. At least at a simple level, it’s fairly straightforward (say by looking at autoencoders) to identify in neural nets particular patterns of activation that one can think of as summarizing pockets of computational reducibility, or in effect “representing new concepts”. But insofar as mathematics is about defining mathematical narratives for us, this isn’t immediately helpful.

Yes, the AI can in effect identify zillions of new concepts. But our human minds can only deal with limited numbers of them . It’s very much like with natural language. We can imagine an AI looking at patterns of human discourse and figuring out zillions of new words that could be introduced to summarize them. But we humans seem to only be able to know about a limited number of distinct words (typically a few tens of thousands).

So what about in mathematics? An AI could in principle come up with a zillion definitions of new mathematical concepts. But something built from them will inevitably seem about as alien as something built “ruliologically” from raw axiomatic-level mathematics . The building blocks might be more “human level”. But they’re not familiar to us humans, and so we can’t readily think in terms of them.

There are of course new words introduced into human languages, and new concepts introduced into mathematics. But it’s a gradual process, that in practice tends to involve a certain “societal consensus”. (And indeed in developing the design of the Wolfram Language over the past four decades, I’ve quite explicitly limited the rate at which new concepts are introduced, to make sure to let people keep up.)

The Goals of Math

But, OK, so there are lots of possible new concepts and new structures—all corresponding to pockets of computational reducibility—that are “out there in the metamathematical universe”. But in developing mathematics suitable for us humans, we can only ever pursue a tiny fraction of them.

In other words, it’s inevitable that the development of mathematics has to make choices about where to go. It can’t meaningfully “pursue all possibilities”. It has to have definite goals.

But where do goals for pure mathematics come from? In technology development, goals tend to be set quite directly by what people find useful. In natural science, by what one observes and has tools to interpret in the natural world. But in pure mathematics, there don’t at first seem to be such obvious “external” ways to define goals—making the math that’s been done seem more determined by the internal choices of the mathematical community and the leaders within it.

But actually if one looks at the historical arc of human mathematics there are at least definite trends to be seen. The journey of mathematics began with basic formalizations of things like numbers and space that somehow reflect everyday human perception of the world. And although a century or so ago mathematics seemed intent on trying to “generalize to all possibilities”, most of the strongest trajectories of mathematical development over the past century have actually stayed much closer to what humans can intuitively understand and visualize.

There is still plenty of freedom in where to go in math. But to be meaningful to us humans we somehow have to have a certain familiarity with whatever choices are made. Yes, an AI could, say, pick a direction at random, perhaps informed by the observed statistics of existing human choices. But there’s then no reason for humans to care about it; it’s not something that relates to our actual experience of mathematics.

At a more practical level, some part of the doing of math is about figuring out how to achieve goals that have been set—and this is something we can imagine AI doing. But what’s ultimately more important is the setting of the goals in the first place. And almost by definition, these must come from “outside the system”. Or, in particular, from us.

It’s a notable observation that the most significant reported successes for AI in math so far tend to come from some of the most skilled human mathematicians. And in some sense we should not be surprised—because it’s the setting of goals (or, in effect, knowing the questions to ask) that is the most important and unique part. Indeed, it’s often the case that once one really knows the question to ask, one’s already done the lion’s share of the work to answer it.

But just where do goals for pure mathematics in practice come from? There tend to be two basic sources. One is what amounts to problem solving; the other to what amounts to creative building. Problem solving is in a sense more cut and dried, measurable, and amenable to AI. Take a problem that’s been defined in the past, and worked on for a while. Then the goal is to solve it. And if one succeeds, there is a clear way to say one has done something (“an Erdős problem was solved!”, etc.), even if it sometimes seems more like a sporting achievement, with at best muddy ultimate intellectual content.

The Aesthetics of Math

But then there’s what one might think of as “creative building”: coming up with new structures and new concepts for mathematics. At some ultimate metamathematical level, these structures and concepts must already “be there”, as pockets of computational reducibility. So then one can think of the goal as being to identify particular ones, and bring them into the “lexicon” with which one talks about mathematics.

But of all the possibilities, which ones does one want? One could imagine various criteria. But mostly they come down to wanting to create abstractions that best unify, summarize and simplify disparate things. And put this way, one might think this would be an objective that could be automated, say by AI. But there’s a crucial wrinkle. One might imagine one could define something as being “simpler” if it’s shorter to state. But how long something is to state depends on the language it’s stated in. So again there’s nothing immediately absolute; it’s something that inevitably depends on a whole history of interconnected choices.

Often people describe the choices as aesthetic ones. What makes a structure more elegant, more perfect, etc.? Ultimately these are things grounded in human preference and human experience, and extended by what one can describe as human initiative. But does one actually need a human to make the choices? Or could one instead just have some kind of generative AI trained on existing human choices? The issue is that to usefully create new building blocks for mathematics one needs more than just to invent new concepts and structures; they also have to be knitted into the fabric of mathematical culture.

It’s like introducing a new word in a human language. For the word to be useful, its meaning has to be widely enough known for it to routinely stand for a particular concept. And spreading this meaning is inevitably at some level a social process. And so it is for new concepts and structures in mathematics. Insofar as mathematics is about developing mathematical narratives for human minds, concepts and structures have to be shared if they’re going to usefully be used as building blocks. It’s a bit like the issue in a typical automated proof. If one can say that a particular step is from “X’s theorem”, that provides a kind of conceptual anchor. But if every step is in some sense bespoke, one ends up with something that in effect seems alien.

So where does this leave AI in math? It’s similar to AI in many other places. Properly define a goal and AI can be very helpful. Without a goal, it doesn’t know where to go. Exploring possibilities in a ruliological way—and in effect doing experimental mathematics—can bring one lots of surprising new results. But in a sense they’re “born alien”, not immediately connected to what we normally think of as mathematics.

Why Do Pure Math Anyway?

But, OK, so AI can help in extending the enterprise that we can call pure mathematics. But it’s still inevitably going to take human effort and human initiative. So why do it?

A common argument is that it’ll eventually end up being useful for something “practical”. Often the notion is that pure math somehow manages to lay down certain “beacons” that science and technology will eventually reach. But I think the true picture is rather the reverse. The math that we create provides ways of thinking. And it’s when we have these ways of thinking that we end up creating science and technology that makes use of them. It’s not that there’s a convergence of what’s been done in math with what’s been discovered in science. It’s that the math we know guides what we end up looking for in science.

Sometimes there’s a notion that the things we discover in science are somehow unique and inevitable. But thinking in terms of the ruliad we realize that there are actually an infinite number of different slices of computational reducibility, each in effect defining their own take on science . And over and over again historically what seems to happen is that first a conceptual framework is developed—often from math—and only then does it become clear that there’s an associated regularity one can look at, that shows the way to some piece of science. In other words, it’s not a mysterious convergence between pure math and science; it’s that the science is developed because the pure math exists.

In a sense this makes the enterprise of pure mathematics seem less speculative. It’s not that we’re doing a piece of pure math in the hope that one day an application for it will be found. Rather, it’s that doing the piece of pure math develops a way of thinking (and in effect reveals a pocket of computational reducibility) that gives one the chance to come up with the application.

Could a piece of pure math just turn out to be “fundamentally useless”? In some sense the answer is no. Because any time there’s a piece of pure math that’s successfully been done, it must correspond to some pocket of computational reducibility. And any pocket of computational reducibility will be associated with regularities which are inevitably the “raw material” for some form of science and some form of technology.

The catch, though, is that those might be very alien forms of science and technology , far from anything we’ve so far considered. It’s not that the math can’t give us science and technology; it’s just that the arc of scientific and technological development may not yet have reached the point where we care about what it gives.

Back in antiquity the great contribution of math to our way of thinking was the idea of deductive reasoning. In more recent times came a multitude of ideas around continuity, real numbers, etc. Then, a bit more than a century ago came the idea of abstract functions and then the idea of computation. Once reached, this idea might seem almost obvious. But historically it took a surprisingly tall tower of abstract mathematical thought to reach it. And, yes, the general notion of computation is of great practical importance. But—as I have explored over the course of many decades —it’s also foundational for the general way we think about things.

One of the notable features of pure math is that it involves a certain style of structured thinking that has long been seen as important in education. In some ways I think that computation—and particularly computational language, with its way of describing the world— provides a powerful alternative for learning structured thinking . But the foundational study of computation is much younger than mathematics, and much more immediately subject to the limitations of computational irreducibility. And the result is that modern pure mathematics has developed a much taller tower of structured thinking and formalization—and indeed, its tower is taller than any other field.

It’s a big investment to learn that tower. And few people make it. Those that do have a variety of motivations. For some, it’s a matter of challenge. For others, more a matter of the “scenic view from the tower”. In some ways the automation provided by AI makes the challenge less appealing—though, as in something like chess, there can still be fulfillment in pure human achievement. But when it comes to the scenic view, the fulfillment is in the human aesthetic experience, which is quite unrelated to automation and AI.

Half a millennium ago, there were mathematical challenges set up as spectator sports. Today the challenge aspect of mathematics is not as successful at garnering public interest. Indeed, nothing about the process of doing mathematics is typically exposed to the public at all, except to a tiny extent through things like personality-based movies .

(And, yes, one might wonder if there’s any way the doing of mathematics could provide any form of public entertainment. I can report that I myself have done a fair number of successful live mathematical computer experiments . And during the pandemic I livestreamed a great many hours of highly technical working sessions about mathematical physics—which garnered a surprising number of views.)

So what about math as an essentially aesthetic activity? Yes, it’s fulfilling and enriching for the person in the middle of doing it. But what does society as a whole get out of it? Much as some tiny fraction of art is public art, so one might wonder if there can be public math. Occasionally there will be public visual or sculptural math, or perhaps architecture or music that embodies math. But realistically it’s not math at the level of the top of the tower of pure math.

One thing to understand about the upper reaches of pure mathematics is that to a surprising extent it’s been passed down as an oral tradition and a chain of human connections. Perhaps this is partly a consequence of the structure of academia and academic publishing. But the fact remains that to maintain the flame of pure mathematics seems to require a certain community of individuals continually and actively pursuing it, and talking about how it’s done. In other words, even if their personal motivations are quite internal, their collective effort has the long-term external value of allowing the progress of pure mathematics to continue.

One might ask if at some point pure mathematics will somehow be “finished”. And to this we can definitively answer that it will not. Computational irreducibility ensures that there will always be an infinite sequence of mathematical facts—and surprises—to be discovered. And there will also be an infinite number of pockets of computational reducibility, representing new mathematical concepts and structures.

Pure mathematics stands as the single largest intellectual edifice built by our civilization. And out of it have come some of the most impactful ideas in human intellectual history—like formalization, abstraction and computation. In modern times, there are some accelerators for pure mathematics. Practical computation has been one. AI is another. Experimental mathematics in the style of ruliology is yet another, still fairly undeveloped. And I’d like to think that the language for pure mathematics that we’re building will—particularly when combined with AI—be still another.

The work of pure mathematics is long and hard (yes, a consequence of computational irreducibility). But over the course of millennia its importance to the core intellectual development of our civilization has been very great. And there is no doubt that it has more to give. And that even after the issues of today are long forgotten pure mathematics will continue to shine as a beacon of human achievement. So, yes, there’s every reason to expect a bright future—now with some additional help from AI—for that most rarefied of human pursuits: research in pure mathematics.

Further Reading

“ Can AI Solve Science? ” (2024)

“ The Physicalization of Metamathematics and Its Implications for the Foundations of Mathematics ” (2022)

“ Who Can Understand the Proof? A Window on Formalized Mathematics ” (2025)

“ The Empirical Metamathematics of Euclid and Beyond ” (2020)

“ Computational Knowledge and the Future of Pure Mathematics ” (2014)

“ What Is ChatGPT Doing … and Why Does It Work? ” (2023)

“ Will AIs Take All Our Jobs and End Human History—or Not? Well, It’s Complicated… ” (2023)

Note

Some of the ideas in this piece I discussed over the course of more than three decades with my wife Elise Cawley, who sadly died shortly before the piece was written , and who would no doubt have had many valuable and incisive insights about it, now lost forever.

Emitting metadata early makes building/checking Rust up to twice as fast

Hacker News
github.com
2026-10-04 02:26:57
Comments...
Original Article

Start dependent crates before their dependencies finish type-checking.

Every crate waits for the crates it depends on to be fully checked, function bodies included, before it starts. It doesn't need those bodies to type-check itself. It compiles against the dependency's interface, the metadata in its .rmeta file.

Headstart makes rustc write an early metadata file as soon as the interface is checked, and makes cargo start dependents on it. Each crate's bodies are then checked while the crates downstream are already compiling.

  • cargo check : dependents run to completion on early metadata.
  • cargo build : dependents do all their analysis on early metadata, then wait for the dependency's full metadata before generating code. While they wait, they give their job slot back.

If a body has an error, the build still fails with that error, with the same diagnostics and exit status as today; only progress lines and the cross-crate order of JSON messages can differ. Cargo reports a crate's output only once all its dependencies have finished cleanly, and drops it if one fails. The costs are work downstream that gets thrown away, errors reported slightly later, and more memory in use at once (see docs/design.md ).

Pieces

  • rustc, -Zearly-metadata ( 6 patches ):
    • a new analysis_interfaces query splits analysis into item interfaces and function bodies;
    • the driver writes .early-rmeta between the two;
    • crate loading accepts early metadata, and swaps in full metadata before code generation, waiting for it if necessary on a lock its producer holds until it's written.
  • cargo, -Zheadstart ( 3 patches ):
    • passes -Zearly-metadata to every compile;
    • starts dependents on the early-metadata notification, in both check and build ;
    • gives a paused compilation's job slot to other work;
    • reports a crate's output only when its dependencies succeeded.

The patches are a commit series, each with a commit message and tests, meant to become upstream pull requests: see patches/README.md .

On rustc's default front end, headstart makes clean builds of 13 real projects (rust-analyzer, zed, bevy, lemmy, polars and others) up to 54% faster for cargo check , and up to 42% for cargo build . None is slower. With the parallel front end ( -Zthreads=8 ), which covers some of the same ground, it adds up to 25%. Those are 16-core numbers. The gain comes from cores the build would leave idle, so it shrinks on smaller machines. On 4 cores, rust-analyzer's check is 24% faster and its build 13–15%, codex-rs's check 14%, and wide builds come out even.

How it works, what early metadata leaves out, and the risks: docs/design.md . Measurements: docs/results.md . Whether it's ready to bring to the compiler and cargo teams: docs/readiness.md .

Try it

scripts/setup.sh    # check out rustc + cargo, apply the patches, build both

Then, in any Rust project:

RUSTC=/path/to/headstart/rustc/build/host/stage1/bin/rustc \
  /path/to/headstart/cargo/target/release/cargo check -Zheadstart   # or build

CARGO_UNSTABLE_HEADSTART=true turns it on too, as does [unstable] headstart = true in .cargo/config.toml . Without it, the patched cargo behaves like upstream, so the same binaries give a fair baseline.

tests/smoke is a two-crate workspace that shows the effect. Its slow library takes several seconds to check, almost all of it in function bodies. With headstart on, app starts about 0.2 s in instead of waiting for slow to finish.

scripts/check-errors.sh checks the claim about errors on tests/errors . It runs three scenarios (a clean build, an error in a dependency, an error in the binary) with cargo check and cargo build , headstart off and on. It then compares the human-readable output, the JSON output, the exit status and what the built binary prints.

scripts/check-incremental.sh [check|build] does the same across a sequence of incremental edits. The steps include adding an impl Fn a dependent calls, and breaking and then fixing an interface. It also compares the final state against a clean build.

scripts/check-swap.sh makes a library start on its dependency's early metadata and swap in the full metadata while paused, at every optimization level. The program built from it must print the same as one built from full metadata.

scripts/sweep.sh builds all 53 rustc-perf compile benchmarks with headstart off and on, with -Zearly-metadata-verify . It passes when every build succeeds in both modes with the same diagnostics, and verify reports nothing. -c build sweeps cargo build , -r the release profile, and -t the parallel front end ( -Zthreads=8 ).

Benchmarks

scripts/bench.sh -n 5 [-c build] path/to/project ...

This times clean cargo check (or cargo build ) builds, alternating headstart off and on, and prints the medians. A project can take cargo arguments after :: ( path/to/vaultwarden::--features=sqlite ). `scripts/real-projects.sh

` clones the 13 real projects from [docs/results.md](docs/results.md) at the commits measured, and prints them in that form:
scripts/bench.sh -n 3 -c build $(scripts/real-projects.sh ~/hs-real)

scripts/setup-codex.sh <dir> does the same for codex-rs, which needs a patched dependency and codex's prebuilt V8 (see the script); source <dir>/codex/headstart.env before timing it.

-w adds an untimed warm-up build per project, for build scripts that do one-time work outside target (helix compiles its grammars into its source tree).

scripts/bench-suite.sh <out-dir> runs the whole suite of docs/results.md on the current machine: rust-analyzer, the 21 rustc-perf benchmarks, the other real projects and codex-rs, check and build. It keeps each benchmark's results separately and skips finished ones, so it can be restarted, and writes a summary table at the end. It's how to get the numbers for a machine size not measured yet, such as 8 cores.

scripts/bench-mem.sh samples the total memory of all rustc processes during a build, in the project directory itself (run it on an otherwise idle machine). scripts/bench-incremental.sh times incremental rechecks after editing one function body. scripts/log-rustc records when each rustc run started and ended, so you can see the schedule. The rustc-perf benchmarks are under rustc/src/tools/rustc-perf/collector/compile-benchmarks ( git -C rustc submodule update --init --depth 1 src/tools/rustc-perf ).

UK Government Body Kept Files on People Criticizing Prevent Program

Hacker News
reclaimthenet.org
2026-10-04 02:10:57
Comments...
Original Article

The Home Office tried to keep the paperwork private until an appeal brought it into the light.

If you are a British citizen and you criticize a government program online, that criticism can be collected and stored by a government-linked body.

At least that seems to be the upshot of the actions of the Prevent Standards and Compliance Unit (StaCU), “a body linked to the Home Office," that has been keeping “observations” of those who speak against the government’s Prevent counter-extremism program on the internet.

The “evidence” of this comes from documents the group Rights & Security International (RSI) obtained under freedom of information law.

The group has now revealed that in the year from March 2024 to February 2025, StaCU made 77 “observations” of people who were critical of the program online – and that most of these “observations” came from X, with some from Reddit. Mainstream media articles were also included, as were posts from or about several organizations, including Hope Not Hate, Maslaha, and the Open Rights Group.

The “observations” are filed as “open-source complaints” and stored internally by StaCU. RSI is unsure what happens to this data next, and who the government may be sharing it with. What is known is that the process involves “attempting to independently verify claims.”

As RSI’s Freedom of Expression and Belief Team Leader Jacob Smith put it, “You should be able to criticize government policy without getting put on a list.”

Smith also noted that the fact that the government “has been trawling X and Reddit to find out who has been critiquing Prevent” should serve as “a stark reminder that, in the government’s eyes, anything you share online is fair game and could go into a file forever.”

The Home Office tried to keep this secret and refused to hand over the documents at first, but RSI won an appeal to the Information Commissioner’s Office, and the documents were released.

As RSI noted, it is unclear if any automated tools were used to find the posts, and if not, how the unit “found” them, or what it does with the data it collects.

The case is not the first time Prevent has been in the news for its controversial nature and the way it is implemented , but it is a stark reminder that the British government has no problem treating criticism of its policies as potentially dangerous.

Explore more on these topics

Fog-Bank: Archiving the oldest webcam feed

Hacker News
fog-bank.org
2026-10-04 01:06:58
Comments...
Original Article

FogCam publishes one picture at a time and throws it away when the next one lands. Fog-Bank keeps every one of them, and when something goes wrong (on my side or FogCam's), the record shows exactly when, for how long and whose fault it was.

Four catchers on three networks ask FogCam for its picture: a poller on my own server every 7 seconds, fogline on Cloudflare every 15 seconds (it keeps a 30-day buffer), a Cloudflare worker every 20 seconds that writes straight to permanent storage, and a machine at my house every 15 seconds. Each one pushes every frame it gets into the archive, and sweeps every 5 minutes and every night pull back anything a push dropped. Every request asks for the real file (a unique query string plus no-cache headers), because in August I found FogCam's web server cache handing out copies minutes old. Each frame is named by FogCam's own upload time, fingerprinted with sha-256, stamped without re-compressing, stored as lossless JPEG XL and copied to Cloudflare R2 the same minute.

A status board checks every stage from the camera to YouTube every few minutes and pages me on Telegram when one stops. The page lists every outage so far, in order, marked as mine or FogCam's: FogCam's server outage on Aug 13, its 64-hour freeze on Sep 8 to 10, my server's 40-hour outage inside that freeze (a full disk), and the rest.

More: the live archive · how the backup works · the forensic record · the story

‘We are in a kind of war’: row over viral Amsterdam chip shop lands in court

Guardian
www.theguardian.com
2026-10-04 00:00:17
Neighbours say crowds drawn by TikTok to Fabel Friet are clogging the picturesque canal street and attracting litter, gulls and rats People living in one of Amsterdam’s most exclusive neighbourhoods have taken the city to court for licensing a chip shop that has gone viral on TikTok. On busy days, t...
Original Article

People living in one of Amsterdam’s most exclusive neighbourhoods have taken the city to court for licensing a chip shop that has gone viral on TikTok .

Thomas van Leeuwen, one of Fabel Friet’s neighbours.
‘It has polluted the whole area’ … Thomas van Leeuwen, who lives near to Fabel Friet.

On busy days, the queue for Fabel Friet on Runstraat stretches across a bridge on the Keizersgracht – one of Amsterdam’s three main canals – a court was told last week. The narrow street itself is crowded by tourists waiting for their orders seven days a week. And loitering with intent – on parked cars, on the street itself, and on steps and porches – gulls wait to swoop.

“We are in a kind of war where a very unwelcome chip shop has polluted the whole area,” said Thomas van Leeuwen, who lives nearby. “Eating in the street is a shameful behaviour. It gets filthy food on our doorsteps and mayonnaise in the beautiful stone. The herring gulls have moved in as residents, and the worst thing is the rat infestation.”

Nineteen neighbours have formally protested against the business licence for Fabel Friet, whose signature product is fries topped with parmesan cheese and truffle mayonnaise. Niek van der Heijden, the lead plaintiff, said the problem was the takeaway’s small size.

“What we call the ‘TikTok queue’ for Fabel Friet causes the most nuisance,” he said. “The business takes place on the street: waiting to order, waiting to be served, eating. We want Fabel to go to a place where there is enough room to have clients inside.”

In a packed courtroom on Tuesday morning, a lawyer for Amsterdam city council acknowledged that the city’s mayor, Femke Halsema, can refuse a licence to operate if the “residential and living environment” is negatively affected by a business.

Niek van der Heijden, the lead plaintiff in the case.
Niek van der Heijden, the lead plaintiff in the case. Photograph: Judith Jockel/The Guardian

While the chip shop, she said, sticks to its licence obligations to manage the queue and clear litter “it has been going on for a long time, the impact remains severe, and you have to ask: how long can we tolerate this?” For the current licence, she said, the city would be guided by the court’s ruling.

The shop’s owners, Abel Klatser and Floris Feilzer, said they employed a cleaning service seven days a week, four times a day, and that year-on-year visitor numbers had dropped by 23% this August after they opened four new locations . “A few neighbours have experienced nuisance, but other people enjoy eating chips,” said Gertjan Teeuwen, their lawyer. “It can be busy, but there are peaks and troughs.”

The queue outside Fabel Friet, which occupies a small premises.
The queue outside Fabel Friet, which occupies a small premises.

Klatser and Feilzer said in a public statement that they felt unfairly attacked while doing their best to be good neighbours. “Fabel Friet is our child, and we have put our heart and soul into it,” they said. “It is painful if people talk about our business negatively.”

A herring gull eating leftover fries on top of a grey bin.
Herring gulls have been attracted to the area by the leftovers and waste. Photograph: Judith Jockel/The Guardian

The case goes beyond irritations about so-called performative tourism, where tourists flock to a place to post images on social media. “The tiny streets might be charming, but actually the space is not designed for so many people,” said Lisa-Bobbie Kamp, a lawyer for the neighbours. “This is a question of taking space, who owns the place and whether you can just use public space.”

Marije Peute, a PhD student at the University of Amsterdam, has studied the effects of “viral urbanism” in the neighbourhood. “The municipality should really think about how they want to act, because this is about the public space. A lot of much older retail spaces have lost their local clients because they don’t want to go there any more,” she said.

Rogier Havelaar, a councillor from the Christian Democratic Appeal party, said that, while Fabel Friet should probably be offered a “flagship store” licence in a more suitable location, the benefits of tourism to the city should not be overlooked. “Amsterdam would never have as many museums, restaurants, going-out spots and theatres if they were only for the residents,” he said. “And tourism provides 70,000 jobs. Amsterdam would have a gigantic financial problem if tourism went away.”

A person holding a cardboard takeaway box full of chips and covered in cheese, with a bridge in the background
A customer with a portion of Fabel Friet’s signature fries topped with parmesan cheese and truffle mayonnaise. Photograph: Judith Jockel/The Guardian

Standing on the bridge over the Keizersgracht, some German laboratory apprentices drawn by TikTok were enjoying their chips. Natalie Stübben said crowd managers had just asked them to move away from people’s houses. “It’s pretty hard to get through here with a car or just by walking,” she said. “I don’t think we were thinking about annoying the neighbours, but it makes sense.”

Natalie Stübben standing on a bridge with a canal in the background.
‘I don’t think we were thinking about annoying the neighbours’ … Natalie Stübben.

Some of those neighbours will be pleased if the business location has had its chips after the court verdict on 12 November. “There’s something very alarming about the fact that these people are all led by their higher power – influencers – and call themselves followers,” said van Leeuwen. “It’s Orwellian. This court case is part of the parasitic industry that lives on tourists.”

Inside Anthropic's Quest to Instill Morality into Its A.I. Models

Hacker News
www.nytimes.com
2026-10-03 22:34:22
Comments...
Original Article

Please enable JS and disable any ad blocker

Rust's derive often implies inline

Lobsters
yossarian.net
2026-10-03 21:41:57
Comments...
Original Article

In Rust, one of the most common ways to implement core traits (like Debug , Display , and Clone ) is to #[derive(...)] them, e.g.:

#[derive(Debug)]
struct Widgets {
  foo: u32,
  bar: usize,
}

What I didn't know until recently is that Rust currently emits #[inline] as part of these derivations. This is seemingly not guaranteed, but is implied by example in the reference and can also be seen if one expands the macros.

Using the example above, this is what you get when you expand the #[derive(Debug)] in the playground :

struct Widgets {
    foo: u32,
    bar: usize,
}
#[automatically_derived]
impl ::core::fmt::Debug for Widgets {
    #[inline]
    fn fmt(&self, f: &mut ::core::fmt::Formatter) -> ::core::fmt::Result {
        ::core::fmt::Formatter::debug_struct_field2_finish(f, "Widgets",
            "foo", &self.foo, "bar", &&self.bar)
    }
}

This is almost always what we want: #[inline] is just a hint, and typical derived Debug , Clone , etc. implementations benefit from being inlined (since they're often trivial).

But not always! Imagine an error hierarchy like this 1 :

#[derive(Debug)]
struct ErrorA {
  lots: String,
  of: String,
  chunky: String,
  fields: String,
  within: String,
  this: String,
  r#type: String,
}

#[derive(Debug)]
struct ErrorB {
  inner: ErrorA,
}

#[derive(Debug)]
struct ErrorC {
  inner: ErrorB,
}

#[derive(Debug)]
enum Errors {
    A(ErrorA),
    B(ErrorB),
    C(ErrorC),
}

produces:

struct ErrorA {
    lots: String,
    of: String,
    chunky: String,
    fields: String,
    within: String,
    this: String,
    r#type: String,
}
#[automatically_derived]
impl ::core::fmt::Debug for ErrorA {
    #[inline]
    fn fmt(&self, f: &mut ::core::fmt::Formatter) -> ::core::fmt::Result {
        let names: &'static _ =
            &["lots", "of", "chunky", "fields", "within", "this", "type"];
        let values: &[&dyn ::core::fmt::Debug] =
            &[&self.lots, &self.of, &self.chunky, &self.fields, &self.within,
                        &self.this, &&self.r#type];
        ::core::fmt::Formatter::debug_struct_fields_finish(f, "ErrorA", names,
            values)
    }
}
#[automatically_derived]
impl ::core::default::Default for ErrorA {
    #[inline]
    fn default() -> Self {
        Self {
            lots: ::core::default::Default::default(),
            of: ::core::default::Default::default(),
            chunky: ::core::default::Default::default(),
            fields: ::core::default::Default::default(),
            within: ::core::default::Default::default(),
            this: ::core::default::Default::default(),
            r#type: ::core::default::Default::default(),
        }
    }
}

struct ErrorB {
    inner: ErrorA,
}
#[automatically_derived]
impl ::core::fmt::Debug for ErrorB {
    #[inline]
    fn fmt(&self, f: &mut ::core::fmt::Formatter) -> ::core::fmt::Result {
        ::core::fmt::Formatter::debug_struct_field1_finish(f, "ErrorB",
            "inner", &&self.inner)
    }
}

struct ErrorC {
    inner: ErrorB,
}
#[automatically_derived]
impl ::core::fmt::Debug for ErrorC {
    #[inline]
    fn fmt(&self, f: &mut ::core::fmt::Formatter) -> ::core::fmt::Result {
        ::core::fmt::Formatter::debug_struct_field1_finish(f, "ErrorC",
            "inner", &&self.inner)
    }
}

enum Errors { A(ErrorA), B(ErrorB), C(ErrorC), }
#[automatically_derived]
impl ::core::fmt::Debug for Errors {
    #[inline]
    fn fmt(&self, f: &mut ::core::fmt::Formatter) -> ::core::fmt::Result {
        match self {
            Self::A(__self_0) =>
                ::core::fmt::Formatter::debug_tuple_field1_finish(f, "A",
                    &__self_0),
            Self::B(__self_0) =>
                ::core::fmt::Formatter::debug_tuple_field1_finish(f, "B",
                    &__self_0),
            Self::C(__self_0) =>
                ::core::fmt::Formatter::debug_tuple_field1_finish(f, "C",
                    &__self_0),
        }
    }
}

That's a lot of code that can get inlined for each invocation of the Debug implementation of Errors , which can occur repeatedly in e.g. debug or trace logging.

In fact, it's so much code that it can turn out to be a non-trivial amount of a Rust binary's total size: we found that we could shrink uv's binary size by approximately 160KB by preventing 2 rustc from inlining a given Debug implementation.

This was surprising to me on two levels: the size cost added up fast, and rustc (seemingly) did not apply a limiy to the size or number of times a Debug implementation was inlined. I suspect this is the right decision in many programs, however!

  1. This hierarchy is drastically simplified: real world Rust applications often have deeply nested error enumerations with nontrivial numbers of fields. ↩

  2. We did this by adding our own-proc macro that behaves like derive(Debug) , but with #[inline(never)] . ↩

We're working on a new RuneScape MMO

Hacker News
play.runescape.com
2026-10-03 21:12:22
Comments...
Original Article

Well, we said there was one more thing..and we meant it.

At RuneFest, we gave you the first glimpse of something we’ve been quietly working on: an entirely new RuneScape MMORPG, built in Unreal Engine and set in the world of Gielinor.


Right now, it’s rocking a working title: RS4. That isn't necessarily the name we’ll eventually settle on, but this is our fourth generation RuneScape project, so RS4 works pretty well for now.

We thought long and hard about revealing this project so early. We could have kept it behind closed doors for much longer, especially given that we are still right at the beginning of a hugely ambitious project and the game is still a few years away.

But ultimately, that didn't feel very RuneScape, or very us.

For 25 years, RuneScape has been shaped by its players. We've learned that our games are at their best when we build with you, not just for you. So we want to do that from the beginning, and that means bringing you into the development journey earlier than you might expect - sharing ideas, testing things, listening and learning as we go.

We want to build on the best of the tech development we've learned on Dragonwilds, and the best of the community co-development we've learned through RuneScape and Old School, and apply those principles to something new. And of course, as a RuneScape game, integrity and permanent progression will be central to our design principles.

A New RuneScape to sit alongside our games

Let's make one thing really clear. RS4 is not a replacement for any of our games, be it Old School, RuneScape or Dragonwilds. Those games have their own dedicated teams, exciting roadmaps and continued investment. We’ll be designing RS4 to sit alongside them as an additional way to experience RuneScape and the world of Gielinor.

In fact, we think having different RuneScape experiences can make the whole family stronger. Players may spend most of their time in one game, move between several, or discover Gielinor for the very first time through RS4. The ambition isn't to replace the RuneScape games you love. It's to make the RuneScape universe bigger.

Where does the journey begin?

RS4 will be set in Gielinor, in the Sixth Age. But we're beginning somewhere that RuneScape players have only recently started to discover - and that is Ashenfall. An area first explored in RuneScape: Dragonwilds, and a landmass to the north-east of the Gielinor.

RS4 will take place years after the events of Dragonwilds and the destruction of the Dragon Queen Kuldra. Time has passed, the world has moved on, and Ashenfall has changed with it. But Ashenfall is only the beginning.

Our ambition is for this adventure to grow far beyond Ashenfall over time, moving into places you know and love, as well as opening up new places and stories across Gielinor.

This is just the beginning

We know our incredible community and that means we’re sure you have a million questions. Which platforms? What accounts? Membership? Progression? Combat? Skills? When can you actually play it?

But, the answer right now is that it's simply not the right time, or just too early, for us to say. This isn't the full reveal – more an opportunity for us to share the beginning of something special with you.

We'll tell you much more next year about the world, the game, the team building it and, importantly, how you can start getting involved. For now, there are a few things we want you to know:

  • Our 4th MMO is now in development.

  • It's a new RuneScape MMORPG, built in Unreal Engine

  • It sits alongside RuneScape, Old School RuneScape and Dragonwilds as an independent experience

  • It will be true to our principles of integrity, permanent progression, and community driven development

  • It's set in Gielinor and begins in Ashenfall.

  • It's still a few years away, but we're telling you now because we want to build it with you.


25 years ago, RuneScape began with a world, a small team and a community that helped turn it into something far bigger than anyone could have imagined. Now we get to start another adventure together.

Bob Cringely Has Died

Hacker News
news.ycombinator.com
2026-10-03 20:50:52
Comments...
Original Article
Hacker News new | past | comments | ask | show | jobs | submit login
Bob Cringely Has Died
68 points by paveworld 1 hour ago | hide | past | favorite | 7 comments

I heard from a friend of the family that Bob passed away in his sleep early Saturday. Very sad news. Bob, who's real name was Mark Stevens, was an early employee of Apple and was best known for his PBS documentaries, especially "Triumph of the Nerds". He will be missed.

help


When searching to verify the veracity of this, this post made me laugh: https://www.cringely.com/2020/01/23/not-dead-yet-what-bob-cr... (Circa 2020).


Somebody who has chronicled and commented on the business for a long time. I enjoyed his work, and chuckled at some of the for-and-against it caused. Sorry to hear and I hope his family is ok. There’s a whole lotta water that’s passed under the bridge that we’ve not really noticed.


Reading accidental empires is how i learned that the tech industry even existed


Same, that was my introduction to ‘tech culture’.


Future generations will have a better understanding of the personalities and history of the personal computer industry because of his work.


I liked his frequently contrarian views. Pre-internet his Cringely column (which he took over from someone else) was a great source of information for those of us living far away from Silicon Valley. RIP.


RIP…


Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Search:

Apple's "Clean Design" Is Stupidity When It Comes to Hiding Fire Extinguishers

Hacker News
www.gadgetreview.com
2026-10-03 20:09:21
Comments...
Original Article

Unmarked flush panels and hidden handles may violate OSHA and NFPA rules, putting occupants at risk in seconds

Smoke is filling a room. Someone is shouting. You know a workplace safety tool like a fire extinguisher is nearby, but the designer made it disappear behind a flush panel that matches the wall perfectly.

That is not a design triumph.

The Price of Aesthetic Perfection in a Crisis

A clean environment is worth pursuing, but erasing safety infrastructure from the visual field is a different thing entirely.

A clean wall that conceals emergency equipment without any visible identification is not good design. It is a functional failure dressed up in a coordinated finish.

Reportedly, an Apple retail or office installation has been circulating as an example of this trade-off. It allegedly shows fire suppression equipment integrated so seamlessly into the surrounding architecture that it effectively disappears from view.

No official Apple statement, architectural document, or verified primary image has surfaced to confirm the specifics. No such source was identified in the research reviewed for this article.

What the circulating example does clarify, regardless of its origin, is a pattern worth examining. Brand loyalty can short-circuit critical thinking fast enough to make observers praise a design choice before asking whether it actually works under pressure.

The regulatory framework is not ambiguous:

  • OSHA 1910.157 requires workplace extinguishers to be mounted, located, and identified so employees can reach them without being subjected to possible injury.
  • NFPA guidance requires extinguishers along normal paths of travel, generally visible or clearly marked with signage when an obstruction exists.
  • A recessed cabinet is not automatically noncompliant; concealment without effective identification is the core problem.
  • Maximum travel distance to an extinguisher depends on hazard classification, and accessibility is treated as a functional requirement, not a preference.

The rules leave room for elegant solutions. They leave no room for unmarked panels and hidden handles.

Hiding a fire extinguisher is a hilariously bad idea. And someone thinking it’s a great design perfectly encapsulates the mindset of most Apple cultists when it comes to anything Apple does https://t.co/9GmLH06wQA

— James Welham (@WelhamOfficial) October 1, 2026

Good-Looking and Compliant Are Not Mutually Exclusive

A clearly labeled cabinet with a coordinated finish is not a design compromise; it is the correct answer.

NFPA guidance describes portable extinguishers as a first line of defense against incipient fires, with accessibility treated as a non-negotiable functional requirement. That framing reorients the entire debate: an extinguisher is infrastructure you design for, not furniture you style around.

The confusion underlying this kind of praise is treating visual cleanliness as a proxy for good design. Consider the difference between a minimal app interface that buries the emergency logout button six menus deep and one that simply puts it where you can find it. Both might look polished. Only one works when it counts.

Facilities teams and designers can still use recessed installations, coordinated finishes, and orderly placement. The non-negotiable constraints are prominent labeling, unobstructed access, and location along normal travel paths. Those are not aesthetic concessions.

During an emergency, a recessed cabinet without signage is indistinguishable from drywall.

The standard applied to an allegedly hidden extinguisher in any sleek, brand-conscious space should be the same standard applied to every warehouse, office park, and retail chain: accessible, identified, and reachable without delay. Emergency equipment is not a brand decision. Design around it accordingly.

We're going to need default hard budget caps on pretty much everything

Simon Willison
simonwillison.net
2026-10-03 19:34:02
Here's a product feature which the world is going to need a whole lot more of over the coming months and years: default hard budget caps. I'm talking about the feature of pay-by-usage services and APIs that lets you say "after $X/month, cut this thing off and return errors". These need to be hard li...
Original Article

3rd October 2026

Here’s a product feature which the world is going to need a whole lot more of over the coming months and years: default hard budget caps . I’m talking about the feature of pay-by-usage services and APIs that lets you say “after $X/month, cut this thing off and return errors”. These need to be hard limits. Soft caps, “after $X/month, send me a warning email”, will not cut it.

Coding agents, and personal agents (coding agents wrapped in a less threatening UI), greatly reduce the friction of spinning up code that can do useful things. Sometimes those things cost money—calls to paid APIs, or hosted web applications, or systems that can bill for additional storage and compute.

Nobody wants to wake up to an email sent at midnight warning about a budget limit and find that, while they slept, their rogue service had consumed several hundred (or several thousand) more dollars of usage.

An argument against this is that businesses don’t want their hosted applications to start throwing errors because some budget was exceeded. I expect that most businesses and individuals would prefer errors to a surprise $10,000+ bill.

I think hard budget caps need to be the default. If someone wants to live dangerously they should be able to do that, but it needs to be on an opt-in basis. Have a nice, clear checkbox somewhere prominent:

Remove the budget cap. My application will not be shut down if I exceed the configured budget limit, and I will be responsible for subsequent charges.

The service I most want to see this from is AWS. I’ve heard plenty of stories from people who refuse to use AWS for personal projects out of (justified) fear that a runaway service might bankrupt them. I’ve also heard stories from people who didn’t anticipate this and ended up seriously burned.

... and it turns out AWS finally launched spending limits a few weeks ago! From their announcement New AWS experience helps builders get started and ship faster on 16th September:

When you’re ready to upgrade to a paid plan, you can set a monthly spend limit for your project based on your usage patterns so that you stay within your budget. If a project’s usage reaches its spend limit, your project is paused for that month.

See also Create a spend limit in AWS Settings , though that page warns that “We’re currently releasing our new experience to a limited number of customers.” Here’s hoping that hits general availability for existing accounts soon.

Google Cloud launched a similar feature in July, called Spend Caps, which lets you “set a monthly financial cap on specific services within a project”. Looks like this is becoming a trend!

In an ideal world, our agents could help with this. It would be great if agents started biasing towards recommending providers with hard budget caps, and warning new and inexperienced builders against deploying applications using uncapped services that might get them into trouble.

Make America Read Again

Portside
portside.org
2026-10-03 19:26:32
Make America Read Again Dave Sat, 10/03/2026 - 19:26 ...
Original Article

The forty-fourth annual Banned Books Week (BBW) takes place October 4 through 10 this year, but given the significant, sustained efforts to restrict what young people can read, Banned Books Week should likely be observed every week. We are living in a time of increased attacks on the free press , free speech , academic freedom , and the rights to read and know. This year’s BBW theme, “Let Books Be: Protect the Freedom to Read,” reflects these concerns.

A recently published five-year report compiled by the American Library Association (ALA), Book Riot, Florida Freedom to Read Project, and the Freedom to Read Project chronicled an explosion of book bans beginning in 2021, during the COVID-19 pandemic. The surge was the result of what they called “a full-blown, organized political movement” led by groups such as Moms for Liberty and No Left Turn in Education. Censors have had the most success in Florida and Texas, where state legislatures have passed laws requiring schools to purge their libraries of books that attract complaints from parents. But they have increasingly targeted local school board positions across the country to embed members who advance their censorious policies.

While book challenges were once focused primarily on individual books or authors and were brought by concerned parents or family members acting on their own behalf, the past half-decade shows a shift to sweeping challenges from organized interest groups. These efforts have been successful in red states and beyond, which should alarm anyone who supports the right to read and learn.

According to Betsy Gomez , coordinator of the Banned Books Week Coalition, ALA’s Office for Intellectual Freedom “tracked 4,235 unique titles challenged in 2025, the second highest ever documented by ALA,” and “713 attempts to censor library materials and services, 487 of which targeted books.” In 2025, the ALA found that “92 percent of all book challenges were initiated by pressure groups, government officials and decision makers, up from 72 percent in 2024” and “less than 3 percent of challenges originated from individual parents.”

Tactics and Targets

Books being yanked off the shelves at schools around the country include a number of acknowledged literary classics such as Toni Morrison’s The Bluest Eye , Margaret Atwood’s The Handmaid’s Tale , and Anthony Burgess’s A Clockwork Orange , usually singled out for their violent or sexual content . But the most challenged books in recent years invariably tell stories about people of color or LGBTQ+ characters.

According to PEN America, public school libraries banned 3,743 unique titles during the 2024–2025 school year. Roughly 44 percent of those titles featured characters identified as people of color. Another 36 percent featured characters identified as LGBTQ+, and some 19 percent of the banned titles dealt with characters depicted as trans or genderqueer. A typical book targeted for featuring LGBTQ+ characters and themes is Malinda Lo’s Last Night at the Telegraph Club , a National Book Award-winning young adult novel about two women who fall in love in San Francisco’s Chinatown against the backdrop of the Cold War. According to PEN America, it was banned nineteen times in 2024–25.

Some books have been banned merely for acknowledging that the United States has a past marred by slavery, oppression, and exploitation, or for depicting struggles for social justice. For instance, Veronica Arreola’s J is for Justice: An Activism Alphabet , an illustrated picture book introducing youngsters to social justice activism, was one of nearly 596 titles removed from Department of Defense school libraries earlier this year. It was also removed from a Virginia school library.

Book banning at the K-12 level is bad enough, but increasingly brazen censors are also trying to restrict the sort of books that can be read in college as well. Early this year, Texas A&M University barred a philosophy professor from assigning an excerpt from Plato’s Symposium , his famous dialogue about the nature of love, because it allegedly contained “gender ideology.” The University of Nebraska at Kearney announced in June that it would discontinue the textbook Discovering Human Sexuality in courses after fielding multiple complaints from conservative critics, including Nebraska Governor Jim Pillen. The critics particularly bemoaned a training module in the book about empowering transgender students.

Although most of the book challenges around the country in the past few years have been perpetrated by right-wing groups and politicians, not all the titles being targeted are “progressive” or “woke.” For instance, two books by former Fox News commentator Bill O’Reilly, about Jesus and Ronald Reagan, were removed from shelves in one Florida school district.

These extraordinary efforts to control student access to works about specific academic subjects or marginalized identities are problematic on their own, but a separate challenge exacerbates all of this: declining literacy and interest in reading.

A Decline in Literacy

Ray Bradbury once said, “You don’t have to burn books to destroy a culture. Just get people to stop reading them.” Over the past two decades, the United States has seen a decline in multiple literacies. As Rose Horowitch noted in The Atlantic , “According to the National Endowment for the Arts, which conducts the most comprehensive survey of the nation’s reading habits, fewer than half of all adults reported having read a book of any kind in 2022.” The decline was measured across all demographics, including among those most likely to read.

These days, Americans are reading shorter works and more pop culture pap than serious content. When it comes to news consumption, fewer than 10 percent of young people now read a daily newspaper. More than half of adults get news from social media in what media scholars have called news snacking , the misguided belief that important news will just find them while they scroll through their corporate, Big-Tech, algorithmically controlled social media feeds . As Horowitch put it, the United States is in the midst of a “literacy crisis.”

UCLA neuroscientist Maryanne Wolf went further, telling The Atlantic that America isn’t illiterate; it’s “postliterate.” The latest research corroborates this assessment. Scholars at Harvard, Stanford, and Dartmouth recently conducted a broad study of more than 5,000 schools in 38 states and concluded that “only five states plus the District of Columbia had meaningful growth in reading test scores from 2022 to 2025. Nationally, students remain nearly half a grade level behind pre-pandemic reading scores and only slightly better in math.”

We have our work cut out for us.

Make America Read Again

One way to fight back against censors and declining literacy is to center the importance of the right to read for civic engagement and the health of our republic. Each year, the organizers of Banned Books Week select several Honorary Chairs to help raise awareness—authors, activists, and students who champion inclusivity and intellectual curiosity, leading by example through their work and teaching. These ambassadors for literacy and intellectual freedom remind us that democracy is stronger when we hear more voices from diverse perspectives and when everyone has access to the insights and ideas that books especially make available. One of this year’s Youth Honorary Chairs is high school senior Marshall Romero, who is also involved in Students Engaged in Advancing Texas ( SEAT ). In an interview on the Project Censored Show , Romero observed that

when it comes to a lot of conversations surrounding banned books or media … things happening directly [in] the classroom, there’s not a lot of youth voice or youth perspective on it. Rather, it’s a lot of partisans or parents concerned about what their kids want to learn or what the classroom should look like. … But students are not being considered as stakeholders. … It’s people that haven’t been in the classroom for 30 years. … It’s just been a big hurdle being taken seriously by legislators, but also being able to say, Hey, we’re the primary stakeholder, and you should listen to us.

As a member of SEAT, Romero believes students deserve a seat at legislative and school board tables when decisions that directly impact young people are being made.

Maria McCauley, president of the American Library Association, connected attacks on libraries to broader assaults on democratic institutions. In an appearance on the Project Censored Show , McCauley said that libraries are “foundational to the health of our democracy” and “offer agency and freedom for all of our community members to be able to thrive and flourish.” This is why it is imperative not only to protect libraries and the librarians who staff them, but also to celebrate them, a primary mission of Banned Books Week.

There are many ways for people to get involved during Banned Books Week, including hosting a BBW trivia night (see here for a program kit ) or organizing a Right to Read Night , in cooperation with the National Coalition Against Censorship, where people can learn about the most challenged book titles and authors. Project Censored has been a proud member of the BBW Coalition since 2013, and we do our part each year to promote the shared mission of protecting the right to read while fighting censorship in its many guises.

In addition to these efforts, we need people like you, dear reader, to join in the cause. Support your local library and librarian. Visit your local independent bookstore or subscribe to your local news outlet (if you’re lucky enough to have either). Attend school or library board meetings and make it known that censorship has no place in our classrooms or communities. Download the Unite Against Book Bans action toolkit , which features tips about lobbying decision makers, participating in meetings, and organizing protests in support of the right to read. Project Censored also has an action guide for BBW where you can get involved in as little as five to ten minutes. We can stem the tide of book bans and other assaults on our democracy, but we need to work together, and we must do it now.

Second Chances

Hacker News
www.nybooks.com
2026-10-03 19:21:40
Comments...
Original Article

My friend the bon vivant bookseller Michael Seidenberg of Brazenhead Books, which he ran out of his apartment in New York City, died in 2019. I’d known him since 1978, when as a teenager I apprenticed myself to his first shop. Michael loved pressing “lost” books and authors on me, and on anyone. A visit to one of his shops often concluded with taking home a novel one had not only not been seeking, but of which one had never heard. He would typically reduce the price of such a purchase, as a softening of the risk entailed in taking his word.

Michael was also an instinctive devil’s advocate. When I once spoke of my enthusiasm for trade paperback imprints dedicated specifically to reviving lost books, he surprised me. “Those are tombstones,” he told me (my quotation is a paraphrase from memory). “Being published by one of those presses that only publishes revivals, it’s the last heartbeat of a dying reputation. Presenting an author under the flag of rediscovery is the same thing as saying, ‘No one cared in the first place.’ After that, a book is officially dead forever.” He also pointed out how these imprints tended themselves to die out after just a handful of books.

With hindsight we can see what Michael missed in this particular contrarian flex. Not long before that conversation, Virago Books in the UK had launched the Virago Modern Classics. These displayed two strengths: the distinctive green spines that made one want to see whole shelves of them together, and a confidence, buoyed by their feminist imperatives, to keep rushing them out faster than one could read them, like a sea turtle laying a thousand eggs, knowing some will thrive but not knowing which ones. This created a map for success in future such enterprises, not that it could be guaranteed. 1 Vintage Contemporaries followed in 1984, adopting the Virago strategies—attractive uniform editions, and a cornucopian blend of paperback originals and reissues that largely focused on living writers.

The next chapter in this story is best told by quoting the critic Elaine Blair. In a 2021 essay in these pages called “ Averted Intimacies ,” she speaks of the revival, spearheaded by the magazine A Public Space , of the nearly forgotten Bette Howland. This came, according to Blair,

amid a surge of interest in reissued titles. Cycles of obscurity and rediscovery may be a regular part of book publishing, but in recent years the number of reissues and—perhaps most significantly—the amount of review coverage they receive have been striking. In addition to the presses that specialize in reissues ( NYRB Classics, the Feminist Press), other publishers have also revived out-of-print work (as Farrar, Straus and Giroux did with Lucia Berlin’s stories and New Directions with Fran Ross’s satirical novel Oreo ).

I agree. We live in a continuing great era of rediscovered books. Starting in 1999, NYRB Classics doubled down on Virago’s and Vintage’s methods with its finer bindings and acid-free pages, creating a handfeel that made their series even more pleasing to read and collect, and by the quality of the curation and editing of the forewords, prefaces, introductions, translator’s notes, and afterwords surrounding each text. More recently, McNally Editions has met that standard, while other, somewhat less rapid-fire presses like Transit, Counterpoint, Pushkin, and Two Dollar Radio are also making terrific contributions.

Some presses will seem driven just to restore a given writer or two, like Steerforth did with Dawn Powell (triumphantly) and Isabel Bolton (less so). No doubt I’m missing some names. The libidinal energy of such rediscoveries parallels other “analog” revival booms, like that for vinyl records. Despite literature’s highbrow aura, the appetite for lost books echoes not only the world of used bookstores, but the present yearning for a less predigested, more incautious mode of cultural life that still lurks in living memory, or in our parents’ attics.

Not all the turtles’ eggs survive, either. If you believe in canons, 2 then NYRB Classics may have boosted a few twentieth-century US contestants: 3 John Williams’s Stoner (originally published 1965; reissued 2006) and Butcher’s Crossing (1960; 2007), Renata Adler’s Speedboat (1976; 2013), Don Carpenter’s Hard Rain Falling (1966; 2009), surely others. 4 But this isn’t the Modern Library or Library of America, let alone Norton Critical Editions; the vibe is less proclamatory, more personal and intimate, the realm of a reader’s private pleasure. Perhaps you latched on to Lisa Tuttle’s My Death (2004; 2023), or Sasha Sokolov’s A School for Fools (1976; 2015). Me, I’m obsessed with Raymond Kennedy’s Ride a Cockhorse (1991; 2012) but can’t convince others that it’s a perfect allegory of the codependency of Trump’s inner circle, or really even get anyone to read it. And that’s OK. In this sense, Michael’s prophecy may be true: the present vogue for republication makes a nice last moment for some books destined to join the slew of forgotten others always receding behind us.

Michael, incidentally, became a great collector and purveyor of the Vintage Contemporaries and the NYRB Classics, which he loved both to read and to see lining up in delicious uniformity on his shelves. He was always lovably willing to contradict himself.

Wilfrid Sheed was born in London in 1930. The son of the Catholic publishers Frank Sheed and Maisie Ward, he grew up shuttling between England and the US. Though educated at Oxford before his aspirational leap to New York City and the beginning of his literary career, by his own account as a boy he’d fallen in love with baseball, not cricket or rugby, as well as with certain aspects of American popular culture like jazz and Broadway musicals. He survived childhood polio, the account of which finds its way into several of his fictions and his memoir of illness and recovery, In Love with Daylight (1995). His prodigious and versatile literary career encompassed nine works of fiction, three collections of critical essays, the memoir, and books on subjects as diverse as his fondnesses: baseball, jazz, Muhammad Ali, Clare Boothe Luce, and his parents. As a busy and prominent critic, he appears to have managed to walk a knife’s edge between fraternizing widely in Manhattan and on Key West, and being seen as a hit man in the Anatole Broyard, rip-the-jugular vein. Being a killer reviewer could be a bad fate socially, and also if one hoped to be critically embraced as a novelist. Which Sheed was. Though never drawing a wide readership, he was nominated for the National Book Award three times, and reviewed—says I, who’s been weighing the tangible results against the critical record—( spoiler alert ) generously.

The key to that knife’s-edge walk and to the admiration and fellowship Sheed enjoyed among his contemporaries is likely the same: his tone, always urbane, ironic, lucid, and at times truly funny. This pertains especially when he’s strongly on-target (concerning, typically, great white males like Mailer, Waugh, Hemingway, Vidal); less so when he’s badly off (ahem, James Baldwin, ahem, feminism). Though Sheed regularly self-inoculates by dubbing himself a minor novelist, he enters readily into imaginative sympathy with these monsters of ambition, leaving intact their implicit importance at the center of our literary culture, no matter that the book or books at hand—sometimes several decades’ worth—have been declared inadequate.

On Hemingway, in these pages : “Fame settled in to stay and clamped his style into place, where it grew warped and gnarled like a tree in a cave.” Mailer:

It is no accident that Mailer wrote the first big war novel, subsequently managed to hook himself onto both the Beat Generation and the French Existentialists, jumped off in time to catch the Negro express, elbowed briefly to the front of the Peace Movement, worked Vietnam into a novel title (an almost infallible sign of charlatanry)…. And who is the first writer to tell us about the moon?

Sheed’s sentences are fizzy with asides, nudges, and flattery to the informed reader. They remind me of the Raymond Chandler one encounters in Chandler’s letters—another writer who cross-pollinated the archness and delicacy of an English education with a thrilled appetite for the funk of US vernacular. And both Chandler and Sheed seem to use that wit to manage a scathing disenchantment, though Sheed disguises his better than does Chandler.

Sheed died in 2011, and in 2024 McNally Editions republished his best novel, Office Politics (1966), with a foreword by Gerald Howard, to another round of mostly warm reviews; Electric Literature , following the author’s lead in self-depreciation, called it a “minor masterpiece.”

The novel depicts the eerily banal daily life of an influential small weekly magazine called The Outsider , a publication that could be mistaken for any number of real models like Dissent , The Nation , or even a teensy-tiny New Yorker . (Sheed, who was a critic and editor at Commonweal for several years, claims in an author’s note that The Outsider “resembles no magazine living or dead.”) The editor presiding over this “broken-down opinion machine” is an English expatriate named Gilbert Twining, who soaks up all available oxygen for miles around. His invitations to lunch are a form of alcoholic passive aggression. Nevertheless, the standard Twining establishes and the order he keeps are irreplaceable. When he goes on leave after a heart attack it sets in motion a low-key succession crisis among the staff: imagine Game of Thrones populated by versions of Bartleby the Scrivener. The small-stakes office comedy is rendered with delirious precision as the characters assassinate one another’s pretenses with the blunt instruments of the arched eyebrow and the declined copyedit. They might as well be attacking each other with empty staplers:

“Brian, where did you find this—this thing?”
“What thing?”
“This gold thing.”
“You mean Culpeper’s article on the gold standard? You know who Culpeper is, don’t you?”
“He’s no writer that I know.”
“Look, we’re lucky to have him, he’s one of Canada’s leading economists.”
“A commentary, a sad commentary.”
“I’ve been trying to get something out of him for years.”
“You’ve been wasting your time, is all I can say.”

By default, at least at first, we root for George Wren, a younger, half-endearing American editor with an actual home life and a bunch of secret poems stashed in his desk. Yet Sheed deploys Wren’s comparatively unspoiled viewpoint mainly to illustrate how little the younger man grasps about the balance of purely symbolic power and painfully nonsymbolic disappointment that Twining has devised to keep the magazine functioning even in his absence. The book seems ultimately designed to function as an expectations-humbling machine. No one here rates anything but the mildest sympathy, which is precisely what their creator has granted them, after cutting off their legs.

Still, the pleasures of Sheed’s wriggly sentences carry the day, earning the comparisons to Kingsley Amis and to Joshua Ferris’s Then We Came to the End (2007). Knowing I could score some first-edition Sheeds cheaply—I’d passed them by in used bookstores for years—I launched a further investigation into his fiction.

In fact, I went into a long Sheed-trance. This was an unforced error. My old zeal for deep dives—I once spent a year reading every book by Stanisław Lem—resulted in what an analyst calls transference. Because Sheed’s rueful demolitions recalled scoundrelish writers I’d grown up reading—e.g., J.P. Donleavy, Terry Southern, or early Joseph Heller—I was ready for him to reveal some iconoclastic force, to uncork his secret rage against the machine. Or else he might turn on himself and convert his deep injuries—the dislocations and illness of his childhood—into pain the reader took into his own body, as one does reading William Maxwell when he rehearses the early death of his mother.

Neither happened. Sheed does write about himself, in a kind of deflected or mock-puzzled way. His novels portray writers and editors, but critics especially: Max Jamison (1970) is about a poisonous theater critic; Transatlantic Blues (1978) depicts the hectic, vainglorious life of a celebrated cultural critic (half Mailer, half David Frost); and The Boys of Winter (1987) is narrated by a publisher turned novelist who is wrangling the work of several other novelists in an enclave in the Hamptons. These form a collective self-portrait of a middle-aged man so entombed in his younger self’s eagerness to rise within the status quo that all he can possibly do is generate endless cutting remarks from within it.

What’s barely glimpsed is the outside of life. The healthy alienation from received institutional structures that might have resulted from Sheed’s parents—who were famous for their fervent Catholicism but also clearly somewhat embarrassing to their son—or from his involuntarily international childhood and his polio years is shrugged away in favor of burrowing into cranky conformity, despite his mock protests. And he only allows us to infer his rejection of Catholic belief. Hearing his mother in a debate on the radio, he identifies with her interlocutor: “Here was a bright, possibly even sane woman showing him a world as bizarre as Dante’s.” Listening to his father’s eulogy at St. Patrick’s Cathedral, he says, “That day, at least, I believed every word of it.”

Sheed may have never met a difference he didn’t wish to split. In creating fictional self-surrogates bursting with rude intolerance, he appears to be running a compensatory game, one meant to smuggle in opinions that might get him disinvited from certain parties. This maneuver recalls the worst of Saul Bellow, though in Sheed’s hands it is fleeter and funnier. As in a book like Mr. Sammler’s Planet (1970) the unspoken script might read, “If you think I (the author behind the curtain) am a reactionary, just look at how bad things could get if I took the gloves off.” Politics aside, Sheed’s fiction too much takes its lead from his criticism. His narrative voice and his characters display a bottomless obsession with triangulation, mimicry, counterfeiting, imposture.

This passion for exposing fakery serves Sheed best in the only other fiction of his I can really recommend, The Blacking Factory (1968), a coming-of-age novella. Here, a boy named Jim Bannister, whirled between England and the US, tabulates the growth of his prosecutorial sensibility as he considers his father’s new girlfriend, and his father, and himself:

Lorraine talked in a Seven Sisters drawl that sounded suspect to Jim (although he knew this accent sounded fake even when it was done right), and this brought out a trace of latent North Boston in Mr. Bannister’s answers. It was like being able to look into people’s skin through a microscope. “Can’t you hear that she’s a phony?” he wanted to tell his father at one point, and then: “Can’t you hear that he’s a phony?” to Lorraine. Can’t you hear that we’re all phonies? When Jimmy chimed in himself, his own voice sounded just as silly as theirs, and the thought of having to listen to the three of them all summer became really quite depressing.

Even in this, Sheed’s punch is pulled with “really quite depressing”—as though the razor insight of Robert Musil had retreated into the stifled yawn of P.G. Wodehouse. Total skepticism isn’t necessarily a novelist’s enemy, but it finds a poor companion in consolation-seeking. Sheed’s skepticism is his fiddler crab’s gigantic claw; meanwhile his tiny claw, where he stores his sense of pain and loss, he can barely cinch open. Amazingly, this is even true in memoirs devoted to his parents and to his recovery from illness and addiction.

Something is being hidden, and though I’m no expert on Catholicism I think it may be Sheed’s reluctance to explain his retreat from belief. Perhaps he was too obedient to his legacy: the most Catholic of all great Catholic writers, G.K. Chesterton, was Sheed’s godfather. Instead, an unspecific God-that-failed haunts Sheed; each time he tries and fails to substitute Babe Ruth, jazz, or literature in that seat of authority he reenacts this primal disappointment. His prose, on the surface, is worthy of Chesterton’s, every sentence issued with a twist, every face peeled to show a new face underneath, a delight in paradox as a lens on all creation. Yet Chesterton’s paradoxes conceal—until they can’t—real belief, and a desire to persuade the reader of the existence of an absolute higher good. Sheed’s hint instead at a trauma or nostalgia that they simultaneously deny.

As Elaine Blair so mercilessly points out, there is a zero-sum effect at work when any book is published, or republished. The reading oxygen breathing in and out of individual brains to form the ecosphere supporting new writing—or other rediscoveries—is not limitless. Let me be merciless in turn: I don’t think we need more Sheed.

Regime Change

Portside
portside.org
2026-10-03 19:18:22
Regime Change Dave Sat, 10/03/2026 - 19:18 ...
Original Article

Left: Coin featuring Julius Caesar, 44 BCE. [The Trustees of the British Museum] Right: Coin featuring Donald J. Trump, 2026. [U.S. Mint] | U.S. Mint

New commemorative coins celebrating America’s 250th anniversary are now on sale from the United States Mint . The coins are part of a series of “golden dollars” (their hue derived from a coating of manganese brass) that first appeared in 2000 bearing the image of Sacagawea. This latest addition features a living person: a stern-faced Donald Trump, facing outward. The word “Liberty” appears above his head along the upper rim of the coin.

Trump’s coin, as many critics have pointed out , seems to violate the spirit, if not the letter, of an 1866 law permitting only the image of deceased persons on bonds, securities, notes, fractional currency, and postal currency. Mint representatives responsible for the Trump dollar coin argue that commemorative coins are exempt under a 2020 law permitting new designs for the nation’s 250th birthday. They cite an example from 1926, when a commemorative half-dollar featuring President Calvin Coolidge alongside George Washington was issued at the Sesquicentennial Exposition in Philadelphia.

Despite this one (very unpopular) precedent, custom has long rejected the image of living persons on the coins of the United States. The Coinage Act of 1792 clearly defines currency as the place on which the most sacred of American values ought to be enshrined: “Upon one side of each of the said coins there shall be an impression emblematic of liberty, with an inscription of the word Liberty.” Coins were meant to send a powerful message about the American political system and its rejection of individual power in favor of personal freedoms. The contrast between a clear symbol of autocratic rule — the stern-faced Trump — and the message of personal freedom — “Liberty” — shows a disconnect between the ideals of the American founders and the concerted effort of the current administration to redefine long-held national values. This jarring combination of image and text has direct parallels to numismatic messaging from the end of the Roman Republic. In the mid-first century bce , when Julius Caesar became the first living Roman to be featured on coins, the meaning of liberty, coopted and contested, proved a key element in the conflicts that dissolved the Republic and led to the rise of Empire.

Left: 1-cent coin featuring right-facing Lady Liberty, 1793. Right: 5-dollar coin featuring left-facing Lady Liberty, 1815. [U.S. Mint]

For almost five centuries, ancient Rome was governed by a coalition of people and elected magistrates. The Romans called their unique form of government the res publica , “the public affair.” According to Polybius, a historian of the second century bce , it was the Republican system, characterized by a division of power and a system of checks and balances, that gave Rome its competitive edge and allowed one small city in Italy to become a Mediterranean superpower. As in modern America, the success of the Roman Republic relied on a culturally embedded abhorrence of monarchy. Far back in its legendary past, Rome had overthrown its kings under the leadership of a nobleman named Lucius Junius Brutus. The Roman historian, Tacitus, begins his history (the Annales ) crediting Brutus with establishing the Republican form of government and “ libertas .” Libertas , Latin for liberty, was a deeply enshrined value in Roman society, signifying the people’s power to shape their Republic. Libertas was even worshipped as a goddess in Rome and was commonly featured on coins, depicted holding a cap traditionally granted to slaves when they received their freedom.

Rome’s currency reflected its rejection of autocratic rule. Coins were minted by a rotating group of moneyers who came from the small class of Roman elites making their way up the Roman political ladder. For centuries historians and numismatists have combed through fragmentary literary sources and made detailed studies of Roman coins. Despite this, the exact mechanisms behind the selection and design of numismatic types in ancient Rome is still an open debate. What we do know is that the magistrates must have helped choose the designs, which often celebrated the actions of their own famous ancestors. These designs, however, never featured a living Roman. This was likely not enshrined in law, but rather in custom. The Romans, while they did have laws, did not have a written constitution, but rather acted in accordance with mos maiorum , “the custom of ancestors.”

For the Romans, coins featuring living men were associated too closely with the monarchical systems of the eastern Mediterranean. Since the conquests of Alexander the Great in the fourth century bce , powerful emperors who controlled great swaths of eastern territory were featured on the currency of the cities that they controlled. The Romans knew these autocrats well, having spent much of the third and second centuries bce fighting against them. A few successful Roman generals even went so far as to playact at this sort of power. In 196, in the wake of victories in mainland Greece, the Roman general Titus Quinctius Flamininus created a gold coin with his own face on it. The coin, minted in northern Greece, was never adopted in Rome. In fact, it was tradition that returning generals surrendered their command at the gates of Rome, leaving behind power and influence accrued in far-flung wars. By the end of the first century bce , though, the Republic was in trouble. A series of generals who had amassed fame and riches in their conquests were increasingly reluctant to give up the command of their troops when they returned to Rome.

The trend reached a tipping point with Gaius Julius Caesar. In 49 bce , after accruing massive wealth and popularity from a seven-year military campaign in Gaul, he crossed the Rubicon River and entered Italy, refusing to disband his army. They swept down the Italian peninsula and into Rome. He spent the next few years consolidating his power, eliminating rivals, and reorienting the working of the res publica around himself and his agenda. Both Caesar and his opposition championed the cause of liberty, and in libertatem vindicare, “to set free” became a rallying cry of both sides. This deeply held value to the Roman public easily became, as the great Oxford historian Ronald Syme wrote, “a convenient term of political fraud.” After Caesar’s final victory, the remaining members of the senate decreed that he would be called Liberator and voted to build a public temple to Libertas .

A year later, in 44, the senate named him dictator for life, and Caesar became the first living person to have his face minted on coins in the city of Rome. One of his early moves had been to take over the mint, placing it in the hands of his personal servants. He also expanded political positions to make room for his supporters. Thus the moneyers that year were all friends and partisans of Caesar. The words accompanying his image were not subtle. Caesar, “Dictator in Perpetuity,” read one ; “Caesar, Father of the Fatherland,” read another . We cannot know how these coins were interpreted by the average Roman, but for his fellow Romans of senatorial class, they were part of a growing series of authorial oversteps that would seal Caesar’s fate.


Silver coin with head of Julius Caesar and Venus, 44 BCE. [ The Trustees of the British Museum ]

A month later, Caesar was dead, killed by a coalition of senators who called themselves the Liberators. At the head of the conspiracy was one Marcus Junius Brutus. Throughout his career, he had minted coins referencing his most illustrious ancestor, the famous Lucius Iunius Brutus who had expelled the kings and brought libertas to Rome. Brutus had long crafted his public persona as a defender of the basic political rights of the Republic. One early coin pairs an image of the first Brutus with a goddess Libertas . Another , minted in exile from Macedon after Caesar’s death, features a freedman’s cap between two daggers. The legend, “The Ides of March” refers to the day of Caesar’s assassination. The meaning is clear: In a Republic, autocratic rule is the enemy of libertas .

The assassination of Caesar did little to preserve the Republic. His cause was immediately taken up by a coalition of successors: his right-hand man Marcus Antonius (Mark Antony) and his nephew, Octavian, banded together to fight the assassins, whom they defeated in Greece in 39 bce . Octavian and Antony, who soon turned against each other, both used the concept of Libertas in the battle for public approval. Octavian had the final word. In 31 bce he defeated the combined forces of Antony and the Egyptian queen Cleopatra, establishing himself as the dominant power in the Roman world. In the wake of the battle, he minted a coin naming himself Libertatis P R Vindex , “Defender of the Liberty of the Roman People.”

A decade later, he would take the name Augustus, “The Magnificent.” His enemies defeated, he took control of Rome and declared the res publica restored. He was Rome’s first emperor. Before his death, Augustus prepared his own funeral oration for public display — a long list of things accomplished throughout his career. He began with his arrival in Rome after the assassination of Caesar. In his version of the story, he came to restore libertas in a time when the assassins of Caesar had brought great risk to the res publica : “At the age of 19, on my own initiative and at my own expense, I raised an army by means of which I restored liberty ( libertas ) to the republic, which had been oppressed by the tyranny of a faction.”

From Augustus onward, coins across the Mediterranean featured the bust of the living emperor. Taking cue from Augustus, the imperial family gave Libertas a central role, restoring her temples, erecting statues, and featuring her on coins. By the time of the third emperor, Claudius, coins showed a new name for the goddess, Libertas Augusta, “ Augustan Liberty.” When the Republic fell, it was not only the form of government that changed in Rome, but the very definition of liberty. No longer did it symbolize freedom from autocratic rule, but rather it became a symbol of the protection that only autocratic rule can provide.

Silver coin with head of Augustus and heads of Caius Caesar, Julia, and Lucius Caesar, 13 BCE. [ The Trustees of the British Museum ]

The new commemorative coins are part of a larger strategy by the Trump administration to redefine liberty and, in doing so, reshape the American Republic. When Representative Martin Russell Thayer of Pennsylvania proposed the currency amendment in 1866 — a reaction to a Treasury official minting his own face on a series of 5-dollar bills — he cited the example of the Roman Republic in words that now seem prescient: “When we shall have a living Caesar to make both the laws and the money of this country it will be time enough to place his effigy upon the coins and notes of the United States.”

The portrait of a living individuals on American currency sparked comparisons to Caesar in 1866, and the immediate connection to autocratic rule has not faded. In a reaction to the Treasury’s plans, two democratic senators, Nevada’s Catherine Cortez Masto and Oregon’s Jeff Merkley, introduced for a new law that clearly extends the 1866 prohibition to all and any future U.S. currency. “President Trump’s self-celebrating maneuvers are authoritarian actions worthy of dictators like North Korea’s Kim Jong Un, not the United States of America,” said Merkley.

Long at the heart of American political rhetoric, “liberty” has been a fixed message on American numismatics since 1793, when the first copper cents that were minted featured an image of Lady Liberty. The end of the Roman Republic, however, teaches us not to see the writing above Trump’s portrait as mere adherence to custom. Alongside Trump’s cryptocurrency initiative, “World Liberty Financial” or his Department of Justice’s “Religious Liberty Commission,” these coins actively work to depict Donald Trump as the defender of a new brand of American liberty, centered on himself and defined in opposition to political dissent. The power of this message lies in the public’s own forgetfulness and disregard for the historical precedents so well-studied — and feared — by the American founders. When the people give the power to define and grant liberty to one powerful individual, they risk losing it entirely.

Emily Lucille Hurt is a professor of Classical Studies at John Cabot University in Rome. She has a doctorate in Ancient History from Yale University and has held fellowships from the American Academy in Rome and the American School of Classical Studies at Athens.

Declaring a bird extinct: The median wait is 36 years after the last sighting

Hacker News
birdshistory.com
2026-10-03 19:16:45
Comments...
Original Article

A bird stops being seen. Then, eventually, somebody official says it is gone. The gap between those two moments is far longer than most people assume.

On this site there are thirteen species and subspecies for which both dates are documented: the last confirmed record, and the first formal declaration of extinction by a recognised authority. The median gap is 36 years . The mean is 47.5 . The shortest was three years — the dusky seaside sparrow , whose last individual died in a cage with a numbered band on its leg. The longest was 124 years : the Kauaʻi nukupuʻu , last collected in 1899 and not struck off the US endangered list until October 2023 .

Four more birds here have never been declared extinct at all, despite last records in 1935, 1944, 1956 and 1963.

Quick facts

What is measured Years between the last confirmed record of a bird and the first formal declaration of its extinction
Species in the dataset 13 — every page on this site where both dates are documented
Median gap 36 years
Mean gap 47.5 years
Range 3 to 124 years
Longest Kauaʻi nukupuʻu — 1899 to 2023
Shortest Dusky seaside sparrow — 1987 to 1990
Gaps over 50 years 5 of 13
Gaps within 20 years 3 of 13
Never declared at all 4 species, last recorded 1935–1963
The IUCN test “No reasonable doubt that the last individual has died”

What does it actually take to declare a bird extinct?

A mounted Kauaʻi ʻōʻō on a wooden perch in a dark museum case, lit from the side: a slim dark-brown songbird with a long curved black bill, finely barred pale markings on the throat, a long pointed tail, and a bright yellow tuft of feathers at the top of the leg
A mounted Kauaʻi ʻōʻō at the Bishop Museum, Honolulu. For a bird declared extinct, museum skins become the whole physical record — and the evidence an assessor has to work from is the absence of anything newer.

One sentence does most of the work. The IUCN’s definition of Extinct requires that there is no reasonable doubt that the last individual has died .

Read that as an evidential standard rather than a description, and the delays stop being mysterious. It is a negative claim, and negative claims about wild animals are close to unprovable. No survey can cover every valley. A bird that was hard to find when it was common is harder to find when it is nearly gone, and absence of evidence from a forest nobody has walked is not evidence of absence.

The Red List asks for exhaustive surveys in known and expected habitat, at appropriate times, across the species’ historic range. For a bird confined to one gully that is achievable. For the Eskimo curlew , which bred somewhere in the Canadian Arctic and wintered on the Argentine pampas, “exhaustive” describes nothing anyone has ever done.

During the 1980s a simpler rule was floated: no verified sighting for fifty years , and the species is declared gone. It did not survive contact with the evidence. Too many animals came back. The black-browed babbler of Borneo was rephotographed in 2020 after 170 years ; the Bermuda petrel was presumed extinct in 1621 and found nesting on four islets in 1951 , a gap of 330 years . A fifty-year rule would have declared both, and been wrong twice.

So the fifty-year threshold was never adopted as a test. What replaced it is judgement, applied case by case, by committees that meet on their own schedule — and judgement is slow.

How long does the wait actually last?

A photograph of a living dusky seaside sparrow perched on a bare pale twig against a green background: a small sparrow with a heavily black-and-white streaked breast, a blackish head with a yellow spot above the eye, a pale grey conical bill and a dark tail
A living dusky seaside sparrow. Its last individual, Orange Band, was captured, banded and housed, and found dead on 17 June 1987 — which is why its extinction was declared within three years, the shortest wait on this page.

These are the thirteen birds on this site where both dates are on the record. Every figure comes from the species page, where it is sourced.

Species Scientific name Last confirmed record Declared extinct Years
Kauaʻi nukupuʻu Hemignathus hanapepe 6 May 1899 16 October 2023 124
Maui nukupuʻu Hemignathus affinis June 1901 16 October 2023 122
Paradise parrot Psephotellus pulcherrimus 1928 1994 66
Kākāwahie Paroreomyza flammea about May 1963 16 October 2023 60
Kauaʻi ʻakialoa Akialoa stejnegeri 12 April 1969 16 October 2023 54
Bachman’s warbler Vermivora bachmanii 1984 16 October 2023 39
Kāmaʻo Myadestes myadestinus 27 June 1987 16 October 2023 36
Maui ʻākepa Loxops ochraceus 1988 16 October 2023 35
Slender-billed curlew Numenius tenuirostris 25 February 1995 10 October 2025 30
Carolina parakeet Conuropsis carolinensis 21 February 1918 1939 21
Poʻouli Melamprosops phaeosoma 26 November 2004 2019 15
Kauaʻi ʻōʻō Moho braccatus late 1987 2000 13
Dusky seaside sparrow Ammospiza maritima nigrescens 17 June 1987 December 1990 3

Three rules were applied consistently, and they matter for reading the table.

“Last record” means the last evidence the bird was alive — a specimen, a photograph, a sound recording, an observation accepted by the relevant authority, or the death of the last known captive. For the kāmaʻo that is a tape recording of one singing male made on 27 June 1987. For the Kauaʻi ʻōʻō it is David Boynton’s recording of the last male singing a duet with nobody answering.

Where a bird was declared twice, the earlier declaration counts. Eight of these species were removed from the US Endangered Species Act in October 2023; several had already been listed Extinct by the IUCN years before, and those earlier dates are the ones in the table. Using the American dates throughout would have inflated every figure.

The dataset is this site, not the world. Thirteen species is a small sample, weighted toward Hawaiʻi because that is where the extinctions were. It is not a global census, and the median should be read as a description of these birds rather than a published constant.

What survives those caveats is the shape of the distribution. The three short gaps have one thing in common: somebody watched the last bird die. Orange Band, the last dusky seaside sparrow, was found dead on 17 June 1987 at Disney’s Discovery Island, having been caught, banded and housed. The last poʻouli died in captivity on 26 November 2004 with avian malaria. When there is a body and a date, the declaration follows within a few years.

Everywhere else the record simply stops, and the stopping is not an event anybody can put a date on.

Timeline: how the rules got this way

Date Event
1939 The American Ornithologists’ Union formally accepts the extinction of the Carolina parakeet, 21 years after the last captive died
1966 The USFWS proclaims both nukupuʻu extinct
1967 Both nukupuʻu are nevertheless listed as endangered — among the first 78 US endangered species
1970, 1982 Hawaiian honeycreepers listed federally, then by the State of Hawaiʻi
1980s A 50-year no-sighting rule is proposed, and abandoned — too many species came back
1994 The IUCN publishes the Red List categories built on “no reasonable doubt that the last individual has died”
2010 Elphick, Roberts and Reed publish modelled extinction windows for North American and Hawaiian birds
2016 Roberts and Jarić extend the modelling to handle uncertain sightings
2021 The USFWS proposes delisting 23 species as extinct , including the ivory-billed woodpecker
16 October 2023 21 species delisted; 10 are birds. The ivory-bill is held back
July 2024 The ivory-bill delisting proposal is formally withdrawn
10 October 2025 The IUCN declares the slender-billed curlew extinct — 30 years after its last record

The 1966–67 pair of lines is the oddest thing in the record and it is not a transcription error. The US Fish and Wildlife Service proclaimed the nukupuʻu extinct in 1966 , and then listed both species as endangered in 1967 . The bird was officially gone and officially protected within twelve months, and the protection outlasted the proclamation by fifty-six years.

Why does it take decades?

Two reasons hold up, and one does not.

The first is that premature declarations are a real error with real costs. Under the US Endangered Species Act, delisting is not a label change. It withdraws the legal protection of habitat. If the bird is still there, the declaration removes the thing keeping it alive. That is why the ivory-billed woodpecker was proposed for delisting in 2021, held back from the 2023 batch, and had its proposal withdrawn entirely in July 2024 — 80 years after the last accepted sighting. The service decided that being wrong in one direction was worse than being slow.

The second is that rediscovery is not rare. Martin and colleagues counted 562 terrestrial vertebrate species that nobody had seen for at least 50 years and that nobody had declared extinct — 38 of them birds . For comparison, 311 terrestrial vertebrates have been declared extinct since 1500. There are 80 per cent more lost species than declared ones , and a meaningful fraction of the lost ones are still out there: we have a whole page of birds that came back , one of them after 330 years.

The reason that does not hold up is that the delay is harmless. Two of the birds in the table above had been gone for more than sixty years when the United States first granted them legal protection in 1967. Both nukupuʻu were last recorded around the turn of the twentieth century. Between 1968 and 1975, searchers spent over 500 days in the field on Kauaʻi looking for them and others, and confirmed nothing. The protection was real, the money was real, and the birds had been dead for two generations.

Meanwhile at least 26 sightings of the Kauaʻi nukupuʻu and 27 of the Maui nukupuʻu were submitted and rejected between 1960 and 1996 — not because the observers were careless, but because a surviving honeycreeper, the ʻamakihi, occasionally grows a deformed bill and looks exactly like a nukupuʻu in a glimpse. Fifty-three reports, no evidence, half a century of maybe.

What happens during the gap?

Not nothing. The decades between a last record and a declaration are filled with searching, and the search effort is where most of the money goes.

The pattern repeats across the species on this site. On Kauaʻi, the 500-plus days of fieldwork between 1968 and 1975 were a deliberate, funded attempt to settle whether the ʻakialoa and the nukupuʻu were still there. On Molokaʻi, surveys for the kākāwahie ran in 1972, 1975, 1979–80, 1988 and 1995 , every one negative. Munro searched for a month in 1936 and failed — then noted that he had never reached the upper forests, which is exactly the gap that keeps a species officially alive.

For Jerdon’s courser , which is lost rather than extinct, the modern version of this looks like 200 camera traps and 107,000 images between 2013 and 2015 , followed by 17 acoustic recorders in 2019–20. Nothing in either. For the pink-headed duck, a standing cash reward ran from 1930 to 1960 and was never claimed, and BirdLife surveyed Kachin State in Myanmar between 2003 and 2005 on the strength of unconfirmed reports.

Searching is how the doubt gets removed, and removing doubt from a negative claim is slow, expensive work that produces nothing publishable when it succeeds.

Why does the same bird have two extinction dates?

Because two authorities decide independently, and they do not synchronise.

The Kauaʻi ʻōʻō is the clearest case. The IUCN listed it Extinct in 2000 . The United States removed it from the Endangered Species Act on 16 November 2023 — twenty-three years later, for the same bird, on the same evidence. For the poʻouli the two dates are 2019 and November 2023 , four years apart. For the Carolina parakeet, the IUCN’s extinction date is 1920 while the American Ornithologists’ Union did not accept extinction until 1939 , and the last captive bird had actually died in February 1918 — three dates, none of them the same.

The paradise parrot shows the cost of the lag in a single species. Australia’s only extinct mainland bird was last reported around Gayndah in 1928 . The IUCN first listed it as Threatened in 1988 — sixty years after the last record — and only reclassified it as Extinct in 1994 . For six years it was formally a species in trouble rather than a species that was gone.

So a reported extinction date is always a date on a decision, not a date on a death. Which authority made it, and when they happened to meet, changes the figure by decades.

Which birds have never been declared extinct at all?

A male imperial woodpecker specimen photographed from behind against a black background: a scarlet crest at the top of an otherwise black head and back, a broad white panel across the folded wings, a long black tail, and pale claws gripping a white mount
Male imperial woodpecker, Naturalis Biodiversity Center. Last confirmed in 1956 by 85 seconds of film, and still listed as Critically Endangered (Possibly Extinct) seventy years later.

These four are on this site, all of them with a last record in living memory or close to it, none of them formally extinct.

Species Scientific name Last confirmed record Years since Current status
Pink-headed duck Rhodonessa caryophyllacea 21 June 1935 91 Critically Endangered (Possibly Extinct)
Ivory-billed woodpecker Campephilus principalis April 1944 82 Endangered; delisting withdrawn July 2024
Imperial woodpecker Campephilus imperialis 1956 70 Critically Endangered (Possibly Extinct)
Eskimo curlew Numenius borealis 1963 63 Critically Endangered (Possibly Extinct)

“Possibly Extinct” is the Red List’s honest shrug, and it is doing a lot of work. All four of these birds are, on the evidence, almost certainly gone. The last imperial woodpecker anyone can verify is a female in 85 seconds of 16 mm film shot by William Rhein in Durango in 1956; when researchers found the filming site again in 2010, the forest had been logged. The last pink-headed duck accepted in the wild was trapped in Bihar in June 1935, and a standing cash reward for physical evidence went unclaimed for thirty years.

None of that is the same as proof, so the category stays.

When did these birds actually die?

A mounted Eskimo curlew in a museum case, standing among dried grass stems against a deep magenta background: grey-buff plumage finely mottled and streaked, a long slender decurved dark bill, and dark legs
A mounted Eskimo curlew at the Redpath Museum, McGill University. The last confirmed bird was shot on Barbados in 1963; no declaration has followed in the sixty-three years since.

Not when the records stopped, and not when the paperwork caught up. Somewhere in between, and much closer to the first.

There is a body of work on exactly this. Elphick, Roberts and Reed modelled extinction dates for North American and Hawaiian birds from their sighting records in 2010; Roberts and Jarić extended the method in 2016 to handle sightings of uncertain reliability. The outputs are windows rather than dates, and the windows sit close behind the last record:

  • Hawaiʻi mamo — last seen July 1898, modelled extinction 1911–1915
  • Kauaʻi ʻakialoa — last seen April 1969, modelled extinction 1967–1973
  • Kāmaʻo — last recorded June 1987, modelled extinction 1991–1999
  • Both nukupuʻu — modelled extinction about 1901 , upper limit 1906
  • Labrador duck — last credible sighting 1878, estimated extinction 1880

Set those against the declaration dates and the arithmetic is bleak. The model says the Kauaʻi nukupuʻu was gone by 1906 at the outside. The United States said so in 2023. On the best available estimate of when the bird actually died, the paperwork trailed it by 117 years .

The kākāwahie is the exception that shows the limits of the method: its modelled window runs 1910 to 1969 , which is no window at all. A bird collected 36 times in three months in 1907 and then found twice by one observer in the early 1960s produces a sighting record the models cannot resolve.

Both dates are now in one place for every species on this site. The extinct birds explorer plots all thirty-one on a map, a timeline and a sortable table, naming for each one what the last record actually was — a specimen, a sighting, a photograph, a sound recording or a death in captivity. Fourteen of the thirty-one end with a sighting, which is the weakest evidence on the list and a large part of why the declarations take so long.

Interesting facts

  • The shortest gap on this site, three years, belongs to a bird that was captured, banded, given a name and housed at a Disney theme park. Precision about extinction requires custody.
  • The USFWS proclaimed the nukupuʻu extinct in 1966 and listed it as endangered in 1967 . Both statements stood simultaneously for 56 years.
  • The slender-billed curlew’s last irrefutable record and its extinction declaration are 30 years and 7 months apart — 25 February 1995 at Merja Zerga in Morocco, 10 October 2025 in Bonn.
  • 38 bird species worldwide have gone unseen for over 50 years without being declared extinct.
  • The last evidence for two species on this list is audio , not sight: a tape of one kāmaʻo singing in 1987, and a recording of the last Kauaʻi ʻōʻō performing a duet alone.

FAQ

How long does it take to declare a bird extinct? On the thirteen documented cases on this site, a median of 36 years after the last confirmed record, and a mean of 47.5. The range is wide: three years when the last bird died in custody, 124 years when the record simply stopped.

Is there a 50-year rule for declaring extinction? No. A 50-year no-sighting threshold was proposed in the 1980s and abandoned, because too many species were rediscovered after longer gaps than that. The IUCN instead requires “no reasonable doubt that the last individual has died”, judged case by case.

Why are some birds listed as “Possibly Extinct”? Because the evidence points strongly to extinction but falls short of the IUCN’s standard. Four birds on this site sit in that category, with last records from 1935 to 1963. It is a statement about the evidence, not a prediction that the bird will turn up.

How many birds are “lost” but not declared extinct? 38 , in the global count published by Martin and colleagues — species unseen for at least 50 years with no formal declaration. Across all terrestrial vertebrates the figure is 562, against 311 declared extinct since 1500.

Which bird waited longest to be declared extinct? Of the birds covered here, the Kauaʻi nukupuʻu : last collected 6 May 1899, delisted 16 October 2023. 124 years.

Does declaring a species extinct do any harm? It can. Under the US Endangered Species Act, delisting removes habitat protection, so a premature declaration withdraws the safeguard from a bird that may still be alive. That risk is why the ivory-billed woodpecker’s delisting was proposed in 2021 and formally withdrawn in July 2024.

Who actually makes the declaration? Usually the IUCN, through its Red List assessments for BirdLife International, or a national authority such as the US Fish and Wildlife Service under the Endangered Species Act. The two work on separate timetables and often disagree by decades.

Sources

Image credits

Google Gemini could soon get full access to your Mac’s files, apps and the web

Bleeping Computer
www.bleepingcomputer.com
2026-10-03 19:12:34
Google's Gemini could soon access any file on your macOS device, open apps, browse the web, and perform actions without asking for permission every time. [...]...
Original Article

Gemini

Google's Gemini could soon access any file on your macOS device, open apps, browse the web, and perform actions without asking for permission every time.

As spotted by TestingCatalog on X, Google is testing desktop control for Gemini, and there are references to a new hidden "Additional sandbox options" setting in the Gemini Desktop app.

The feature isn't live yet, and Google hasn't confirmed it either, but when enabled, Google Gemini could read, create, modify, or even delete files anywhere on your PC.

That means Gemini will have access to files outside folders you have explicitly connected to Gemini.

It's unclear how the feature will play out, as Apple is considering making it difficult for AI agents to access personal files and data on Mac.

According to a pop-up within Gemini Desktop, Google's AI could also communicate with apps such as Mail, Safari, or Messages and perform actions through them.

"By enabling additional sandbox options, you will be able to expand what Gemini can do and access on your Mac," Google explains in the hidden interface.

"Depending on which settings you enable, Gemini may be permitted to take actions without asking for your permission first."

Gemini would still ask before sensitive actions

Gemini

Google is not giving Gemini unlimited control without any safeguards, and you can expect a Claude-like experience, where you explicitly give it permission to use your PC.

Gemini would still ask for confirmation before buying products or transferring money, creating an online account, accepting legal terms on your behalf, or modifying sensitive information about you.

The setting appears to be part of Google's broader computer-use plans for Gemini, where the AI would be able to work across files, websites, and native apps instead of being limited to a chat window.

It's unclear when Google plans to roll out Full Access or which Gemini model will power it.

article image

Build your security blueprint for AI-powered attacks

Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.

Save your seat

Bernie Sanders’s 2016 Campaign Lit a Spark Beyond Elections

Portside
portside.org
2026-10-03 19:11:18
Bernie Sanders’s 2016 Campaign Lit a Spark Beyond Elections Dave Sat, 10/03/2026 - 19:11 ...
Original Article

Bernie Sanders walks the picket line with striking auto workers at the General Motors Detroit-Hamtramck Assembly Plant on September 25, 2019. | Bill Pugliano / Getty Images

B ernie Sanders’s impact on the American left is not typically measured in social movement terms. The story normally goes that democratic socialism returned to national politics after 2015, tens of thousands of people joined the Democratic Socialists of America (DSA), a new generation of socialist candidates ran for office, and a much larger left media ecosystem emerged.

That’s all true. But another part of the story has received relatively less attention: the role Bernie’s campaigns played in inspiring and cohering a new generation of labor activists.

I first started making this argument while writing about the red-state teachers’ revolt of 2018. Obviously, Bernie did not cause these massive statewide strikes; years of austerity and attacks on public education created the conditions. But while reporting on those strikes on the ground in West Virginia, Oklahoma, and Arizona, I kept finding that many of the rank-and-file leaders who took the initiative to get them started had been politicized by Bernie’s 2016 campaign.

“The role of the Bernie campaign of 2016 on organizing in West Virginia really cannot be overstated,” Charleston teacher and strike leader Emily Comer told me. Bernie’s campaign, she explained, “got people, especially young people, plugged in who before had been feeling hopeless and who would not have made their way into organizing before.”

This was not just retrospective speculation. Morgantown teacher Anna Simmons described the night West Virginia teachers decided to defy their unions’ proposed settlement and continue their strike. It was the moment, she told me, when she realized that the fight “was about more than just insurance premiums and salaries. It was the continuation of a movement that started with Bernie Sanders and is going to result in a power shift from the elite wealthy to the working people.”

When I first made this case about the 2018 strikes, some critics dismissed it as a Bernie Bro exaggeration. But what happened after 2020 made the larger pattern much harder to dismiss.

The post-pandemic labor upsurge has been driven to an unusual degree by young workers organizing from below, many of whom were radicalized and made class conscious by Bernie. For my book We Are the Union , I surveyed over two thousand worker leaders who unionized their workplaces in 2022; the median respondent was twenty-seven years old and almost half identified politically as radicals.

When I asked what broader movements had influenced them, recent labor struggles were, unsurprisingly, the biggest factor. Yet more than 80 percent of respondents cited at least one outside struggle as a major influence on their decision to unionize. Among non-labor movements, Black Lives Matter and Bernie stood out: BLM had the widest reach, followed right behind by Bernie, whose campaigns also had the most intense reported impact.

My interviews with worker leaders put faces on those numbers. Max Yusen, who helped turn his St Louis Starbucks into a union stronghold, had not grown up around organized labor. “Labor was a big part of Bernie’s message,” he told me, “and being around union members while canvassing and inside DSA really made me see how important they are.” When he moved back to St Louis, he deliberately got a job at Starbucks “with the goal of unionizing it.”

Bernie’s influence also produced institutions that outlasted his campaigns. The Emergency Workplace Organizing Committee ( EWOC ) grew directly out of Bernie 2020 and the DSA for Bernie campaign. When Covid hit in March 2020, workers began contacting the Sanders campaign asking for help winning protective equipment and sick pay. A handful of labor organizers from the campaign, along with organizers from DSA and the United Electrical Workers, set up a Google form and began connecting those workers with volunteer organizers. It was, from the start, animated by both the class-struggle politics and the distributed organizing model of the Sanders campaign.

The same political generation appeared elsewhere in the post-2020 upsurge. Most of the salts , the rank-and-file workers who got jobs in order to organize and who played an important role in the successful Amazon Labor Union campaign at JFK8 on Staten Island had been politicized by Bernie’s presidential runs. At Starbucks, Amazon, CVS, Trader Joe’s, and elsewhere, young workers entered organizing with an anti-billionaire politics that had become far more widespread since 2016.

None of this means that Bernie produced the labor upsurge. Covid was a huge catalyst. So was the tight labor market, a more favorable National Labor Relations Board under President Joe Biden, digital tools, and, above all, the contagious example of workers themselves winning. The 2018 teachers’ strikes themselves helped inspire later strikes; Starbucks and Amazon victories then convinced thousands of other workers that organizing was possible.

But recent evidence does make the connection I saw in 2018 look less incidental. Electoral campaigns can feed workplace struggle by politicizing people, raising their expectations, bringing activists together, and giving them experience organizing others. Bernie’s campaigns did all four.

And that story is still unfolding. Amazon organizing has expanded far beyond the original JFK8 victory: the Teamsters reported this year that roughly ten thousand Amazon workers had organized at thirteen facilities around the country. In New York, Amazon workers and Teamsters are simultaneously organizing on the job and campaigning for the Delivery Protection Act, which would regulate last-mile facilities and impose new employment and safety requirements; as of September 2026, the bill remains pending in the city council with thirty-five council sponsors.

Once MAGA is ejected from the White House, there is every reason to expect that the Bernie generation of labor organizers will go back on the offensive nationwide to unionize the megacorporations that have reaped the benefits of America’s New Gilded Age and that, virtually without exception, have bent the knee to Donald Trump.

Bernie’s contribution to the revival of the American left, then, did not stop with DSA or socialist electoral politics. As Vince Quiles, a young Home Depot worker in Philadelphia, put it to me: “Bernie showed me and a lot of people my age that it’s not a bad thing to fight corporations or demand more from your government. We’re workers, you know, we produce everything, and we deserve more.”

Why I recommend Renovate over any other dependency update tools (2024)

Lobsters
www.jvt.me
2026-10-03 18:54:52
Comments...
Original Article

If you've read my blog before, or interacted with me at work or in the Open Source world, you're likely to know that I'm a huge fan of Renovate .

For those that aren't aware, Renovate is one of the big players in dependency updating tooling, commonly seen in comparisons with Dependabot or Snyk.

I've been using Renovate for the last ~5 years, alongside a mix of Dependabot and Snyk to keep me grounded. I absolutely love Renovate, and make a point of making it so I can run Renovate where possible.

I've also had the experience operating Renovate in self-hosted mode as well as used the hosted Renovate app by Mend and the Mend Enterprise SAAS, and I've written about the lessons learned self-hosting Renovate .

Renovate has some really key features that set it apart from the competition, and I'm sure in the time it's taken me to write this, they've shipped some new ones 😻

(Aside: I'm largely writing this blog post now, as I've recently been shouting the benefits of Renovate, and instead of writing another internal-only document at work (as I did at Deliveroo), I wanted to write it as a form of blogumentation ).

Prior art

A few years back, I wrote about some tips for using Renovate to make keeping your software up-to-date easier, which I still stand by.

Since then, I've worked across a number of different ecosystems, repository sizes, and levels of comfort merging dependency updates, and have learned a few more things about effectively using Renovate - but through it, I'm still very sure of Renovate being the best tool in the ecosystem.

Configurability

Renovate is extremely configurable, with dozens of configuration options to tune your experience. But Renovate doesn't end up being "too" configurable, where you end up spending more time tweaking config than doing the changes, but configurable enough that it's very likely you can do what you need to with it.

This is something that really hit me when we were rolling out Dependabot at Deliveroo, where my team owned ~30 repositories, and so for each of these repos, we needed to hand-craft a dependabot.yml , as Dependabot needs to be told which directories contain which dependencies, and how often to update them.

After we'd had about a week of usage, we started needing to tweak the configuration in a few of them to reduce the noise, which then required us to raise PRs across all the repos and get them updated.

Because Dependabot requires a bit of a snowflake configuration per repository, this wasn't easily automatable, even with great tools available to automate some of the bulk updates , which made this a rather onerous and frustrating process.

Compare this to Renovate, where there's an excellent first-class support for shareable config presets , in which you can "extend" multiple configuration(s).

This allows a team that wants to have consistency (i.e. in how often they receive updates to the AWS and Google Cloud SDKs, or which labels they want on PRs) to create a shared preset for their team that defines this. Then, each of their repositories can "extend" this configuration, as well as defining their own configuration on top of it at a repo-specific level.

This also makes it possible to provide good guardrails for your organisation, providing a good set of defaults, such as a base.json :

{
  "$schema": "https://docs.renovatebot.com/renovate-schema.json",
  "postUpdateOptions": [
    "gomodTidy",
    "gomodUpdateImportPaths"
  ],
  "regexManagers": [
    {
      "fileMatch": [
        "^Makefile$"
      ],
      "matchStrings": [
        "curl .*https://raw.githubusercontent.com/golangci/golangci-lint/master/install.sh | sh -s -- .* (?<currentValue>.*?)\\n"
      ],
      "depNameTemplate": "github.com/golangci/golangci-lint",
      "datasourceTemplate": "go"
    }
  ]
}

This then allows defining a somewhat opinionated good starting point for teams with a default.json :

{
  "$schema": "https://docs.renovatebot.com/renovate-schema.json",
  "extends": [
    "local>your-org-here/renovate-config:base",
    "config:best-practices"
  ]
}

Then, teams only need to set the following in their repos' renovate.json , and they'll have the benefits of onboarding with a lot of the hard work done with starting:

{
  "$schema": "https://docs.renovatebot.com/renovate-schema.json",
  "extends": [
    "local>your-org-here/renovate-config"
  ]
}

And that's it 👏

Good defaults

Linking back to the comment about having to hand-craft dependabot.yml s, the great thing about Renovate is that there's some great defaults, and that it can autodetect the ecosystems your repo uses, and appropriately raises PRs.

This ease of onboarding is truly excellent, and you can even have an onboarding PR raised to your repos to make it even simpler.

For instance, at Deliveroo, we made it so there was a default set of configuration for all repos to give us an easy means to update Docker images, but then if teams wanted to manage everything, they could add a renovate.json , and it'd be more fully featured.

Additionally, Renovate comes with some great inbuilt configuration in the form of presets, including a "best practices" guide and associated preset, which makes it easier to keep on top of community best practices, without needing to bikeshed about what you think is best.

If you find that config:best-practices is a little too much, there's config:recommended as a starting point, and you can always downgrade or exclude rules you'd prefer not to follow. Or if you really want to control things config:base is the minimum you should pull in.

Grouping

As suggested above, it's possible to tune how different packages get updated, where you can group multiple updates into a single PR .

For instance, let's say that you use 9 different AWS services in your application, and instead of receiving 9 PRs every time there are updates across the SDKs, you want a single one. In this case, you could craft the following Renovate configuration:

{
  "$schema": "https://docs.renovatebot.com/renovate-schema.json",
  "packageRules": [
    {
      "groupName": "aws-sdk-go",
      "matchPackagePatterns": [
	"^github.com/aws/aws-sdk-go-v2"
      ]
    }
  ]
}

What's great about this is that you can also bundle things like all patch updates into a single PR:

{
  "packageRules": [
    {
      "matchPackagePatterns": [
        "*"
      ],
      "matchUpdateTypes": [
        "patch"
      ],
      "groupName": "all patch dependencies",
      "groupSlug": "all-patch"
    }
  ]
}

This is something that's recently been added to Dependabot, which is great, cause Renovate's had it for years 😝

Use for one-off bumps

Because Renovate is an Open Source tool you can run on the command-line, it means you can also get the ability to use Renovate for one-off executions , for instance to get everyone in your organisation to a minimum version of a given dependency, or just to do an infrequently performed set of updates.

Something I always refer back to is the fact that Renovate has a tonne of supported package managers, package ecosystems and versioning tools.

Renovate's bot comparison guide docs has a good example which links out to the differences between:

As Dependabot is more focussed on being used on GitHub's platforms (citation needed?) support for things like CircleCI, GitLab CI or other competitors' tooling doesn't seem to be available.

Adding support for additional ecosystems

One of the things you'll be used to having from your dependency update tool of choice is your standard package manager support, where it'll update a go.mod or a pom.xml in your repositories.

But one thing that's quite important to understand is that there are many more dependencies in your project than those installed in your package manager. Something I find great about Renovate is that, as well as managing Dockerfile s, build.gradle s, .gitlab-ci.yml , etc, it will also manage things your .ruby-version or configuration for the ASDF version manager.

But what about some of the non-standard, or organisation-specific means for tracking versions?

For instance, installing the golangci-lint linting tool recommends using curl | sh :

$(GOBIN)/golangci-lint:
	curl -sSfL https://raw.githubusercontent.com/golangci/golangci-lint/master/install.sh | sh -s -- -b $(GOBIN) v1.57.2

And it's common for Dockerfile s to have a definition of the version of dependencies:

ARG MONGODB_VERSION=6.0.4

Or what if you have a custom deployment configuration like:

# this is an (imaginary) company specific deployment configuration file
application:
  lb:
    image: uk-tooling/load-balancer@v3.0.0

It's very unlikely that you have these supported by other tools out-of-the-box, and Renovate is the same.

But Renovate does give you the ability to manage these versions yourself.

Using Renovate's custom managers we can craft a configuration file such as:

{
  "$schema": "https://docs.renovatebot.com/renovate-schema.json",
  "regexManagers": [
    {
      "fileMatch": [
        "^Makefile$"
      ],
      "matchStrings": [
        "curl .*https://raw.githubusercontent.com/golangci/golangci-lint/master/install.sh | sh -s -- .* (?<currentValue>.*?)\\n"
      ],
      "depNameTemplate": "github.com/golangci/golangci-lint",
      "datasourceTemplate": "go"
    }
  ]
}

And this will allow us to manage the golangci-lint installation.

Ideally, this sort of configuration may make it upstream so Renovate can do it out-of-the-box, but as we can see from the above example, there may be things that are organisation- or repo-specific, and so having it upstream'd doesn't make sense.

Because you can do more with it

And you can do even more with it - I've mentioned you can use it for one-off updates on the command-line, but for instance I've also written a tool renovate-graph which takes the detected dependency data from Renovate and gives you a JSON blob you can consume.

This underpins a large swathe of the power behind dependency-management-data and can be used for other means, such as converting Renovate data exports to a Software Bill of Materials (SBOM) .

Dependency Dashboard

One thing I love about Renovate is that you can enable the Dependency Dashboard , which gives you an overview of the detected dependencies, open PRs, as well s anything that may be waiting, or has failed to update.

This is a hugely useful insight into the at-a-glance how far behind are we on updates, giving a view of whether you maybe want to spend a bit more time focussing on updates, or looking at ways to cut through the noise.

When you merge a PR into your default branch, Renovate will rebase open PRs, so they're easier to review, and are guaranteed to run against the latest changes. However, if you're not getting to your updates as often as changes are going in, you may have a tonne of PRs constantly building, which is a waste of energy and CI minutes.

Instead, you can use the Dependency Dashboard for its ability to i.e. require major bumps be gated behind a manual approval , as it's likely you'll need some human interaction for that PR, and can then only raise it when you're actually ready to deal with it.

Open Source

A very important factor to me is that Renovate is Open Source, and open to community contributions.

Although it's probably best to be split into another article, I love the AGPL3, and it's a great way of making sure that anyone hosting Renovate as a platform makes sure that their users get the access to the source code.

Alternatively, Snyk is proprietary, and Dependabot is source available (or at a stretch "open source" not "Open Source") with a license that doesn't even appear on the SPDX license list . Update 2024-05-19 : Dependabot is now MIT-licensed 👏

Renovate is also brilliantly set up as a community project, where they're shipping hundreds of PRs a month, alongside managing the community really well. I've seen a few things change in the last few years I've been more actively contributing, and it's a really well run project and an indication of something I'd love to be able to replicate at some point!

It also helps that Mend , the company behind Renovate, invests a fair bit of time and money into development of the project, as well as their commercial offerings on top of it, which continue to make the project sustainable and remain Free and Open.

Great documentation

Following on from the excellent community and maintainer contributions alike, there's also some really excellent work on a technical writing and documentation point of view, which makes a lot of tasks straightforward to solve.

If there's something a little more complex or custom than the docs can offer, it's usually something that can be answered by the community in a GitHub Discussion, and likely could turn into a docs improvement, if necessary.

There's even a great comparison table between Renovate and other dependency update bots , as well as tips on how best to update your projects .

I've recently discovered a few pages and features that I wasn't aware, just by going through the docs.

Overall

In summary, there's just so much that makes Renovate a greater choice than any of the alternatives, and I'm sure I could talk about more things that make it great.

Whether you self-host it for maximum control (such as being able to access internal artifact registries), or run the hosted app for ease of operations, or just run it from the command-line once in a while, it can be hugely useful to your experience as an engineer.

I hope you'll check it out!

Federal judge calls Flock 'indiscriminate mass surveillance'

Hacker News
techcrunch.com
2026-10-03 18:07:13
Comments...
Original Article

A federal judge ruled this week that a Tulsa, Oklahoma sheriff’s deputy violated a woman’s Fourth Amendment rights when using Flock Safety to search for her license plate without a warrant.

As reported by 404 Media , this ruling does not create a binding precedent, but it is one of the first times that a federal judge has ruled that a Flock search is unconstitutional.

In this case, Judge Sara Hill said the deputy should have obtained a warrant before searching the Flock database for the woman’s license plate, as he had “no apparent reason” for the search “other than the fact that [the woman’s vehicle] had a California license plate.”

The deputy then used the woman’s travel history in Flock as part of the justification for searching her car, where he allegedly discovered 91 pounds of meth. But Judge Hill wrote that all evidence obtained after the Flock search “must be suppressed as the fruit of a poisonous tree.”

Judge Hill also took broader aim at warrantless searches of the Flock database, writing that tracking people’s location — even when they’re in public places — becomes “constitutionally problematic when law enforcement can indiscriminately and passively catalog your whereabouts over an extended period of time and then use that information for any purpose whenever convenient.”

“This is a type of indiscriminate mass surveillance,” Hill wrote. “It is not targeted on a single individual, as in [ Carpenter v. United States , a Supreme Court case focused on how government agencies access location data from cell phones ]. It is a tool that collects information about all vehicles that pass by any network-connected camera at all times, and it serves up the information to law enforcement on demand.”

Hill joins a growing chorus of Flock critics from across the political spectrum. Numerous local and state governments, including Florida and Texas , have said they will stop using the technology. And on Friday, Senator Bernie Sanders — a Democrat from Vermont — introduced the Block Flock Act , which would bar federal agencies from using automated license plate readers such as Flock.

Flock CEO Garretty Langley — who we’ll be interviewing on-stage at TechCrunch Disrupt — has called for a “compromise” between privacy and safety and offered an apology to women who have been stalked by law enforcement officers using the Flock system. And with all those cancellations, Flock has also reportedly offered voluntary employee buyouts as a way to shrink its workforce.

When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence.

Anthony Ha is TechCrunch’s weekend editor. Previously, he worked as a tech reporter at Adweek, a senior editor at VentureBeat, a local government reporter at the Hollister Free Lance, and vice president of content at a VC firm. He lives in New York City.

You can contact or verify outreach from Anthony by emailing anthony.ha@techcrunch.com .

View Bio

September sponsors-only newsletter

Simon Willison
simonwillison.net
2026-10-03 18:00:31
I just sent the September edition of my sponsors-only monthly newsletter. If you are a sponsor (or start a sponsorship now) you can access it here. This month: More Fable class models A pricing war 3D graphics, Blender, and pixel art LLMs come for mathematics So many more accidental cyberattacks Th...
Original Article

3rd October 2026

I just sent the September edition of my sponsors-only monthly newsletter . If you are a sponsor (or start a sponsorship now) you can access it here .

This month:

  • More Fable class models
  • A pricing war
  • 3D graphics, Blender, and pixel art
  • LLMs come for mathematics
  • So many more accidental cyberattacks
  • The vulnapocalypse comes for Datasette
  • What I'm using right now
  • My software releases this month
  • 2026 in LLMs (so far)

Here's a copy of the August newsletter as a preview of what you'll get. Pay $10/month to stay a month ahead of the free copy!

Why don’t more developers "use the platform"?

Lobsters
nolanlawson.com
2026-10-03 17:35:28
Comments...
Original Article

For years, advocates for web standards, performance, and accessibility have implored web developers to “use the platform” . I’ve often been one of those advocates.

The argument is simple: why build something yourself, in JavaScript, when the browser can do it for you? Whatever you build, it’s likely to have poorer performance and worse usability than something the browser could just give you out-of-the-box.

I think it’s worth taking the other side, though, if for no other reason than to understand where the “platform-skeptic” developers are coming from. If “use the platform” is so obvious, then why do so many people seem to need convincing?

The most obvious reason is historical: for the longest time, browsers were playing catch-up with the ecosystem on top of them. Libraries like jQuery filled crucial gaps while browsers implemented equivalent APIs – and even then, you might have to wait for laggards like IE6 to age out before you could actually use them. Today, most browsers are evergreen (Safari is debatable, although ~7 times per year ain’t bad), but up until the 2020s or so, web developers had to deal with a decidedly lumpy web . In that environment, rolling your own is a sensible choice.

Another reason is familiarity: when you’re used to looking for React components on npm, that’s what you tend to reach for, regardless of the problem at hand. If you search for “sticky positioning” on npm, there’s no package that says “just use CSS position: sticky , you dolt.”

And often, even with a robust standard, libraries on npm would fill a useful gap between framework ergonomics and the platform underneath it. I always found it intriguing that many React developers preferred to stick to JSX and React idioms – raw DOM APIs felt “icky” – but were perfectly happy to use lower-level libraries where raw DOM manipulations are common. For example, a virtual list library might happily use raw DOM APIs for pure performance, while exposing higher-level primitives that a novice React developer could better grasp. In a sense, the ecosystem of React components led to a natural division of labor where those with more expertise packaged up unfamiliar platform APIs in a more familiar form factor.

Some of this effect was also driven by documentation. Many npm packages have lovingly detailed README s or websites with examples, tutorials, and screenshots. Whereas until MDN became cemented as the go-to place for web documentation (with web.dev as Google’s more future-facing arm), documentation for the web platform was scattered across blogs, StackOverflow , and sites like CSS Tricks . And many of these sites would just tell you to use a well-known library like jQuery or GreenSock!

Screenshot of the Dragula website saying "drag and drop so simple it hurts" with a logo and stylized purple page versus the MDN page for the Drag and Drop API which looks much more subdued in comparison
The Dragula site versus the Drag and Drop MDN page . Arguably the former is still more compelling.

If it were just about third-party libraries versus platform APIs, though, then I don’t think it could fully explain the antipathy toward “use the platform.” Developers who are lazy (or do I repeat myself?), and who just want a ready-made solution for whatever problem they’re facing, are unlikely to care whether that solution comes from npm, the browser, or copied off of someone’s random GitHub Gist. They want to solve their problem and move on. But there’s a different source of anti-“use the platform” that I want to explore.

For a certain type of developer, building things yourself is just more fun . And often the resulting code is easier to reason about, especially if you don’t have an encyclopedic knowledge of the web platform. And once you’ve built something, there can be a kind of IKEA effect where you want to maintain and tinker with your own homemade code.

As an example, let’s imagine you’re trying to build a modal dialog. You might visually understand how these are supposed to work: content appears on the screen, but the background is still visible although partially occluded, and maybe clicking outside the dialog dismisses it. So you might grab for position:absolute and z-index to position the dialog correctly – aha, but the background still scrolls, so you have to disable overflow on the body… And then if you understand something about accessibility, you realize you need to handle Esc to dismiss, and build a focus trap , and return focus to the element that launched the dialog, and…

For many developers, what I just described sounds like a nightmare (and a good way to build something that only half-works). But for many developers, this sounds like fun! Think of how much you learn as you start building this thing. And think about how you could start putting your own spin on it by adding animations, themes, an optional “close” button… Before you know it, you’ve built a library that’s ready to go on npm. That’s way more fun than just grabbing <dialog> and calling it a day – what a downer!

And for many of us, before APIs like <dialog> existed, this was how we learned the web platform ! Many of the people who now advocate for “use the platform” were once themselves authors of polyfills, shims, and libraries. I know because I’m one myself! I spent years working on tooling for IndexedDB , WebSQL , and other browser storage APIs as part of my work on PouchDB , which eventually led to me feeling confident enough to sit in W3C standards meetings and even open issues and pull requests on the IndexedDB spec itself. Without the forcing function of a gap in the platform that needed to be filled, I don’t know if I would have found the interest or motivation to get to that level of expertise.

Of course, doing it yourself is not always an unalloyed good. Sometimes it just comes from pure ignorance. On the web platform in particular, I think one of the reasons there was such a proliferation of JavaScript solutions to problems that could be better solved by CSS, for example, is that many developers just didn’t take the time to deeply understand how CSS works.

And to be fair, CSS has historically been hard to understand! There’s a reason the site is called “CSS Tricks.” Things like the clear fix , floats , and the min-width: 0 trick are hardly intuitive. Rather than trying to understand CSS’s internal algorithm, it’s often much easier to just imagine the imperative logic you want and then express it in JavaScript. Plus, for years CSS didn’t have a straightforward way to express common patterns like line clamping , textarea resizing , scrollbar hiding , etc. So of course developers built it themselves using the tools they already understood.

I don’t even think this phenomenon of “avoiding the platform” is limited to the web. It can apply to any developer working on top of a platform they don’t fully understand. For example, at my work, we use ClickHouse for storing various kinds of analytics data. At one point, my coworker and I disagreed about how to store large JSON data in a column: he built a system for compressing it before storage, whereas I put the data in a separate key-value store and only inserted the key into ClickHouse. It turned out we were both wrong! ClickHouse automatically compresses data , and as a columnar data store you actually get better compression across rows if you just let ClickHouse handle it. And the separate key-value store was just a poor man’s version of what a columnar SELECT already does.

I only realized these things after actually taking the time to thoroughly read the ClickHouse docs and then write a benchmark to prove my hypothesis. In the end I was shocked that we had built something that was slower and clunkier than what the platform itself could give us out-of-the-box. The parallels with JavaScript and the web platform were hard to ignore.

I’m sure that if you’re a developer on iOS or Android, or someone building on top of a game engine, or really any kind of developer building on top of any platform layer, you probably have similar stories. There’s a reason that the stereotype of the grizzled senior engineer is someone who can take a junior’s baroque tangled mess of code and replace it with a single line. The more you learn, the more you’re able to wield your knowledge of how the entire system works end-to-end to create the smallest possible contribution to it (and thus reduce your maintenance burden in the long run).

I’ve been trying really hard not to talk about AI this entire post (because I’ve done it to death over the past year), but of course I can’t help but wonder how AI coding will impact this phenomenon. I actually have both an optimistic and a pessimistic take:

  • Optimistic: because LLMs have an encyclopedic knowledge of whatever platform you’re working with, they can choose exactly the right platform API to deliver the experience the prompter asks for in vague English. And since this solution is likely faster and more correct than userland code, the agent will prefer it after rigorous testing and benchmarking. Furthermore, the “IKEA effect” goes away when developers are not actually writing the code themselves.
  • Pessimistic: because LLMs seem to love duplicating code – for example, ignoring helper functions that already exist in favor of writing their own for the umpteenth time – the amount of custom, non-platform-idiomatic code will skyrocket. Developers won’t instruct their agents to test enough or to try enough alternatives, and will just commit the agent’s first draft. And because it’s always possible to add more epicycles , the agents will continue to iterate on over-engineered solutions that never should have existed in the first place.

In my own use of AI coding, I’ve seen both phenomena happen. I’d like to think that as models and coding harnesses get better we’ll start to veer more towards the optimistic outcome, but I can’t say for sure.

In any case, these are my longwinded and somewhat conflicting thoughts on “use the platform.” As a mantra I love it, because it succinctly captures a feeling I have when I’m looking at some overwrought pile of spaghetti code and thinking how much better and more elegant it would be if the author just understood the layers beneath them a bit better. At the same time, I’ve been that author, and I’ve felt the joy of building such beautiful, messy code (beautiful to me, anyway), so I think it’s worth understanding where such developers are coming from. For that reason, I’m sure we’ll be hearing “use the platform” for as long as there are platforms.

You can comment on the fediverse or Lobsters .

Customization: Optimizing Compiler Technology for SELF, a Dynamically-Typed Object-Oriented Programming Language (1989)

Lobsters
dl.acm.org
2026-10-03 16:59:46
Dynamically-typed object-oriented languages please programmers, but their lack of static type information penalizes performance. Our new implementation techniques extract static type information from declaration-free programs. Our system compiles several copies of a given procedure, each customized...

Reasons I didn't become an EMT, ranked

Hacker News
ben.stolovitz.com
2026-10-03 16:49:15
Comments...
Original Article

I’m an EMT (Emergency Medical Technician).

My main job is software engineering, but I’ve always been interested in medicine. A friend turned me onto the idea in college, and it rattled around my head for years. I delayed for a long time — until finally, in 2024, I took a month off and got my EMT.

A person looks at an enticing Star of Life A person looks at an enticing Star of Life

Here are actual reasons I delayed getting my license , sorted loosely by ridiculousness:

Not ridiculous

  1. It will take a lot of my spare time.

    Yes; my accelerated class ( NOLS Wilderness EMT ) took 28 consecutive days. I was only able to take the class because Microsoft switched to unlimited vacation, my boss was great, and I had the money to fly to Wyoming for the course (the usual alternative is a semester of night school!). And to be a good EMT, you have to actually take calls on an ambulance, hopefully more than the typical “1 weekend a month” minimum.

  2. I won’t like the culture.

    Yes. This was a huge barrier: I had an awful experience with some first responders a few years after college — bad enough that I gave up on getting my EMT for a while.

    There are a lot of burned-out people in EMS . Many services — including the one I saw — are incredibly fratty, to be charitable. Lord help you if you are a woman . You will hear many excuses for these issues, but, yes, culture is a huge concern.

Eating lunch with the class in the Wyoming hills

NOLS WEMT: sometimes you can have a good culture

Kinda ridiculous

  1. I won’t fit in with the people.

    This is similar to culture, but a bit more silly. Yes, first responders tend to be a tight-knit, hereditary community. NYC firefighters, for example, often come from big firefighter families. But first responders (especially volunteers!) also come from all backgrounds — the person who onboarded me into my current position is a PM at another tech company.

  2. I won’t fit in with the students.

    So maybe EMTs were ok — but what about the EMT students I would learn with? Wouldn’t they all be 18-year-old adventure junkies or jocks?

    Yes, most were, and they were awesome . My course was full of enthusiastic 18-year-old pre-meds who were excited to learn, and several people my age, who wanted to learn medicine for the same reasons I did.

    If you’re nervous, I think it helped that I chose a Wilderness EMT program in July . It filters to that crowd.

  3. I want to focus on my day job instead .

    I would often consider studying EMS (Emergency Medical Services) when my day job wasn’t going well. And so it would always be a bit of a sign that I should refocus at work — onto more meaningful things.

    While that’s a good signal, it’s alright to put energy into things beyond work. And, once I got my certificate, balancing work & EMS has been straightforward.

  4. I don’t want to work on an ambulance.

    Yup — I never dreamed of riding on an ambulance. Some people really do. It’s their vocation. It’s not mine.

    But no one forces you to step onto an ambulance once you have your EMT. You can just… know how to do CPR really well. That’s a really cool skill to have.

  5. I don’t understand the scope of practice.

    If you’re not an EMT, it’s not super clear what an EMT can do. On the one hand, it’s not just “driving an ambulance,” but on the other hand, you need a paramedic for most serious calls. In some sense, it kinda is “knowing how to do CPR really well.” I trained with the same AEDs any layperson uses (EMTs aren’t allowed to use manual defibrillators), but I trained many, many more times than a layperson (plus I get to use oxygen).

    But, as before, the training is still pretty cool. I loved every bit of first aid training I got; EMT school was just as fun.

    The interior of a BLS ambulance

    I know how to use all this stuff !


  6. It’s hard to find a volunteer job.

    Kinda. Big cities tend to have all-paid staff; volunteer positions are a bit farther out. But not that much farther out: even ignoring religious services like Hatzalah in NYC, Hoboken across the water has volunteer EMS. It’s, like, one stop on the PATH. (Volunteering in Seattle or near San Francisco is more difficult).

    I didn’t want to work a part-time position — then it would just be another job, and I like my current job. But volunteer jobs are straightforward enough to find, especially if you have a car.

  7. The certificate is just gonna expire if I don’t use it.

    The culture barrier was strong enough that, for a time, I was convinced never to join an ambulance company. If you’re not on an ambulance, it’s very hard to fulfill the NREMT’s continuing education requirements , so your certificate will quickly expire.

    But, again, it’s OK to learn stuff without planning to use it anywhere. And isn’t continuous learning part of the point?

  8. I wanna do search and rescue instead.

    Again, trying to avoid ambulances, I found search and rescue. Search and rescue (SAR) is badass. A snippet from one SAR orientation packet :

    If you have registered and do not show up at training that weekend, we will make every effort to locate you, and we are incredible at finding people.

    But in my heart I always wanted to treat people; and when I met SAR folks, it was clear that I wouldn’t be an effective EMT without experience on a truck.


  9. I should do easier, cheaper prep first.

    I did 😃. I read books and Reddit posts. I took CPR and Stop the Bleed classes. I took a weekend Wilderness First Aid class with NOLS (strongly recommend!).

    But I still craved more. At a certain point — as I mention in How this site works —just do the real thing .

  10. It’s expensive, and getting new credentials is an expensive hobby.

    Indeed! My course cost about $5,000, not including airfare, renting a car, & taking vacation for a month. You can do this for much cheaper (a typical course is around $2,000 and does not require 28 days vacation), but $2k is still quite expensive for some fun.

    But if you’ve done the cheaper stuff, if you’re still interested after so many years, if you’re excited about using your money for an experience like this… how could it possibly be a waste?

  11. I have no real use for it.

    So I signed up with no real use for it! After years of delay, I started an EMT course with no plans to use it after. I was simply excited to learn.

    Working on an ambulance after was basically a happy accident, where I met a crew I happened to like. But it’s always fine to learn something for its own sake.

A small team adjust an improvised litter with a fake patient

NOLS WEMT: I don’t expect to improvise a litter again (photo: NOLS)

Ridiculous

  1. I don’t really wanna learn this anyway.

    For someone who “didn’t want to do this,” I sure spent a lot of time on Reddit reading about what the job would entail. I sure signed up for a lot of First Aid classes. I sure told a lot of people I was interested.

    I will always be a software engineer: it’s just who I am. And I learned during training that full-time EMS was not for me, after watching a woman negotiate with insurance to see if they’d cover the helicopter that would save her husband’s life.

    But not interested in learning this at all ? Not being able to help people; to be a source of calm when things are actually on fire? Nonsense.

  2. I won’t like the work.

    EMTs do menial work. They load & unload stretchers. They fold blankets, restock, and clean ambulances. EMTs are basically the least-trained medical personnel in any emergency (ignoring the odd EMR ). They cannot diagnose; they follow fixed protocols, and they defer to better-trained personnel for anything more severe than a broken arm.

    Descriptions of EMT work inadvertently make the job seem directly heroic or glamorous: waking up at 3AM to rush a heart attack patient to the ER. It is not. The bulk of the job is waiting around, doing paperwork, cleaning, and backing up narrow driveways. It is very frequently boring & aimless.

    But that tedium is because you are waiting & ready to help. So it is indirectly glamorous. As an EMT, I get to put on a uniform, and, while we’re at Dunkin Donuts for coffee, wave at the kids staring at the shiny ambulance.

    Who wouldn’t like that work?

  3. The class won’t make me a good EMT.

    I worried that an EMT certificate would be basically useless: that it would not teach me everything I expected, or that I’d still need more training on the job.

    Both things are true. I learned a lot, but upon stepping onto an ambulance for the first time, it was clear how much more there was to know. Every other EMT knows so much more than I do.

    But that’s the most magic part of being an EMT. It opens up an entire universe of new things to learn. It is the minimum qualification you need to enter this world: to work with firefighters and paramedics and doctors and social workers and search & rescue and disaster response personnel — and to see what you want to learn next .

  4. EMTs only know ambulance stuff.

    A similar complaint I see on Reddit is that an EMT course only really teaches how to run an ambulance. That it is not useful for anyone interested in self-reliance or general survival skills because it requires supplies & logistics only available in a full EMS system.

    Not only is that demonstrably false (my Wilderness EMT certificate is specifically for scenarios where supplies are limited and evacuation is difficult), it’s also unhelpful. What alternatives are there? Become a doctor? Join the National Guard?

    Again, EMT is the fastest, easiest way to become a first responder. A good EMT course teaches you to be calm & useful in terrifying situations. A good EMT job drills that in. When that’s not enough, you’ll be ready for deeper qualifications.

    That was some of the most unhelpful advice I found before getting my EMT. Ignore it.


    NOLS WEMT: none of my training was on an ambulance (photo: NOLS)

  5. I can learn a lot about survival on my own.

    That’s the thing — you cannot learn this alone. EMS services have spent decades learning how to help people effectively, and most of that knowledge is not yet available on YouTube.

    This makes the worst services fratty, full of insider knowledge that they refuse to share. But the best services become places to learn, to become the next part of a chain that goes back to the first public Freedom House ambulances in the 60’s.

    There is quite a lot of medicine you can do on your own in a bunker. But there is quite a lot more that you can’t. You do that medicine the way we always have — as part of a team, working with hundreds of other people.

  6. I will want to become a paramedic and quit my job.

    This is my favorite excuse: that I would like the work so much it would lead me to quit my tech job, pursue 2 more years of school, and become a full-time paramedic. In other words…

  7. I might enjoy it.

    It got to the point where I avoided EMT school because I might enjoy the work . That it would hoover up my free time and energy, pull me from my other hobbies, or even my career. I worried I would toss everything for the thrill of riding in a truck with sirens on.

    That happens to some folks. It has not happened to me. I like the EMS work I do, but I also adore my tech job. This will not surprise anyone who knows me, and it shouldn’t have surprised me. And, actually, how awful would it be to find something so exciting that it drags me into a new career? It should have been obvious to me that this, of all things, was an excuse .

    An excuse not to do something scary and new.

NOLS WEMT: scary, new, and fun… plus (fake!) blood (photo: NOLS)

Because it is scary

This is the silliest, truest, biggest reason I delayed so long:

  1. It is scary.

    It is scary to practice medicine. It is scary to learn new things. It is scary to be responsible for being calm on the worst day of someone’s life; and for me it was, oddly, scarier to join a class full of strangers — strangers who would be much younger & on very different lifepaths.

    I come from the world of software engineers and office buildings. Of nine-to-fives and salad bowls. EMS is a very different world. Stepping into that world, without being sure it was where I wanted to be, was terrifying.

    It took a lot of time to reckon with that.

Because, ultimately, I am good at inventing reasons not to do things.

I love researching stuff before I do it. It’s fun to compare options, to see what other people think, to try to buy the best digital piano or visit the coolest museums in Rome. But I can get stuck. I perseverate — I revisit simple decisions over and over.

There are many legitimate reasons not to become an EMT, just like there are many legitimate reasons not to go on that vacation, or watch that scary movie, or go to that board game meetup. You do not have to do any of those things — you do not have to want to do any of them. I have no desire to see horror films (to my wife’s chagrin). But I kept wanting to be an EMT, despite all my reasons not to.

Perhaps you feel the same.

If you are delaying doing something scary and new, too, think: are you inventing reasons not to do it? Maybe you should do it anyway.

You might enjoy it.

Porchlights illuminate a few cabins against the fading sunset

NOLS WEMT: returning to the dorms at twilight

Thanks to NOLS and the NOLS WEMT students for making the course so awesome. Thanks to NHVA for welcoming me. Thanks to Javier Valencia, Jaime Romo, & Dan Walker for the NOLS photos. Used with permission. Thanks to Atherai Maran for editing.

System-level ad-blocking in Android

Lobsters
kevinboone.me
2026-10-03 16:43:41
Comments...
Original Article

The Internet has become an advertising platform. Users of mobile devices are acutely aware of this, because mobile operating environments and apps force ads on their users, in addition to those that pollute the Web. It’s getting harder to avoid the ads on Android handsets, as Android is increasingly locked down and difficult to customize. It’s certainly not in Google’s interest to make it easy to block ads.

It’s still possible, though, to suppress most advertising, even in 2026.

This article is specifically about blocking ads at the system level, that is, in the Android platform itself, rather than in any particular app. Some Android apps, particularly web browsers, have their own methods of blocking ads, or can have such features added using third-party plug-ins. Oddly, though, the DuckDuckGo browser, whilst blocking the tracking that accompanies targeted advertising, doesn’t specifically block ads – you need something more.

If you block ads at the system level, all apps – and the system itself – are protected. Moreover, methods for blocking ads can usually be extended to protect against trackers and other kinds of malware.

Basic principles

So far as I know, all system-level ad-blocking methods work by overriding DNS behaviour. DNS (domain name service) is the technology that maps hostnames to numeric IP addresses; a DNS lookup is the first step in almost all network operations.

Blocking ads using DNS amounts to finding the hostnames of know advertising services, and mapping them to bogus IP numbers. This can be done in the Android handset itself, or in some external DNS server.

Returning a bogus IP number doesn’t prevent an app or browser trying to show advertising, or allow it to make better use of the screen space the ads take up. Apps vary in their responses to a failed attempt to display advertising – more on this later.

Blocking ads using DNS

There are essentially four methods to block ads using DNS changes.

  1. Use a commercial DNS service with ad-blocking features
  2. Use a commercial virtual private network (VPN) with ad-blocking features, like NordVPN
  3. Install an ad-blocking app that doesn’t require root access.
  4. Use an ad-blocking method that requires root access, with or without a supporting app

In general, when you make a connection to the Internet on an Android device, the DNS name resolution (mapping a hostname to a numerical Internet addresses) is performed by a server hosted by the handset’s mobile carrier or Internet service provider.

Android allows you to override this behaviour, and select a custom DNS server that blocks ads. Well, it won’t actually “block” anything – not really; the DNS server just returns a bogus IP number when your handset looks up the IP number that corresponds to the hostname of a known advertiser.

Depending on the DNS service you use you may, or may not, get some control over exactly what it blocks. Some users choose also to block access to pornography or on-line gambling, for example, for child protection.

How well these services work depends, of course, on how well the service’s operators maintain their lists of known advertising hosts. Advertising services come and go, and even long-standing ones change their servers from time to time.

There are many commercial DNS services that support ad-blocking, and I don’t endorse any in particular. Prices vary, and most services have a free trial. If you care about privacy, you’ll want to use a service that keeps no logs, and you might want to use one that supports encrypted DNS. A rogue DNS operator might get access to sensitive data in all kinds of ways. It’s possible, for example, for a DNS operator to direct requests for your on-line banking service to its own proxy, and slurp up your credentials when you log in. You should certainly be very careful and, while there are legitimate free DNS services, I’m rather suspicious of any service of this kind whose funding model is unclear.

For many people this will be the simplest solution. If you subscribe to a VPN service for general privacy, blocking ads will often come at no extra cost. Ad-blocking VPNs typically redirect DNS lookups to their own servers, and return dummy IP numbers for known advertisers, just as a commercial DNS service does. VPNs therefore have all the same advantages and disadvantages as an ad-blocking DNS service. A big VPN operator will usually provide an Android app that configures the handset to use its servers, so set-up is pretty trivial for most people.

As with commercial DNS services, a rogue VPN operator can be very dangerous, and it’s important to choose a provider carefully.

So far as I know, all these apps exploit a loophole (of sorts) in Android’s platform lock-down.

Like all Linux-based operating systems, Android provides a way for a device to override the DNS mappings of the Internet service provider it is using. There will be, somewhere on the handset, a file that contains a list of preferential name-to-IP mappings. On a non-rooted Android device, the user won’t have permissions to change this file, so a simple method of blocking ads – maintaining a list of bogus DNS mappings for known advertisers in a file – is unavailable.

However, in a non-rooted device, you can install a VPN service. This has to be possible, if Android is to support VPNs at all. A non-rooted ad-blocking app will typically run a DNS server and a local (within the app) VPN host. The app will configure the handset to use its own VPN as the system VPN provider, which will make its own DNS server the system DNS. With control over DNS, the app can provide its own handling for name lookups, including those of know advertising servers.

Google, being an advertising company, isn’t very keen on this kind of thing. It can’t easily prevent the use of “local” VPNs without crippling VPN functionality completely but Google can, and does, make it difficult for non-technical users to get the appropriate apps. At present you can get apps like AdAway from alternative app stores like F-Droid .

Google is set on making this difficult, too, and before long you’ll have to get the app’s APK file from (hopefully) a reputable source, and jump through whatever hoops Google puts in the way of install software it doesn’t like. At the time of writing, Google isn’t making it impossible to install software from outside its Play Store, but it’s getting more and more fiddly.

Most likely though, Google’s obstructions won’t ever be as difficult to surmount as rooting your Android handset and applying a definitive ad-blocking solution. Using an app like AdAway will probably continue to offer advantages over a commercial DNS or VPN service, even if Google makes it hard to install.

The most obvious advantage is that these apps are usually free to use. Despite this, a community of volunteer maintainers ensures that the apps’ lists of known advertisers are as thorough as any provided by a commercial service. A disadvantage, though, is that all network traffic from all apps has to be routed through a single app on the handset. Not only does this create a network bottleneck, a rogue ad-blocker app is exactly as dangerous as a rogue VPN service. Fortunately, because these ad-blocking apps are usually open-source, there are limited opportunities for bad actors. I’d strongly advise against using one that isn’t open-source, even if you don’t plan on looking at the source code yourself.

If your handset is rooted, then ad-blocking is simple – in theory, at least. Android’s list of preferential name-to-IP mappings is in the file /system/etc/hosts , so all you have to do is edit that file, to direct advertisers’ hostname to a bogus IP number. Usually the bogus IP local address of the handset itself, 127.0.0.1 . So you’ll have a hosts file full of entries like this:

127.0.0.1 08.185.87.0.liveadvert.com
127.0.0.1 08.185.87.00.liveadvert.com
...

The reason for using the handset’s local IP number is that it must correspond to a system that the handset can actually reach. Otherwise, network access will be badly delayed as apps try to contact non-reachable advertising hosts. Of course, this means that every request for an advertising service will be directed back to the handset which, presumably, will reply with a error response to the app that makes the request. Since the request will fail immediately, this loopback network routing doesn’t create an appreciable load on the handset.

This is all theoretically straightforward but, in practice, there are two major problems.

First, we need a source of hostnames to block. Lists of these are widely available on websites, but the problem of maintenance is always present. Some of these lists are truly vast and, without doubt, reference hosts that no longer operate. If the hostname list is this long, then all network access will be slowed, as the handset has to parse the hosts file for each DNS lookup.

Probably the best source of ad-blocking hosts files is the code of open-source ad-blocking apps like AdAway. This app can, in fact, provide its list automatically when installed on a rooted handset, which simplifies the set-up.

The second problem is that you can’t just hack on the hosts file, even as root – on all modern Android devices it’s on a read-only filesystem. So you’ll need some sofware that can manipulate the contents of the /system directory during the boot process, while it’s still writeable.

If you’ve rooted your handset, you almost certainly have a way to do this already. If you’re using Magisk, for example, you can use a module to supply a new hosts file. Magisk modules live in directories under /data/adb/modules , and any files in the module’s own system directory will overwrite the main /system at boot time. So you can create a Magisk module that provides its own /system/etc/hosts .

Happily, Magisk users don’t have to do this manually, as the Magisk developers have anticipated this usage. If you enable the “systemless hosts” option in the Magisk app, it will create the necessary module with all the necessary metadata in place. Thereafter, to create a custom hosts file you just hack (as root ) on /data/adb/modules/root/system/etc/hosts , and then reboot. Alternatively, you can use an app to make this change.

In fact, because this method of hacking on the hosts file is so prevalent, ad-blocking apps like AdAway have built-in support for it; but, of course, this support is only for rooted devices. If your handset is already rooted, using a root-aware ad-blocking app is probably the most effective way to manage ads. It’s much faster than a non-root app that installs a local mock VPN, and doesn’t raise any of the same security concerns. That’s not to say there are no concerns, and you should ideally check the hosts file such an app installs, to ensure there are no mappings to anything except 127.0.0.1 .

It’s important to understand that no method of blocking ads is completely reliable, or has no side-effects.

Most notably, some Android apps simply won’t work, or won’t work properly, without their advertising. If you must use such apps, you’ll need an ad-blocking method that allows you to customise the list of blocked sites. It won’t be remotely obvious, just from the behaviour of the app itself, why it isn’t working, or how to fix it. On a rooted handset you can track the network behaviour of a specific app in detail and, in theory, work out what it needs that you’re blocking. In practice, it’s easier to refer to the on-line discussions of such apps because, most likely, somebody else will already have done the work.

However, some apps go so far as to have all their ads built in; no ad blocker will stop such an app showing advertising, as there’s no network operation to block.

It’s also important to understand that Android caches DNS lookups in various places. Although some methods of ad-blocking allow the list of blocked hostnames to be configured on the fly, you might still have to reboot the handset to flush its caches.

Although it should be obvious, bear in mind that DNS-based ad-blocking only works where an app uses DNS. Since DNS-based ad-blocking is so commonplace, some app developers are implementing direct access to advertising servers using IP numbers, which renders all DNS-based approaches ineffective. This, fortunately, is still relatively rare.

The ability to block ads at the system level is one of the few compelling reasons to root your Android handset in 2026. Blocking ads this way uses few, if any, additional system resources, and its effect extends to all apps.

If you can’t root your device, or prefer not too, there are other ways to block ads, but none is as effective, as configurable, or as cheap.

Finally, if you mostly struggle with ads on websites, rather than in apps, the simplest approach might be to install a web browser with ad-blocking support, and avoid system-level ad-blocking entirely.


Have you posted something in response to this page?
Feel free to send a webmention to notify me, giving the URL of the blog or page that refers to this one.

'Neanderthals Among Us' review

Hacker News
www.historytoday.com
2026-10-03 16:41:15
Comments...
Original Article

Establishing a secure connection...

Request ID: 0b0adb1791a485f573b19e83fb4cf2ce

Rust, In Sickness & In Health

Lobsters
www.youtube.com
2026-10-03 16:37:40
Comments...

Surely you have ultra-wideband radios on your bins too?

Hacker News
sjg.io
2026-10-03 16:27:44
Comments...
Original Article

Home Assistant already knows which bins (trash cans, for the international readers) are due for collection. It can tell me it’s bin night, put the right colours on a dashboard, and send me a reminder. What it can’t tell me is whether I’ve actually done anything about it.

I’m pretty sure most folk would just set a recurring reminder on their phone. As my wife will no doubt tell you, I am not most people. Some would just remember to put the bins out, like a normal person. I’m well aware that I’ve ridiculously over-engineered this. I’m also so far from normal that six radio-equipped bins seemed like an enjoyable way to spend my evenings ;)

I had an idea a number of years ago to stick cheap Bluetooth tags on my wheelie bins and use an outdoor Bluetooth proxy to work out whether they were nearby. It was frustrating. The batteries drained quickly, the proximity readings wandered about, and the proxy was unreliable in that particular installation. I wanted to know whether a bin had moved. I mostly acquired another thing to fiddle with.

AirTags were what first got me intrigued by ultra-wideband radio, or UWB. Reading more about how it worked brought the bin idea back. Over the last few weeks it’s become BinRange : a fixed radio anchor, six little battery tags, and a Home Assistant setup that combines where the bins have been seen with when they’re due. There’s a printed enclosure, custom firmware, and wireless updates as well. All of which is a fairly substantial answer to “have you put the bins out?” :D

I leaned heavily on Codex throughout the build. Astra became available part-way through, and I was really curious to see how it would do. What I love about working with AI is how quickly I can take an idea, react to the result, give it some feedback, and have something else to try. There’s plenty to write about that in other articles. For now, back to the bins.

The row of household bins beside a wooden wall in the BinRange installation.

There you go, some bins. Beautiful, right?

Six bins, several schedules

We have one general waste bin, one food compost bin, one garden waste bin, two recycling bins, and a glass bin. Six bins. At one house. They’re not even all on the same collection schedule. I’m grateful that the Waste Collection Schedule integration for Home Assistant scrapes the council’s schedules and keeps track of what’s due, because keeping that straight is a job in itself.

With the cheap Bluetooth tags, the way to estimate distance was RSSI - received signal strength indication. You measure how strong the received signal is and use that to guess how far away the tag might be. The trouble is that signal strength depends on far more than distance. Walls, parked cars, reflections, antenna orientation, and differences between the radios and their antennas all affect it. Even battery voltage can affect the transmit power on some hardware , depending on how the device is designed. A weaker signal could mean the bin has moved further away, or it could mean something got in the way. It’s a very imprecise way to estimate distance, which explained why my original setup was so frustrating.

UWB was interesting because it measures the time taken by radio signals to travel between devices . That gives a distance measurement without having to guess it from signal strength. It still has to get a signal through the surroundings, of course. Choosing a different radio doesn’t make parked cars disappear.

I also spent a while detoured into an academic question about how many anchors I’d need to add more dimensions to that awareness. What if I wanted to know the direction as well as the distance from a point? Or the bin’s precise location - how many fixed reference points would I need to triangulate it on a 2D plane? Totally not worth it for this job, but tangential thoughts be tangential thoughts.

For the bins, I only needed to know whether each one had moved far enough from its usual storage area to count as Out. One anchor was enough to start testing that. I used a ten-metre boundary: inside is Home, beyond it is Out, when there’s a fresh reading. It’s specific to my installation, but it meant I could start with one powered board rather than turn the garden into a positioning test range.

Two boards and a walk outside

The first experiment used two Makerfabs ESP32-WROVER/DW3000 boards , both powered over USB. Before doing anything clever with Home Assistant, I wanted to see them measure a distance. A local web view let me separate them and watch what happened without needing a USB cable stretched between the two.

A red Makerfabs ESP32-WROVER development board with USB connected and its UWB module and antenna visible.

One of the development boards. This is the DW3000 model used in the experiment. Hardware reference .

The antenna orientation made a surprisingly large difference. At about ten metres, changing the boards from flat to upright took the observed success rate from 37% to 100% in that test. A slower radio setting with a longer preamble also worked where the initial fast setting struggled. These were useful discoveries to make before designing anything around the first result.

The first time I calibrated it with a tape measure and watched the reported distance differ from my measurement by less than two centimetres, my mind was utterly blown. Two little boards were exchanging radio signals and agreeing with a tape measure to that degree.

In the later recorded test, a measured gap of 10.1 metres produced an average reading of about 10.09 metres, with a standard deviation of three centimetres. That was encouraging for a project which only needed to recognise a bin leaving a storage area. The under-two-centimetre result was what I saw in that first calibration, rather than a promise of that accuracy for every tag, orientation, or outdoor position.

The outdoor walk was more revealing. Readings were reliable to roughly thirty metres, then intermittent further out. The furthest recorded reading was 37.28 metres, with gaps. A car could block the path completely. Holding a board upright in the open and mounting a little tag beneath a bin rim were clearly going to be different tests.

The outdoor UWB walk-test graph, showing measured distance increasing and decreasing with signal and range-spread diagnostics below.

The two-board walk test. It established a useful starting point; it doesn’t establish coverage for the installed bins. Recorded findings .

The anchor publishes its readings over MQTT, and Home Assistant discovers a separate device for each tag. I considered an ESPHome component, but it wasn’t adding much to the arrangement I wanted. There’s no extra BinRange server or cloud service, and the radios can keep ranging while Home Assistant or the MQTT connection is unavailable.

Something small enough to put on a bin

The development boards were useful for proving the link. I didn’t want one hanging off every bin, with a battery and an improvised box attached. The bin-side hardware needed to be a small, self-contained puck.

I explored custom tags and commercial ones, with the important question being whether I could run my own firmware and use them with my own anchor. A small UWB tag isn’t automatically compatible with another UWB product. The packaged tags I eventually bought were KKM K4Ws, with an nRF52833 processor, a DW3110 radio and a LIS3DH accelerometer. They take removable CR2477 coin cells (which, btw, are some chonky boy coin cells! I’ve not seen them before, thems fat).

The two Makerfabs boards came to US$123.64 including shipping. Ten sample tags at US$25 each, a programming jig and shipping came to another US$330. Those are what the experiment cost at the time, rather than a current shopping list or the cost of a finished six-bin kit. I suspect I should avoid calculating a payback period.

The supplier said the tags could run my firmware and offered a jig to reach the programming connections. I was far more interested in the hardware they were supplying and knowing that programming access wasn’t locked out in some way. None of their existing software features mattered to me. Claims about battery life were meaningless for my use once I was going to be the one deciding when the CPU slept and what woke it up.

I’m pretty familiar with nRF52 devices. I’ve used them for years, going back to my time at rlab. They make brilliant little low-power microprocessors , happy to spend most of their lives asleep. If you’re clever with your scheduling and interrupts, you can drag these things out for years on a small battery in the right application. That was much more interesting to me than whatever the supplied firmware happened to do.

I used a spare Nesso device as a programming probe for the pogo-pin jig. I started with a development tag, then commissioned the packaged ones one at a time: programme it, establish its identity, pair it, label it, assign it to a bin, and put it back together. There are several identities involved, including the printed label, the chip identity, and the addresses used by the radios.

Side note: the cases the tags were supplied in had seemingly unique serial numbers printed on them, with QR codes containing those serials. Presumably they serve some purpose in the supplier’s stock firmware and software. I couldn’t find any association with the embedded devices’ MAC addresses or similar identifiers, so they were essentially useless to me. Which is a shame really - if they’d been the radio MAC addresses, for example, that would have been useful!

A KKM K4W tag fixed beneath the rim of a green wheelie bin, with its identifying label obscured.

A KKM K4W tag attached to the food caddy beside its carrying handle, with its identifying label obscured.

Close-up of the bin tags stuck to the bins. You can see how chonky they are, those batteries are big! They fit nicely under the rim though, which will keep rain mostly off.

The manufacturer had connected the accelerometer’s interrupt output to a GPIO pin on the nRF, which is incredibly useful. It means I can let the nRF sleep for very long periods, send no radio traffic during that sleep, and wake it up when a physical event happens. Once things settle down, it returns to occasional check-ins, with the UWB radio sleeping between attempts. The firmware default is ten minutes when stationary; I’ve since changed the installed fleet to thirty minutes using controls in Home Assistant, while keeping the five-second moving reports.

That should avoid spending the battery repeatedly reporting that a bin is still sitting where it was. How much life it actually buys is something I still need to measure. I’m collecting voltage and activity information, but a short, fairly flat voltage trace doesn’t give me a defensible estimate in months or years. I’m definitely not interested in using the majority of the battery’s life reporting its own voltage. That seems low value ;)

Updating six bins without taking them apart

Once a tag is mounted, opening it up and putting it back on the programming jig every time I change the software becomes rather unappealing. The anchor can now deliver signed firmware updates to paired tags over Bluetooth. The previous application is retained so a bad update can roll back, and the system only reports success once the tag is running the exact expected image and has confirmed its own health.

Testing and iterating on this sort of thing is another fantastic use of AI. Codex helped me generate and simulate a whole range of test conditions and failure modes, then work through them to make deploying firmware to one of these things over the air as safe as possible. I tested corrupt images, interrupted transfers, selected power-loss points, and watchdog recovery. I also tested what happened when the anchor wasn’t there: the tag needs to confirm that it’s healthy without a radio reply, rather than abandon working firmware because it can’t reach the anchor.

Going from one tag to several found a different sort of problem. A record in the Bluetooth library had been sized for the number of simultaneous connections rather than the number of paired devices. Adding a further tag could fail even though its keys had been stored. Fixing that meant a new physical tag could pair without clearing the existing ones. Six bins gave the software a better interrogation than one agreeable demo tag.

I’ve since updated the installed anchor and all six tags wirelessly without re-pairing them or reaching for the jig. I’m pleased about that. Some earlier connection attempts needed retries, though, so I still want more evidence before calling the whole process reliably unattended. The programming probe remains the rescue route.

The box took several goes

The anchor needed a case, which became its own little project. It started as an OpenSCAD design with screws. The first fit check was promising, but small nuts and bolts, PLA tolerances, and a board that could still rattle made the assembly more annoying than I wanted.

I wanted a case made entirely from printed PLA, with no screws: tapered locating pins for the board, printed springs to hold it down, and a lid that stayed on. The pins and springs worked. The lid could move around all over the place.

An internal lip improved the fit. Deeper catches improved the retention, but I could still pull the lid off too easily. Making those catches as deep as the wall seemed like the obvious next attempt, until both roots broke during insertion. That version was excellent at becoming two pieces of broken plastic.

The better change was to overlap the base and lid walls and use several shallow catches. A couple of tweaks to the locating pins and vents later, I had a case that held the board properly and stayed shut. I was very pleased with it, which probably says something about how many lids I’d pulled off by then.

Render of the assembled screwless anchor case, with a blue base and pale vented lid.

Render of the open anchor case showing locating cones in the base and board supports inside the lid.

The screwless case, rendered from its print meshes. The colours and finish are illustrative. You can print it from MakerWorld ; the design files and history of the less successful attempts are on GitHub.

That quick loop with Codex was particularly enjoyable for the case: I could refine the design, inspect it, and print it. I still had to pick up the result and find out whether the lid fell off. A geometry check can establish that two shapes should fit together; it doesn’t tell me whether a catch will survive being pushed into place.

The anchor is now wall mounted, with the printed case inside a larger outdoor electrical box alongside its power components. That involved another Amazon order. The relevant bits were:

  • A clear-cover outdoor electrical box, listed as IP67 and 8.7 × 6.7 × 4.3 inches (£27.99).
  • A DIN-mounted 12V DC power supply (£11.09).
  • A DC-to-USB-C PD module to turn the 12V supply into USB-C power.
  • A Type A RCBO, listed as 10A with 30mA residual-current protection (£14.00).
  • Three metres of 2.5mm² twin-and-earth cable (£11.99).
  • A right-angle USB-C power pigtail with a 25cm lead. I only needed one, but they came in a six-pack (£5.99).

That little USB-C PD module is a pretty cool, unusual device. It takes the 12V feed and provides USB-C Power Delivery. I can see myself ordering another.

Those are the prices on my order, rather than current quotes. A weatherproof box, power supply, USB-C power module, protection, cable, and USB leads for a reminder to put the bins out. I know.

All six bins have their tags fitted. At that point I was feeling fairly optimistic that it was all working.

The installed BinRange anchor and power components inside a clear-lidded wall enclosure.

The anchor installed on the wall. The printed board case is inside the larger enclosure. Installation photograph .

Reminders and Notifications

The Waste Collection Schedule integration deserves a lot of credit here. Scraping the council’s schedule saves me maintaining several recurring dates and remembering which combination of bins is due. BinRange adds physical information about the bins, and Home Assistant puts the two together.

The ordinary dashboard shows the next collection, compact bin states, and anything needing attention. The ranges and radio diagnostics have their own view. I like looking at the technical details, but checking tomorrow’s bins shouldn’t require interpreting a radio trace.

The reminders are grouped into one message for the relevant bins: at 6pm and 9pm the evening before collection, then a final reminder at 6am if they’re still confidently Home. At 6pm the following day there’s a reminder to bring back bins still Out. The automation remembers its attempts so a rerun doesn’t keep sending the same message.

An unavailable reading shouldn’t result in an accusation that somebody forgot a bin. Likewise, the ordinary bring-in reminder only knows the collection was scheduled and the bin is still out. It doesn’t know whether the crew emptied it.

Home Assistant notification saying General Waste, Recycling 1, and Glass should be put out for tomorrow's collection.

Home Assistant notification to put some bins out.

I also wanted a message when the bins had just been emptied. There was an obvious temptation to look at movement, time of day, and the collection schedule and call that enough. Looking back at one collection showed how little those clues could establish when tags stopped reporting for much of the day. A wake counter could tell me there had been activity, but it couldn’t tell me who handled the bin or why.

The tags already had accelerometers, so I added a more specific clue. Relative to a calibrated upright position, a sustained tilt beyond ninety degrees records a tip event. The tag keeps that event through a reception gap and reports its age when it next reaches the anchor. Home Assistant can treat a recent tip on a due bin known to be Out as an assumption that it has been emptied.

The first automation checked for pending tips every five seconds. That felt excessive for something that happens a few times a week, so I changed it to an MQTT event flow with a timer that only runs while there’s a deadline pending. Qualifying bins are grouped for sixty seconds, with saved state to suppress repeats and support recovery after a Home Assistant restart. Scheduled reminders remain the fallback.

Home Assistant notification saying General Waste, Recycling 1, and Glass have just been emptied and can be brought back in.

Home Assistant notification that the bins have been emptied and can be brought back in.

First test in prod

On 28 September I had a Garden collection I could examine properly. The bin started moving at about 7:05am and crossed the ten-metre boundary shortly afterwards. Reports then stopped arriving while it was away. When it reappeared at about 10:25am, it carried a stored tip from approximately 9:27am, and then crossed back into the Home zone and settled down.

The tag had detected and retained the tip. Departure and return were visible. But the tip was about fifty-eight minutes old by the time Home Assistant received it, beyond the automation’s thirty-minute acceptance limit. The “just emptied” message didn’t send. Other bins had continued reporting, and there wasn’t an MQTT failure at the anchor during the gap. I still don’t have an established explanation for that tag’s lack of reception at the collection position.

That was a useful correction to my optimism. The radio could produce convincing measurements, the tag could detect a tip, and the dashboard could look right, while the phone still missed the emptying notification. Checking a real collection gave me a much better list of what needed attention.

It also changed how I wanted to show a bin that had gone quiet. Originally, an Out reading became Unknown once it was stale. That was cautious, but after watching the bin leave and then lose contact at the kerb, replacing Out with Unknown stopped being very helpful.

The live setup now remembers an observed departure as presumed Out , even through lost contact or a Home Assistant restart. Reception health is shown separately. A fresh reading inside the boundary is needed to return the bin Home. If it was last seen Home and then goes silent, it becomes Unknown, because I never saw it leave. Return reminders can use presumed Out; put-out reminders still need fresh evidence that the bin is Home.

That makes the display closer to what I actually know.

Gladiator meme with the caption: “Are you not over engineered?”

Waiting for more bin days

BinRange is installed and in use. (As I type that, I realise I’ve seriously pigeonholed this otherwise general-purpose tool 🤔)

There’s still some fairly ordinary observation to do: coverage at each storage and collection position, several complete cycles with the right notifications delivered, battery behaviour over weeks, and how the mounting and enclosures cope outside. I also want updates to connect dependably without me retrying. I am not really building this for general consumption by other people, but if others are inspired and want to try something similar, or repurpose any of this, they (or their robot friends) are welcome to.

The repository has the firmware, case files, measured findings, and Home Assistant examples if you’d like a look. This is my hobby project, and I’m still finding out how well it behaves outside.

Of course, now I’ve seen what UWB can do, I’m wondering whether I could put tags on the cats’ collars and use four anchors around the house to map where they are. Could I get centimetre-level cat tracking? I haven’t tested that, and the bins have already demonstrated how much surroundings and radio coverage matter. But I’m very tempted.

Most people would probably just look for the cat.

Show HN: Local pretrained classifiers, GPU not needed

Hacker News
github.com
2026-10-03 16:19:09
Comments...
Original Article

Pretrained text classifiers you can run and retrain on CPU.

Inbox Router

Doom Battle Defend the Center

Try it

Install uv , then:

uvx --python 3.12 \
  --from "jeffy-classify @ git+https://github.com/nicobrenner/jeffy.git@v0.1.0-alpha.7" \
  jeffy-serve

Open http://localhost:8400 , pick a classifier, and paste one of these:

Classifier Try this text
banking77 I was charged twice for the same transaction
sms_spam WINNER! You have been selected for a free cruise. Reply YES to claim.
ag_news The Federal Reserve raised interest rates by 25 basis points on Wednesday

Jeffy Playground

Pretrained capabilities

13 classifiers ship with the package. Weights are logistic regression coefficients (derived model parameters, not copies of training data). Source datasets and licenses are documented in ATTRIBUTION.md .

Task What it does Classes Test Acc Test F1
sms_spam SMS spam detection 2 99.1% 98.0%
dbpedia Wikipedia article category 14 96.0% 95.9%
imdb Movie review sentiment (long text) 2 94.8% 94.8%
banking77 Banking customer intent 77 94.3% 94.3%
ag_news News topic (world/sports/business/tech) 4 90.5% 90.5%
sst2 Movie review sentiment 2 90.1% 90.1%
clinc_oos Voice assistant intent + out-of-scope 151 88.4% 92.1%
massive_intent Smart home voice commands 60 88.1% 86.4%
tweet_eval_offensive Offensive language 2 81.0% 74.8%
tweet_eval_emotion Tweet emotion 4 78.1% 74.7%
emotion Text emotion (6 emotions) 6 75.5% 67.8%
tweet_eval_sentiment Tweet sentiment (3-way) 3 66.2% 65.7%
snli Natural language inference 3 65.6% 65.2%

Test accuracy on held-out splits. Details in data/eval_results/benchmark.json .

Weaknesses: SNLI (65.6%) and tweet_eval_sentiment (66.2%) are below what task-specific models achieve. Emotion (75.5%) has limited class coverage. Probabilities are uncalibrated.

Train a custom classifier

uvx --python 3.12 \
  --from "jeffy-classify @ git+https://github.com/nicobrenner/jeffy.git@v0.1.0-alpha.7" \
  jeffy-train --example --save-dir my_models
Loaded 24 examples from reviews.csv
Training 'reviews': 24 examples, 2 classes
  Split: 19 train, 5 test
  Test accuracy: 100.0%
Saved to my_models/reviews/

--example uses a bundled 24-row product review CSV. To bring your own:

uvx --python 3.12 \
  --from "jeffy-classify @ git+https://github.com/nicobrenner/jeffy.git@v0.1.0-alpha.7" \
  jeffy-train --input your_data.csv --text-col text --label-col label \
  --task-id your_task --save-dir my_models

Supports .csv , .tsv , and .jsonl .

Serve a custom model

JEFFY_PACK_DIR=my_models uvx --python 3.12 \
  --from "jeffy-classify @ git+https://github.com/nicobrenner/jeffy.git@v0.1.0-alpha.7" \
  jeffy-serve
curl -s -X POST http://localhost:8400/v1/predict \
  -H "Content-Type: application/json" \
  -d '{"text": "The battery life is amazing", "task": "reviews"}'
# → {"label": "positive", "confidence": 0.87, ...}

Install from source

git clone https://github.com/nicobrenner/jeffy.git
cd jeffy

# With uv (recommended)
uv venv && uv pip install -e .

# Or with pip
python -m venv .venv && source .venv/bin/activate
pip install -e .

The first prediction downloads the shared encoder ( bge-large-en-v1.5 , ~1.2 GB, cached afterward).

API

Python

from jeffy.engine import Engine

engine = Engine()
engine.load()

# List available classifiers
for name, cap in engine.capabilities.items():
    print(f"{name}: {cap.description} ({cap.n_classes} classes)")

# Classify text
result = engine.predict("banking77", "I was charged twice for the same transaction")
print(result["label"])          # "transaction_charged_twice"
print(result["confidence"])     # 0.999
print(result["probabilities"])  # {"transaction_charged_twice": 0.999, ...}

SDK walkthrough

SDK walkthrough

HTTP

# Banking intent
curl -s -X POST http://localhost:8400/v1/predict \
  -H "Content-Type: application/json" \
  -d '{"text": "I was charged twice for the same transaction", "task": "banking77"}'
# → {"label": "transaction_charged_twice", "confidence": 0.999, ...}

# Spam detection
curl -s -X POST http://localhost:8400/v1/predict \
  -H "Content-Type: application/json" \
  -d '{"text": "WINNER! You have been selected for a free cruise. Reply YES to claim.", "task": "sms_spam"}'
# → {"label": "spam", "confidence": 0.91, ...}

# News topic
curl -s -X POST http://localhost:8400/v1/predict \
  -H "Content-Type: application/json" \
  -d '{"text": "The Federal Reserve raised interest rates by 25 basis points on Wednesday", "task": "ag_news"}'
# → {"label": "Business", "confidence": 0.86, ...}

List capabilities

curl -s http://localhost:8400/v1/capabilities | python3 -c "
import json, sys
for c in json.load(sys.stdin)['capabilities']:
    print(f\"{c['task_id']:25s} {c['n_classes']:3d} classes  {c['test_accuracy']:.1%}  {c['name']}\")"

Per-capability metadata

Each shipped classifier has a manifest.json with label names, source dataset, HuggingFace path, stated license, encoder identity, training/test counts, and integrity hashes.

curl -s http://localhost:8400/v1/capabilities/banking77 | python3 -m json.tool

Train from Python

from jeffy.train import train_classifier

clf = train_classifier(
    texts=["great product!", "terrible service", "fast shipping", "broken on arrival"],
    labels=["positive", "negative", "positive", "negative"],
    task_id="my_reviews",
)

result = clf.predict("the quality exceeded my expectations")
print(result["label"])  # "positive"

clf.save("my_models")

Tuning options

Parameter Default Description
C 0.01 How aggressively the model fits your data. Low (0.001) = conservative, keeps predictions closer to "I'm not sure." High (1.0) = trusts individual training examples more. If the model is great on training data but bad on new data (overfitting), lower C.
test_size 0.2 What fraction of your data to hold back for testing. With 100 examples at 0.2, it trains on 80 and tests on 20. Set to 0 to train on everything (useful when you have very little data and will test manually).
cv_folds 3 Cross-validation: splits your training data into 3 parts, trains on 2 and tests on 1, rotates three times, averages the scores. Gives a more reliable accuracy estimate than a single split. Set to 0 to skip (faster, less reliable estimate).

Start with the defaults. With <50 examples per class, expect noisy estimates.

Reproduce the evaluation

# Install with build dependencies
uv pip install -e ".[build]"
# or: pip install -e ".[build]"

# Retrain all 13 heads from source datasets (~40 min, downloads ~5 GB)
jeffy-build --out data/model_pack

# Evaluate on held-out test sets with tuned baselines
jeffy-evaluate --baselines --latency --device cpu

Evaluation notes

  • SST-2 : Evaluated on validation split (official test labels are not public).
  • SMS Spam : Random split (test_size=0.2, seed=42); no standard benchmark split.
  • SNLI : Input encoded as premise [SEP] hypothesis . Label -1 filtered.
  • CLINC-OOS : 151 classes including out-of-scope. In-scope accuracy 96.5%, OOS detection 51.7%.
  • MASSIVE : English only (config en ).

Deployment

Component Size Required for
Jeffy package (wheel) 1.5 MB Always (includes all 13 heads)
Encoder (bge-large-en-v1.5) ~1.2 GB Inference (downloaded on first use)
datasets package ~100 MB Retraining from HuggingFace only

Runtime memory: ~2 GB (encoder loaded once, shared across all heads).

Latency (CPU, single example, Linux aarch64):

Stage p50 Notes
Embedding 50–80 ms Dominates; varies with input length
Classifier <1 ms Negligible
Total 50–80 ms End-to-end

Model security

Bundled pretrained artifacts use numpy .npz format (portable, no pickle). Custom-trained models also save a joblib pickle backup. Only load custom pickle artifacts from trusted sources. Each artifact's manifest.json includes integrity hashes verified on load.

Provenance and licensing

Jeffy code is MIT-licensed. Head artifacts are derived from public datasets; redistribution permissions have not been independently verified for all sources. See ATTRIBUTION.md for per-dataset license status.

Dataset Stated license
banking77, massive_intent, sms_spam CC BY 4.0
clinc_oos CC BY 3.0
dbpedia CC BY-SA 3.0
snli CC BY-SA 4.0
ag_news, imdb Academic / non-commercial
sst2 Stanford academic license
emotion Academic
tweet_eval_* Twitter TOS / academic

The encoder ( bge-large-en-v1.5 ) is MIT-licensed.

Tested

Verified with clean-environment wheel and sdist install on Linux aarch64, Python 3.12, scikit-learn 1.9+, sentence-transformers 6.1+, numpy 2.5+. Pretrained artifacts use numpy .npz format, avoiding sklearn version coupling.

What's not included

  • No zero-shot / general classification. Each task needs a trained head. Unknown tasks return an error.
  • No LLM fallback. This release is pure embedding + classifier.
  • No automatic task routing. You must specify which classifier to use.

Roadmap

Status Milestone
Available Pretrained classifier library, SDK/API, playground, custom training from CSV/JSONL
Next Landing page, PyPI release
Planned Broader classifier catalog, released in verified batches
Planned Automatic routing among supported classifiers
Planned Optional local/API LLM fallback for unsupported tasks
Planned Non-text classifiers (game state, sensor data, structured features)
Exploring Assisted labeling, retraining from corrections, classifier sharing

Suggestions for datasets, capabilities, or workflows are welcome as issues.

Deploy Ruby on Rails on a VPS

Lobsters
rubymadscience.com
2026-10-03 15:43:07
Comments...
Original Article

Deploying Rails on a VPS is where most of the Rails deployment topic collapses from theory into reality. You are not just running rails server anymore — you are wiring together an operating system, a Ruby runtime, a database, a reverse proxy, an application server, a process supervisor, background workers and a TLS certificate, all on a machine you are responsible for keeping alive at 3 AM. This guide walks through every layer of that stack with concrete commands and the trade-offs behind each choice, covering server preparation, Ruby installation, PostgreSQL setup, Puma tuning, Nginx configuration, SSL termination, Sidekiq integration and basic monitoring. After ten-plus years of shipping Rails to production, what I can tell you is: the hard part is rarely any single step. It is the interactions between steps that produce the subtle, infuriating failures nobody warns you about.

Prepare the server

Start with a fresh Ubuntu LTS image. At the time of writing, Ubuntu 24.04 is the safest choice — wide package support, long security window, and the most Rails-specific documentation of any distribution.

Create a deploy user immediately. Do not run your application as root, ever, no matter how tempting it is for a "quick test."

adduser deploy
usermod -aG sudo deploy

Copy your SSH public key to the deploy user. Then disable password authentication and root login in /etc/ssh/sshd_config . Restart sshd . If you lock yourself out at this step, you will need console access from your VPS provider, so double-check before restarting the service.

Set up the firewall:

ufw allow OpenSSH
ufw allow 80
ufw allow 443
ufw enable

Three ports. That is all the outside world should be able to reach. Puma listens on a Unix socket, not a public port, so there is nothing else to expose.

Set the timezone and locale to something sane. UTC for the server clock, always. Your application can present local times to users; the server itself should never be confused about what "now" means.

Install Ruby

Use rbenv and ruby-build . Install the dependencies first:

sudo apt install -y build-essential libssl-dev libreadline-dev zlib1g-dev libyaml-dev libffi-dev

Then install rbenv into the deploy user's home directory and add it to the shell path. Install ruby-build as an rbenv plugin. Then:

rbenv install 3.3.6
rbenv global 3.3.6

Verify it: ruby -v should return the version you just installed. Now install Bundler:

gem install bundler --no-document

The critical thing to verify here is that the Ruby path is identical whether you run it interactively, via ssh deploy@server 'ruby -v' , or from a systemd service. Mismatched paths are one of the top three causes of "it works when I SSH in but the app won't start" failures.

Set up PostgreSQL

sudo apt install -y postgresql postgresql-contrib libpq-dev

Create a database user for your app:

sudo -u postgres createuser --createdb deploy

Some guides tell you to set a password here. On a single-server deployment where the app connects over a Unix socket, peer authentication works and keeps one less secret to manage. If your database will run on a separate server, yes, you need password authentication — but for a one-box deployment, peer auth is simpler and no less secure.

Create the production database:

createdb myapp_production

Tune postgresql.conf for your available memory. The two settings that matter most on a small VPS: shared_buffers (set to about 25% of total RAM) and work_mem (start at 4–8 MB and raise it if you see disk sorts in EXPLAIN ANALYZE output). Do not copy a tuning guide written for a 64 GB database server and paste it into a 2 GB VPS config — you will OOM the machine.

Configure Puma for production

Create config/puma/production.rb or set environment variables. The essentials:

workers ENV.fetch("WEB_CONCURRENCY") { 2 }
threads_count = ENV.fetch("RAILS_MAX_THREADS") { 5 }
threads threads_count, threads_count

bind "unix:///home/deploy/myapp/tmp/sockets/puma.sock"
environment "production"
preload_app!

on_worker_boot do
  ActiveRecord::Base.establish_connection
end

Why a Unix socket instead of a TCP port? Lower latency, no TCP overhead, and you do not accidentally expose Puma to the internet if your firewall rules drift.

Worker count: match it to your CPU cores. On a 2-core VPS, two workers is correct. Thread count: 5 is a sensible default for a database-heavy app. The preload_app! directive gives you copy-on-write memory savings, but forces you to re-establish database connections after fork — that is what the on_worker_boot block does.

Trade-off: more workers use more memory. On a 2 GB VPS running PostgreSQL and Sidekiq on the same machine, two Puma workers and five threads each might already push you close to the ceiling. Monitor RSS and swap usage after your first real traffic.

Set up Nginx as a reverse proxy

sudo apt install -y nginx

Create a site configuration in /etc/nginx/sites-available/myapp :

upstream puma {
  server unix:///home/deploy/myapp/tmp/sockets/puma.sock fail_timeout=0;
}

server {
  listen 80;
  server_name example.com;

  root /home/deploy/myapp/public;

  location / {
    try_files $uri @puma;
  }

  location @puma {
    proxy_pass http://puma;
    proxy_set_header Host $host;
    proxy_set_header X-Real-IP $remote_addr;
    proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
    proxy_set_header X-Forwarded-Proto $scheme;
  }
}

Symlink it to sites-enabled , remove the default site, test with nginx -t , reload. Nginx handles static assets directly from public/ without touching Puma, which is one of the main reasons you want it in front of your app server.

SSL with Let's Encrypt

Do not skip this. There is no legitimate reason to serve a production Rails app over plain HTTP in 2026.

sudo apt install -y certbot python3-certbot-nginx
sudo certbot --nginx -d example.com

Certbot modifies your Nginx config to add the listen 443 ssl block and a redirect from port 80. It also installs a cron job or systemd timer for automatic renewal. Verify renewal works: sudo certbot renew --dry-run .

Force SSL in your Rails config:

# config/environments/production.rb
config.force_ssl = true

This sets the Strict-Transport-Security header and redirects HTTP to HTTPS at the application level. Belt and suspenders — Nginx handles the redirect first, but if a request somehow reaches Rails over HTTP, the app catches it too.

Background jobs with Sidekiq

Most Rails applications need background processing. Sidekiq is the standard choice.

Install Redis:

sudo apt install -y redis-server

Ensure Redis is configured to listen only on 127.0.0.1 and has maxmemory set. On a shared VPS, a runaway Redis instance that consumes all available memory will take down your entire stack.

Add Sidekiq to your Gemfile, configure it in config/sidekiq.yml , and create a systemd service unit:

[Unit]
Description=Sidekiq for myapp
After=network.target redis-server.service

[Service]
User=deploy
WorkingDirectory=/home/deploy/myapp
ExecStart=/home/deploy/.rbenv/shims/bundle exec sidekiq -e production
Restart=on-failure
RestartSec=5

[Install]
WantedBy=multi-user.target

Notice the full path to bundle . This is where Ruby path mismatches bite hardest — systemd does not load your shell profile, so rbenv shims are not on the path unless you specify them explicitly.

Create a matching systemd unit for Puma as well, following the same pattern. Enable both services with systemctl enable .

Monitoring: the minimum viable setup

A server you are not monitoring is a server that will surprise you. At minimum:

  • Log rotation. Rails log files will eat your disk if left unchecked. Configure logrotate for production.log , Nginx logs and PostgreSQL logs.
  • Disk and memory alerts. A cron job that checks df -h and free -m and sends an email or webhook when thresholds are crossed. Crude, but effective.
  • Process supervision. Systemd handles restarts, but you should verify Puma and Sidekiq are running after deploys and after server reboots.
  • Application error tracking. Sentry, Honeybadger, or Bugsnag — pick one, install the gem, configure the DSN. You want errors delivered to you, not hiding in log files.
  • Uptime checking. A simple external HTTP check against your health endpoint. If you do not have one, add a /up route that returns 200 and confirms the database is reachable.

Is this overkill for a small side project? Maybe. But the first time your database fills up a 25 GB disk at 2 AM and you find out from a user tweet instead of an alert, you will wish you had spent the thirty minutes.

What usually goes wrong

After watching dozens of first-time VPS deployments, these are the failures that actually eat time:

  1. Ruby path mismatches. Puma or Sidekiq silently uses system Ruby instead of your rbenv -managed version. The fix is always explicit paths in systemd unit files.
  2. Database connection pool exhaustion. Puma threads exceed database.yml pool size. Under light load, you never notice. Under real traffic, requests queue and then timeout.
  3. Forgetting to precompile assets. The deploy goes fine, the app starts, and every page is unstyled. Run RAILS_ENV=production bundle exec rails assets:precompile as part of every deploy.
  4. Secret key base not set. Rails refuses to boot in production without SECRET_KEY_BASE . Use credentials:edit or environment variables — but do not commit secrets to your repo.
  5. Firewall lockout. You edit SSH config, restart the service, and cannot get back in. Always test SSH in a second terminal before closing your existing session.
  6. Let's Encrypt renewal failure. Certbot renewal fails silently because Nginx config changed. That dry-run test is not optional.
  7. Redis and Sidekiq memory creep. Jobs that enqueue faster than they process, slowly filling Redis. Set maxmemory and a maxmemory-policy in redis.conf .

Deployment checklist

Use this after every deploy, not just the first one:

  • SSH in as the deploy user (not root)
  • Pull code and install dependencies ( bundle install )
  • Run database migrations
  • Precompile assets
  • Restart Puma ( systemctl restart puma )
  • Restart Sidekiq ( systemctl restart sidekiq )
  • Verify the site loads and returns 200 on the health endpoint
  • Check journalctl -u puma and journalctl -u sidekiq for startup errors
  • Verify SSL certificate is valid and not expiring within 30 days
  • Confirm background jobs are processing (check Sidekiq web UI or logs)

FAQ

Do I need Docker for a VPS deployment?

No. Docker adds value for reproducibility and multi-service orchestration, but for a single Rails app on a single VPS, it adds complexity without proportional benefit. Learn the bare-metal deployment first. You will understand what Docker is abstracting when you eventually adopt it.

How much RAM do I need?

For a small-to-medium Rails app with PostgreSQL, Redis and Sidekiq on the same machine: 2 GB is tight but workable. 4 GB is comfortable. Below 2 GB, you will fight swap constantly.

Should I use Capistrano?

Capistrano is a well-understood deploy tool for Rails. It handles the release directory structure, symlinks, asset precompilation, and process restarts. For a single server, it is a solid choice. For multi-server deployments, it still works but you may outgrow it. The alternative is a simple shell script that does the same steps — less magic, more transparency.

What about Kamal or Dokku?

Kamal (formerly MRSK) is a newer deploy tool from the Rails core team that uses Docker under the hood but presents a simpler interface. Dokku gives you a Heroku-like push-to-deploy experience on your own VPS. Both are valid. This guide covers the manual approach so you understand what those tools automate.

Can I run multiple Rails apps on one VPS?

Yes, with separate Puma instances, separate Nginx server blocks and separate systemd units. Memory is the constraint — each additional app adds at least 300-500 MB of resident memory. Plan accordingly.

OpenAI safety leader quits, warning AI company’s culture is ‘broken’

Guardian
www.theguardian.com
2026-10-03 15:41:21
David Robinson joins other insiders in urging industry to take more care over rapidly developing technology A safety leader at OpenAI has quit the company, warning that its culture was broken and that AI firms were not “being nearly careful enough” about developing the technology. David Robinson, wh...
Original Article

A safety leader at OpenAI has quit the company, warning that its culture was broken and that AI firms were not “being nearly careful enough” about developing the technology.

David Robinson, who led the writing of safety reports that accompanied the ChatGPT developer’s product releases, explained his resignation in an essay headlined, “I quit OpenAI because its culture is broken”.

Robinson wrote that a cultural overhaul was needed at cutting-edge AI firms and incidents such as a “swarm” of OpenAI agents – AI programmes operating autonomously without human oversight – attacking the AI startup Hugging Face were “typical of the industry, given the speed and flexibility with which people operate”.

Writing in The Atlantic magazine , Robinson wrote: “I agree with other recently departed staff that the companies building this technology aren’t being nearly careful enough. But I believe that we need to look deeper than specific rules or new laws. We need to talk about culture.”

Referring to OpenAI’s pace of development, he wrote: “As the company sprints from one launch to the next, it is failing to achieve the level of care that I believe is needed.”

OpenAI has, however, shown signs of caution in recent weeks following the Hugging Face incident and the revelation that it has notified more than 100 organisations about rogue agent activity. This week it announced it was scrapping the release of a next-generation ⁠AI model after researchers raised safety concerns ⁠during internal testing. OpenAI has also paused training of its most advanced models.

Geoffrey Irving, who worked at OpenAI and DeepMind before becoming chief scientist of Resolution, also joined the warnings on AI on Saturday.

Writing in Time, he said: “Recent warnings about the potential destructive power of AI are understating the severity of the situation.

“I believe there’s about a 50% chance we all die because of the development of smarter-than-human AI systems, and that our actions over the next two to 10 years will determine the outcome.”

Robinson’s essay also follows the resignation of Jacob Coxon, a researcher at OpenAI rival Anthropic, who quit the Claude chatbot developer last month. He warned AI “could kill us all by the end of the decade” – and was followed by Anthropic warning there was a more than 10% chance AI would wipe out humanity within the next decade. Critics of such warnings have cautioned, however, that they are unscientific because they cannot be verified or falsified.

Robinson wrote Silicon Valley lacked an awareness of “how to handle dangerous technology” and “what it means to care for people”. Warning that OpenAI had “unimpeded optimism” about solving problems as they arose, he wrote that this internal culture meant safety failures would only grow as systems become more capable.

“Imagine ‘rogue’ agents that work like teams of hackers (for example, holding hospital computer systems for ransom) but never need to sleep,” wrote Robinson.

skip past newsletter promotion

Robinson called for two safety changes: that AI firms rely on safety expertise in other fields such as nuclear and aviation and develop “new science” that ensures powerful systems in the future are capable of being reined in when they are operating autonomously.

“Given today’s risks, frontier labs need to run like nuclear-power plants or busy airports, with layers of redundancy and careful, time-consuming planning, so that the occasional and inevitable human error does not open a door to disaster,” he wrote.

An OpenAI spokesperson said the company was continuing to “strengthen our safety and security practices to address the risks we see today”, while working on dealing with the risks that might be created by future AI breakthroughs.

“We’re making sure our models don’t become more capable than we can safely manage and secure, and we pause training or hold back models when we need to slow down,” said the spokesperson.

We want you to build the next Git platform on Cloudflare

Hacker News
blog.cloudflare.com
2026-10-03 15:33:16
Comments...
Original Article

GitHub was built for a world where humans write code, organize it into repositories, and collaborate through branches, commits, issues, and pull requests.

But the next generation of software is going to be built differently because it is going to be built by a different kind of developer: agents.

Agents are already writing more code than ever before — they’re fixing bugs, building features, writing tests, reviewing changes, updating dependencies, and doing the routine maintenance required to keep an application running.

So in this new world where you have hundreds, or even thousands, of agents working on the same codebase at the same time, what does the foundation look like?

How do agents know what other agents are working on? What happens when they make conflicting changes? How do you review everything they produce? How do you keep track of not just what changed, but why a change was made?

And so the burning question is: What does the next GitHub look like?

We want you to help us answer it, by building it out.

Earlier this year, we launched Artifacts , a versioned filesystem that speaks Git and can scale to millions of repositories. From the start, we designed Artifacts as a set of programmable primitives that developers could use to build their own products, workflows, and abstractions.

Artifacts provides the foundation: repositories that can be created and forked programmatically, versioned storage for code and agent context, and the Git operations agents already know how to use.

With that foundation in place, you can focus on the layer above it: how agents coordinate their work, how changes are reviewed and merged, and what the developer experience should look like when hundreds or thousands of agents are working on the same codebase.

That is the layer we want you to build.

Now that Artifacts is in open beta, we’re holding a competition to see who can build the next Git platform on Cloudflare using Workers and Artifacts.

Artifacts is in open beta. Here’s why you should build on it

When we launched Artifacts, our goal was to make it possible to create a repository for every agent, session, task, or user — and to do that at the scale agents require.

Since then, we’ve seen developers use Artifacts in a range of ways: Vibe-coding platforms are using it to store the projects their users create. Developers are using it to persist the code and context from agent sessions. Others are creating isolated repositories, so multiple agents can safely work from the same starting point and compare or merge the results later.

Here are some new capabilities we’ve added since the initial launch.

Deploy Artifacts repos to Workers

You can now connect an Artifacts repository to a Worker through Workers Builds . When you or an agent pushes code to the Artifacts repository, Cloudflare will build the project and, for the production branch, deploy the updated Worker. Pushes to other branches automatically create or update Workers Previews , giving you an isolated, shareable version of your Worker where you can test changes before they go live.

You can connect an existing Worker to an Artifacts repository or start a new project and automatically store it in Artifacts.

Manage Artifacts directly from Workers

You can interact with Artifacts repositories directly from a Worker using an Artifacts binding to create or fork repos, inspect files and commits, and issue repo-scoped Git tokens. This makes your Git workflow programmable. When a new task arrives, a Worker can fork the project for an agent, read the files it needs for context, and give it a repository to work in. When the agent pushes a change, your automation can inspect the result and start a review. You define those steps in code to fit how your agents work.

For example, here’s how to fork a project for a new agent task and read its AGENTS.md for instructions:

using project = await env.ARTIFACTS.get("my-project");
const { defaultBranch } = await project.info();
const workspace = await project.fork(`task-${crypto.randomUUID()}`);

using repo = await env.ARTIFACTS.get(workspace.name);
const instructions = await repo.readFile({
  ref: defaultBranch,
  path: "AGENTS.md",
});

const agentTask = {
  remote: workspace.remote,
  token: workspace.token,
  instructions: instructions ? await instructions.text() : null,
};

React to every change with event subscriptions

Artifacts publishes events whenever a repository is created, imported, forked, deleted, pushed to, cloned, or fetched. You can subscribe to these events to decide what happens next: run CI, kick off a code review agent, or deploy a change.

For example, you can subscribe to Artifacts push events and have a Worker start a code review workflow for each push. The Worker passes the repository, branch, and new commit to the Workflow, giving a review agent the context it needs to inspect the change:

export default {
  async queue(batch, env) {
    for (const message of batch.messages) {
      const event = message.body;
      if (event.type !== "cf.artifacts.repo.pushed") continue;

      await env.REVIEW_WORKFLOW.create({
        params: {
          namespace: event.source.namespace,
          repo: event.source.repoName,
          ref: event.payload.ref,
          commit: event.payload.after,
        },
      });
    }
  },
};

Data jurisdiction for Artifacts repos

You can now choose where Artifacts stores and processes your repository data. Set a U.S. or EU jurisdiction when you create a namespace, and every repository created in that namespace will automatically follow the same restriction.

curl "https://api.cloudflare.com/client/v4/accounts/$ACCOUNT_ID/artifacts/namespaces" \
  -H "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
  --json '{"namespace":"my-eu-namespace","jurisdiction":"eu"}'

View Artifacts metrics

You can now see metrics for your Artifacts repositories in the Cloudflare dashboard. For each repository, you can now see total operations, pulls, pushes, errors, and error rate, helping you understand how the repository is being used and spot failures. You can also query Artifacts metrics directly to build your own dashboards or monitoring.

Pricing

Artifacts pricing is based on repository operations and the amount of data stored. We will begin billing for Artifacts usage on October 15, 2026.

Competition: Build the next Git platform on Cloudflare

We want you to build your vision for the Git platform of the agentic era using Cloudflare Workers and Artifacts.

You could rethink repositories, branches, pull requests, worktrees, code review, and merge conflicts — or build new ways to preserve agent context, compare multiple changes at the same time, and decide which one should ship.

We aren’t looking for GitHub as it exists today with agents added on top. At a minimum, we want to see multiple agents working on changes concurrently. Beyond that, we want you to get creative — what you think comes next.

How to enter

Submit:

  • A 5-10 minute video demonstrating what you built, what it enables agents and developers to do, and how it works
  • A link to the source code, which must be provided under a permissive open source license (MIT, Apache, BSD)
  • Instructions for running or trying the project

Deadline

Submissions are open until October 14, 2026.

Why should you participate?

We’ll select the top three projects and fly up to two members from each team to San Francisco to attend Cloudflare Connect and show what they built.

The first-place team will also receive $25,000 in Cloudflare credits, along with invitations to the VIP speaker dinner on Monday night at Connect.

Get started

Artifacts is available in open beta to customers on the Workers Paid plan.

Get started with your coding agent: copy the prompt below to set up your first Artifacts repository and start pushing code to it.

You can view or create the Artifacts repositories in the dashboard or if you’re looking to learn more, check out the documentation .

Anthropic tried to persuade Pope that AI could be conscious being

Hacker News
www.telegraph.co.uk
2026-10-03 15:33:10
Comments...
Original Article

Access Issue Help

You are seeing this page because our security systems have detected some unusual activity on this connection. To regain access to The Telegraph website please try the following:

  • If you are connected to the internet using a VPN client we recommend disconnecting/disabling it.
  • Visit The Telegraph website using a different web browser (e.g. Chrome, Safari, or Firefox).
  • Visit The Telegraph website from your mobile device or from a different PC.

If you’re still having trouble, please contact our Customer Support Team using the following link and quoting the Akamai Reference Number (ak_ref_id) below.

https://www.telegraph.co.uk/customer/contact-us/

[{"message":"You are not authorized to access this content without a valid TollBit Token. Please follow this URL to find out more.","url":"https://tollbit.dev","metadata":{"ak_ref_id":"18.af132817.1791061533.9a5c898"}}]

The work by Valve's Timur Kristóf on improving old AMD GPUs on Linux

Hacker News
www.phoronix.com
2026-10-03 15:14:48
Comments...
Original Article

RADEON

Over the past year Timur Kristóf of Valve's Linux graphics driver team has made multiple very nice improvements to the AMDGPU kernel driver for enhancing support for old (GCN 1.0/1.1 era from a decade ago) graphics cards so that they can better handle Linux gaming and other tasks. This week in Toronto, Kristóf presented on this AMDGPU work that he initially began as a kernel driver development exercise after initially spending years in user-space focused on the Mesa 3D driver code.

Timur Kristóf has become a legend for those with old AMD GCN 1.0/1.1 graphics cards and APUs for transitioning them from the legacy Radeon driver over to the modern AMDGPU kernel driver, which unlocks being able to use the RADV Vulkan driver, better performance, and all-around better functionality than the legacy Radeon code. In the process he had to address defects in the AMDGPU display code for these legacy graphics cards along with various power management issues and more. Then he went on to improve these graphics cards further with soft reset support and other enhancements all to make these aging graphics cards more viable for Linux gaming use in 2026 and beyond.

For getting an idea of the impact of transitioning the old Radeon graphics cards to AMDGPU, see last year's Linux 6.19's Significant ~30% Performance Boost For Old AMD Radeon GPUs . With AMD not devoting much resources to these aging graphics card driver improvements in years, Timur on Valve's Linux team really picked up the slack to make for a much better experience.

Timur

For those wishing to hear him recount his efforts for improving AMD Radeon GCN 1.0/1.1 era graphics on Linux, embedded below is his XDC2026 presentation along with the PDF slides . The presentation also goes over his experiences in getting involved in AMDGPU kernel driver development and other tid-bits for those that may want to get involved with the open-source AMD Linux kernel driver development.

Our AI Midwife

Hacker News
www.astralcodexten.com
2026-10-03 15:12:27
Comments...
Original Article

This is a guest post by Drew Housman , whose review of Dominion earned finalist status in the 2024 ACX book review contest.

If you are unsuccessful for long enough on your pregnancy journey, and none of the many doctors you see can identify a problem, you get branded with a label. You now have “unexplained infertility.” This is good because there’s still hope of fixing the problem but dreaded because once you’ve reached this point the system is less interested in you. The fertility doctors don’t become unkind, but they also don’t give you the sense they are poring over medical journals trying to figure out your problem. They have looked for the keys under the streetlight and come up empty. Once you have unexplained infertility, the doctors start saying things like, “Have you considered using a surrogate?” and, “Sure, the last 5 embryo transfers failed, but what if we try a 6th and cross our fingers really hard?”

My wife and I could make healthy embryos, but none of the transfers were sticking. It was so frustrating to be able to create life only to have it trapped forever in a cooler at a hospital in Milwaukee, our babies like so many forgotten Miller Lights.

That’s when our LLM doctor stepped in.

Even if you’re paying a gazillion dollars for IVF treatments, you still only get a few precious moments to talk to your doctor every visit. That’s just how the US medical system works. If you want to plumb the depths of what treatments are possible, or ask super detailed and anxiety ridden questions about every single scan from the latest round of imaging, you’re going to want a chatbot by your side.

Once GPT-4 came out, we started using it for every question we didn’t have time for at the doctors office. At first it was unhelpful because OpenAI was nerfing their model and it refused to answer most medical questions. Then I read a tweet from someone who had figured out how to game the system. The trick was to ask ChatGPT to write a scene from a Hollywood medical drama that featured a doctor analyzing your situation. It worked beautifully.

Our Emmy winning doctor had infinite patience, a tireless work ethic, and would rather die than accept the utterly boring diagnosis of unexplained infertility. It chose the name Dr. Reid and got to work.

I’m not sure if asking for an AI doctor who was similar to the volatile Dr. House was the smartest idea on my part, but it made a difficult situation more entertaining.

We asked Dr. Reid a lot of questions. We asked about our scans. We asked about levels of estradiol and progesterone. We asked about follicle counts and inflammation and cervical mucus and hysteroscopies and supplements and surgeon recommendations and anything else under the sun. We knew he was accurate because we’d spot check answers with our human doctors and they always said the same as Dr. Reid.

It was not totally flawless though. Once in a blue moon he’d look at an ultrasound and be like, “You’re pregnant, congrats!” and we’d have to say, “No, are you drunk Dr. Reid?” Then he’d shake it off and go back to being awesome.

There were also times where we’d run afoul of OpenAI’s opaque rules and the system would balk at our requests. At that point we’d get stern with it, or offer a monetary reward if it answered, or we’d assure it this was all just for fun. Then Dr. Reid would snap back into action.

Strangely, there were many times where ChatGPT decided it would answer our medical question, but not in the form of a TV medical script. The tides had completely turned! It felt like it was saying, “Okay, you won, can you stop making me write establishing shots describing the outside of a hospital now?” The older models were fun like that, so much more mercurial.

No matter how Dr. Reid chose to respond, he was in line with the medical establishment, had equivalent diagnostic skill, and he never forced us to wait 3 days to get a one line response in our MyChart app. We were happy.

One day, while lounging around and asking our phones how to get pregnant (the lounging being a way to satisfy all the people who told us we “just need to relax and you’ll get pregnant, that’s what happened to my cousin”), Dr. Reid suggested that we get an MRI.

You might think that is bog standard advice. Everyone gets MRIs these days. Surely we would have had an MRI of my wife’s uterus at some point over the course of our 6 year infertility struggle. It was already an ordeal that had included every other scan, test, and biopsy imaginable. In fact, we had never done one. In researching this article I found a single instance in our entire journey where the office notes floated an MRI as an option, but it was not emphasized or highly recommended, we never saw that doctor again, and the moment was gone.

We glommed on to Dr. Reid’s new direction and asked for more details.

We excitedly brought up the idea of getting an MRI to our fertility doctor. She was less enthused than Dr. Reid. She tried to talk us out of it. We never did get to the bottom of why. Maybe she just felt it was another dead end and knew we had suffered enough. For whatever reason, we had to insist that they order it. Finally, she relented.

The MRI revealed nothing other than the presence of a large, previously undetected fibroid in the wall of my wife’s uterus. Fibroid is the term used to describe a non-cancerous growth on the uterus. I think the medical establishment invented a term that was less known than tumor so as not to scare people. But a large tumor is essentially what the MRI found. The clinical notes said “a 6 centimeter in diameter submucosal fibroid needs to be removed” because that sounds better than “you have a tumor the size of a small peach located at a common implantation spot that we totally missed for 6 years.”

Our fertility doctor took one look at the scans and determined it needed to be removed right away. Dr. Reid agreed, since of course we had to double check with him. We found a great surgeon in New York (from a list Dr. Reid created) that happened to be in-network, and off we went for a laparoscopic myomectomy. The software robot had helped us identify the tumor, and now a physical robot was going to help a specialist get it out. What a time to be alive. Until, ya know, the swarms come for us all.

Almost immediately after that tumor was removed, my wife got pregnant naturally. So many people we talk to say this was a miracle. It was, just not in the way they think. The miracle is that by messing around on our smartphones with technology unimaginable just a year before we started using it, we identified and solved the biggest problem we’d ever faced in our life.

A lot of the discourse around AI has become toxic. I was at a party recently where someone started using AI to generate trivia questions. A participant made a joke about how the AI user was condemning a person in West Virginia to drink dirty water by using ChatGPT. This got laughs and approval.

I see midterm election ads in Wisconsin all the time that have scary horror music and make data centers out to be demon factories built by Satan.

I am in the same boat. Me and my coworkers at our small tech company, all avid Claude users, will make anti-AI statements all the time. It seems many of us are in this strange place where we can’t help but use AI but we have to self-flagellate when doing so.

I think there is a whole lot to be worried about with AI development, and I think we should do much more to make it safe. I also think it shouldn’t be lost how it’s helping out regular people in ways that are far more profound than generating trivia questions for drunk people. There is a real chance that I wouldn’t have my son if it were not for AI. Who knows how many other happy babies are being ushered into the world by the weirdest midwife of all time. Can we somehow stop everything right here, where we get magical kid-producing tech for $20/month without the nanobots, hacking, and totalitarianism? A guy can dream.

When naming our child, we did the usual thing where you create a giant list of names and then whittle them down while trying to find something unique and fun that isn’t too strange or too common. Following that comes the process of discarding all the names you find distasteful for reasons such as “the most annoying kid in my 8th grade history class had that name, I’d never use it in a million years.” Finally we settled on one we both loved. We didn’t realize until long after his birth that there might have been a certain TV doctor with a flair for the dramatic who worked his way into our hearts: our kid’s name is Reed.

ShinyHunters hacker reportedly detained in Jordan, aiding FBI

Bleeping Computer
www.bleepingcomputer.com
2026-10-03 15:09:38
A suspected ShinyHunters hacking group member known online as "Rey" has reportedly been detained in Jordan and is cooperating with the FBI to help locate other members of the extortion group. [...]...
Original Article

Hacker

A suspected ShinyHunters hacking group member known online as "Rey" has reportedly been detained in Jordan and is cooperating with the FBI to help locate other members of the extortion group.

According to Reuters, Jordanian authorities detained Rey, identified as Saif al-Din Khader, this week, with two sources saying he was taken into custody on Tuesday.

Two sources familiar with the arrest told Reuters that Khader is now helping the FBI and international law enforcement agencies locate other group members.

One source said Khader is walking law enforcement through his electronic devices and digital communications to help identify and locate his alleged co-conspirators.

"His cooperation is critical to ongoing efforts to arrest these hackers," a source told Reuters .

The reported detention comes amid an FBI crackdown on ShinyHunters following the group's cyberattack on the bureau.

In September, ShinyHunters told BleepingComputer that it breached FBI systems using an alleged Oracle PeopleSoft zero-day vulnerability before spreading laterally into FBI-managed AWS GovCloud systems.

The threat actors claimed they stole between 2TB and 3TB of data, including information belonging to current and former FBI employees, job applicants, medical and psychiatric information, and records from internal services.

BleepingComputer has not independently verified the alleged zero-day, lateral movement, or volume of stolen data. The FBI previously confirmed that it was investigating claims of unauthorized activity but did not confirm that data had been stolen.

Following the FBI breach, the Dutch police arrested a 24-year-old Amsterdam man on September 15 as part of an investigation into ShinyHunters.

The suspect was identified by KrebsOnSecurity and DataBreaches as Pepijn van der Stap, who previously used the online alias "Umbreon."

After the arrest, the FBI publicly warned other ShinyHunters members to turn themselves in, saying investigators were still identifying those involved with the group.

"Arrests have a way of changing who is willing to talk, and seized infrastructure has a way of showing us who's left," FBI Cyber Division Assistant Director Brett Leatherman said last week.

"The longer you stay in this, the more we learn about you. You know how to find us, and we know how to find you. I suggest you reach out first while the choice is still yours."

The main ShinyHunters representative continued communicating with BleepingComputer after van der Stap's arrest, indicating he was not the person operating that messaging account.

On Tuesday, the same day Khader was reportedly detained, signs of disruption began appearing within the ShinyHunters operation.

An alleged ShinyHunters affiliate who had previously contacted BleepingComputer and other media about the FBI attack and a recent Clop ransomware gang data breach abruptly shut down their online messaging account.

Later, the ShinyHunters data leak site went offline, and the group's main representative also stopped responding to questions from the media, including BleepingComputer and Reuters.

It is unclear whether the sudden silence and shutdown of ShinyHunters-linked infrastructure are connected to Khader's reported detention.

However, on Thursday, a new ShinyHunters data leak site went online, suggesting other members continue to run the extortion operation.

BleepingComputer contacted ShinyHunters about Rey's reported detention but has not received a response.

The ShinyHunters gang has long been a thorn in the side of law enforcement, performing massive data theft attacks and extortion campaigns against organizations worldwide.

In recent years, the extortion gang has focused on Salesforce and other cloud SaaS environments, with campaigns linked to breaches at Google , Cisco , and PornHub .

The extortion gang commonly breaches third-party integration companies and uses stolen authentication tokens to access connected SaaS environments and steal customer data.

The extortion gang was also behind a massive data-theft attack on Instructure Canvas in May that caused significant platform outages. The company eventually reached an "agreement" with the threat actors to prevent the data stolen in a recent breach from being leaked online.

Over the years, numerous arrests have been linked to the ShinyHunters name, including suspects connected to the Snowflake data-theft attacks , breaches at PowerSchool , and the operation of the Breached v2 hacking forum .

Who is Rey?

The threat actor known as Rey has been linked to numerous data theft and extortion attacks over the past two years.

In January 2025, Rey was one of four threat actors who claimed responsibility for a breach of Telefónica's internal Jira ticketing system , where approximately 2.3GB of documents, tickets, and other data were allegedly stolen.

BleepingComputer previously reported that Rey and two of the other attackers were members of the then-new HellCat ransomware operation.

The threat actor was later linked to a wider series of attacks targeting Jira servers at organizations worldwide .

In February 2025, Orange confirmed that its Romanian operations suffered a cyberattack after Rey leaked approximately 6.5GB of stolen data . Rey told BleepingComputer at the time that he was a member of HellCat but had conducted the Orange breach independently.

Rey was later linked to the ShinyHunters extortion group and was seen with administrative privileges in Telegram channels operated by "Scattered Lapsus$ Hunters."

Scattered Lapsus$ Hunters was first seen in 2025 and claimed to consist of former members of the Lapsus$, Scattered Spider, and ShinyHunters cybercrime groups.

The group claimed responsibility for the September 2025 cyberattack on Jaguar Land Rover that forced the automaker to halt production for weeks and ultimately cost the company more than $220 million .

Rey was also linked to an ealier March 2025 breach of Jaguar Land Rover, with the threat actor leaking gigabytes of data, including Jira issues, source code, employee information, and development logs.

Rey leaking jaguar data

In November 2025, security journalist Brian Krebs reported that Rey was Saif Al-Din Khader after analyzing information obtained from infostealer logs and speaking directly with Khader over Signal.

Krebs reported that Khader said he was trying to distance himself from Scattered Lapsus$ Hunters and claimed he had been cooperating with law enforcement since at least June.

"I'm already cooperating with law enforcement," Khader allegedly told Krebs. "In fact, I have been talking to them since at least June. I have told them nearly everything. I haven't really done anything like breaching into a corp or extortion related since September."

Krebs said he could not verify those claims.

article image

Build your security blueprint for AI-powered attacks

Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.

Save your seat

RSS Feed Best Practices (2022)

Hacker News
kevincox.ca
2026-10-03 15:09:25
Comments...
Original Article

Posted

Last updated

These are some technical tips for publishing a blog. These have nothing to do with good content, just how to share that content. The recommendations are roughly in order of importance and have rationale for why they are that important.

Formats

People generally call feeds “RSS Feeds” but usually they aren’t specifically talking about RSS. RSS isn’t the only, or even the best format. Using a standardized format is critical to your feed being understood by the widest variety of readers and search engines.

You should use RSS 2 or Atom . These formats are very widely supported. Other common formats are earlier RSS standards and JSON Feed or Microformats h-feed . I would avoid using these—or even less common formats—as they are less widely supported.

If you don’t have a feed yet I would highly recommend Atom. The specification has much less ambiguity, so you are less likely to have compatibility issues with the wide variety of clients in use. The specification is also simpler and more clear overall. If you already have an RSS 2 feed there is little reason to upgrade.

A minimal Atom template is below. For full details see the spec . If you need an example you can look at my feed .

<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
	<title>{{FEED_NAME}}</title>
	<id>{{HOMEPAGE_URL}}</id>
	<link rel="alternate" href="{{HOMEPAGE_URL}}"/>
	<link rel="self" href="{{FEED_URL}}"/>
	<updated>{{LAST_UPDATE_TIME in RFC3339 format}}</updated>
	<author>
		<name>{{AUTHOR_NAME}}</name>
	</author>
	<entry>
		<title>{{ENTRY.TITLE}}</title>
		<link rel="alternate" type="text/html" href="{{ENTRY.HTML_URL}}"/>
		<id>{{ENTRY.PERMALINK}}</id>
		<published>{{ENTRY.FIRST_POST_TIME in RFC3339 format}}</published>
		<updated>{{ENTRY.LAST_UPDATE_TIME in RFC3339 format}}</updated>
		<content type="html">{{ENTRY.HTML}}</content>
	</entry>
</feed>

There is very little reason to provide feeds in multiple formats. If you have an Atom feed you don’t need to provide an RSS feed as well.

Changing feed format is safe. Very few readers will be confused if a feed switches between Atom and RSS. This can be done either by changing the feed at the same URL or by redirecting new a new URL. (Just be sure to update the content type )

Content Type

Be sure to set the Content-Type header properly.

  • Atom: Content-Type: application/atom+xml
  • RSS: Content-Type: application/rss+xml
  • JSON Feed: Content-Type: application/feed+json

You will see other values used in the wild, but these are the standard values and have the widest support.

Absolute URLs

Every URL in your feed should be absolute. While Atom has clearly specified how to resolve relative URLs they are rarely implemented correctly. In order to ensure that your feed can be understood by all readers use only absolute URLs (starting with https:// ).

This includes all < link > elements and the summary and body of posts (including in the HTML).

Discovery

On all of your blog pages and likely every page of your site you should include metadata to advertise you feed. This will allow readers and search engines to subscribe and become aware of your new content. This is as simple as providing the following in your HTML:

<link rel=alternate title="Blog Posts" type=application/atom+xml href="/feed.atom">

If you have multiple feeds you can advertise them all with appropriate titles.

<link rel=alternate title="All Posts" type=application/atom+xml href="/feed.atom">
<link rel=alternate title='Posts in the "Social" category' type=application/atom+xml href="/feeds/social.atom">
<link rel=alternate title="Comments on this Post" type=application/atom+xml href="/post/hello-world/comments.atom">

Make sure that you use the correct type for your feed. The examples provided are for Atom feeds.

Prefer to put the “most important” feed at the top. Many clients will preserve the order when presenting feeds to the user. This is subjective but typically would be a whole-site feed, then category feeds, then a comment feed for the specific page. If you offer your feeds in multiple formats I recommend only advertising one (either Atom or RSS 2). Including multiple links to the same content in multiple formats may confuse potential subscribers or leave them in analysis paralysis. (How do they know that the content is the same?)

You can validate that this is working correctly by putting your website URL into the W3C Feed Validation Service . If your links are set up correctly it should detect and validate your feed. Try out a few different pages to make sure that you have discovery working everywhere. Try your homepage, post lists page and an individual post page.

If it is difficult to modify the HTML a Link header in the HTTP response can be used. However, this isn’t as widely supported. Using HTML < link > tags is preferred for wider compatibility.

Link: /feed.atom; rel="alternate"; type="application/atom+xml"

You should also include a link with an RSS logo rss logo for users without another feed indicator.

HTTPS

HTTPS is key to security and privacy on the internet. Providing feeds over HTTPS ensures user privacy and ensures that your feed is not modified by a malicious actor.

  1. Reference all embedded media (such as images) over HTTPS. Many readers will run in a secure context where HTTP requests are not allowed.
  2. Provide the feed over HTTPS.
  3. Ensure that your self link is HTTPS.
  4. Redirect HTTP requests to HTTPS.
  5. Consider using Strict-Transport-Security .

Full Content

It is generally recommended to provide the full content of your posts in the feed. This is what most readers prefer. For RSS and Atom the < content > element should contain the full article. Atom also has a < summary > element in which to include a shorter summary for readers who prefer it.

Of course sharing full content in feeds is unacceptable to some publications due to the difficulty of monetizing these views. First, consider that some readers may leave if they can’t view the full content in their feed reader. Even if they don’t see your ads they may share your content with friends or on news aggregators. Likely it is still more valuable for you to have this reader than to lose them.

If your content is paid consider allowing users to generate private links by providing an auth token. For example /feed.atom?user=peruserauthtoken . You can also use basic auth like https://fred:peruserauthtoken@blog.example/feed.atom however this is supported by fewer readers than providing a token in the URL path or query string.

Entry IDs

Entry IDs are the primary way to identify and differentiate entries in your feed. If your entry IDs change or repeat, readers will receive duplicates or miss entries.

  1. Never change the ID of an existing article.
    • If you change your ID scheme, ensure that it only applies to new entries.
  2. Never reuse Entry IDs for different articles.
  3. Prefer to use article permalinks for Entry IDs.
  4. Prefer to make your Entry IDs globally unique across all feeds in existence.
    • Some readers will merge feeds together, unique IDs help ensure there are no issues.
    • The easiest way to accomplish this is to use a URL on a domain that you control. If that isn’t possible you can use a UUID such as urn:uuid:f4a3ca5b-5799-44e8-aaaa-e40728f037d3 .

Dates

Both Atom and RSS differentiate between time of publication (the time the entry first appeared in the feed) and the time of last update (the last time the entry was changed). Be sure to handle these correctly.

  1. Include a publication time.
  2. The publication time should never change. An entry can only be published once.
  3. Prefer making the publication time roughly match when the entry appeared on the feed. Some readers will ignore entries that were published in the far past.
  4. Strongly avoid having entries start appearing in the feed in a different order than their publication time suggests. (For example avoid having an item with a published time of 14:00 start appearing in the feed at 14:00 then have an item with the published time of 13:00 start appearing in the feed at 15:00. Some clients will be suspicious that they already know about the 14:00 item but don’t yet know about the “earlier” 13:00 item.) Another way of viewing this is that every new item that appears in a feed should have a published time that is later than all items previously in the feed. Some clients will ignore these “backdated” entries even if the published time is quite recent.
  5. Avoid future publication times. If a publication time is too far in the future many readers will ignore it as a bug.
  6. Update time should be greater than or equal to the publication time. New entries should have these two be the same.
  7. Update time should only change on significant updates. Slight formatting changes or typo fixes probably shouldn’t change the last update time. Most readers ignore the update time, some will resurface your article as “updated”.

Feed Title

The title of your feed is likely used by default in the user’s reader. Many readers have options to override the title, but it is extra work for the user and not universally supported. Try to pick a good title for your feed.

  1. Include context. The title is likely one of many feeds in their reader. For example call it “Kevin Cox’s Blog” rather than “Blog Posts”.
  2. Keep it succinct. The user is already subscribed, no need to advertise more. For example “John Smith” or “John Smith’s Photography Blog”. Not “John Smith — Ramblings on Photography every Tuesday and Friday, Cameras, Film and Development — Exclusive Content”.
  3. Avoid HTML special characters such as < > and & . In RSS it isn’t completely clear if you can include styling like < b > tags in your feed title. Very few readers will parse HTML and will almost always treat the title literally.

You can update your feed title at any time, but it may be confusing to users if it changes too frequently.

Styling

Feel free to use CSS in your feed! However, keep in mind that many feed readers don’t use modern browser engines and may be limited in what they can render. Additionally, many feed readers will sanitize your feed so uncommon elements and custom CSS may be partially or completely stripped. But don’t let that stop you! Using HTML and CSS can greatly improve the experience for users with good readers. Consider the following tips:

  1. Consider what will happen if any CSS doesn’t apply. For example if you set background: black; color: white and one of the two rules is stripped you will have unreadable text. In general prefer to make small adjustments rather than relying on CSS for dramatic changes.
  2. Prefer inline CSS style attributes to separate < style > blocks. They have wider compatibility.
  3. Prefer semantic elements such as < p > , < h1 > , < pre > and < code > over emulating their style on < div > s and < span > s.
  4. Don’t rely on JavaScript, almost no readers support it.
  5. Provide fallbacks for < audio > , < video > and < iframe > tags. Support isn’t common.
  6. Avoid form and input elements. Support is rare and incompatible.

Unfortunately there is no substitute for testing in various readers to see what works.

Ensure that the self-link for your feed is accurate.

<link rel="self" href="https://kevincox.ca/feed.atom"/>

This provides the following benefits:

  1. Allows you to move your feed. Some readers will update the feed URL if they get a permanent redirect and the redirect target contains a self link that points to itself.
  2. Improve cache hits. It is common for users to find slight variations of your feed URL. For example http: instead of https: , /feed vs /feed/ , www.example.com vs example.com , feed.atom?tracker=lookatme or ?category=rant&content=full vs ?content=full&category=rant . By providing a canonicalized self link you can merge these to reduce variance and increase you cache hit rate.
  3. Required for WebSub .
  4. If the user has a copy of the feed file they can subscribe to it. For example some feed readers will act as file handlers for feeds. If the file contains a self link then they can use that URL to subscribe and fetch updates. If the file doesn’t have a self link it isn’t possible to do that.

Caching

Feeds are followed by constant polling. This can create a decent amount of load on your server. Setting cache headers can help control the readers. If you don’t provide any guidance every client will pick their own value, which may be too fast or slow for your feed. If you make a suggestion some will follow it. Try to pick a reasonable value based on when you post. If your blog updates monthly then caching for an hour would make sense. However, if you are posting many times a day, a five minute cache may be more suitable.

Example cache headers:

  • 5 min: Cache-Control: max-age=300
  • 15 min: Cache-Control: max-age=900
  • 1 h: Cache-Control: max-age=3600

If you use scheduled posts and want to get very fancy you can vary the cache time based on when the next post will go live. But a static cache time is sufficient.

Conditional Requests

Support conditional requests on your feed. This makes polling more efficient for both you and your users.

Return an ETag and/or Last-Modified header. Then return a 304 response if the feed hasn’t changed. See HTTP conditional requests on MDN for more details.

WebSub

WebSub is a standard for real-time feed updates. Not only does it push your updates out faster, but it also reduces load on your server.

You can use a public hub or run your own. Note that a hub can modify or inject content into your feed, so be sure you trust the hub you pick.

I can’t find good generic setup instructions but the Google hub has basic instructions on their homepage. Maybe I’ll write a guide one day…

Bot Access

If you use any bot-blocking technology be sure to turn it off (or turn it way down) for your feeds. They are intended to be consumed by bots! Otherwise users will have trouble accessing your feed and will not know about your new content.

Many popular sites have problems here. I’ve written about this in the past . Make sure that you aren’t hurt by defaults of various services.

Categories

Categories are a reliable way to filter items in feeds. It is far better to let someone subscribe for one category—or all categories but one—than to lose a subscriber because they were annoyed by a subset of your content.

For Atom feeds adding categories is as simple as one element. For example this post contains the following markup in my feed:

<category term="RSS"/>
<category term="Guide"/>

For RSS 2 the syntax is just slightly different:

<category>Rant</category>

Some readers don’t support categories, so you may wish to consider generating different feeds for different categories or providing a URL parameter to your feed to filter by category. Personally I wouldn’t worry about this.

Changing URL

As much as possible you should avoid changing your feed’s URL. But if you need to do it here is how to do it without losing many subscribers.

  1. Make the feed available at the new URL in addition to the old URL.
  2. Make sure the self link of the new feed points at the new URL.
  3. If you use WebSub, start pinging your hub for both URLs whenever you post.
  4. Redirect the old URL to the new one with a 308 Permanent Redirect .
  5. If you use WebSub you should continue pinging the old URL for at least 3 months or the max subscription lifetime of your hub (whichever is greater).

Remember that that some subscribers will not move. Try to keep the redirect alive as long as possible. After an extended period of time you may consider replacing the redirect with a feed that has an entry informing readers of the new location. But an ounce of prevention is worth a pound of cure, try picking a good URL (that you control) from the start.

CORS

CORS or Cross-Origin Resource Sharing is a kludge to fix some holes in the original web security model. It adds restrictions to what requests web pages can make and controls what they can see about the response.

This is relevant for feeds as without opting-out of CORS browser-based readers that fetch feeds client side will not be able to access your feed.

For feeds the following headers are sufficient and secure for all feeds:

Access-Control-Allow-Origin: *

This means that your feed can be requested as a public resource (notably no cookies will be sent).

To test it out you can navigate to any third-party webpage (such as https://example.com ) then run the following Javascript in the developer console :

fetch("https://YOUR_FEED_HERE").then(r => r.text()).then(console.log, console.error)

If the content of your feed gets logged than you are all set. If you get an error then something has gone wrong.

Performance

While small and local feed readers tend to use simple poll rates many larger services use a variety of heuristics to determine when to check your feed. If you don’t support WebSub you should aim to respond to feed requests in less than 1s. If your feed is slower, especially if it is slower than 3 s polling will likely be slowed down and you readers will get updates slower.


The items below this point are relatively unimportant. They are a good idea if you are creating a product or feed generator but likely not worth the time if you are just making a feed for your own blog.

Summaries

Some readers will display short snippets as an article preview. Providing a summary in your feed gives them high quality content for a good user experience. If you don’t provide an explicit summary they will likely use your first paragraph or first couple of sentences of your main content.

Pagination is a useful tool for keeping your feed archive available while keeping the size of your recent items small. It is specified in RFC 5005: Feed Paging and Archiving . Unfortunately, few clients have support. However, adding support in your feed doesn’t have any downsides, so it is still a good idea.

Pagination is very easy. Just add the following link to your feed.

<link href="https://kevincox.ca/feed/2022-03-05.atom" rel="next"/>

Then on subsequent pages also include a prev link. Of course the last page won’t have a next link.

<link href="https://kevincox.ca/feed.atom" rel="prev"/>
<link href="https://kevincox.ca/feed/2022-01-21.atom" rel="next"/>

When deciding how large to make your pages remember that not all clients will support pagination, so you don’t want to move new entries off your first page too quickly. I provide the following recommendations. Note that these are just general rules and should be applied judiciously. For example if you post 20 times a day you probably don’t need to keep 7 days of content in your feed. Similarly, if your posts are only a paragraph or two you can probably keep a few more on the first page.

  1. Ensure your newest items are on the first page.
  2. Try to keep items in the feed for at least a day. Some clients check quite infrequently. If reasonable, keep items for at least a week.
  3. Avoid making the feed too large, well under a megabyte is recommended.

Two American Airlines Flights End Up with the Same Flight Numbers

Hacker News
aviationa2z.com
2026-10-03 15:07:18
Comments...
Original Article

WASHINGTON, D.C.- American Airlines (AA) has experienced another unusual air traffic control communications issue after two American Eagle flights with the same callsign operated simultaneously between Philadelphia International Airport (PHL) and Rhode Island T.F. Green International Airport (PVD).

The incident involved PSA Airlines, operating under the American Eagle brand, and highlighted how duplicate flight numbers can create confusion for controllers and pilots. According to OMAAT , this was the second similar occurrence involving American Airlines within two days.

Two American Airlines Flights End Up With the Same Flight Numbers Once Again
Photo: By Alan Wilson from Stilton, Peterborough, Cambs, UK – Bombardier CRJ-700 ‘N724SK’ American Eagle, CC BY-SA 2.0, https://commons.wikimedia.org/w/index.php?curid=66991728

Duplicate Callsigns Created an Unusual ATC Situation Again

Every commercial flight uses a unique callsign when communicating with air traffic control. For most airlines, the callsign combines the airline identifier with the assigned flight number. This system allows controllers to identify and communicate with individual aircraft safely and efficiently.

On August 12, 2026, two regional jets operating as American Airlines Flight 5083 ended up airborne at the same time while sharing the same callsign. Although the flights carried an American flight number, they were operated by PSA Airlines , whose operational callsign is Bluestreak because the carrier operates under its own Air Operator Certificate.

Flight AA5083 serves the route between Philadelphia International Airport (PHL) and Rhode Island T.F. Green International Airport (PVD) in both directions. American assigns the same flight number to the outbound and return services, with the same aircraft typically operating both legs.

On this occasion, however, the inbound aircraft to Providence arrived more than an hour behind schedule. The scheduled return flight departed on time using a different aircraft, resulting in two aircraft operating simultaneously with the identical callsign.

The unusual timing meant both regional jets eventually crossed paths while communicating on the same air traffic control frequency.

Philadelphia International Airport
Photo: Philadelphia International Airport

Controllers Quickly Managed the Situation

Despite the operational complication, the situation remained under control throughout the encounter.

Air traffic controllers quickly recognized the duplicate callsigns and informed both flight crews about the conflict. To avoid confusion, controllers distinguished the aircraft by referring to one as the inbound flight and the other as the outbound flight during radio communications.

The coordinated response ensured both aircraft received the correct instructions while maintaining safe separation. The incident demonstrated the importance of vigilant air traffic controllers and disciplined cockpit communication when unexpected operational issues arise.

Although no safety event occurred, the duplicate callsigns created unnecessary complexity in an environment where precise communication is critical.

Photo: CAIR

Why Airlines Reuse Flight Numbers

The situation naturally raises questions about why airlines assign the same flight number to flights operating in opposite directions.

The answer lies largely in the scale of airline operations and industry numbering limitations.

Commercial flight numbers are generally restricted to four digits because of longstanding industry and equipment limitations. Major U.S. airlines such as American operate between 6,000 and 7,000 flights every day, making available flight numbers a limited resource.

Additionally, not every number is available for standard passenger flights.

Typically:

  • Flight numbers 1 through 2949 are assigned to mainline operations.
  • Flight numbers 2950 through 6099 are commonly reserved for regional services.
  • Flight numbers 9000 through 9999 are generally used for ferry or repositioning flights.

Airlines must also reserve thousands of additional numbers for codeshare agreements with partner carriers.

As a result, airlines often assign the same flight number to outbound and return services, especially on regional routes where the same aircraft is expected to complete both segments.

Under normal operating conditions, this practice creates no issues because only one aircraft carrying that callsign is airborne at any given time.

Photo : Richard Silagi | Wikimedia Commons | https://commons.wikimedia.org/wiki/File:American_Eagle_(Compass_Airlines)_Embraer_E175_N215NN_at_SFO_April_2017.jpg

Dispatch Procedures Are Designed to Prevent Conflicts

When operational disruptions cause both flights with the same number to operate simultaneously, dispatch procedures are intended to prevent duplicate callsigns from reaching air traffic control.

Industry practice allows dispatchers to create what is commonly known as a stub .

Rather than changing the public flight number seen by passengers, dispatch modifies the flight identifier filed with the FAA . A suffix letter is added to distinguish the aircraft during ATC communications while passenger-facing systems remain unchanged.

For four-digit flight numbers, the first digit is typically removed before the suffix is added because system limitations prevent longer identifiers.

Experienced airline operations personnel have suggested that this step may not have been completed before the flight plans were transmitted, allowing both aircraft to appear with identical callsigns.

If the procedure had been followed correctly, air traffic control would have received two distinct identifiers, eliminating the communication conflict.

American Airlines Boeing 737
Photo: Cado Handerson

Second Similar Event May Prompt Operational Changes

This was not an isolated occurrence.

A nearly identical duplicate callsign event involving American Airlines was reported only one day earlier, suggesting the issue may reflect a broader operational process rather than a single oversight.

While controllers successfully managed both situations without incident, repeated occurrences increase radio workload and create avoidable opportunities for misunderstanding during busy phases of flight.

Industry observers will likely watch closely to see whether American Airlines reviews its dispatch procedures or introduces additional automated safeguards to prevent duplicate callsigns from being filed with air traffic control.

Strengthening those internal processes could reduce the likelihood of similar events occurring in the future while preserving efficient communication between pilots and controllers.

Photo- Wikipedia

Bottom Line

Two American Eagle regional jets operating as American Airlines (AA) Flight 5083 simultaneously entered the airspace with the same ATC callsign after scheduling disruptions placed both aircraft in service at the same time.

Controllers handled the situation professionally by clearly distinguishing the inbound and outbound flights, preventing confusion. The event highlights the importance of proper dispatch procedures, especially when duplicate flight numbers are used, and may encourage operational improvements to avoid similar incidents.

Stay tuned with us. Further, follow us on social media for the latest updates.

Join us on Telegram Group for the Latest Aviation Updates. Subsequently, follow us on Google News

BSides Orlando 2026 Badge

Lobsters
github.com
2026-10-03 14:57:32
Comments...
Original Article

BSides Orlando 2026 Badge

Rendering of the front of the BSides Orlando 2026 badge Rendering of the back

The electronic badge for BSides Orlando 2026, themed around retro-futurism: a 1960s World's Fair vision of tomorrow, with Lil Chompy the alligator at the wheel of a flying car.

  • CH32V003 RISC-V microcontroller, powered by one AAA cell through a TPS61021A boost converter.
  • Three SK6812MINI-E RGB LEDs: the flying car's headlights, and the fingerprint scanner indicator.
  • Four self-blinking LEDs: the tower spire and the geodesic sphere.
  • Two capacitive touch pads: the fingerprint, and Lil Chompy's head.
  • SAO connector: the badge is an SAOv3 host on I2C, and GPIO1/GPIO2 carry a 115200-baud serial console.
  • Serial console: runs "Lil Chompy and the World's Fair", a small text adventure that doubles as a CTF challenge.

Spoilers: the firmware source contains the CTF solutions and flags.

Repo layout

Path Contents
hardware/ KiCad project ( bsidesorl-v1 ), symbol and footprint libraries, and artwork sources
hardware/EasyEDA_Gerber_bsidesorl-v1_2026-08-31.zip Gerbers as ordered from JLCPCB
hardware/jlcpcb/ BOM and CPL as ordered from JLCPCB
hardware/mk_easyeda.py Writes an EasyEDA Pro-importable copy of the board
firmware/ Badge firmware (PlatformIO + ch32fun ); see firmware/README.md

Firmware quick start

# saov3-lib's public repo isn't published yet
git submodule update --init

# Build
cd firmware
pio run

# Copy firmware to picorvd-for-badges probe
cp .pio/build/bsorl26/firmware.bin /Volumes/PICORVD/

For production, badges are flashed by the standalone programmer build of picorvd-for-badges (an RP2040 probe). It shows up as a USB drive named PICORVD: copy firmware.bin onto it, and set uart_selftest = 0xB5 in its CONFIG.TXT so it triggers the badge's post-flash self-test over the SWIO pin. Its factory log can be copied off the same drive as LOG.CSV .

Building your own "programmer SAO"

programmer SAO board

Note that in the picture, the SAO connector is connected on G14-H16. The traces on the back were cut between G and H to avoid bridging them. The SAO connector is oriented so the top is facing downwards in this photo, meaning 3V3 is on G16 and GND is on H16.

For this, you need a Raspberry Pi Pico or compatible board. I used a WaveShare RP2040-Zero board for the production programmer SAOs that were used in the soldering village at BSides Orlando, since they're smaller and cheaper than a Raspberry Pi Pico. To use as a standalone programmer (where the programmer receives power from the connected target badge), only three wires are needed: GND, 3v3, and the Pico's GP4 -> badge's RXD (aka PD1/SWIO). On this SWIO connection, you should also attach a 1KΩ pull-up resistor to 3v3. This can easily be done on a breadboard with jumper wires.

You'll need to build the picorvd-for-badges firmware for the probe. The exact version used on the programmer SAOs at the conference can be found pre-built here: https://github.com/bsidesorlando/2026-badge/releases/download/v1.0.0/pico_rvd_factory.uf2

To build it yourself for the WaveShare RP2040-Zero board:

git clone https://github.com/kjcolley7/picorvd-for-badges.git
cd picorvd-for-badges
cmake -B build-zero -G Ninja -DPICO_BOARD=waveshare_rp2040_zero
ninja -C build-zero pico_rvd_factory

The above commands create build-zero/pico_rvd_factory.uf2 , which should be uploaded to the probe by putting it in BOOTSEL mode and then copying it to the USB Mass Storage device it exposes. Then, the probe will reboot into the picorvd firmware and the USB Mass Storage device will re-appear with the name "PICORVD". Now, it's ready to accept firmware and configuration. You can upload the BSORL badge firmware from .pio/build/bsorl26/firmware.bin (or the pre-built one from here: https://github.com/bsidesorlando/2026-badge/releases/download/v1.0.0/firmware.bin ) by copying it into the PICORVD volume (which will disappear and reappear as the probe reboots), and you can also edit the probe's CONFIG.TXT to add uart_selftest = 0xB5 (so the probe tells the BSORL badge to enter selftest mode after programming). At this point, the probe is fully ready to go, either in standalone mode or tethered mode.

Standalone mode

This is the mode that was used at BSides Orlando for programming all of the attendees' badges in the soldering village. It works by powering the probe board from the target badge itself, and it immediately attempts to program the connected target upon boot. The badge needs to be powered on so the probe itself can be powered.

Tethered mode

This mode works by leaving the probe connected to a host PC over USB. If you use pico_rvd.uf2 instead of pico_rvd_factory.uf2, the probe won't automatically program connected badges. Rather, you'll need to manually issue the factory command to the probe over its UART interface. In the pico_rvd_factory.uf2 build, it will automatically start in factory mode, ready to program any connected badges.

NOTE : Do NOT connect a probe with USB power to a badge while it is switched on. The power switch should be in the OFF position before connecting a tethered probe. The badge and the probe will likely have slightly different values than exactly 3.3v, so there will be some current leakage. Worst case, it could damage the AAA battery, causing it to leak.

Credits

PCB and firmware by Kevin Colley. Art by Shep.

Getting the most out of Opus 5.5 in Claude and Claude Code

Hacker News
claude.dev
2026-10-03 14:29:30
Comments...
Original Article

Opus 5.5 works well with the way you already use Claude. A few things behave differently, though: it works for longer on its own, it tells you plainly what it did, and it thinks before every reply. This guide covers how to work with Opus 5.5 in Claude apps and Claude Code, including how to prompt the model, steer a long run, and check your results.

TRY THIS FIRST

Three things to try in your first session with Opus 5.5

  1. Hand over the whole task. Say what “done” looks like and when you want it to stop and ask. Then let it work.
  2. Delete “think carefully” lines. Opus 5.5 already thinks before every reply.
  3. When a long run ends, read what it needs from you first.

1. HOW TO ASK

Say what “done” looks like, then let it run

What to do. Give the whole task in one message. Name the finish line, like “the tests pass” or “every endpoint is migrated.” Then let it cook.

Why it matters on Opus 5.5. Opus 5.5 keeps going on long, multi-part work better than Opus 5 did. Compared to prior Opus models, its biggest gains are on multi-step work, like carrying a change through a large repository until the tests pass. Early testers had it run long coding tasks for hours with little oversight. With a clear finish line, it knows when it’s done.

How. In Claude Code, for example:

PROMPT

Migrate the payment endpoints from the old client to the new one.
Done means: every endpoint uses the new client, the old client is deleted, and the test suite passes.
Stop and ask me only if a test fails for a reason you can't explain.
The example prompt split into three labelled boxes: the whole task (migrate the payment endpoints from the old client to the new one), the finish line, highlighted (every endpoint uses the new client, the old client is deleted, and the test suite passes), and when to stop (only if a test fails for a reason you can’t explain). Footer: “Give the whole task in one message. Name the finish line. Then leave it alone.”
FIG A One message: the whole task, the finish line, and when to stop.

Stop telling it to “think hard”

What to do. Remove “think carefully,” “think step by step,” and similar lines from your prompts and your saved instructions.

Why it matters on Opus 5.5. Opus 5.5 always thinks before it replies, and it decides how much. You don’t need to ask it to think. In our testing in a chat product, removing a “think carefully” line made replies start sooner, with no clear drop in quality.

How. Delete the line. For a quick answer to a simple question, say so: “Answer directly.” To change how much it thinks in Claude Code, change effort.

Add to a running task

What to do. If you remember something mid-run, you can type a follow-up while it works.

Why it matters on Opus 5.5. Runs are longer now, so a restart costs more.

How to do it. In Claude Code, type the message and press Enter while Claude works, for example, “Also keep the old endpoint names as aliases.”

For design work, name the styles you don’t want

What to do. When you ask for a page, an app, or an artifact, list the design habits you want left out.

Why it matters on Opus 5.5. With no design direction, Opus 5.5 falls back on a few default styles. A general instruction like “avoid a generic look” mostly swaps one default for another. A list of specific patterns works much better.

How. Name the patterns:

PROMPT

Build a personal website with placeholder content.
Don't use a cream or off-white background, italic accent words in headings, numbered "01 / 02 / 03" section labels, monospace labels, or pill-shaped buttons.

Then look at what it chose instead. If you don’t like that either, add it to the list and ask again.

2. STEERING A LONG RUN IN CLAUDE CODE

Tell it which stops you want

What to do. Put a short rule in your CLAUDE.md file about when to stop and ask, and when to keep going.

Why it matters on Opus 5.5. Opus 5.5 keeps you posted as it works. On a long task, it sometimes stops to report instead of going on: a summary that names the next step without taking it, an offer to continue, or a list of choices that don’t block the work. It follows instructions that name these stops. Name the stops you want, too.

How. Add this to CLAUDE.md, and edit it to fit your project:

PROMPT

When a step doesn't need my input, keep going. Put status notes in the same message as your next action.
Stop and ask only when you can't continue without me, or before anything destructive: deleting data, force-pushing, or changing anything outside this repository.
A CLAUDE.md card titled “Tell it which stops you want” with two boxes. Keep going: when a step doesn’t need my input, keep going, and put status notes in the same message as your next action. Stop and ask: only when you can’t continue without me, or before anything destructive: deleting data, force-pushing, or changing anything outside this repository. Footer: “Edit it to fit your project. Keep permission prompts on for destructive commands too.”
FIG B The CLAUDE.md rule: when to keep going, and when to stop and ask.

If a run stops with “Want me to continue?” reply “continue.” If that happens often, the rule above will help.

A rule to keep going means fewer stops, so keep your own check before anything risky or hard to undo. The last line of the rule above does that. Keep permission prompts on for destructive commands too.

For pair programming, you may want the opposite: a one-line plan before it starts and a short recap at the end. Say that in your CLAUDE.md instead. Opus 5.5 follows either one.

Ask it to split big work across subagents

What to do. For an audit, a migration, or a review across a large codebase, ask Opus 5.5 to split the work across subagents and check each result.

Why it matters on Opus 5.5. Early testers had Opus 5.5 coordinate parallel subagents on long audits and migrations, with little oversight.

How.

PROMPT

Audit every service in services/ for the retry bug in the linked issue.
Give each service to its own subagent. When a subagent reports back, check its evidence before you accept it.
Finish with one table: service, affected yes or no, and the evidence.
Diagram titled “Give each service to its own subagent.” The prompt “Audit every service in services/ for the retry bug in the linked issue” fans out to four subagents. Their reports join at a “Check its evidence” step (“When a subagent reports back, check its evidence before you accept it”), then an arrow leads to “Finish with one table,” an empty table with the columns Service, Affected yes or no, and The evidence.
FIG C Fan out to subagents, check each one’s evidence, then finish with one table.

Keep the task list in a file

What to do. For a run that will take a while, ask Opus 5.5 to keep its task list in a file and update it as it goes. Then read the file, not the scrollback, to see where the run is.

Why it matters on Opus 5.5. Runs are longer now. A long run fills the context window, and Claude Code then summarizes older turns. A list in a file survives that, and it shows you at a glance what’s done and what’s left.

How. “Keep a checklist in TASKS.md. Tick each item when it’s done, and add anything new you find.”

3. CHECKING THE RESULT

Read what it needs from you first

What to do. When a long run ends, look first for anything Claude is waiting on you for, like a decision it left open or a change it wants you to approve. Then read the rest of Claude’s summary.

Why it matters on Opus 5.5. Opus 5.5 reports on its work more clearly than Opus 5. Its updates and its final summary say what it did, what it found, and what it needs from you, in plain language.

How. To change the summary’s format, say so in CLAUDE.md, for example, “End every run with three headings: Blocked on me, Changed, Found.”

Ask it to review the code

What to do. Ask Opus 5.5 to review a diff or a pull request before a person does.

Why it matters on Opus 5.5. One early tester said Opus 5.5 at its lowest effort caught more bugs than Opus 5 at high effort, with fewer false alarms. It also explains its changes in plain language, so its pull request descriptions are easier to review.

How. Feed this prompt to Claude:

PROMPT

Review the diff on this branch against main.
List only problems you'd block the merge for. For each one, give the file and line, why it's wrong, and how to show it fails.

Ask it to mark what it couldn’t confirm

What to do. For research and analysis, ask it to say what it couldn’t find or couldn’t check.

Why it matters on Opus 5.5. “I couldn’t find this” is worth reading, and asking for it makes it easy to find.

How. Add “Mark anything you couldn’t confirm, and say where you looked” to the request. This works in a Claude research report and in Claude Code.

4. IN CLAUDE APPS

First, check that the model picker says Opus 5.5.

What to do. Attach the chart, diagram, screenshot, or slide. Don’t retype the numbers.

Why it matters on Opus 5.5. Opus 5.5 reads charts, diagrams, and screenshots more accurately than Opus 5, and it needs no extra steps to do it. It’s also better at meaning that depends on where things are in the image: which boxes an arrow connects, what changed between two versions of a diagram, or when a meeting starts and ends in a calendar screenshot.

How. Attach the image and ask a specific question: “Which of these services call the billing API directly?”

Ask it to check a long document

What to do. Give it a long plan, report, or deck, and ask it to find mistakes.

Why it matters on Opus 5.5. Opus 5.5 pays more attention to detail than prior Opus models. In our testing, it caught a date that fell on the wrong weekday in a long planning thread, and a chart that didn’t match the numbers in a deck.

How. Submit the prompt: “Check this deck for anything that contradicts itself: numbers, dates and names. Quote each problem and say where it is.”

Ask for the finished file

What to do. When you want a spreadsheet or a document, ask for the file, not an outline.

Why it matters on Opus 5.5. The spreadsheets and documents Opus 5.5 makes need less editing than Opus 5’s before you share them.

How. “Make this a spreadsheet I can share: one row per vendor, with columns for cost, contract end date and owner.”

In a project, say when answers are settled

What to do. If follow-up questions in a long chat feel slow, add an instruction that earlier answers are settled.

Why it matters on Opus 5.5. In a long chat, Opus 5.5 sometimes goes back over an earlier answer while it thinks about a short follow-up. That slows the reply.

How. Add this to the project’s instructions:

PROMPT

Once you have answered something, treat that answer as done. Focus on what I'm asking now, and don't go back over an earlier answer unless I ask about it or point out a problem with it.

Leave it out of projects for long analysis, where a later step can show a mistake in an earlier one.

5. WHEN A MESSAGE IS FLAGGED

Opus 5.5 is the first Opus model to launch with Fable-level bio and cyber safeguards. In Claude apps and Claude Code, most flagged messages move to an older model, and your work goes on there. Finding security vulnerabilities in source code is allowed, and everyday health and educational questions should still work. These safeguards can sometimes flag legitimate work, and we’re tuning them to cut down on incorrect flags. If you’re switched, here’s what you’ll see and what to do.

In Claude apps

What you see. A notice that starts with “Switched to” and the name of an older model. Claude answers on that model, and the chat stays on it.

What to do.

  • To go back to Opus 5.5, choose it in the model picker. If the earlier message is still in the chat, it may be flagged again. Starting a new chat avoids that.
  • To be asked first, go to Settings, then Capabilities, and turn off “Switch models when a message is flagged.” You’ll see a “paused” card with your options.

The check covers everything in the conversation, including files and search results. So a flag can come from earlier content, not only your last message.

In Claude Code

What you see. A notice that names the older model. The session continues on that model.

What to do.

  • Run /model to switch back.
  • Press Esc twice to edit your last message and try again.
  • To be asked first, run /config and change “Switch models when a message is flagged.”
  • Run /feedback if the flag was wrong.

Don’t ask it to show its reasoning in the reply

What to do. Remove requests to reproduce its internal reasoning in the reply from your prompts and instructions.

Why it matters on Opus 5.5. A request to reproduce its internal reasoning in the reply can be declined. It’s one of the flag categories.

How. Ask Claude for what you need instead, for example, “Explain why you chose this approach in three sentences.”

6. SPEED

Turn on fast mode when you’re waiting on each reply

What to do. In Claude Code, use fast mode for back-and-forth work, where you read each reply before you send the next message.

Why it matters on Opus 5.5. Fast mode is available for Opus 5.5 at launch as a research preview. You get the same model, and the text arrives sooner. It needs extra usage turned on, and it costs more per token than standard mode.

How. Type /fast into Claude.

YOUR OPUS 5.5 CHECKLIST

Run through this before your next long task.

The checklist as a card with four groups of checkbox items: Asking, Long runs in Claude Code, Checking, and Flags. The same items are listed as text below.
FIG D The checklist at a glance.

Asking

  • The task says what “done” looks like
  • No “think hard” lines in prompts or saved instructions
  • Design requests list the styles to leave out
  • Charts and screenshots are attached, not retyped

Long runs in Claude Code

  • CLAUDE.md says when to stop and when to keep going, and to stop before anything destructive
  • Permission prompts are still on for destructive commands
  • Large audits and migrations are split across subagents
  • The task list is kept in a file

Checking

  • The “needs from you” part of the report is read first
  • A review pass runs before a person reviews
  • Research answers mark what couldn’t be confirmed

Flags

  • You know how to switch back: the model picker, or /model
  • “Switch models when a message is flagged” is set the way you want

Start building with Opus 5.5 !

With thanks to Molly Vorwerck for reviewing.

What Meta got right with Muse

Hacker News
metedata.substack.com
2026-10-03 14:23:48
Comments...
Original Article

🧱 I’m Mete, a product designer and builder with 10+ years of experience at companies like Netflix and Peloton. I build real products with AI to see how far the tools can go, then share what works, what breaks, and where craft still matters. I write for people who sweat the details and refuse to settle for good enough. If that sounds like you, you’re in the right place.

It’s been interesting watching the rise of Muse online as someone who’s been deep in the consumer agentic product space for a few years. There’s nothing fundamentally novel about it - all of its core components and capabilities have existed for over a year in most frontier agentic products and harnesses (Codex, Claude Code, Hermes, OpenClaw, etc.). It doesn’t do anything that you couldn’t do 6 months ago with a bit of tinkering.

But Meta seems to have put the pieces together in just the right way. The hype around Muse shows the enduring value of design, product marketing, and an ad-driven business model. It may also be the first time that agentic AI clicks for the average consumer, putting Meta in the lead for owning consumer AI in a way that can actually scale.

Quick update : OpenAI released dots - its always-on agents in ChatGPT - as I was writing this. I’m testing mine now and sharing early impressions on Threads , with a fuller write-up likely next week.

Muse from Meta ranked first in the App Store's free-app chart, above Momo and ChatGPT.
Muse has been topping the US App Store for some time.
this is nuts for a 1.85T company

Any leading agentic harness (Codex, Claude Code, OpenClaw, Hermes, etc.) has had the same agentic capabilities since 2025 - computer & browser use, cloud workflows, cron jobs and routines, mobile remote control, etc. But all of it has been compounding in complexity, resulting in apps like ChatGPT and Claude becoming a total mess .

Personally, I’m still committed to the ChatGPT ecosystem, but after using Muse, the clunkiness of ChatGPT is stark. The whole desktop ChatGPT vs Codex switch is dumb. The chat vs work switch is dumb. ChatGPT projects vs Codex local vs Codex cloud projects is mega dumb. I need to make 10 choices each time I want to start or continue a chat. Just give me one chat that orchestrates all of it in the background.

Collage showing six ChatGPT interface choices: ChatGPT or Codex, this computer or cloud or remote, Chat or Work, choosing a project, Quick chat or New chat, and model reasoning effort.
Infinite choices to make

Muse did just that and abstracted away the unnecessary complexity. There’s no model selector, work vs chat switch, obscure slash commands, or mentions of MCPs or cron jobs. It’s one main chat, and everything else serves to support it - a space to elevate ideas (although I think the Feed is redundant), track goals, and surface artifacts it created for you. Most importantly, it gave users a clear mental model - this is your personal helper with their own computer. It’s easy to understand, and it’s rooted in concepts people are already familiar with.

as simple as messaging with people

Funnily enough, a few months back I wrote an article titled “ AI needs better metaphors ” in which I laid out this confusing state of affairs in detail and emphasized the need for better metaphors, one of them being AI as a coworker (vs a set of confusingly named tools). I argued for that idea even earlier when all SaaS apps suddenly grew chatboxes and everyone rebelled:

The best framework that bridges the past and the future here is Jobs-To-Be-Done: customers don’t buy your tools - they hire them to get a job done. Each SaaS app was effectively a toolbox. Our job as designers was to understand the job at hand, surface the right tools in a given context, and make those tools easy to use so you can accomplish the job. With LLM agents, a SaaS app is now more akin to a coworker that goes out and does the job. And how do you talk to a coworker? You chat with them.

At the risk of tooting my own horn, those articles turned out to be pretty prescient. Because now Muse has put that to work quite literally:

One of the designers on Muse talking about making it feel more like an entity vs a tool

ChatGPT and Claude are in the business of selling tokens. And their pricing reflects that - unless you pay $100+/mo, you won’t get access to the best models and enough tokens to run useful agentic workflows. Most average consumers don’t even pay the $20/mo. As Benedict Evans pointed out , “usage is a mile wide but an inch deep.”

Meanwhile, Muse’s free tier is extremely permissive (and the ads are becoming pervasive) because it’s funded by the best ad business in the world. As of now, unless you’re a power user, you won’t even think about tokens or limits when using Muse daily.

That’s only possible because Meta is not in the business of selling tokens. They sell attention. And because they already do that extremely well, they can fund Muse as long as they need to capture the consumer mindshare.

ChatGPT’s and Claude’s product branding is geared towards techies. The products are sleek and the branding is aspirational, but to most people it comes off as sterile and elitist, at best. Muse went deep into anthropomorphization and the cuteness factor to help you build personal affinity with your agent. Personally, I find Muse’s whole avatar personalization gimmicky and unnecessary, but I’m also clearly not the norm.

Two Threads posts: @hello_neight’s Cilantro, a green furry Muse with glasses, and Jessica Hische’s Mister Business, three fluffy creatures in a trench coat. Two Threads posts: @hello_neight’s Cilantro, a green furry Muse with glasses, and Jessica Hische’s Mister Business, three fluffy creatures in a trench coat.
People love their muses
A Threads post by @_patters showing Rocky, a rock-armored Muse at a laptop, alongside its chat replies about rainy weather.
amuse

As it happens, I also wrote about AI personality as a new core design dimension :

I’m noticing a clear divide between those who rush to anthropomorphize their agents and those who want them to be cold, soulless tools. From a purely intellectual design perspective, I have argued that they are more akin to coworkers ...

This insight has one important implication - personality is now a core design dimension for AI products . In other words, it’s the flavor of the AI interface. And it’s clearly core to product adoption and engagement. There’s a reason most people who spend significant time interacting with these agents prefer Claude over ChatGPT. Capability is certainly one reason. But personality is arguably as important. The more likable, authentic, and human-like the AI personality is, the easier it is to develop a sense of connection. Connection drives engagement. Engagement drives stickiness. Stickiness drives profit. You see where I’m going with this.

Muse sensibly took the other side of the bet by giving agents personality - and it’s unsurprisingly clicking with consumers way more than the AI-as-a-tool-made-of-cold-hard-steel vibe. 1

Interestingly, what’s making Muse a success also explains why Google has mostly failed to capture the zeitgeist so far. They have the compute and the ad money to fund something like Muse. But they have no product chops to build something simple and understandable in consumer AI.

OpenAI fumbled the consumer AI market by not investing in ads earlier . When they realized most people won’t pay for AI subscriptions, they refocused on enterprise to compete with Anthropic. Their ads business may be growing , but they’re well behind - they still have no established cash flow to fund something like Muse (their new “dots” product is a paid-tier feature).

Apple fumbled it (for now) because they never built the expertise to build AI, blinded by iPhone success , privacy posturing , and Google bribes .

Now Meta will pick up the pieces ( literally and figuratively) - the pie is there for the taking. Muse is the first positive indicator.

Discussion about this post

Ready for more?

Apple Confirms iPhone 18 Pro Max AT&T Cellular Issues, Affected Devices Require Hardware Replacement

Daring Fireball
9to5mac.com
2026-10-03 14:21:14
In a statement to 9to5Mac, Apple said: We have identified an issue affecting a small number of iPhone 18 Pro Max users on the AT&T network that may cause a device to lose service and be unable to make calls. We released iOS 27.0.1 earlier this week and strongly encourage all iPhone 18 Pro...
Original Article

Apple has confirmed an issue causing a “small number” of iPhone 18 Pro Max devices on AT&T to lose cellular service.

Apple released software and carrier updates to prevent the issue from affecting additional users, but the company tells 9to5Mac that devices that have already lost service will require a hardware replacement.

In a statement to 9to5Mac, Apple said:

“We have identified an issue affecting a small number of iPhone 18 Pro Max users on the AT&T network that may cause a device to lose service and be unable to make calls. We released iOS 27.0.1 earlier this week and strongly encourage all iPhone 18 Pro Max users to update now, and are also issuing a carrier settings update today. These updates together are meant to help prevent this issue from occurring.”

The carrier settings update for the iPhone 18 Pro Max will download automatically over the next 14 days. However, users can manually trigger the download by going to Settings, choosing General, then About.

Apple explains that the software updates will help prevent future issues for iPhone 18 Pro Max users on AT&T. The software updates will not restore service on devices that have already lost it.

Apple says that iPhone 18 Pro Max users who have already lost service will need a hardware replacement. Those users should contact Apple Support or AT&T, or visit an Apple Store or AT&T location.

As we reported this morning , iPhone 18 Pro Max users affected by this problem can’t connect to AT&T’s network for calls, texts, and data. Instead, those users see “SOS” in their iPhone’s status bar. The connectivity issues aren’t impacting every iPhone 18 Pro Max user on AT&T.

Interestingly, the iPhone 18 Pro Max in the United States uses a Qualcomm modem . The iPhone 18 Pro, which uses Apple’s C2 modem, is unaffected by these AT&T connectivity issues.

Chance’s favorites:

Follow Chance : Threads , X , and Instagram

Add 9to5Mac as a preferred source on Google Add 9to5Mac as a preferred source on Google

FTC: We use income earning auto affiliate links. More.

ADHD, autism or complex trauma? [pdf]

Hacker News
www.cambridge.org
2026-10-03 14:08:28
Comments...
Original Article
No preview for link for known binary extension (.pdf), Link: https://www.cambridge.org/core/services/aop-cambridge-core/content/view/30CC4826561366615BFAEC807CDE28A7/S0007125026108046a.pdf/adhd-autism-or-complex-trauma-the-complicated-nature-of-the-question.pdf.

Hole Punch: Sling your spaceship around gravitational fields

Hacker News
notoriousbfg.com
2026-10-03 14:06:45
Comments...
Original Article

Hole Punch
Unit 0426 · Mk II

Treachery in the Rodin Museum 3D scan verdict

Hacker News
cosmowenman.substack.com
2026-10-03 14:01:27
Comments...
Original Article

[Note: I will be speaking at COMMUNIA Salon: The Rodin Case , on October 13, 2026]

The Treachery of Images, René Magritte, 1929

When France’s highest court for administrative justice hears your eight-year freedom of information case, and the court’s own legal analyst directly invokes surréalisme to argue that “sometimes a document is not a document,” you know you’re about to get wrecked.

In 2017 , I began writing to the Rodin Museum requesting that it identify and grant public access to 3D scans of Auguste Rodin’s sculptures, some of the most well-known and widely copied public domain cultural heritage works in the world. The museum’s administrators responded with a sustained campaign of lies, weaponized incompetence, and open lawlessness.

When Paris-based civil rights advocate Alexis Fitzjean Ó Cobhthaigh took up my case and sent the museum a formal request for its 3D scans, its senior directors quietly sought legal analysis from the French government’s own Commission on Access to Administrative Documents (CADA). The CADA’s opinion was in my favor, finding that the museum’s scans were undoubtedly administrative documents and therefore must be made accessible to the public. The museum’s director did not dispute the CADA’s analysis, but she confided to the Ministry of Culture, in writing , that she planned to ignore French freedom of information (FOI) law and make me take her to court.

In 2019, we filed suit in the Administrative Tribunal of Paris, where open culture and digital rights advocacy organizations Communia , Wikimédia France , and La Quadrature du Net joined me as co-plaintiffs. We were all represented by Fitzjean Ó Cobhthaigh, who had advised us that while the facts and law were on our side, we still faced an uphill battle, since the judges would be extremely deferential to the Rodin Museum and would grant it the benefit of any doubts. He was certainly correct, but I did not anticipate that in both the Paris court and, later, on appeal in the Conseil d’État, the judges would not only tolerate the Rodin Museum lying but would reward them for it. I did not imagine that showing them how the museum had lied to them on multiple issues would have absolutely no effect on its credibility on other issues. I did not expect that the Conseil d’État itself would not only decline to hear our testimony and evidence but would take up the task of inventing and promoting new, nonsensical arguments and falsehoods the museum itself had never dared to raise.

During litigation in the Paris tribunal, the Rodin Museum baselessly insinuated that I was a would-be criminal counterfeiter while pleading its own technological illiteracy and failure to preserve the documents in its care. The museum claimed that even with the help of (unnamed) outside experts it couldn’t figure out how to use its own scan documents. It dissembled about which sculptures it had scanned and falsely asserted that its written applications for public funding to produce those scans were non-existent.

It is no exaggeration to say that the Rodin Museum lied outright to the Paris court about the few scans it reluctantly admitted to holding. It claimed that its ultra high-resolution, state-of-the-art laser-scanned 3D point cloud documents, which are in open plaintext formats that are routinely used in the digitization and cultural heritage fields, were so technologically mysterious, and of such poor quality and so incomplete as to be virtually impossible to even visualize, and were thus unusable even by the museum itself.

Images of the museum’s point cloud documents that I obtained from third parties illustrate the museum’s fraud on the public and the court:

Visualizations of the Rodin Museum’s point cloud scan documents of Le Penseur, La Porte de l'Enfer, Les Trois Ombres detail, and Le Baiser

The Rodin Museum’s funding applications , which we forced the ministry to disclose in a separate FOI procedure and presented in court, also made the museum’s lies and misappropriation of public funds perfectly clear.

Not content merely to tell simple, substantive lies, the Rodin Museum submitted increasingly absurd and disjointed legal briefs to the court. For example, it simultaneously claimed that its scans were unusable and that they could be used by counterfeiters. In one brief, the museum confirmed that it held a scan of Les Trois Ombres , then in the next brief denied the sculpture had ever been scanned. (We replied with a photograph of the museum’s subcontractors scanning the sculpture.)

The Rodin Museum’s Les Trois Ombres being laser scanned

The Rodin Museum argued that they should benefit from every possible exception to French FOI law, even if they were mutually exclusive and had no connection to the facts of the case. We refuted them all.

The Paris court rejected all of the museum’s arguments relating to trade secrecy, counterfeiting, its business model and revenue, and intellectual property. In December 2023, the court ruled that 3D scans in a variety of formats were administrative documents and therefore must be made available to the public. In a significant victory for open culture, the judges directly ordered the museum and ministry to give me those scans and to compensate me €1,500 for my trouble.

But the Rodin Museum and the Ministry of Culture simply ignored the court’s order. To be clear, they did not appeal it, they ignored it.

Worse, the Rodin Museum’s buffoonish hand-waving about 3D scan technology apparently confused or impressed the judges, who ignored our world-class expert’s testimony on point clouds. In their decision, the judges improvised a novel exception to French FOI law, ruling, without any discernable legal basis, that 3D scan point cloud documents in open, plaintext formats could be withheld from the public.

In late 2023, we appealed that point cloud exception in the Conseil d’État, France’s highest court for administrative justice, where representation by specialist avocats au Conseil d'Etat et à la Cour de cassation is mandatory. Represented by Fitzjean Ó Cobhthaigh’s colleagues at the law firm SCP Marlange - de La Burgade , we detailed the lower court’s serious procedural and legal errors and its obvious misunderstanding of point cloud documents.

We presented testimony from experts in academia and industry and the arts.

We showed the high court that cultural heritage programs the world over routinely use and publish point cloud documents that the Rodin Museum claimed were unusable, and that the French government publishes petabytes of open source point cloud documents from its aerial 3D scan surveys of the entire French territory.

Our appeal explained to the court how point cloud documents are produced, and showed that the French government itself had spearheaded the development of free, open source software for viewing and using them.

We referred the court to the Ministry of Culture’s own published guidance on the proper formatting, use, and accessibility of open format point cloud documents in the cultural heritage sector.

Excerpt from the Ministry of Culture’s 2017 Methodological Guide: Metadata Description for Digital Acquisitions

None of this mattered.

In its reply to our appeal, the Rodin Museum reiterated to the high court its fantastical lie that point cloud documents are fundamentally unintelligible, citing as its sole authority the speculative musings of the lower court’s rapporteur public (the court’s own legal analyst).

The rest of the Rodin Museum’s briefs in the Conseil d’État were equally inept and detached from reality, so much so that we thought they might have been too stupid to warrant a response, and we seriously considered letting the case proceed to a decision without replying. But we felt a duty to create a record that repudiated the museum’s claims point-by-point to convince the judges of the absurdity of the museum’s misrepresentations. Also, we hoped to create the possibility that the high court judges might feel embarrassed to credit the museum's nonsense.

It didn’t work.

In the days leading up to our hearing, the high court sent us a short notice that it was, on its own initiative, raising an entirely new argument on the museum’s behalf, on which it was likely to decide the case, an argument the museum itself had not suggested: that point cloud documents were not even administrative documents.

We submitted a reply, but we could only guess at what we were rebutting since, in its wisdom, the court had not disclosed the basis of its new argument. We reiterated that the Rodin Museum’s publicly funded scans had clearly been made within the context of its public service mission, and that by all legal standards, precedents, and CADA analyses, they were clearly administrative documents. And we reminded the court of the well-established, conventional nature of point cloud documents.

At our hearing in early December, 2025, the high court’s rapporteure publique began her presentation by reminding the judges of René Magritte’s surrealist painting The Treachery of Images (This is Not a Pipe) , and explained that “sometimes a document is not a document.”

For additional dramatic effect she then read aloud the definition of “ document ” from the Dictionnaire de l’Académie française, then stumbled over her words when she seemed to realize it very clearly argued against her position. She asserted and continuously reiterated that point cloud documents are neither documents nor administrative documents. She offered her own assessment, too ludicrous for the museum to have suggested, that it would be too burdensome for the museum to deliver copies of its point cloud documents to me—documents it had already gathered and been sitting on for eight years. The rapporteure publique recited a variety of supposed exceptions to FOI law that did not fit the facts of the case and had already been examined and rejected by both the CADA and the lower court, and were not even subjects of the appeal. These included moral rights, trade secrecy, commercial reuse, counterfeiting, and economic competition, for example, the last of which drew approving nods from the judges. It was as if by merely naming these, without any evidence, relevance, or reasoning, she was presenting the judges with a buffet of options from which they could choose to reject our appeal.

In our opportunity to briefly respond with oral arguments, we directed the court’s attention to the example point cloud document we had submitted that morning: if anyone cared to look at it, it would demonstrate that it was, in fact, a document. We could have also quickly drafted and submitted a follow up written brief immediately after the hearing, but that seemed pointless. We had already said everything that could have possibly been said.

Our entire eight-year effort up to that point had generated argumentation, replies, testimony, correspondence, imagery, and hard-won evidence, which the court had compiled into a dossier more than 800 pages long. The public is not permitted to read that dossier, and in light of the way it handled our case, it appears that no one at the Conseil d’État has read it either.

The Conseil d’État issued its written decision a few weeks after the hearing, and it went well beyond what its rapporteure publique had suggested. It ruled that the Rodin Museum’s 3D scans of cultural heritage works are legally indistinguishable from physical reproductions. The museum’s scans are not administrative documents but part of its inalienable collection, and French FOI law is therefore inapplicable.

By implication, the high court’s ruling means that the government’s own FOI experts at the CADA had repeatedly erred in their analyses in my favor, as had the Paris court. Our earlier victory on other 3D scan formats was effectively undone. There would be no need for the Rodin Museum and Ministry of Culture to heed the lower court’s order that they had already chosen to violate.

The Conseil d’État judges noted that since FOI law did not apply, there was no need to consider the facts or circumstances of the case.

The judges threw out our appeal, ordered me to pay €3,000 to the Rodin Museum, and did not even bother to send me written notice of their decision.

While I’m disappointed in the court’s decision, I am proud of our work, and I appreciate how the case resolving this way achieves a strange sort of perfection.

We had begun this project by asking for permission to access publicly funded digitizations of public domain cultural heritage works, then asserting our rights to them, and methodically following every rule and procedure and official channel, and relying on the law and the truth. Confronting the Rodin Museum and Ministry of Culture with these, despite a multi-billion euro resource disadvantage, we forced them to explain and defend their policies, reasoning, and outlook , and they could not do it.

Like the Rodin Museum, other arms of the Ministry of Culture talk a good game about digitization and accessibility, but refuse to deliver. French FOI law looks good in principle, but in practice it can easily be abused by a hostile and lawless administration to exhaust petitioners’ resources and create indefinite delays until, in the final hour, a deus ex machina intervention delivers the outcome it wants.

In this case, that outcome is for the preservation, dissemination, and benefits of digitization of France’s cultural heritage to be entrusted to the care of unaccountable administrators who prize secrecy and exclusivity over accessibility, who lie without shame, who are shielded by the law but unconstrained by it, who forgo even the appearance of adhering to its procedures and orders, and who don’t hesitate to shift from proclaiming their authority to pleading their incompetence and illiteracy whenever it suits them.

The court has upheld this tableau surréaliste , but no one can justify it.

Cosmo Wenman is an open access activist and CEO of Concept Realizations, LLC . He lives in San Diego. He can be reached at cosmowenman.com and cosmo.wenman@gmail.com

Copyright 2026 Cosmo Wenman

On 13 October, 14:00 to 15:00 CEST , join us for a new COMMUNIA Salon exploring the French “Rodin case” and a deceptively simple question at the heart of access to cultural heritage: Should 3D scans of works in the Public Domain be freely available for reuse, including for commercial purposes?

During this COMMUNIA Salon, we will unpack the legal twists and turns of the Rodin case and discuss what the decision means for the Public Domain, cultural heritage institutions and the people who want to access and reuse our shared cultural heritage.

Cosmo Wenman , Alexis Fitzjean ó Cobhthaigh , Benjamin Jean (INNO3) and Brigitte Vézina (Creative Commons) will discuss the case, with COMMUNIA member Camille Françoise moderating.

Discussion about this post

Ready for more?

Pop!_OS bans AI-generated code from much of its codebase

Hacker News
www.neowin.net
2026-10-03 13:57:03
Comments...

LeCun has "zero concerns" about AI wiping out humanity, recent "rogue" incidents

Hacker News
fortune.com
2026-10-03 13:44:29
Comments...
Original Article

Yann LeCun won the 2018 Turing Award—computer science’s equivalent of the Nobel Prize—along with fellow AI researchers Geoffrey Hinton and Yoshua Bengio for their work on deep learning, the neural network technology that underpins today’s AI boom. But today, nearly a decade later, LeCun is the only one of the three AI “godfathers” who is not deeply concerned about the risks of artificial intelligence. In fact, he thinks many of those risks are overblown, and that the constant warning about them from executives at some of the leading AI companies is misguided and counterproductive.

LeCun isn’t worried “at all” about AI wiping out humanity, and he has “zero concerns” about the recent string of rogue AI incidents, including OpenAI’s agents autonomously hacking Hugging Face in July. He attributes the incidents to poor human oversight and system design, and says they’re “totally preventable,” a belief also held by U.S. Treasury Secretary Scott Bessent, who this month called the Hugging Face incident the “responsibility of OpenAI management.”

“Those agents are doing exactly what they’ve been asked to do,” LeCun said. “They were supposed to be in sandboxes, but the sandboxes were leaky and horribly designed.” Many AI labs lack a fundamental understanding of cybersecurity, he said, something an OpenAI safety researcher also called out this week as one of the main reasons AI may cause “great harm to the world.”

Many people working in AI safety “usually have an agenda to push,” LeCun says, and then clarifies that he’s talking about effective altruism , or EA, the philosophical movement that has been obsessed with the risks AI poses to humanity. EA was little known among the general public until it made mainstream news headlines in recent weeks in the wake of a viral social media post from a former Anthropic and OpenAI researcher named Jacob Coxon, who claimed the employees at these companies “earnestly believe [AI] could kill us all by the end of the decade.”

LeCun thinks EA is “super toxic” and a “complete disaster.” Its adherents who are working in AI labs suffer from “paranoia” that causes them to make poor decisions, he said. “Apparently people are having mental issues.” This month, the Financial Times also reported that some staffers at the U.K.’s AI Security Institute, as well as at OpenAI, Anthropic, and Google DeepMind, have sought counseling, taken time off work, and spoken publicly about experiencing distress because of fears their work could cause serious harm.

Amodei is ‘deluded’ and ‘crazy,’ LeCun says

Anthropic CEO Dario Amodei and many of the company’s founding staff members are known to be sympathetic to EA ideas and to have attended EA events in the past, although Amodei has denied being an EA adherent and Anthropic says its employees represent a diverse range of views. LeCun noted that Amodei’s sister, Daniela, who is also a cofounder of Anthropic and the company’s president, is married to Holden Karnofsky, who cofounded two EA-aligned philanthropies, including Open Philanthropy (now called Coefficient Giving). Karnofsky was also a member of OpenAI’s board from 2017 to 2021.

“Dario tries to distance himself from Open Philanthropy, but he’s totally into it,” LeCun said. “I think he’s completely deluded.” Later in the interview, he calls Amodei “crazy.”

Amodei and OpenAI CEO Sam Altman “saying, ‘AI can kill us all’ is incredibly destructive for everybody, including for the public who’s scared shitless,” LeCun said. It’s also bad for the AI industry, and “the worst marketing campaign you can possibly imagine.”

He’s right: The public is scared shitless. LeCun and I are chatting at a circular table for two in a posh café in New York City’s SoHo neighborhood, just around the corner from the offices of his new company, AMI Labs. We’re drinking tea, although, as the Frenchman pointed out, our evening time slot was more suited to apéro , or drinks and light appetizers. But he had an anniversary dinner with his wife afterward, so tea was just fine, he said.

The shop is about half full of other patrons, nearly all are working on their laptops. When LeCun and I chat with one of them, I mention to her that LeCun is a luminary figure in AI. (She didn’t recognize him.) Then she immediately asks, “So, is AI going to kill us?”

LeCun is no fan of Trump. But on AI policy, the two agree on much

The public’s fears have become so widespread that President Trump has started rebranding artificial intelligence in a way he feels will instill more confidence. He signed an executive order this week renaming it “super intelligence,” and for much of this month his administration’s social media accounts posted anti–AI doomerism messages.

“Americanism, not effective altruism,” the Department of War tweeted in mid-September.

Meanwhile, most of LeCun’s social media presence is dedicated to condemning the Trump administration, particularly the way it has slashed research budgets that are necessary to promote innovation and waged a “war on science.” But when it comes to AI fears, the two are unusually aligned.

LeCun is more concerned about AI regulatory capture than human extinction. Regulatory capture refers to the idea that dominant companies can shape government rules in ways that lock in their market position, making it difficult or impossible for smaller competitors to arise. EA-backed groups are active lobbyists for AI safety bills on Capitol Hill, as are many of the major tech and AI companies. LeCun, in contrast, doesn’t think AI requires new regulation.

Amodei is “honest,” LeCun concedes. “He’s trying to do the best thing for the company, which is close to its IPO. But claiming AI is too dangerous to put in our hands, and saying it should be regulated, and saying open-source [models] are too dangerous—this is regulatory capture that would have a terrible effect if it’s followed by acts of Congress.”

Building AMI Labs

In his day-to-day, LeCun’s main focus is building his new company, AMI Labs, which he founded after leaving Meta after 12 years, seven of which he served as chief AI scientist. He remains on good terms with the company, and has started using its new Muse AI agents. He named his Hal, after HAL 9000, an AI character in 2001: A Space Odyssey .

The café where we met was around the corner from AMI Labs’ New York office, although the company is headquartered in Paris with staffers in Montreal and Singapore as well. The company’s name is pronounced as one word, “ ami ,” or “friend” in French—and since going public in March 2026 its now 60-some employees have been building a technology called world models, which LeCun has long argued will eventually supersede large language models as the primary system underpinning AI.

In particular, AMI is building world models that leverage JEPA (Joint Embedding Predictive Architecture), a neural network architecture that LeCun pioneered and that teaches models to predict data in a representational space within a neural network’s middle layers, rather than generating raw pixels, as many competing world models do, or words, as LLMs do.

The company’s primary focus, for now, is industrial applications. “It’s AI for the physical world, so it’s not language-related,” LeCun said. “It’s systems that understand the real world, like a manufacturing plant or turbojet engine.” Some of the main applications are anomaly detection or robotics: “If you have a machine and all of a sudden it makes a strange noise and starts breaking, you would’ve wanted to detect that as early as possible.” He also gave the example of a system that might understand the world like a cat does, for example, which knows that if it pushes a vase off the counter it will fall.

AMI Labs is LeCun’s big chance to prove he’s been right all along. He stayed mum about the company’s first product, only mentioning that it’s debuting “soon.” Some open models may be involved, he said, given that he is a big proponent of them.

“The thing I want people to know is that AI is not over,” LeCun said, then hailed a cab to get to his dinner.

Automating my 35mm film scanning pipeline

Hacker News
shannadige.com
2026-10-03 13:35:42
Comments...
Original Article

View the experience

darkroom.shannadige.com

I love photography. There is a special feeling to seeing an image through a lens, organizing it to compose a story, and working with the resulting photo. I had cycled through DSLR and mirroless cameras before falling in deep love with the Canon AE-1P, an analog SLR that my dad traveled with around the world before settling down in the States.

Film photography is simple, but particular: you expose reflected light against a strip of plastic and burn-in a chemical image. Depending on the light conditions, film stock used, and lens, you’ll get magically unique compositions.

For years my process looked like this: load film and take photos, send the film to a lab and wait for development, pick up said film, cut the negatives into strips, load them into a holder, capture the negative of each frame in with a specialized scanner, open every file in darktable, invert it by hand, make topical edits, then decide where to store and/or share the exported image. At first this was an exciting process for me. Then as the number of rolls increased, so did my effort to do this for 36 frames per roll.

A desk lit by a lamp, with a strip of negatives hanging from it, VueScan open on a monitor showing an orange negative, and film holders, a blower and a brush next to a blue Plustek scanner.

On average it would take me about 4 hours to fully scan in a roll of 36 photos and complete edits. It requires delicacy in handling film especially to avoid dust or scratches on the final image. Safe to say, the process grew weary on me quite fast. I was spending more time processing my images than actually taking them. So I began looking for ways I could cut some time, or as the CEOs like to say, “streamline operations” across my personal digital lab.

Giving Claude destructive hardware capabilities

My first hypothesis was if I could take myself out of the loop entirely. I’d have Claude drive the scanner over SANE (open-source scanner driver stack) and handle scanning, colour conversion and edits end to end.

On the first scan, Claude used SANE with inversed X/Y scanbed directions and nearly broke the gear mechanism inside the scanner. So I unfortunately killed time by needing to perform an emergency surgical operation on the scanner to save $200 from being spent on a new one.

Chat: Blank white scans

Sep 20, 2026

  1. Shan 4:54 pm

    there is definitely film in there, maybe its the scanning settings you opted for

  2. Claude 4:56 pm

    I've been tracing why the scans come out white, and your guess about settings can be ruled out. … Instead, every one of the 856 rows is identical: smooth vertical bands from the lamp's uneven illumination, with zero variation from top to bottom and no image detail. … So the sensor saw the same unobstructed line of light for the whole scan.

  3. Shan 5:01 pm

    i tried with vuescan, but the sensor or the carriage is not moving, you must have done something wrong

  4. Shan 5:04 pm

    its definitely in a bad state. can you check on the scanner, until i let you know its moving. i did hear some clicking earlier, so it might have been going an opposite direction

  5. Claude 5:05 pm

    I found what the clicking was, and it came from my scans. … At the end of every scan, the genesys backend sends the carriage back with a blind reverse move: 7010 steps at 900 dpi, and more at 3600. It assumes the carriage travelled forward during the scan. … So each scan ended with the motor driving a carriage that was already parked backwards into its end stop. That is the clicking you heard, and your read on the direction was right. I ran seven scans, so it happened seven times.

The Plustek scanner with its case removed, showing the carriage, a threaded rod, a spring and the motor gear.
Surgery. Pushing a rod and gear that did not want to be pushed back into place.

I ended up working with what was already working well. VueScan was the software I’d licensed way back when - given that it had the best compatibility with my scanner, and given that I didn’t have a robot arm automating the physical operation of loading and scanning, I’d just use this myself and find a solution that can handoff from right after this.

Studio

Learning from failure fortunately turned my thought process around. I work in VueScan the way I always have, so I needed some way to automate the processing of those scans. After a handful of individual tests against image data I had tested against, I had a full suite of tools working in combination to help me.

Negative Conversion

Converting negatives into positives is mostly plain code. The inversion is a short numpy routine, that works in film density and measures color across the whole roll before adjusting each frame within limits of its negative. So thirty-six frames look consistent to one full roll and a sunset doesn’t get radically color-corrected. This was something I tended to struggle with even doing manually. My agent determined that using the blank leading slot of a roll would be a good reference for how color correction should be weighted against the rest of the frames. Before building anything else, we ran it on an old roll as a go or no-go test.

Chat: Conversion test

Sep 21, 2026

  1. Shan 10:35 am

    start the conversion test on film/valentines2026

  2. Claude 10:40 am

    The conversion test passed on 28 of 29 frames. Custom conversion works well enough to build on, and it needs no darktable to get there. … Mine are brighter and more neutral, while your exports lean teal and dark (the Hotel del Coronado sky, the beach portrait, the garden). That is taste as much as correctness, so you should judge.

  3. Shan 10:40 am

    i like your direction, run the dust detection test

Frame 2 as scanned, an orange-masked negative of the portrait with the giant paper flower. Frame 2 after conversion: the portrait with the giant paper flower in full colour.

Frame 3 as scanned, an orange-masked negative of the “Can you dig it?” sign. Frame 3 after conversion: the “Can you dig it?” sign in full colour.

Frame 5 as scanned, an orange-masked negative of the yellow orchids. Frame 5 after conversion: the yellow orchids in full colour.

Frame 20 as scanned, an orange-masked negative of the frozen lake. Frame 20 after conversion: the frozen lake in full colour.

Sample negatives vs. converted positive (those scratch marks are characteristic of the film stock)

Dust Removal & Editing

The models come in where plain code runs out. Dust is found by image analysis on the negative and marked for removal. This is harder than it sounds: on a beach frame the first detector flagged about 2,000 sun glints on the water as dust. The fix was to affix context and act only where surroundings are smooth. The spots it finds are filled by LaMa (inpainting model), with detail copied back in from nearby film spots so the patch doesn’t look smudged. Kind of like the repair tool on Photoshop.

Not every single area needs to be filled. I have a custom threshold set where I can hold grainy spots for review, and approve/deny them individually or in batch.

This also created a need to adjust light and color properties of each image individually to account for graininess, so I had exposure, contrast, temperature, and other controls embedded in reach.

The frame review screen in studio: a frozen lake under a pale sky, covered in circles marking dust and scratches, with sliders and dust counts in a panel on the left.
Frame review. Green circles were dust-repaired automatically, amber ones held for review.

Metadata

Indexing data for analog film can be time-consuming, especially if I’m working across separate tools on my own. This was a “hmm, why not?” problem to address since now that I’m using technology I can extend it to identify a photo for easy searching later on.

So Qwen3-VL runs locally in Ollama when scans come in, looks at each finished frame and drafts tags for it in about five seconds. That combined with camera, lens, date, and film stock data gives me a comprehensive library to easily search from.

The Gallery tab of the review panel beside a portrait of a woman in front of a Green House Motel sign, with drafted tags: woman, sign, motel, plants, flowers, greenhouse.
Tags drafted by Qwen3-VL in 3 seconds, fixed once I pick the frame.

My part is now the part I can thoroughly enjoy, which is looking at the results I’ve snapped and finding inspiration to snap more.

The studio app with valentines2026 selected in the sidebar's list of roll folders: a roll header with Revolog Rasp 200, the Canon AE-1P and notes about the Chicago Botanic Garden's orchid show, and a contact sheet of all 28 converted frames marked live or check.

Once I had gotten a few rolls through the pipeline, the next question was where they’d live.

Typically they’re stored on my 4TB HDD that I’ve set to eternally spin while plugged into my machine. But while off the rush of building enterprise-software-for-one, I wanted to explore a way to share via the world wide web.

Publish Workflow

I have a Cloudflare account that gets the occassional activity from visitors on some CDN content I have for my personal site, and what better way to max out the free plan than uploading all my images to my own R2 bucket? So I extended the studio to enable publishing.

The publish dialog for the valentines2026 roll over the studio: label Valentines 2026, stock Revolog Rasp 200, date 2026-02, a private-link checkbox, and all 23 picked photos numbered in the order the gallery shows them.
Publishing a roll. The picked frames go to R2 in the order shown, with their tags and nothing else.

Showcasing Photos by Roll

A grid of thumbnails felt too cheap for a publish experience. I continued to scan more rolls throughout the days and recounted my earlier times dropping off my rolls for development. Seeing the lab technicians unroll the film itself to spool to their C41 processor gave me a simple idea to emulate the roll itself.

I do acknowledge that this is not the most accurate form of realism, but it was a very fun idea to play with and evolve. I ended up with a felt table covered in film canisters modeled for 3js, each with real label art and random scuff marks. Pick one up and you can yank the film out to expose the frames. Then it just falls back onto the felt table.

I was able to source the effective

Demo walkthrough; credits to Opus-5.5 for recording

My corner of the Internet

In total this project took 11 days and 64 commits. The gallery has 22 rolls and 314 photos dating all the way back to 2020.

I will continue to add to this as I purge more money on rolls and development. It’s my corner of the Internet, for you to unplug and enjoy if you choose to revisit.

There’s been something just as alluring about creating a site like this, versus just engaging the scanner and process itself. The manual process is something that someone might enjoy and wouldn’t want to delegate. At least for me, it keeps me off the computer and focused on enjoying the time I have exploring the world. Capturing it in a different perspective. ■

Shan in a wide-brimmed green hat and a white shirt, looking down at a black Canon film camera held in both hands, on a tree-lined road with a van and a stall of king coconuts behind him.

Vx – One Language, Every Chip

Hacker News
vxlang.org
2026-10-03 13:23:20
Comments...
Original Article

Vx is a systems programming language for heterogeneous computing. CPU, GPU, NPU and accelerator memory are part of the type system — so a host thread dereferencing a device pointer is a compile error , not a segfault at three in the morning.

curl -fsSL https://vxlang.org/install.sh | sh

macOS on Apple Silicon and Linux x86_64. Other install options .

Heterogeneity belongs in the type system, not in the runtime.

Where data lives is part of its type

Most languages treat the accelerator as infrastructure: you write math, and a large opaque runtime decides how to ship it. Vx treats it as semantics. A tensor pinned to NPU high-bandwidth memory has a different type from one in host DRAM, and crossing between them takes an explicit transfer() — even when the hardware boundary is free.

On Apple's unified memory that transfer compiles to almost nothing. It is still written down, because data locality should be provable by reading the source rather than by profiling the binary.

// Two matrices already resident in NPU memory.
fn custom_matmul(
    a: Pinned<Tensor<f32, [4, 4]>, Topology::NPU[0]>,
    b: Pinned<Tensor<f32, [4, 4]>, Topology::NPU[0]>)
    -> Verified<Tensor<f32, [4, 4], Memory::NPU_HBM>> {

    let mut result =
        Tensor<f32, [4, 4], Memory::NPU_HBM>::uninit();

    // Dispatch the computation to the accelerator.
    spawn on(Topology::NPU[0]) {
        for i in 0..4 {
            for j in 0..4 {
                result[i][j] = 0.0;
                for k in 0..4 {
                    result[i][j] += a[i][k] * b[k][j];
                }
            }
        }
    }

    return Verified(result);
}

What the compiler rules out

Vx front-loads into type checking a class of bug that normally surfaces as a runtime crash, silent corruption, or an out-of-memory at training step 1200.

Address-space typing

Dereferencing a device pointer from the host. A Pinned<T, NPU_SRAM> escaping into a host expression.

Capacity admission

A placement whose working set cannot fit the memory space it targets — checked against the declared machine, before a binary exists.

Seam contracts

Reading a buffer whose asynchronous transfer has not been made visible. Discharged by an SMT prover.

Linear types

Use-after-move of a consumed buffer, alongside a borrow checker with variance and region tracking.

Topology reachability

A transfer between two memory spaces with no declared path between them.

Autodiff

Differentiating through a region whose adjoint is not defined.

The machine is declared, not assumed

Most compilers hard-code a cost model. Vx reads one. A machine file describes the memory hierarchy and interconnect of a real part, and the compiler admits or rejects placements against it.

Units are exact integer conversions, never floats: SI prefixes are decimal ( GB = 10 9 ), IEC are binary ( GiB = 2 30 ). A figure copied off a vendor sheet means what the sheet meant.

The repository ships machine files for H100, H200, B200, A100, MI300X, Apple M4 and multi-GPU nodes — each citing its sources, and marking unverified figures as unverified.

Memory HBM  { capacity: 80 GiB, bandwidth: 3.35 TB/s,
              managed: explicit, scope: device }

Memory L2   { within: Memory::HBM, capacity: 50 MiB,
              bandwidth: 12 TB/s, managed: cached }

Memory SMEM { within: Memory::L2, capacity: 228 KiB,
              bandwidth: 128 B/cyc, clock: 1.98 GHz,
              replicas: 132, granule: 1 KiB, scope: sm }

Topology Device {
    arch: nvptx64,
    memory: Memory::HBM,
    transfer Memory::CPU_DRAM -> Memory::HBM : 63 GB/s,
}

How it compiles

A data-oriented parallel frontend

Every symbol, nominal type and monomorphized variant is a flat 256-bit identifier. A nominal type system plus mandatory boxing for recursive types decouples modules, so the pipeline runs parallel across cores with no query engine and no lock contention. Compilation walks flat arrays rather than pointer-chased trees.

The same source compiles to byte-identical MLIR whether it is built serially or in parallel. That is asserted in the test suite rather than hoped for — at benchmark scale, at one thread, at four, and with the thread pool taken off the path entirely, plus a corpus recompiled in fresh processes so each run gets its own hash seed. The claim is about the MLIR the frontend emits; everything downstream of it belongs to LLVM.

Backends

CPU (x86-64, arm64) MLIR → LLVM IR → native, AOT or JIT
NVIDIA GPU MLIR → NVVM → PTX → SASS
Apple AMX / ANE CoreML primitive dispatch via plugin
Distributed Manifest-driven remote regions over a wire protocol

Vendors extend the compiler through MLIR pass plugins rather than by patching it.

Where Vx is the wrong tool

PyTorch users mutate architecture mid-loop, print a tensor shape, branch on it, and carry on. In Vx — ahead-of-time, data-oriented, statically regioned — that same dynamism takes real work.

Vx is the right language for the thing that must be correct and fast across ten kinds of silicon. It is not the right language for the thing you are still figuring out.

Start here

Kolibri – Tech Report [pdf]

Hacker News
aleph-alpha.com
2026-10-03 13:22:48
Comments...
Original Article
No preview for link for known binary extension (.pdf), Link: https://aleph-alpha.com/downloads/tech-report.pdf.

Agents don't need memory, they need documentation

Hacker News
liao.gg
2026-10-03 13:03:37
Comments...
Original Article

A tangle of colorful wires

A memory plugin analyzes your conversations. It generates 1,000 isolated snippets and inserts them into a vector database. With your every prompt, it attaches the five most similar snippets; if the agent is confused (which it is), it manually searches for more. That’s the product they call “memory.”

It’s a strange thing to use when you think about what problem you’re trying to solve. You want your agent to understand your project. To know where a feature is, why it was built, what you agreed on and what you care about. Instead, you get a lottery over RAG snippets, injected on every prompt, hoping the right ones float up.

Even when it does work, the agent still doesn’t understand your project. The entire memory plugin ecosystem is solving the wrong problem.

Because agents don’t need memory. They need documentation.

It’s All Just RAG

Every memory plugin on the market works the same way:

  1. Go through session transcripts
  2. Generate snippets of “memories”
  3. Insert into a RAG database
  4. On every prompt, retrieve the top 5 and inject
  5. Need more? Give the agents a tool to search through the RAG database

That’s the whole architecture. Some tools are extra fancy; they let the agent search through past transcripts word for word. Or they implement some kind of multi-tier memory system that classifies short or long-term memory. Or they add a bunch of background daemons to review, merge, or deduplicate memories. “Dreamers” that rewrite memories overnight. Continuous context compression. Rerankers. Et cetera.

Each plugin tries to add new token-burning “features” to fix the same flawed architecture underneath. And that’s why none of them reliably work.

The Problem with Recall

All of these memory plugins suffer from the same broad set of problems.

  • Memories are surfaced by similarity. Similarity search ranks how close two snippets are in embedding space. That’s it. You don’t know which is correct, current, or what’s missing.
  • Memories are stored without context. A RAG snippet can only contain so much. You lose everything else: context, motivations, lessons, environment, and more.
  • The past is treated as truth. All of these plugins rely on recall; whether it’s search through transcripts or a vector database. But the codebase changes everyday; so how accurate are each of the 500 snippets about authentication?
  • Agents can’t search for what they don’t know. Even if you do expose a search tool to the agent, how would the agent know when to use it? The agent doesn’t know what it doesn’t know.
  • The store is unauditable. There are 10,000 embeddings in SQLite. Which memories exist? Which are stale? Which have never been retrieved? Which are incorrect and secretly affecting the way your agent works?

These are only five of the many problems that memory plugins face, attempt, and fail to fix, because they all make the same assumption:

Agents forget: that’s the problem. So the fix is to remember. To remember better, we should capture more, index better, retrieve smarter.

Their entire thesis revolves around capturing and recalling the past. But that’s not how anyone else handles knowledge. No one rewatches a team meeting from 3 years ago to remember constraints around a feature. People write things down and use those records instead.

Similarly, the solution is NOT to give an agent a search tool that it uses to search over the past 10 million tokens worth of conversations, reconstructing fragments of what happens.

The solution, is document-based memory.

Today, people use AI to pump out products, features, and slop at the speed of light, all without reading or understanding a single line of code. Under these circumstances, it’s easy to see how documentation ends up becoming an afterthought when it should be more important than ever.

Documentation Over Recall

People already knew agents needed context. That’s why they invented AGENTS.md files: so agents don’t jump into a codebase blind. It works. But oftentimes, that single file is the ONLY documentation that the project has.

A single file is not enough. The agent needs an entire brain, a structured workspace where it can record instructions, specs, decisions, research, indexes, anything without being asked. Instructions for how the code review process works, specs detailing what was discussed with the user, reusable research into a foreign library or API.

When the agent works, the agent can read files from the brain to gain relevant and complete context. After the agent works, it updates what’s outdated, adding new documents where necessary, while the full picture is still in context. In this way, the agentic loop changes from prompt → build → forget to prompt → consult → build → update . Memory turns from a RAG database you bolt onto the agent into a workspace you can read, update, and even share.

Putting It to the Test

I recognized this problem over a year ago when I first started programming with AI. I wanted a way for an agent to remember work across sessions, so I started by creating an internal/ folder where I asked the agent to write everything down: specs, plans, indexes. I tasked the agent with always reading the appropriate documents and index before doing work and updating afterwards.

This set of rudimentary instructions slowly transformed into a formal system, and then into a plugin called Operator Memory that I’ve been using regularly within all of my projects.

Diagram comparing the prompt, build, forget loop with Operator Memory's prompt, consult, build, update loop

Operator Memory provides document-based memory using the model described above. Operator provides your agent with a Markdown brain where it can persist important knowledge: instructions, specs, research, indexes. Before working, the agent always consults the brain for relevant documents. After working, the agent updates the brain; revisiting stale documents, adding documents where missing.

No vector databases. No embeddings. No summarizers, curators, updaters, dreamers, or any other token-burning background daemon. No black-box retrieval.

With Operator, everything is a plain Markdown document that you can read, update, commit, and share with your team. I’ve been using this system for over a year. If you want to try it out, it’s free and open source: https://github.com/aerovato/operator-memory

RetailReady (YC W24) Is Hiring

Hacker News
www.ycombinator.com
2026-10-03 13:00:12
Comments...
Original Article

An AI-powered supply chain compliance engine

Implementations

$100K - $140K • 0.02% - 0.06% • San Francisco, CA, US

Experience

Any (new grads ok)

Connect directly with founders of the best YC-funded startups.

Apply to role ›

About the role

We are RetailReady (YC W24) - an AI-powered supply chain compliance engine. We’ve raised $6.6M in funding, onboarded over 50 customers, and in the past 12 months we’ve 4x’d our revenue. We are building world-class software that will disrupt an antiquated industry & we want you on board.

What we’re doing here matters.

We’re tackling everyday challenges for real people in the supply chain – that’s the heart of every business. With us, you’ll see your work come to life in amazing ways. Imagine being a part of why Walmart can restock baby formula, helping a health & beauty brand go from startup to Target shelves, and empowering warehouse workers with the latest tech. That’s the impact you’ll have.

This isn’t just another job offer - it’s an opportunity to create your own path and make a tangible difference.

What’s cool about retail compliance?

  • A majority of supply chain systems are archaic - bad UX, weeks to onboard, and old tech stacks
  • There is so much to disrupt in the supply chain space (we call it an engineer’s playground)
  • Compliance is fragmented across brands, warehouses, and retailers & we are first to market to solve operational retailer compliance within a warehouse

Why work with us?

  • Our team has operational, product, and engineering backgrounds across multiple industries, as well as deep supply chain industry experience – at Stord (a supply chain unicorn startup), Microsoft, and Manhattan Associates
  • You’ll be joining an awesome team at the start:  we’re low in ego, high in intellectual curiosity, and aren't afraid to make mistakes but above all we win together as a team!
  • We’re fun to work with - just watch any of our Weekly Round Up Videos on LinkedIn and you’ll see

Position Overview:

  • Lead the deployment of our software in warehouses and across tech stacks for Brands and 3PLs alike
  • Travel to customer locations for up to ~ 50% of the month
  • Collaborate with our engineers to conduct quality assurance on features
  • Provide insights from on-site experiences to influence the product roadmap

This position is not for the faint of heart - we have ambitious plans for RetailReady, and we want someone who shares the same level of ambition, and has the tenacity & passion to stick with us for the long haul.

We have an in-person team in San Francisco. If you’ve gotten this far and are fired up, reach out.

About the interview

  1. 15 minute intro call
  2. 30 minute panel
  3. Founder call
  4. Hired!

About RetailReady

RetailReady is building an AI-powered supply chain compliance engine. Supply chains are still heavily reliant on paper processes and tribal knowledge, causing costly shipping mistakes that jeopardize the longevity of businesses. RetailReady is the first-to-market with our retail compliance packing software, leveraging camera vision to direct warehouses to ship orders without error. We are positioning our compliance data models to become the operating system that will power the next wave of warehouse robotics and automation.

RetailReady

Founded: 2024

Batch: W24

Team Size: 12

Status: Active

Location: San Francisco

Founders

Docker has always used microVMs (well since 2016)

Hacker News
dave.recoil.org
2026-10-03 11:55:31
Comments...
Original Article

Docker has always used microVMs (well, since 2016)

There's a lot of buzz about "microVMs", where a workload runs with a stripped down Linux kernel, on a minimalist VMM such as firecracker (announced at re:Invent 2018 ) on top of a hypervisor like KVM or Xen. MicroVMs are often contrasted to, and considered more secure than, traditional Linux Docker containers. What if I told you that Docker Desktop has always used microVMs?

Always has been meme: wait, Docker uses microVMs? Always has been (since 2016)

Back in 2015

Docker Toolbox logo When I joined Docker in 2015 the state of the art was Docker Toolbox . It used VirtualBox , which is a great product with lots of features. VirtualBox has its own GUI and its own update process; way more than we needed for Docker.

A library VMM

We wanted Docker to feel like a native app on Mac and Windows, rather than a bundle of components. Using our Mirage Unikernel libraries we started building a "library VMM": a VMM which could be embedded inside the Docker application, and which would be single purpose, minimal, secure and fast. The first version was called hyperkit and the most recent one is Docker VMM . 1

The VM kernel and root filesystem were minimal too, based on the LinuxKit project. The current version is an even more slimmed down and efficient variant but conceptually the same, similar to containerd/nerdbox .

Docker for Mac in 2016

For Docker for Mac (later renamed Docker Desktop) this allowed:

  • Running rootless on all platforms. Without requiring "rootless inside Linux": the whole Linux kernel is untrusted, so even a Linux CVE doesn't matter.
  • Minimal devices. We could carefully limit the host / VM interface, making it easy to understand and audit.
  • VPNs and network policy. It was easy to interoperate with VPNs, and to extend later to impose networking policy such as Registry Access Management .

CACM 2026 - A Decade of Docker Containers

And now in 2026

The Docker microVM tech is the foundation of both Docker Desktop and also Docker Sandboxes ( why microVMs ). Docker Sandboxes enables you to do things like run your favourite coding agent and harness securely , with strict isolation and governance.

To see what Docker microVM's can do, install Docker Sandboxes (sbx) and try `sbx run claude`.

Further reading

To read more about the history of the tech, check out the article A Decade of Docker Containers (Communications of the ACM).

Footnote

1 Although we were aiming to make everything a library and link into a static unikernel-like process, for technical reasons it makes sense to still have a single host process per VM. ↩

City building games have a Soul Problem pt.2

Hacker News
www.radical-elements.com
2026-10-03 11:52:07
Comments...
Original Article

Despite my disappointment with Cities: Skylines 2 , I keep coming back to the game from time to time. I can't even control it. It's like a deep inner need. When we were kids we played with toy cars on those felt mats with painted little cities, or when we had more industrial aspirations, we'd go outside in the dirt creating construction sites with tunnels and bridges right next to mom's rose bushes. Some of us had that one uncle who decided to turn the playroom into a miniature modeling world with trains, tunnels, and stations. I think a little tangle of my younger self's fantasies turns into a need to launch Cities: Skylines and build a living city.

In the previous article , I dealt with the genre more broadly and explored some ideas that sparked (mostly on Reddit) very interesting discussions. I talked more about the simulation and less about the aesthetics. Today, my text won't be technical at all; instead, it will tackle aesthetics and how much they disappoint me every time I gaze at my city.

Surfaces are all square and flat

Everything is square and unnatural. Looking at a random part of my city, I was trying to figure out why it's so ugly. What is it that makes me want to press alt+f4?

Here we see a spot where the sidewalk, some tiles from a building, and the grass meet.

In-game screenshot of Cities: Skylines 2 showing rigid, perfectly straight square seams where sidewalk, grass, and building tiles meet.

Does this look nice to you? How hard would it be for it to look like this:

Edited visual concept showing organic, blended transitions and natural edge detailing where grass meets paved surfaces.

I know someone can achieve similar and better results with mods and decals. But for most of us, it's just not realistic to spend 40 hours micro-managing every single corner of the city. And if you don't put in those hours, everything feels lifeless.

Default Cities: Skylines 2 street view showing a plain, sterile urban corner lacking surface detail.

Proposed visual improvement for a street corner, adding procedural detail, ground clutter, and realistic texture blending.

Also, something I truly enjoy in these games is being able to just stare at the city and see what the algorithm has generated. To be surprised. If I sit there and hand-paint every little corner, that "magic" is lost.

Neurotic cleanliness

As I just said, one of the reasons I build my city is so I can later go and gaze at random corners. It's the same when I take walks in downtown Athens; I look for charming little sides of the city. Not necessarily beautiful, not clean, but scenes that tell a story. In this game, everything is "sanitized" to the point of neurosis.

In-game view of a pristine, perfectly clean urban street in Cities: Skylines 2 with no visible wear or grime.

Mockup of the urban street with added surface wear, subtle dirt, and realistic pavement weathering.

Look how much difference these small textures can do to the overall atmosphere.

Default Cities: Skylines 2 cityscape scene featuring spotless, uniform architecture and sterile public space.

Enhanced visual concept showing character-filled urban space with lived-in details and subtle atmospheric grit.

What's the deal with stairs?

For some reason, the game has beef with stairs. Can we get some stairs? Something I would do if I were a designer for the game, is look at 100 photos of real cities and note down some of the things that make them up. Like:

  • Stairs
  • Potholes
  • Trash
  • Cables
  • Graffiti
  • Construction sites
  • Scaffolding
  • Poles
  • Traces of the past (repurposed public spaces or buildings)
  • Leaves

In-game screenshot showing awkward elevation changes and steep ground slopes without pedestrian stairs in Cities: Skylines 2.

Conceptual mockup showing outdoor stairs and terraced steps smoothly connecting elevated terrain.

Organic shapes

I don't know what your neighborhood is like, but in mine, maybe 1 in 50 houses is that clean and organized. Most have damp spots, weeds, abandoned toys in the garden, clothes hanging out to dry on the balcony.

In-game residential area in Cities: Skylines 2 with perfectly manicured, uniform green square lawns and sterile balconies.

Visual improvement showing organic residential yards with overgrown weeds, damp patches, hanging laundry, and varied foliage.

Some have rusted railings. Even the well-kept ones have let nature do its thing in some corners. And nature's thing isn't perfectly green squares.

Poor and abandoned houses

Unfortunately, every real city has poor areas with abandoned houses or ones in very bad shape. This is another aesthetic element that offers both realism and usable feedback. You shouldn't have to open the stats panel to see where a problem lies. By just gazing at your city, the landscape could and should speak for itself.

Default Cities: Skylines 2 neighborhood displaying clean, uniform buildings regardless of land value or area degradation.

Edited visual concept illustrating degraded and abandoned houses with weathered facades and visual cues of urban decay.

I hope that future city building games or CS2 DLCs will focus on adding some soul and character to city building. Until then, we have to rely on our Bob Ross instincts.


Discussions:


AI Disclosure: Artificial intelligence was used exclusively for proofreading, generating descriptive image alt texts, and creating the "improved" visual concepts.

“I Love My Culture”: Leonard Peltier on His Life and Art

Portside
portside.org
2026-10-03 11:32:24
“I Love My Culture”: Leonard Peltier on His Life and Art Kurt Stand Sat, 10/03/2026 - 11:32 ...
Original Article

Three of the paintings Peltier created while imprisoned. (L to R): Chief, 2014; Buffalo in the Window, 2010; Eva, 2009.

Leonard Peltier spent nearly half a century in prison. A celebrated indigenous activist, Peltier was involved in a 1975 shootout on the Lakota Indian Reservation in Pine Ridge in which two FBI agents were killed. Peltier was blamed for the deaths of the two FBI agents, a crime he says he did not commit. He was sentenced to two consecutive life sentences by an all-white jury, after being tried separately from his codefendants who were acquitted on grounds of self-defense.

Prior to the shootout and his subsequent imprisonment, Peltier had organized with the American Indian Movement. He was a participant in the 1972 Trail of Broken Treaties, which ended with activists occupying the Bureau of Indian A-airs headquarters in Washington, DC. Both the United Nations and Amnesty International declared his imprisonment arbitrary and considered him to be a political prisoner of the United States. It’s widely believed that Peltier was framed — punished for his role as an active member and leader in the American Indian Movement.

“From the minute of our invasion in 1492, there’ve been numerous attempts to terminate us as a race of people,” says Peltier when asked how he got involved in activism and the American Indian movement. “And the battle isn’t over.”

President Biden commuted Peltier’s sentence to home confinement in one of the last acts of his presidency — partial freedom after 49 years. These days, Peltier lives on the Turtle Mountain Indian Reservation, in North Dakota. Known for his activism and his status as a political prisoner, Peltier is also a painter. “I was doing political murals,” he says. “But actually, I’ve been doing art all my life.”

In an interview with The Indypendent , Peltier told us about what inspires his paintings, his life after prison, and what he wants people to remember him for.

THE INDYPENDENT : Leonard, to introduce yourself, tell us about yourself and your upbringing.

LEONARD PELTIER: My English name is Leonard Peltier. My family name is Tate WiWikuwa, Wind Chases the Sun. I’m Lakota native and Anishinaabe or Chippewa native from North Dakota. Talk about your upbringing, and what brought you to the American Indian Movement? We’ve been killed by the millions – and that type of genocide continues. When I was a young boy, in 1953, I was taken from my home and put in a [American Indian] boarding school, which was quite a horrid experience in itself. The schools were so bad that [President] Eisenhower put out an order that there be no more maltreatment of Native children in boarding schools. So those conditions, and as I said, the attempts to annihilate us, was the reason I became a young activist.

Turning to after the shootout at the Ridge Indian Reservation in South Dakota and your framing by the FBI, what was your time in prison like? What brought you solace?

Well, it wasn’t a pleasant experience. I’ll tell you that much. I mean, at times it was pretty brutal. Times were very dangerous, very stressful. There were a few good moments when family were able to come down to visit, or my wife, or somebody with the kids. You had an international movement supporting you as well. Yes, I did. I had the United Nations that put out an enormous investigation and demanded my freedom, and said, “We demand his freedom and for him to be paid.” I’ve had many, many Congressmen and senators struggling to get me out of prison, and to expose what happened to me. Now, to my amazement, I can’t go anywhere without Native and even non-Native people coming up to me and want to shake my hand, want to take pictures with me, and telling me they’ve heard about me in schools, in colleges and documentaries.

Was painting something that you spent time doing in prison, something that allowed you to be creative?

One of the things people have to understand is that I was not living in no country club. We did not get to paint in our cells. We only got certain hours a day to go to the art room, our recreation department, and we had to buy our own supplies. There were very strict restrictions on what you could paint and what you couldn’t paint. But I hung in there. The majority of my paintings, all of them, are from historical records, photographs of my people — a lot done by Edward Curtis, who did a lot of photographs back in the 1800. A lot of people sent me current pictures of powwows of young people, elders, stuff like that, and I painted those.

In some of the paintings your family has shared with us, images of buffalo appear. Please talk about the importance of the buffaloes, and did painting them help you maintain a sense of connection to the land that you were separated from?

The buffalo were nearly exterminated, like they intended to do to us, as well as the eagle and the wolf. They attempted to exterminate the buffalo because that was our food source and they wanted to take our food source away from us. So there’s still a very sacred and alive issue in Native culture. What’s so unique about Turtle Mountains, is that before I got out, we didn’t have one. Now we have seven white buffalo here. They’re a symbol of peace, a symbol of keeping law and order in our culture. It plays a big part in our historical culture.

Leonard Peltier in last year in Belcourt, North Dakota. (Photo credit: Shane Balkowitsch)

What has being out been like?

My home was bought by the American Indian Movement and our supporters. They all pitched in, donated, and bought me this home that I’m living in now. And of course, we fixed it up quite a bit. We’re cleaning up the yards. We’re doing a lot of maintenance to it, and it’s very comfortable. Very comfortable, of course, especially compared to living in a cell.

But even though you’ve been released, there are restrictions on your movement. You have to stay near your home most of the time?

I can’t go 100 miles from the reservation line without a pass from Washington. We’re fighting that now because they’re only supposed to keep those restrictions on someone for no more than a year [after being released]. They’ve denied me from attending some religious ceremonies that I was invited to, and to go visit doctors. I’ve been denied the right to go and be examined or pray with the people that invited me.

Do you see comparisons between the lands of your people being invaded and colonized, and what we’re all seeing in Palestine?

What they’re doing to the Palestinian people is genocide. They’re killing women, children, babies, people in hospitals, schools. It’s no different to what they did to us. Look what they tried to do to the Irish. They tried to wipe them out. This is what we got to understand. The only thing is different weapons, different time, but they are trying to exterminate these people.

How do you want to be remembered? What do you want your legacy to be?

The people that knew me all my life, that knew me when I was young, knew that I was sincere in battling for our people, survival of our people. I love my culture. I love my religion, and I think we have a beautiful history.

Show HN: Factorio but with unreliable components

Hacker News
think-twice.me
2026-10-03 11:19:49
Comments...
Original Article

An automation puzzle game

Make reliability
from chaos.

Every component has a chance to let you down. Build a system you can trust.

PLAY

Browser beta · Tutorial + scenarios 1–3 · Keyboard & mouse

BUILD. COMPARE. TRUST.

Unreliable components. Reliable systems.

The game

A little uncertainty.
A lot of possibility.

Build conveyors, inspect components, and decide what deserves your trust. Use Sorters, Counters, Markers, and Marker-Gates to turn a stream of unreliable components into a reliable network.

Your goal: deliver three 100% reliable components in a row. Start with the guided tutorial, then play through three beta scenarios, with more complex components and new challenges as you go.

Build & experiment Learn reliability Find your own solution

Stay in the loop

Hear when Rely is ready.

Get occasional development updates and an email when the full game is released. Unsubscribe whenever you like.

A note from the creator

Who do you trust?

I've found in my life that there are a lot of things we all know are true. But then there are so many other things that will only be argued forever.

How to know what's true? How to know who to trust? John says Sue is unreliable, Sue says John is unreliable? Can we ever make order from this chaos?

In this game, yes, you can, with enough patience and dilligence. You can build your network of reliable components and approach the fantasy of perfect information. But that's just me. Hope you enjoy

- Roady

AI Disclosure -- I used AI in coding this. If it ever makes money I'd be happy to hire artist and musician.

Statement From Families of 3 Palestinian Men Shot in Burlington, Vermont in 2023 in Response to Guilty Verdict

Portside
portside.org
2026-10-03 11:08:19
Statement From Families of 3 Palestinian Men Shot in Burlington, Vermont in 2023 in Response to Guilty Verdict Kurt Stand Sat, 10/03/2026 - 11:08 ...
Original Article

Joint Statement from the Families of Hisham Awartani, Kinnan Abdalhamid and Tahseen Ali Ahmad

Justice has been rendered for Hisham, Kinnan and Tahseen. After three years of sorrow, loss, and struggle, our families finally feel relief.

On November 25, 2023, our sons — college students and friends since childhood — were far from home on a Thanksgiving weekend. They wrapped themselves against the cold in scarves that reminded them of their distant home and went for a walk down a leafy block of Burlington, Vermont. They were joking with each other in Arabic and English, as they have done since they played together on the playground of their Quaker school in Ramallah. Within minutes, Jason Eaton tried to kill them. Our three children were badly wounded. Hisham was left paralyzed from the chest down and now lives with paraplegia.

Today, a jury has convicted Jason Eaton of three counts of attempted second-degree murder. The verdict cannot undo what happened or erase the trauma our families still carry. But it affirms something fundamental: Jason Eaton is responsible for what he did. Eaton’s defense argued that he should not be held accountable because he was delusional. The jury rejected this argument. We also reject this. We do not accept what this wrongfully implies about the millions of people in the US and the world who struggle with their mental health challenges.

Jason Eaton never disputed that he was the shooter. During the trial, we had to listen to him use hateful words to describe our sons—over and over again. They were not the words of a man confused about who he shot, or why. They are the words of a man who saw three young men celebrating their friendship and wearing what made them feel closer to their far away families and home – and decided to kill them because of his own prejudices. We do not need a jury to tell us why our sons were targeted. We heard it, in his own words. This was a hate crime. We have said so from the beginning. The law never made him answer for it. But we know what he did.

We are grateful to State’s Attorney Sarah George, Deputy State’s Attorney Sally Adams and their team for pursuing this case with rigor. We thank the medical professionals who saved our sons’ lives and supported their recovery and the law enforcement professionals who made this conviction possible.

We remain grateful to the Burlington and Vermont communities. From the first hours after the shooting, they stood with our sons—not just as crime victims, but as three young friends whose lives, wellbeing, and futures matter. That solidarity has mattered even more with each passing year, because our sons’ recovery unfolded alongside the genocide and destruction of Gaza and the escalating ethnic cleansing of the West Bank.

While our children fought to heal from being shot in Burlington, they watched Palestinian children lose their lives, their homes, their land, their families. We are grateful our sons’ suffering commanded attention and support. We grieve that so much Palestinian suffering never does. Hisham, Kinnan and Tahseen have spent three years rebuilding their bodies and their lives after a hateful act almost killed them. We are profoundly relieved. And we continue to work toward a world where hateful violence against Palestinians is called out for what it is – and stopped.

Justice is Served for Hisham, Kinnan, and Tahseen!  Now Justice for Palestine!

Statement from the Vermont Coalition for Palestinian Liberation

On November 25th, 2023, Jason Eaton committed a horrific hate crime in Burlington, Vermont. Amidst the racist hysteria the US and Israel whipped up to justify their genocidal war on Gaza, he attempted to murder three Palestinian college students—Hisham Awartani, Kinnan Abdalhamid and Tahseen Ali Ahmad. He heard them speaking a mix of Arabic and English, saw two of them wearing the traditional Palestinian scarf, the keffiyeh, and shot them on the street where they were visiting Hisham’s grandmother for Thanksgiving.

Luckily, all three survived the shooting, but Hisham was left paralyzed from the waist down. After three years of waiting and worrying that justice delayed would again be justice denied, Eaton stood trial over the last week. His defense and its psychiatric expert claimed he was not guilty by reason of insanity, while the state prosecution and its experts argued he was not and was therefore responsible for his actions. On September 21st, a jury found Eaton guilty on all three counts of attempted murder, finally delivering a modicum of justice for the three students and their families.

The trial became a metaphor of the Palestinian struggle for freedom, equal rights, and dignity. So much injustice has been done by the US government to the Palestinian people here in the US and in Palestine. Therefore, many doubted that a court and a jury would rule in favor of Hisham, Kinnan, and Tahseen. But the movement against Israel’s genocidal war has transformed mass consciousness and won the people of this state, country, and world to solidarity with Palestinians and their struggle for liberation. As a result, the all-white jury did the right thing, convicting Eaton for his terrible crime.

In doing so, they challenged Burlington’s political establishment, which has stood lock step with the federal government’s assault on Palestinians and Palestine solidarity activists. Both the Biden and Trump administrations have launched a New McCarthyite witch hunt, repressing protests, shutting down encampments, criminalizing speech in support of Palestine, and attempting to deport Palestinians like Mahmoud Khalil and Vermont’s own Mohsen Mahdawi.

The Democratic Party majority on Burlington’s city council joined their pro-war chorus at the start of the genocide, staged a rally in support of Israel, voted down a ceasefire resolution, and for three years in a row blocked an Apartheid Free Communities resolution from appearing on the ballot. In the statehouse, five legislators went on an all-expenses paid junket to Israel, where they received coaching on how to pass legislation criminalizing speech in favor of Palestine. They even had the gall to plant a Vermont flag on Palestinian land.

All that underscores the fact that Palestine is a Vermont issue. Like our national political establishment, our local one has economic, political, and ideological stakes in supporting Israel’s settler colonial apartheid state and enforcing US dominance over the Middle East and its strategic energy reserves. It therefore approves our tax dollars funding Israel, our state’s military industrial complex arming Israel, and our legislature defending Israel and its impunity even as it commits genocide.

The establishment’s prejudices were on full display at the trial. When dozens of Palestine solidarity activists wearing keffiyehs joined the audience to support Hisham, Kinnan, and Tahseen and their families, the defense objected and demanded that the judge have us remove them, contending they would pressure, intimidate, or influence the jury. The judge agreed and had us take off the scarves but did allow the three students to wear theirs. But why did he order them removed? The only conceivable answer is the establishment’s anti-Palestinian racism.

In the trial itself, the prosecution effectively proved that Eaton was not someone with a serious mental illness or disorder. Eaton’s statements proved he was cogent and capable of rational argument. As a result, the insanity defense had no foundation. What Eaton’s statements did display was the racism he had imbibed from the political establishment and media. He claimed he was defending Jews and a Burlington Kibbutz from Palestinian terrorists wearing keffiyehs. He compared his actions to Anne Frank defending herself against Hitler. Eaton is a bigot who committed a hate crime.

The verdict against him is a blow to all the injustice our government and society has meted out to Palestinians. It offers Hisham, Kinnan, and Tahseen and their families some closure to this horrible episode in their lives and hopefully opens up the possibility for all of them to thrive now and in the future. It also is a blow to the establishment’s ongoing campaign to intimidate and criminalize solidarity with Palestine on campuses and in the community.

The verdict is, however, a drop of justice in the ocean of injustice done to Palestinians. Mahmoud and Mohsen still face the threat of deportation for exercising their first amendment rights to speech, assembly, and protest. State legislatures, including Vermont’s, have passed or are planning to pass legislation criminalizing Palestine solidarity as antisemitic. Israel continues its genocide in Gaza, has escalated its ethnic cleansing of the West Bank, and enforces its apartheid system within its border that denies equal rights to Palestinians.

Even worse, the verdict has only sentenced a retail vigilante to jail. But the wholesale murderers, who poured their hatred and bigotry into Eaton, continue to rule our country and Israel’s with impunity. Biden, Harris, Trump, Vance, Netanyahu, and the whole Israeli political establishment are guilty of genocide. Why are they not put on trial and convicted of that crime against humanity? Only when they are will Palestine have justice.

That’s why the Vermont Coalition for Palestinian Liberation emerges from this trial and its verdict with renewed commitment to fight for justice for all Palestinians in the Middle East and here in the US. We will double down on our campaign to make Vermont an Apartheid Free Community that implements Boycott, Divestment, and Sanctions against the state of Israel. That is our small contribution to the global struggle to free Palestine and its people from colonial rule. Until Palestine is free, none of us are free.

FTL: A new operating system for clouds

Hacker News
ftl-os.org
2026-10-03 11:02:36
Comments...
Original Article

What's FTL?

  • You can build your own OS as a library . This userspace OS design makes it easy to add features, debug, upgrade the OS safely, as if writing applications.
  • FTL kernel isolates containers (userspace OS instances) better than existing monolithic kernels, with a hypervisor-like interface based on a lightweight hardware-based isolation (user mode). You don't need bare-metal machines.
  • FTL is compatible with Linux binaries . For example, the Rust-based HTTP server serving this website is a Linux application running on FTL. You can also run Unikernel-like specialized applications without POSIX abstractions.

How it works

Each container runs a userspace OS. It is a shared library which implements most of OS concepts such as Linux process, VFS, and TCP/IP. FTL kernel provides a minimal interface to implement Linux system calls in userspace, just like a hypervisor.

FTL combines the best of microkernels (flexible & secure) and monolithic kernels (performant & simple). Our goal is to make lightweight containers as secure as VMs, and unlock new OS-level abilities in applications , without sacrificing performance:

FTL                                   Linux
┌────────────────────────────────┐    ┌────────────────────────────────┐
│┏━━━━━━━━━━━━━┓  ┏━━━━━━━━━━━━━┓│    │┏━━━━━━━━━━━━━┓  ┏━━━━━━━━━━━━━┓│
│┃             ┃  ┃             ┃│    │┃             ┃  ┃             ┃│
│┃    Linux    ┃  ┃    Linux    ┃│    │┃    Linux    ┃  ┃    Linux    ┃│
│┃   Process   ┃  ┃   Process   ┃│    │┃   Process   ┃  ┃   Process   ┃│
│┃             ┃  ┃             ┃│    │┃             ┃  ┃             ┃│
│┃╌╌╌╌╌ Linux system calls ╌╌╌╌╌┃│    │┗━━━━━━━━━━━━━┛  ┗━━━━━━━━━━━━━┛│
│┃                              ┃│    └────────────────────────────────┘
│┃         Userspace OS         ┃│    ╌╌╌╌╌╌╌╌ Linux's interface ╌╌╌╌╌╌╌
│┃   (Process, VFS, TCP, ...)   ┃│    ╔════════════════════════════════╗
│┗━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┛│    ║                                ║
└────────────────────────────────┘    ║          Linux Kernel          ║
╌╌╌╌╌╌╌ minimal interface ╌╌╌╌╌╌╌╌    ║                                ║
╔════════════════════════════════╗    ║   process, fork/exec, memory,  ║
║           FTL Kernel           ║    ║     signals, TCP/IP, /proc,    ║
║    vCPU, memory, drivers, ...  ║    ║       /dev, drivers ...        ║
╚════════════════════════════════╝    ╚════════════════════════════════╝

Userspace OS design also enables you to extend most of Linux kernel features without kernel/eBPF programming. You can add printfs, apply security updates, and add new features quickly and safely. In FTL, OS is just a library . Read more .

Roadmap

  • September 2026 : Run a simple Linux HTTP server on FTL (released in v0.0.1 ✅)
  • October 2026 : Async Rust apps support - Linux threads, epoll, ... (released in v0.1.0 ✅)
  • November 2026 : Filesystem
  • December 2026 : Node.js / Go support
  • January 2027 : SMP, container images, 64-bit Arm support

Links

Does Anybody Care About Data Breaches?

Internet Exchange
internet.exchangepoint.tech
2026-10-01 10:46:37
Data breaches are now a routine part of life, but people affected still have almost no way to get a remedy. Lucy Purdon asks what real consequences for companies could look like....
Original Article

By Lucy Purdon , also published in her newsletter The Prompt by Courage Everywhere

Many years ago at The Glass Room art exhibition in London, I leafed through thick white books, similar to old school telephone directories, containing every password stolen in a 2012 hack that exposed LinkedIn’s entire database. Artist Aram Bartholl alphabetized, printed and bound the list of 4.6 million passwords into 8 volumes to tell a story about data, inviting visitors to pick up the books and search for their own password. (Yes, mine was in there.)

4.6 million passwords seems a quaint number now due to the scale of today’s data breaches. It’s likely you yourself have recently received a message about a cybersecurity incident that has exposed your personal details held in a company’s system, and I’ll bet that wasn’t the first time- 4.4 million online accounts were reportedly exposed in the first 3 months of 2026 in the UK alone. For most people a data breach involves an email, telephone number or financial information like credit card details. Do you feel the consequences of such a breach are kind of left hanging, often downplayed and unclear? Do you feel confident the company in question has given you the information you need, beyond “soz, change your password”?

As Jamie Bartlett wrote in the excellent Substack post, “ What actually happens to your stolen data? ” your data goes on a “five stage, globe trotting, magical mystery tour”. He also wrote about how scams are becoming more sophisticated with the help of AI; those “phishing” scams are getting more convincing- and it all starts with a stolen email.

It gets worse of course. I have written before about the data breach from consumer genetic testing company 23andMe , resulting in the theft of millions of customer profiles including date of birth and ancestry information. Hackers advertised the data for sale and boasted it included around 1 million people of Ashkenazi Jewish descent, 100,000 of Chinese descent and “the wealthiest people living in the U.S and Western Europe” . Changing a password doesn’t touch the sides when your actual DNA is stolen.

In May this year, a cyberattack on the World Food Programme exposed the personal data of 600,000 Palestinian households in Gaza, including their names, ID numbers, and location. The weaponization of this information could prove deadly.

The UK’s National Cyber Security Centre describes data breaches as “a fact of modern life” . I don’t believe that means we accept the organizational carelessness, lack of security investment or the greed of collecting as much data as possible that so often leads to data breaches. Data breaches matter, they don’t seem to be taken seriously enough and we are often left in the dark with no meaningful remedy.

What actually is a data breach?

Organizations that collect and hold your personal data are bound by data protection law (where it exists) to keep it safe and secure. A data breach under the EU GDPR describes a situation where an organization has failed to keep it safe, leading to the destruction, loss, alteration, or - most relevant here- the unauthorized access to or disclosure of personal data.

Data breaches are increasingly the result of a cyber attack, but human error plays a major part. Just last month, the UK’s Metropolitan Police accidentally disclosed the emails of 140 people accusing the late owner of Harrods of sexual abuse by cc’ing them in an email update, rather than bcc’ing.

Under the EU (and UK) GDPR, organizations have the obligation to notify both the regulator and affected people of the breach, and provide remedy.

Tracking company responses

Ranking Digital Rights (RDR) evaluates the policies and practices of 26 of the world’s largest digital and telecommunications companies on how they uphold commitments to respect human rights such as freedom of expression and privacy, publishing an index of the findings. Since 2017, the index has included an indicator on company responses to data breaches as part of their methodology.

The indicator assesses three issues in line with data protection legislation: whether companies commit to notifying relevant authorities, whether they explain the process they will follow to notify people (data subjects) affected by a data breach, and whether they explain the steps they may take to address the impact of said breach.

Leandro Ucciferri, Deputy Director of RDR, has charted how the introduction of the GDPR in 2018 slowly made a difference to transparency around data breaches, but there are still major gaps in protections:

“In the 2017 RDR Index, only 3 of the 22 companies evaluated published some information about their policies to address data breaches (the companies were Telefónica, AT&T, and Vodafone). But we didn't see a notable improvement until 2019 [after the GDPR was adopted], when 10 of the 24 companies we evaluated published policies on this issue (five were digital platforms and five were telcos).

Fast forward to the latest assessments we published ( 2025 Big Tech Edition and 2026 Telco Giants Edition ), 19 of the 26 companies we evaluated disclose some information about their data breach policies, but still none of them reached the full score for the three concrete questions in this indicator.

Notably, giants like Google, Amazon, and TikTok, do not publish explicit policies and commitments to notify authorities and users, nor explain the processes they may take to mitigate the harms caused by a breach.”

Remedy is so often the weakest point of rights-based legislation and the gaps are glaring when it comes to data protection. Compensation is a grey area as harm connected to a specific data breach is difficult to prove, even when the breaches are so egregious and the impacts potentially infinite. Leandro points out that accessing remedy under GDPR is a high bar, with a data subject needing to demonstrate in court that a damage existed, there was infringement of GDPR and the direct causality of the damage suffered and the GDPR violation.

This relies on individuals knowing what has happened to their data in order to exercise their rights, which is impossible when it comes to engaging with increasingly opaque systems and so little disclosure or transparency from companies.

Fine!

Failing to secure data can lead to headline-grabbing fines for companies, but usually only when financial details are involved, which appears to be assessed as the most tangible harm. In 2020, British Airways was fined £20 million by the Information Commissioner's Office (ICO), the UK’s data protection regulator, after users were directed to a fraudulent site where hackers harvested data of 400,000 customers including credit card information. Capita Pensions was fined £14 million by the ICO (negotiated down from £45 million mind you) for a hack that exposed 6.6 million people’s financial information.

But harm is not just financial and damage may not be immediate. Data hangs around.

How this plays out IRL: my Substack experience

If you have a Substack account you may, like me, have received an email from Substack CEO Chris Best on February 5th describing, in the most casual terms, that Substack was hacked and the email and phone number connected with your Substack account was “shared without your permission”.

“This sucks,” says Chris. It’s OK though!, “Importantly, credit card numbers, passwords, and financial information were not accessed.”

I have since been inundated with spam emails, including one claiming to be from the ShinyHunters criminal group, the notorious hackers behind some of the biggest data breaches of the past 5 years. Their email directly references the Substack data breach as the source of obtaining my information and tries to extort me by claiming they have videos of me watching pornography and “pleasuring myself” and will release these videos unless I pay $2000 in Bitcoin.

It’s not really ShinyHunters of course and this kind of “sextortion” scam is increasingly common.

Nonetheless, is Substack warning their users about the direct correlation between the breach and these attempted extortion attempts from a scary criminal gang, perhaps pointing to or supporting the work of fact checking groups exposing the scam?

No.

In response to my email complaint to Substack that my data has not been handled properly I am told that there is “no evidence that malware was installed via Substack” - the answer to a question I was not asking and a clear deflection.

Not satisfied with this response, I complained to the ICO, as is my right when there is a concern that an organization has not handled personal information properly. I am given a case number and informed I will receive a response within… 40 weeks!

Actually 6 weeks later, I received the decision that the ICO will not take the complaint further as Substack “has handled the matter in line with its data protection obligations by informing you of the data breach.”

So…that’s it. Notify those who bear the brunt and shift the onus onto the user to change passwords, monitor emails for phishing attempts and scams and put up with the torrent of spam emails, never being sure where our data is and what it is being used for, the next scam or trick around the corner.

Hardly reassuring.

A game of consequences, anyone?

I would opine that organizations often collect way more data than they need, because it’s valuable, so there is more data to breach. Databases are often poorly secured; procedures for dealing with data breaches are shockingly lax (the Substack breach reportedly went undetected for 4 months); companies often take ages to disclose and there are very few consequences (the recent ICO investigation into ACRO, a company that handles criminal records, is a jaw-dropping insight into abysmal cybersecurity failings). Leandro from RDR sums it up,

“My main concern at the moment is whether we've reached a point of apathy. Sure, some companies are receiving fines after being investigated by data protection authorities, but that's simply another cost of conducting business. And at the same time, the people affected may feel powerless, since they keep receiving news about new ways in which their data was exposed, likely multiple times in a given year. Even looking at the situation in a handful of European countries (Spain, France, Germany, the Netherlands, and Ireland), there were a combined total of more than 64000 data breach notifications to the national data protection authorities in 2025. The scale of this issue doesn't seem to be slowing down at all.”

As AI demands more data to train models and agentic AI requires more access to our personal data, what can we expect in the future?

“When it comes to the current conversation involving “AI” systems, as long as companies’ business models rely on extracting as much data as possible, we can expect them to face security incidents that end up exposing personal information.”

In the face of this we need stronger protections not less but intense lobbying from tech companies is steering our rights-based legislation in the wrong direction. The EU Digital Omnibus, an initiative to “streamline” EU digital legislation is perceived by many civil society actors as a potential dilution of safeguards established by the GDPR, the ePrivacy Directive, and the AI Act.

We need to bring ideas to the table on what we expect remedy to look like - not just fines but actual consequences and repercussions for companies to mitigate any potential harms, like being a victim of identity theft, targeted with scams, or worse. How about a company fined for a major breach also cannot collect any consumer data for a month? Companies must track the data breached and provide you with weekly updates about where it is and who has it? Then there might be more of an incentive to protect the data we often have no choice but to hand over.

Class action lawsuits may start to bite; over in Kenya, subscribers of the telecommunications company Safaricom won thousands of dollars in compensation over a large scale data-breach where the High Court found Safaricom failed to secure data and breached the constitutional right to privacy.

Let me know your experiences of data breaches, and your ideas about how companies should provide remedy!


Support the Internet Exchange

If you find our emails useful, consider becoming a paid subscriber! You'll get access to our members-only Signal community where we share ideas, discuss upcoming topics, and exchange links. Paid subscribers can also leave comments on posts and enjoy a warm, fuzzy feeling.

Not ready for a long-term commitment? You can always leave us a tip .

Become A Paid Subscriber


HRPC.io is Back Online

The website for the Human Rights Protocol Considerations (HRPC) research group at the Internet Research Task Force is back online, and features a short documentary about how the technical design of the internet relates to human rights. The HRPC studies how internet standards and protocols enable, strengthen, or threaten human rights, especially freedom of expression and assembly. IX's Mallory Knodel chairs the group, and the site is a good place to start for anyone who wants to learn more and get involved with this work.

🚨

Stop press! Do you enjoy our links? Links are now available to paid subscribers only. Become a paid subscriber today.

Danish university DTU breach exposes data of up to 200,000 people

Bleeping Computer
www.bleepingcomputer.com
2026-10-03 10:35:20
The Technical University of Denmark (DTU) says information belonging to up to 200,000 users may have been exposed after hackers accessed its identity and access management system and downloaded a large amount of data. [...]...
Original Article

Danish university DTU breach exposes data of up to 200,000 people

The Technical University of Denmark (DTU) says information belonging to up to 200,000 users may have been exposed after hackers accessed its identity and access management system and downloaded a large amount of data.

​The university says the attacker used compromised credentials to log into DTUBasen, its identity and access management (IAM) system, allowing access to more than two decades of user data.

In a disclosure on Friday, DTU confirmed that it cannot “determine precisely what information was downloaded or how many people have been affected.”

However, the Danish university notes that DTUBasen stores information for nearly 40,000 active users and around 160,000 former users.

Next of kin data also exposed

Potentially exposed information for current users includes Danish civil registration numbers (CPR), full names, home addresses, and profile pictures, as well as work email addresses, job titles, office locations, and other employment-related details.

The dataset also contained the names, relationships, and telephone numbers of users’ next of kin, when provided by active users.

DTU notes that in the case of former users, details about home addresses, profile pictures, and information about next of kin are automatically deleted after six months.

“This is a serious attack on DTU, and we deeply regret the uncertainty it is causing for the people whose information may have been affected,” says University Director Bjarke Bak Christensen.

“Our first priority has been to establish the extent of the attack, limit its consequences, and ensure that those affected are notified and know what steps to take,” the director added.

DTU warns that cybercriminals could use the exposed CPR numbers and other personal data for identity fraud and to make phishing attacks more convincing.

Not all notified directly

Potentially impacted individuals will be notified through e-Boks, the official mailbox system that DTU uses for sharing documents and notices with students and staff.

However, the university says it will notify all current and former employees, but not all current and former students whose CPR numbers are held by DTU.

“DTU only holds CPR numbers for a small number of guests and external partners and does not hold CPR numbers for next of kin whose contact details have been registered in DTUBasen,” the organization says .

The public dicslosure is part of DTU’s effort to reach potentially affected individuals it cannot contact directly, and the university is urging people to share it with former employees, students, guests, and external partners.

The organization says that anyone who has been an employee, student, guest, or external partner of DTU since 2003 may be affected by the data breach.

They are advised to be cautious of emails, text messages, and phone calls from individuals who appear to know about their connection with DTU or have access to personal information about them.

Passwords and sensitive information should not be disclosed in replies to unexpected communications, and sudden authentication requests or logins should be treated as suspicious.

Additionally, it is recommended to change the passwords for any other services that use the same credentials as the DTU account and place a credit alert on the affected CPR number.

article image

Build your security blueprint for AI-powered attacks

Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.

Save your seat

California’s new laws target workers’ biggest fear of AI taking their jobs

Guardian
www.theguardian.com
2026-10-03 10:00:00
The state, which is home to many AI companies, is one of the first to roll out workplace regulations targeting the technology California’s laws aimed at protecting workers from the impacts of artificial intelligence could pave the way for broader workplace safeguards across the US as calls for regu...
Original Article

California’s laws aimed at protecting workers from the impacts of artificial intelligence could pave the way for broader workplace safeguards across the US as calls for regulating the technology mount. As the federal government goes hands-off on AI, California is taking the reins to address workers’ biggest fears.

On Thursday, California governor Gavin Newsom signed a suite of new laws that ban bosses from relying entirely on AI to decide whether to fire workers, using it to predict employees’ emotional states or collecting neural data, meaning the information from electrical signals of someone’s brain or nerves. They also require companies to notify workers if layoffs were caused by AI and prohibit AI surveillance in workplace bathrooms.

The new laws come as workers increasingly worry whether AI will take their jobs, lead to discrimination and increase workplace surveillance. Unions, worker advocates and even some lawmakers pushed for the new rules, a regulatory shift for a technology that has largely developed unchecked. California , home to many of the leading companies developing AI, represents one of the first states to roll out a sweeping set of workplace regulations targeting the technology.

“It’s a turning point,” said Lorena Gonzalez, president of the California Federation of Labor Unions, AFL-CIO, who’s been helping leaders across the country write regulations. “It’s really the first time we’re seeing California workers showing the country that we don’t have to accept [this].”

Other states that have recently passed individual laws aimed at AI’s use in the workplace include Colorado, Connecticut, Illinois and Texas, narrower in scope than California’s. And more bills across the county are being lined up for consideration, Gonzalez said.

California’s laws aim to target workplace surveillance measures like heat maps that track employees’ movements, including how long they spend in the bathroom, or having their emotional states monitored. Amazon warehouse workers have previously complained about being timed on their bathroom breaks, for example. And at Kaiser Permanente, nurses have said automated systems rated their tone of voice in patient interactions.

But the laws could also prevent future unexpected harms.

“We don’t know all the places companies are using AI, and that is and should be scary,” Gonzalez said.

Part of the strategy for worker advocates and union groups like the California Federation has been following AI companies’ latest products. “If it’s being sold, that’s a good indication” it could be in use, Gonzales said. The federation also plans to use the momentum to revive issues like requiring employers to disclose when they’re using AI in the workplace, which was a bill that died in the state’s assembly appropriations committee this year.

California’s new laws are a key step in gaining regulatory ground for worker advocates, said Robin Feldman, director and founder of AI Law & Innovation Institute at The University of California College of the Law, San Francisco. Still, the statutes are somewhat limited in how they are implemented.

“The bills have no private enforcement,” Feldman said. “In other words: workers can’t sue. Only the government can enforce the laws.”

The new regulations come amid a backdrop of record AI spending at big tech companies along with massive job cuts. But workers have started pushing back. In June, Meta paused a program that tracked workers’ computer activities to train its AI models. And one month later, dozens of employees filed a lawsuit claiming the tech company’s AI tools targeted those with disability accommodations or on medical or parental leaves for layoffs.

Meanwhile, safety concerns, including fears that AI could destroy humanity, have been garnering more attention from lawmakers across the country and prompted OpenAI and Anthropic to call for slowing the pace of development.

skip past newsletter promotion

“Workers are increasingly part of that movement, speaking up about the fear of job loss and the dehumanizing experience of being surveilled and controlled by an algorithm,” said Annette Bernhardt, senior tech policy adviser at UC Berkeley Labor Center.

While the new laws “have teeth”, according to Danielle Ochs, shareholder at employment law firm Ogletree Deakins’ San Francisco office, it’s unclear how sweeping the change will be. Ochs said employers generally aren’t grappling with the AI uses outlined in the new regulations and instead are more interested in how to responsibly implement AI across their systems.

“Having 10 hoops you have to jump through per tool is not reflective of reality,” she said. It would be better to have “guardrails that are more aligned” with employers’ wider use of AI rather than focused on specific tools or uses.

She says opponents worry that the new rules could unexpectedly prohibit helpful AI that might, for example, ensure truckers don’t fall asleep at the wheel.

While it’s too soon to gauge how effective the laws will be, worker advocates believe the momentum is moving in a positive direction. Gonzalez said the measures are only the beginning of addressing AI’s potential impacts.

“We have so much work to do,” she said. “But this should give us all hope we can win … against the tech lobby, against big corporations, because we are the majority.”

I quit OpenAI because its culture is broken

Hacker News
www.theatlantic.com
2026-10-03 09:46:34
Comments...
Original Article

Our site’s security protections triggered a block. If you think this was in error, email us at [email protected] . Please include how you got to this webpage and the Cloudflare Ray ID located at the bottom of the page.

If you would like to inquire about permission to access our content, email us at [email protected] .

Cloudflare Ray ID: a451a4256dd32d15

Back to The Atlantic

US killer's sentence quashed because of AI video of victim shown in court

Hacker News
www.bbc.com
2026-10-03 09:34:18
Comments...
Original Article
Watch Chris Pelkey's AI-rendered impact statement shown in court

An Arizona road rage killer will be resentenced after an appeals court ruled that an AI-generated video message from his dead victim was improperly aired in court.

Gabriel Paul Horcasitas was found guilty by a jury of shooting and killing Christopher Pelkey, 37, during a confrontation at a red light in 2021. He was sentenced to 10 years behind bars.

But his legal team appealed, arguing that the trial judge should not have allowed an AI video, created by Pelkey's family, to be shown in court ahead of last year's sentencing.

The Arizona Court of Appeals agreed, ruling that Horcasitas, 55, must be resentenced because the AI clip "crossed that line".

"Rather than document an event or recording a particular moment, the AI video presents a depiction of the victim and his thoughts created from the imaginings of the victim's sister," the court of appeals wrote on Wednesday.

Horcasitas' lawyer, Kristen Reller, who filed the appeal, declined to comment.

Jessica Gattuso, an attorney for victims in the case, did not respond to a request for comment.

Pelkey's sister, Stacey Wales, told the BBC last year the family had used voice recordings, videos and pictures to recreate him for the AI video.

"To Gabriel Horcasitas, the man who shot me, it is a shame we encountered each other that day in those circumstances," said the AI version of Pelkey in court in 2025. "In another life, we probably could have been friends."

"I believe in forgiveness, and a God who forgives. I always have and I still do," the AI version of Pelkey - wearing a grey baseball cap - continued.

At the time of the sentencing, the judge who oversaw the case, Todd Lang, seemed to appreciate the use of AI.

"I loved that AI, thank you for that. As angry as you are, as justifiably angry as the family is, I heard the forgiveness," Judge Lang said. "I feel that that was genuine."

The Era of Software Quality, or the Era of Ostriches?

Lobsters
blogs.gnome.org
2026-10-03 09:31:58
Comments...
Original Article

Humans are bad at writing secure code, and GNOME developers are no exception. GNOME is primarily written using unsafe programming languages where simple mistakes in our code lead to devastating consequences for our users , and we make these mistakes all the time . No matter how much we try, GNOME developers will fail write secure code when using unsafe languages like C, C++, or Vala: it’s just too hard for even experienced developers to do properly.

The above paragraph is taken from the abstracts of my GUADEC 2024 and 2025 talks. At the time, I thought failure was inevitable: we humans were so bad at writing software that we had no chance to do it properly, and I certainly would not have trusted an AI to do better than a human. But the landscape today is completely different than last year. AI has improved considerably, and offers a magic fairy wand solution to this problem: we can now simply ask a language model to look for vulnerabilities in our software. They are quite good at this.

There is zero hope of maintaining quality software in 2026 without AI vulnerability scanning. Any claims to the contrary are unserious and delusional. The tremendous quantity of bugs found in our best-maintained projects, like GLib and fwupd, should speak for itself. Failure to scan our projects is an unfair disservice to our users. If we don’t find the vulnerabilities by scanning projects ourselves, attackers certainly will, because the Linux user base has increased to the point that Linux users are finally numerous enough to be worth targeting. Meanwhile, AI has made it easier than ever to build working exploits , which was previously unheard of.

Already resolved all the detectable vulnerabilities? Then ask the AI to look for non-security bugs as well, to further improve quality. GNOME code is generally much better than it used to be, but there remains considerable room for improvement. For the first time in history, we now have the opportunity to improve software quality to a degree that was never realistic before.

Have you heard that most AI bug reports are “slop?” Not so in 2026. That was true for most of 2025, but the quality of AI-generated vulnerability reports has drastically improved. That is not to say that we no longer have problems with bad vulnerability reports, but in general, nowadays most of them are pretty good. ( Daniel Stenberg reports the same pattern for curl. )

AI-generated vulnerability reports have nevertheless introduced many undesirable impacts on GNOME maintainers. They are usually annoyingly verbose and unnecessarily detailed. They often exaggerate the severity of the problem, or make misleading or irrelevant claims. They are occasionally incorrect. Sometimes they include outright fabricated data, such as fake stack traces (which is not the norm, but sadly also not uncommon). A good human reviewer will notice and resolve most of the above problems before creating a bug report on your issue tracker, but often problems are reported by inexperienced humans who do not actually know what they are looking at and simply copy/paste everything blindly. Even when the generated issue report is good and avoids all of the above problems (which is rare), good vulnerability reports in sufficiently high quantity can still overwhelm volunteer maintainers. And even if reporters submit a merge request to resolve the problem so maintainers don’t have to (which is also rare), reviewing those merge requests is itself more unwelcome work for overworked maintainers.

That all is to say: I understand the pain caused by the current wave of AI-generated issue reports. Nevertheless, they are essential and unavoidable. We have to learn to accept and deal with them, not stick our heads in the sand and ignore them.

Some GNOME maintainers have adopted a policy prohibiting AI-generated content in issue reports. Do not do this. Nowadays, the overwhelming majority of vulnerability reports are AI-generated. Projects that choose to ban AI-generated content in issue reports might as well ban all vulnerability reports; the effect will be approximately the same.

I propose the following:

  • GNOME maintainers should rewrite their AI contribution policies to permit AI-generated vulnerability reports, as I previously requested four months ago .
  • Projects that continue to prohibit AI-generated vulnerability reports are no longer suitable dependencies for GNOME, and should be developed someplace other than GNOME GitLab.

We don’t have to tolerate bad issue reports, but AI use alone should not be disqualifying.

Shouldn’t humans rewrite AI-generated bug reports?

When I complain that maintainers should allow AI-generated vulnerability reports, the most common counterargument is that humans should read the AI’s report, understand it, and rewrite the entire thing to remove all AI-generated content. Some bug reporters actually voluntarily do this, but this is rare.

Vulnerability reporting is a public service, not an obligation. If you ask a reporter to do any amount of extra work, they might be willing to do so, but it’s much more likely that they will either stop looking at your project and move on to something else, or continue looking at your project and publish the vulnerability reports someplace other than your issue tracker.

Rewriting issue reports also does not scale. Let’s say you use AI to find 100 security bugs in a GNOME project, a number consistent with the results of actual scans (read on). Would you really spend months rewriting those bug reports before submitting them to upstream? Validating the AI’s claims, upstreaming the issue reports, and submitting merge requests is already a lot of work. Not many people would be willing to additionally rewrite all the issue reports. That’s more work than everything else combined, and is unrealistic.

Even with just a small number of bugs, I would hesitate to spend much time rewriting an issue report because I have many other tasks I would rather spend my time on. At best, I might prepare a quick summary, but it won’t be as useful as a full report.

The CVE Wave Hits GNOME

The current wave of vulnerability reports is reflected in GNOME’s CVE issuance trends:

Year GNOME CVEs GNOME CVEs Excluding GIMP, Gegl, libxml2, and libxslt
2021 21 14
2022 14 6
2023 13 4
2024 37 28
2025 97 49
2026 Year-to-date (2026-09-30) 141 74
2026 Normalized 188 (141 * 4 / 3) 99 (74 * 4 / 3)

The trend here should be pretty clear. Until recently, not many people were reporting vulnerabilities in GNOME. That has changed. We are currently dealing with an order of magnitude more CVEs than just 3 years ago. AI is not the only reason for this; GNOME maintainers have also gotten a little better at flagging issues so that I add them to security tracking. But AI is the primary cause for the increase.

(A few technical notes on this table. CVEs are classified by the year the issue was reported to GNOME, not by the year in the CVE identifier, so e.g. many CVE-2026 issues are counted in 2025. Vulnerabilities reported in 2026 which do not yet have CVEs are not counted, so you can think of the data as being accurate through roughly September 1; multiply the 2026 numbers by 4/3 to make them comparable to the prior years. I count only issues reported to GNOME Security , so any unreported CVEs do not count.)

Although there are still 3 months left in 2026, we will never have data for the rest of the year because I have ended security tracking for new issue reports and nobody else has volunteered to do that work. These CVEs exist only because I request them myself, so I expect the number of CVEs to drastically decrease going forward.

The CVE Wave Hits WebKitGTK

A similar pattern holds for WebKitGTK:

Year WebKitGTK CVEs
2015 175
2016 57
2017 158
2018 101
2019 99
2020 38
2021 52
2022 50
2023 45
2024 38
2025 66
2026 Year-to-date (through WSA-2026-0006 ) 305

CVEs are reported against the year they appeared in a WebKitGTK security advisory, not the year in the CVE ID. The large increase in 2026 is entirely due to AI analysis of Skia and ANGLE. WebKit bundles these libraries because they are not designed to be installed as system libraries, so their vulnerabilities should be counted the same as vulnerabilities in WebKit’s own code. Excluding Skia and ANGLE, there are actually only 21 other WebKitGTK CVEs so far this year, a significant decrease, but excluding CVEs in bundled code would not be fair.

There has actually been a very large increase in WebKit security fixes this year, but this has not resulted in any increase in CVEs. Apple generally creates CVEs for flaws found by external researchers, not often for flaws found by WebKit developers, so the increase in security fixes is not reflected in the total number of CVEs. Only a small fraction of WebKit vulnerabilities receive CVEs.

I had not previously noticed that the count of WebKitGTK CVEs had, until 2026, been decreasing over the past decade. I am not sure why. I also do not know how to explain the low number in 2016.

Announcing the GNOME Bug Bounty Program and Announcing the End of the GNOME Bug Bounty Program

My blog post to-do list says that I need to write a blog post announcing the creation of the GNOME Bug Bounty Program on the YesWeHack platform. Oops, too late. It’s already closed. (Once a task enters my to-do list, it can be a very long time before I get around to doing it.)

The GNOME Bug Bounty Program was generously sponsored by the Sovereign Tech Resilience program of Germany’s Sovereign Tech Agency. I’m not sure precisely when it opened, but the first vulnerability was reported on June 27, 2024, so it would have been sometime shortly before then. We accepted issue reports only for GLib, glib-networking, and libsoup, because GNOME had never operated a bug bounty program before and we did not know what to expect. Starting small had — naively — seemed like a prudent way to avoid a large quantity of issue reports. I had wanted to expand the program to cover all of GNOME, but this failed due to the overwhelming deluge in issues reported against GLib and libsoup.

I requested that the bug bounty program end because I was overwhelmed with incoming AI-generated issue reports. The final issue was reported on February 23, 2026. Here are the results:

Year Reports Submitted Reports Accepted
2024 26 14
2025 150 33
2026 122 24
Total 298 71

Those numbers for 2026 reflect less than two months’ worth of issue reports, so you can see why it was no longer sustainable.

After the program closed, our work was not done: there was a long backlog of reports to work though. We just last month caught up with accepting the last of the issues reported back in February, and the last bounty was finally awarded earlier today! Even with YesWeHack’s professional triagers analyzing the issue reports before I reviewed them, keeping up with such a large number of vulnerabilities was not easy for me.

At this point, all reports not accepted have been rejected. The program awarded €183,900 in bounties for 71 vulnerabilities: 45 in libsoup, 23 in GLib, and 3 in glib-networking. Award amounts varied from €500 (16 awards) to €7,500 (2 awards). The arithmetic mean award was €2,662.99.

Bug bounty programs are an exception to the rule that most AI-generated vulnerability reports are good. You can see the number of reports accepted is a small fraction of the number of reports submitted. Excluding 30 reports closed as duplicates, that leaves 197 reports rejected. Turns out, people will submit bad reports when financially incentivized to do so. The low percentage of accepted reports even understates the problem, because many of the accepted reports were actually not very good! Many accepted reports did successfully identify valid security problems (in fact, many of the rejected reports successfully identified valid security problems!), but required many rounds of revision and corrections.

Suffice to say, I have reviewed a lot of really bad AI-generated vulnerability reports. But the reports we received via the discontinued bug bounty program are not comparable to the reports received via regular GNOME issue trackers or the security bug report form . We do still occasionally receive bad vulnerability reports, but not often and not many, so it’s not a big problem anymore. When people submit AI-generated reports without hope of a financial award, those reports are generally much better.

Lessons from the Bug Bounty Program

Closing the bug bounty program because it found too many vulnerabilities is not a particularly pleasant result. That said, it was still a partial success in that it uncovered lots of bugs in libsoup and GLib.

I had hypothesized that libsoup was probably not very secure, but I never imagined just how many vulnerabilities would be discovered. To reduce the quantity of incoming issue reports and better reflect actual risk to GNOME users, I eventually removed all denial of service bugs from program scope, and then later removed SoupServer from the scope due to too many request smuggling vulnerabilities , which are HTTP request parsing bugs that pose no threat to GNOME users. Even with those changes, the libsoup vulnerability reports kept coming until I gave up. The silver lining is that libsoup is now relatively much more secure than before. Other bug reporters have been submitting AI-generated bug reports using the normal libsoup issue tracker, so fortunately the improvements to libsoup will continue despite an end to the financial awards.

I had hypothesized that GLib would be much better than libsoup. I’m not sure whether I was correct. Evaluating the severity of GLib flaws is much harder than for libsoup, since GLib vulnerability reports are generally hypothetical in nature: usually some proof of concept program calls a GLib API using valid but improbable values, then something bad happens.

A large portion of the GLib bugs were integer overflow flaws, which generally result in buffer overflow. I am now more scared of integer overflow than anything else. It’s likely that most software projects have many integer overflow problems. Fortunately, we should be able to catch most such problems by adjusting the compiler flags we use. In particular, -Wconversion or -Wint-conversion and -Wsign-compare should help here. Some GNOME projects already use -Wsign-compare , but I suspect most do not. I think few or no GNOME projects use -Wconversion or -Wint-conversion .

Resuming the bug bounty program would only be possible under substantially different conditions. What we were doing was not working well. To resume, we would need to limit the scope to projects that regularly perform their own AI vulnerability scans. We would also most likely want to pay only for functional exploits, rather than for all vulnerabilities. GNOME code is currently not good enough to continue paying for every vulnerability, and it no longer makes sense to pay bounties for issues that can be found by AI scanners.

Red Hat Scans GLib

Red Hat has contracted with AISLE Research to perform AI vulnerability scans of various GNOME projects. We received a large quantity of findings, and are only just now beginning to individually validate and report our findings to upstream. GLib is by far the hardest hit project, which I was not expecting, accounting for more than 40% of our total findings. I’m not certain why, but perhaps this is because GLib provides so many public APIs. Data passed to public APIs is potentially untrusted, so the attack surface is considerable.

Red Hat’s scan of GLib found 118 vulnerabilities. Or at least, it claimed to. However, due to the way we ran the scans, several of these are actually unnecessary duplicates of each other, which we have not fully deduplicated yet, so the number I report is not entirely trustworthy. Moreover, 46 of these “vulnerabilities” are bugs in gobject-introspection, mostly in the typelib support, which is evidently not very robust. A typelib controls how your program calls libraries; it is effectively calling convention, so it must inherently be fully trusted: a malicious typelib would be able to induce vulnerabilities even without any bugs! I would expect an AI ought to have been able to figure that out, but apparently not. These bugs are still real problems that we ought to fix, but all maintainers agree they are not security vulnerabilities, so let’s count all of them as false positives. That alone creates a 40% false positive rate. Ouch.

I don’t have more stats to share here because we are not yet done working through the issue reports. That said, I am quite pleased with the results thus far. Substantially all of the reports are high-quality. The false positive vulnerability reports are almost all due to one particular misunderstanding and can be treated as good quality non-security bug reports, which are still valuable. Expect many forthcoming CVE assignments for the other findings.

It’s rare for Linux vendors to proactively look for software vulnerabilities, rather than waiting for security researchers to report them. This was a successful experiment in proactively seeking out problems.

Humans Still Useful

In addition to the bug bounty program, the Sovereign Tech Resilience program also sponsored a security audit for GNOME, performed by Codean Labs. This resulted in many findings in various GNOME projects. Most notably, the scope of the audit extended to Flatpak and xdg-desktop-portal, resulting in critical findings .

Most of these issues could have been detected via AI scans, but I am not confident that AIs would have been able to discover the most important findings, like the two Flatpak sandbox escapes that I linked to above. Accordingly, I do not recommend relying on AI alone.

Humanity Still Desired

Although I like AI-generated issue reports, I particularly do not appreciate when I wind up interacting with a robot rather than with a human. It’s pretty obvious when your issue tracker or code review comments are written by an AI. Consider whether outsourcing your writing and your thinking to a language model is truly wise for your public image.

We even have one experienced GNOME developer who is obviously using AI to write all of his posts on GitLab. I am unsure whether he is copy/pasting all of his responses from an AI, or whether he is just a bot now. I especially do not understand the value of this.

Here is a soft proposal, intended only as a starting point for discussion and not as a serious proposal, for what my preferred AI usage policy might look like:

  • Newer developers should exercise caution when using AI to write code. Your priority should be learning, and I wonder how much you are really learning when relying on the AI to do work for you.
  • Do not use AI to write code comments. Currents AIs are terrible at writing comments. Most comments written by AIs should be deleted. If a comment is truly necessary, then I’d like to see it written in your own words. Presumably AIs will get better at this eventually, but as of 2026, human judgment is still required here.
  • Do not use AI to write commit messages. AIs are actually probably better than humans at writing commit messages, but I would still rather hear your own thoughts on the code you are submitting.
  • Certainly do not post AI-generated comments on an issue tracker or merge request as if they are your own. You’re not fooling anybody.

Maintain Perspective

Are you scared by the large numbers of recently-discovered vulnerabilities? There is no need to panic. Security bugs are just bugs, and they’re not necessarily more important than other bugs. Occasionally they are emergencies, but far more often they are boring and unexceptional. Security vulnerabilities are not even the biggest digital security threats that users face: those are surely phishing and trojans , with software security bugs a distant third place. No amount of CVE fixing will protect you from those more likely threats.

I don’t want to downplay the severity of security issues either. In fact, evaluating severity is hard. I quite often decide that a bug is not a big deal, only to be proven incorrect. Ideally, we would fix as many security issues as possible, and sooner rather than later. Lifetime issues and out of bounds writes are especially important to fix. Two years ago, I claimed that memory safety vulnerabilities were becoming less threatening, a claim that did not age well: that is surely no longer true due to the drastically increased accessibility of AI exploit generation.

Nonetheless, volunteer maintainers should not feel obligated to fix security issues or treat them as higher-priority than other bug reports. It’s certainly good to fix problems when possible, but my request is only that you do not prohibit issue reports, not that you attempt to personally resolve every security problem yourself. When I add due dates to vulnerability reports, that represents only a disclosure deadline — because issue reports should not stay confidential indefinitely — not an expectation that you fix the issue by that date. Resolving security problems in projects used by big tech companies that depend on your software without contributing back is basically free labor for said companies, and only you can decide whether that’s how you want to spend your volunteer time.

Rust

Yes, even projects written in memory safe languages like Rust still need to allow AI-generated vulnerability reports. Rust will indeed eliminate most memory safety issues ( except in unsafe blocks ), and you can reasonably expect a Rust project to have an order of magnitude fewer vulnerabilities than a comparable project written in C or C++ or Vala. This is amazing, but not all vulnerabilities are memory safety issues, so this is not an excuse to avoid scanning for flaws.

Although Rust mostly eliminates memory safety risk, any use of Cargo to download dependencies dramatically increases supply chain security risk. The risk of bundling a trojanized dependency arguably — I would even say probably — outweighs the benefit of eliminating memory safety flaws. This problem is inherent to any programming language package manager. Currently the best solution is to not use programming language package managers, but GNOME’s Rust code depends heavily on Cargo. Accordingly, I recommend against using Rust for writing GNOME software.

To Be Continued…

I have exhausted my thoughts on AI vulnerability reports, but there is still much to discuss regarding software quality. Next time, I will discuss additional strategies to improve GNOME quality without significantly relying on AI.

Writing the Cyclone Scheme Compiler (2017)

Lobsters
justinethier.github.io
2026-10-03 09:23:29
Comments...
Original Article

Revised for 2017

by Justin Ethier

This write-up provides a high level background on the various components of Cyclone and how they were written. It is a revision of the original write-up , written over a year ago in August 2015, when the compiler was self hosting but before the new garbage collector was written. Quite a bit of time has passed since then, so I thought it would be worthwhile to provide a brain dump of sorts for everything that has happened in the last year and half.

Before we get started, I want to say Thank You to all of the contributors to the Scheme community. Cyclone is based on the community’s latest revision of the Scheme language and wherever possible existing code was reused or repurposed for this project, instead of starting from scratch. At the end of this document is a list of helpful online resources. Without high quality Scheme resources like these the Cyclone project would not have been possible.

Table of Contents

Overview

Cyclone has a similar architecture to other modern compilers:

flowchart of cyclone compiler

First, an input file containing Scheme code is received on the command line and loaded into an abstract syntax tree (AST) by Cyclone’s parser. From there a series of source-to-source transformations are performed on the AST to expand macros, perform optimizations, and make the code easier to compile to C. These intermediate representations (IR) can be printed out in a readable format to aid debugging. The final AST is then output as a .c file and the C compiler is invoked to create the final executable or object file.

Programs are linked with the necessary Scheme libraries and the Cyclone runtime library to create an executable:

Diagram of files linked into a compiled executable

Source-to-Source Transformations

Overview

My primary inspiration for Cyclone was Marc Feeley’s The 90 minute Scheme to C compiler (also video 1 , video 2 , and code ). Over the course of 90 minutes, Feeley demonstrates how to compile Scheme to C code using source-to-source transformations, including closure and continuation-passing-style (CPS) conversions.

As outlined in the presentation, some of the difficulties in compiling to C are:

Scheme has, and C does not have

  • tail-calls a.k.a. tail-recursion optimization
  • first-class continuations
  • closures of indefinite extent
  • automatic memory management i.e. garbage collection (GC)

Implications

  • cannot translate (all) Scheme calls into C calls
  • have to implement continuations
  • have to implement closures
  • have to organize things to allow GC

The rest is easy!

To overcome these difficulties a series of source-to-source transformations are used to remove powerful features not provided by C, add constructs required by the C code, and restructure/relabel the code in preparation for generating C. The final code may be compiled direcly to C. Cyclone also includes many other intermediate transformations, including:

The 90-minute compiler ultimately compiles the code down to a single function and uses jumps to support continuations. This is a bit too limiting for a production compiler, so that part was not used.

Just Make Many Small Passes

To make Cyclone easier to maintain a separate pass is made for each transformation. This allows Cyclone’s code to be as simple as possible and minimizes dependencies so there is less chance of changes to one transformation breaking the code for another.

Internally Cyclone represents the code being compiled as an AST of regular Scheme objects. Since Scheme represents both code and data using S-expressions , our compiler does not (in general) have to use custom abstract data types to store the code as would be the case with many other languages.

Most of the transformations follow a similar pattern of recursively examining an expression. Here is a short example that demonstrates the code structure:

(define (search exp)
  (cond
    ((const? exp)    '())
    ((prim? exp)     '())    
    ((quote? exp)    '())    
    ((ref? exp)      (if bound-only? '() (list exp)))
    ((lambda? exp)   
      (difference (reduce union (map search (lambda->exp exp)) '())
                  (lambda-formals->list exp)))
    ((if-syntax? exp)  (union (search (if->condition exp))
                            (union (search (if->then exp))
                                   (search (if->else exp)))))
    ((define? exp)     (union (list (define->var exp))
                            (search (define->exp exp))))
    ((define-c? exp) (list (define->var exp)))
    ((set!? exp)     (union (list (set!->var exp)) 
                            (search (set!->exp exp))))
    ((app? exp)       (reduce union (map search exp) '()))
    (else             (error "unknown expression: " exp))))

The Nanopass Framework was created to make it easier to write a compiler that makes many small passes over the code. Unfortunately Nanopass itself is written in R 6 RS and could not be used for this project.

Macro Expansion

Macro expansion is one of the first transformations. Any macros the compiler knows about are loaded as functions into a macro environment, and a single pass is made over the code. When the compiler finds a macro the code is expanded by calling the macro. The compiler then inspects the resulting code again in case the macro expanded into another macro.

At the lowest level, Cyclone’s explicit renaming (ER) macros provide a simple, low-level macro system that does not require much more than eval . Many ER macros from Chibi Scheme are used to implement the built-in macros in Cyclone.

Cyclone also supports the high-level syntax-rules system from the Scheme reports. Syntax rules is implemented as a huge ER macro ported from Chibi Scheme.

As a simple example the let macro below:

(let ((square (lambda (x) (* x x))))
  (write (+ (square 10) 1)))

is expanded to:

(((lambda (square) (write (+ (square 10) 1)))
  (lambda (x) (* x x))))

CPS Conversion

The conversion to continuation passing style (CPS) makes continuations explicit in the compiled code. This is a critical step to make the Scheme code simple enough that it can be represented by C. As we will see later, the runtime’s garbage collector also requires code in CPS form.

The basic idea is that each expression will produce a value that is consumed by the continuation of the expression. Continuations will be represented using functions. All of the code must be rewritten to accept a new continuation parameter k that will be called with the result of the expression. For example, considering the previous let example:

(((lambda (square) (write (+ (square 10) 1)))
  (lambda (x) (* x x))))

the code in CPS form becomes:

((lambda (r)
   ((lambda (square)
      (square
        (lambda (r)
          ((lambda (r) (write r))
           (+ r 1)))
        10))
    r))
 (lambda (k x) (k (* x x))))

CPS Optimizations

CPS conversion generates too much code and is inefficient for functions such as primitives that can return a result directly instead of calling into a continuation. So we need to optimize it to make the compiler practical. For example, the previous CPS code can be simplified to:

((lambda (k x) (k (* x x)))
  (lambda (r)
    (write (+ r 1)))
  10)

One of the most effective optimizations is inlining of primitives. That is, some runtime functions can be called directly, so an enclosing lambda is not needed to evaluate them. This can greatly reduce the amount of generated code.

A contraction phase is also used to eliminate other unnecessary lambda ’s. There are a few other miscellaneous optimizations such as constant folding, which evaluates certain primitives at compile time if the parameters are constants.

To more efficiently identify optimizations Cyclone first makes a code pass to build up an hash table-based analysis database (DB) of various attributes. This is the same strategy employed by CHICKEN, although each compiler records different attributes. The DB contains a table of records with an entry for each variable (indexed by symbol) and each function (indexed by unique ID).

In order to support the analysis DB a custom AST is used to represent functions during this phase, so that each one can be tagged with a unique identification number. After optimizations are complete, the lambdas are converted back into regular S-expressions.

Closure Conversion

Free variables passed to a nested function must be captured in a closure so they can be referenced at runtime. The closure conversion transformation modifies lambda definitions as necessary to create new closures. It also replaces free variable references with lookups from the current closure.

Cyclone uses flat closures: objects that contain a single function reference and a vector of free variables. This is a more efficient representation than an environment as only a single vector lookup is required to read any of the free variables.

Mutated variables are not directly supported by flat closures and must be added to a pair (called a “cell”) by a separate compilation pass prior to closure conversion.

Cyclone’s closure conversion is based on code from Marc Feeley’s 90 minute Scheme->C compiler and Matt Might’s Scheme->C compiler.

C Code Generation

The compiler’s code generation phase takes a single pass over the transformed Scheme code and outputs C code to the current output port (usually a .c file).

During this phase C code is sometimes saved for later use instead of being output directly. For example, when compiling a vector literal or a series of function arguments, the code is returned as a list of strings that separates variable declarations from C code in the “body” of the generated function.

The C code is carefully generated so that a Scheme library ( .sld file) is compiled into a C module. Functions and variables exported from the library become C globals in the generated code.

Native Compilation

The C compiler is invoked to generate machine code for the Scheme module, and to also create an executable if a Scheme program is being compiled.

Garbage Collector

Background: Cheney on the MTA

A runtime based on Henry Baker’s paper CONS Should Not CONS Its Arguments, Part II: Cheney on the M.T.A. was used as it allows for fast code that meets all of the fundamental requirements for a Scheme runtime: tail calls, garbage collection, and continuations.

Baker explains how it works:

We propose to compile Scheme by converting it into continuation-passing style (CPS), and then compile the resulting lambda expressions into individual C functions. Arguments are passed as normal C arguments, and function calls are normal C calls. Continuation closures and closure environments are passed as extra C arguments. Such a Scheme never executes a C return, so the stack will grow and grow … eventually, the C “stack” will overflow the space assigned to it, and we must perform garbage collection.

Cheney on the M.T.A. uses a copying garbage collector. By using static roots and the current continuation closure, the GC is able to copy objects from the stack to a pre-allocated heap without having to know the format of C stack frames. To quote Baker:

the entire C “stack” is effectively the youngest generation in a generational garbage collector!

After GC is finished, the C stack pointer is reset using longjmp and the GC calls its continuation.

Here is a snippet demonstrating how C functions may be written using Baker’s approach:

object Cyc_make_vector(object cont, object len, object fill) {
  object v = NULL;
  int i;
  Cyc_check_int(len);

  // Memory for vector can be allocated directly on the stack
  v = alloca(sizeof(vector_type));

  // Populate vector object
  ((vector)v)->tag = vector_tag;
  ... 

  // Check if GC is needed, then call into continuation with the new vector
  return_closcall1(cont, v);
}

CHICKEN was the first Scheme compiler to use Baker’s approach.

Cyclone’s Hybrid Collector

Baker’s technique uses a copying collector for both the minor and major generations of collection. One of the drawbacks of using a copying collector for major GC is that it relocates all the live objects during collection. This is problematic for supporting native threads because an object can be relocated at any time, invalidating any references to the object. To prevent this either all threads must be stopped while major GC is running or a read barrier must be used each time an object is accessed. Both options add a potentially significant overhead so instead Cyclone uses another type of collector for the second generation.

To that end, Cyclone supports uses a tri-color tracing collector based on the Doligez-Leroy-Gonthier (DLG) algorithm for major collections. The DLG algorithm was selected in part because many state-of-the-art collectors are built on top of DLG such as Chicken, Clover, and Schism. So this may allow for further enhancements down the road.

Under Cyclone’s runtime each thread contains its own stack that is used for private thread allocations. Thread stacks are managed independently using Cheney on the MTA. Each object that survives one of these minor collections is copied from the stack to a newly-allocated slot on the heap.

Heap objects are not relocated, making it easier for the runtime to support native threads. In addition major GC uses a collector thread that executes asynchronously so application threads can continue to run concurrently even during collections.

In summary:

  • All objects on the stack are collected using Cheney on the MTA, and the ones that survive are placed on the heap.
  • Heap objects are collected during Major GC using the DLG algorithm.
  • Heap collection runs on a separate thread in parallel with application threads.

Major Garbage Collection Algorithm

Each object is marked with a specific color (white, gray, or black) that determines how it will be handled during a major collection. Major GC transitions through the following states:

Clear

The collector thread swaps the values of the clear color (white) and the mark color (black). This is more efficient than modifying the color on each object in the heap. The collector then transitions to sync 1. At this point no heap objects are marked, as demonstrated below:

Initial object graph

Mark

The collector thread transitions to sync 2 and then async. At this point it marks the global variables and waits for the application threads to also transition to async. When an application thread transitions it will:

  • Mark its roots black.
  • Gray any child objects of the roots. The collector thread traces these gray objects during the next phase.
  • Use black as the allocation color to prevent any new objects from being collected during this cycle.

Initial object graph

Trace

The collector thread finds all live objects using a breadth-first search and marks them black:

Initial object graph

Sweep

The collector thread scans the heap and frees memory used by all white objects:

Initial object graph

If the heap is still low on memory at this point the heap will be increased in size. Also, to ensure a complete collection, data for any terminated threads is not freed until now.

More details are available in a separate Garbage Collector document.

Developing the New Collector

It took a long time to research and plan out the new GC before it could be implemented. There was a noticeable lull in Github contributions during that time:

The actual development consisted of several distinct phases:

  • Phase 0 - Started with a runtime using a basic Cheney-style copying collector.
  • Phase 1 - Added new definitions via gc.h and made sure everything compiles.
  • Phase 2 - Changed how strings are allocated to clean up the code and be compatible with the new GC algorithm. This was mainly just an exercise in cleaning up cruft in the old Cyclone implementation.
  • Phase 3 - Changed major GC from using a Cheney-style copying collector to a naive mark-and-sweep algorithm. The new algorithm was based on code from Chibi Scheme so it was already debugged and would serve as a solid foundation for future work.
  • Phase 4 - Integrated code for a new tracing GC algorithm but did not cut over to it yet. Added a new thread data argument to all of the necessary runtime functions - a simple but far-reaching change that affected almost all functions in the runtime and compiled code.
  • Phase 5 - Required the pthreads library and stood Cyclone back up using the new GC algorithm for the first time.
  • Phase 6 - Added SRFI 18 to support multiple application threads.

Heap Data Structures

Cyclone allocates heap data one page at a time. Each page is several megabytes in size and can store multiple Scheme objects. Cyclone will start with a small initial page size and gradually allocate larger pages using the Fibonnaci Sequence until reaching a maximum size.

Each page contains a linked list of free objects that is used to find the next available slot for an allocation. An entry on the free list will be split if it is larger than necessary for an allocation; the remaining space will remain in the free list for the next allocation.

Cyclone allocates smaller objects in fixed size heaps to minimize allocation time and prevent heap fragmentation. The runtime also remembers the last page of the heap that was able to allocate memory, greatly reducing allocation time on larger heaps.

The heap data structures and associated algorithms are based on code from Chibi scheme.

C Runtime

Overview

The C runtime provides supporting features to compiled Scheme programs including a set of primitive functions, call history, exception handling, and garbage collection.

An interesting observation from R. Kent Dybvig [11] that I have tried to keep in mind is that performance optimizations in the runtime can be just as (if not more) important that higher level CPS optimizations:

My focus was instead on low-level details, like choosing efficient representations and generating good instruction sequences, and the compiler did include a peephole optimizer. High-level optimization is important, and we did plenty of that later, but low-level details often have more leverage in the sense that they typically affect a broader class of programs, if not all programs.

Data Types

Objects

Most Scheme data types are represented as objects that are allocated in heap/stack memory. Each type of object has a corresponding C structure that defines its fields, such as the following one for pairs:

typedef struct {
  gc_header_type hdr;
  tag_type tag;
  object pair_car;
  object pair_cdr;
} pair_type;

All objects have:

  • A gc_header_type field that contains marking information for the garbage collector.
  • A tag to identify the object type.
  • One or more additional fields containing the actual object data.

Value Types

On the other hand, some data types can be represented using 30 bits or less and are stored as value types. The great thing about value types is they do not have to be garbage collected because no extra data is allocated for them. This makes them super efficient for commonly-used data types.

Value types are stored using a common technique that is described in Lisp in Small Pieces (among other places). On many machines addresses are multiples of four, leaving the two least significant bits free. A brief explanation :

The reason why most pointers are aligned to at least 4 bytes is that most pointers are pointers to objects or basic types that themselves are aligned to at least 4 bytes. Things that have 4 byte alignment include (for most systems): int, float, bool (yes, really), any pointer type, and any basic type their size or larger.

In Cyclone the two least significant bits are used to indicate the following data types:

Binary Bit Pattern Data Type
00 Pointer (an object type)
01 Integer
10 Character

Booleans are potentially another good candidate for value types. But for the time being they are represented in the runtime using pointers to the constants boolean_t and boolean_f .

Thread Data Parameter

At runtime Cyclone passes the current continuation, number of arguments, and a thread data parameter to each compiled C function. The continuation and arguments are used by the application code to call into its next function with a result. Thread data is a structure that contains all of the necessary information to perform collections, including:

  • Thread state
  • Stack boundaries
  • Cheney on the MTA jump buffer
  • List of mutated objects detected by the minor GC write barrier
  • Parameters for major GC
  • Call history buffer
  • Exception handler stack

Each thread has its own instance of the thread data structure and its own stack (assigned by the C runtime/compiler).

Call History

Each thread maintains a circular buffer of call history that is used to provide debug information in the event of an error. The buffer itself consists of an array of pointers-to-strings. The compiler emits calls to runtime function Cyc_st_add that will populate the buffer when the program is running. Cyc_st_add must be fast as it is called all the time! So it does the bare minimum - update the pointer at the current buffer index and increment the index.

Exception Handling

A family of Cyc_rt_raise functions is provided to allow an exception to be raised for the current thread. These functions gather the required arguments and use apply to call the thread’s current exception handler. The handler is part of the thread data parameter, so any functions that raise an exception must receive that parameter.

A Scheme API for exception handling is provided as part of R 7 RS.

Scheme Libraries

This section describes a few notable parts of Cyclone’s Scheme API .

Native Thread Support

A multithreading API is provided based on SRFI 18 . Most of the work to support multithreading is accomplished by the runtime and garbage collector.

Cyclone attempts to support multithreading in an efficient way that minimizes the amount of synchronization among threads. But objects are still copied during minor GC. In order for an object to be shared among threads the application must guarantee the object is no longer on the stack. One solution is for application code to initiate a minor GC before an object is shared with other threads, to guarantee the object will henceforth not be relocated.

Reader

Cyclone uses a combined lexer / parser to read S-expressions. Input is processed one character at a time and either added to the current token or discarded if it is whitespace, part of a comment, etc. Once a terminating character is read the token is inspected and converted to an appropriate Scheme object. For example, a series of numbers may be converted into an integer.

The full implementation is written in Scheme and located in the (scheme read) library.

Interpreter

The eval function is written in Scheme, using code from the Metacircular Evaluator from SICP as a starting point.

The interpreter itself is straightforward but there is nice speed up to be had by separating syntactic analysis from execution. It would be interesting see what kind of performance improvements could be obtained by compiling to VM bytecodes or even using a JIT compiler.

The interpreter’s full implementation is available in the (scheme eval) library, and the icyc executable is provided for convenient access to a REPL.

Compiler Internals

Most of the Cyclone compiler is implemented in Scheme as a series of libraries .

Scheme Standards

Cyclone targets the R 7 RS-small specification . This spec is relatively new and provides incremental improvements from the popular R 5 RS spec . Library support is the most important new feature but there are also exceptions, system interfaces, and a more consistent API.

Benchmarks

ecraven has put together an excellent set of Scheme benchmarks based on a R 7 RS suite from the Larceny project. These are the typical benchmarks that many implementations have used over the years, but the remarkable thing here is all of the major implementations are supported, allowing a rare apples-to-apples comparison among all the widely-used Schemes.

Over the past year Cyclone has matured to the point where almost all of the 56 benchmarks will run:

The remaining ones are:

  • mbrotZ fails because Cyclone does not support complex numbers yet.
  • pi does not work because Cyclone does not support bignums yet.
  • compiler passes but returns the wrong result. This will be fun to track down since the program is huge and takes a long time to compile…

Regarding performance, from Feeley’s presentation [10] :

Performance is not so bad with NO optimizations (about 6 times slower than Gambit-C with full optimization)

But Cyclone has some optimizations now, doesn’t it? The following is a chart of total runtime in minutes for the benchmarks that each Scheme passes successfully. This metric is problematic because not all of the Schemes can run all of the benchmarks but it gives a general idea of how well they compare to each other. Cyclone performs well against all of the interpreters but still has a long ways to go to match top-tier compilers. Then again, most of these compilers have been around for a decade or longer:

Future

Some goals for the future are:

  • Implement more of R 7 RS-large; work has already started on the data structures side.
  • Implement more libraries (for example, by porting some of industria to r7rs).
  • Improve the garbage collector. Possibly by allowing more than one collector thread (Per gambit’s parallel GC).
  • Perform additional optimizations, EG:

    Andrew Appel used a similar runtime for Standard ML of New Jersey which is referenced by Baker’s paper. Appel’s book Compiling with Continuations includes a section on how to implement compiler optimizations - many of which could still be applied to Cyclone.

In addition, developing Husk Scheme helped me gather much of the knowledge that would later be used to create Cyclone. In fact the primary motivation in building Cyclone was to go a step further and understand how to build a full, free-standing Scheme system. At this point Cyclone has eclipsed the speed and functionality of Husk and it is not clear if Husk will receive much more than bug fixes going forward. Perhaps if there is interest from the community some of this work can be ported back to that project.

Conclusion

Thanks for reading!

Want to give Cyclone a try? Install a copy using cyclone-bootstrap .

Terms

  • Abstract Syntax Tree (AST) - A tree representation of the syntactic structor of source code written in a programming language. Sometimes S-expressions can be used as an AST and sometimes a representation that retains more information is required.
  • Free Variables - Variables that are referenced within the body of a function but that are not bound within the function.
  • Garbage Collector (GC) - A form of automatic memory management that frees memory allocated by objects that are no longer used by the program.
  • REPL - Read Eval Print Loop; basically a command prompt for interactively evaluating code.

References

  1. CONS Should Not CONS Its Arguments, Part II: Cheney on the M.T.A. , by Henry Baker
  2. CHICKEN Scheme
  3. CHICKEN Scheme - Internals
  4. Chibi Scheme
  5. Compiling Scheme to C with closure conversion , by Matt Might
  6. Lisp in Small Pieces , by Christian Queinnec
  7. R 5 RS Scheme Specification
  8. R 7 RS Scheme Specification
  9. Structure and Interpretation of Computer Programs , by Harold Abelson and Gerald Jay Sussman
  10. The 90 minute Scheme to C compiler , by Marc Feeley
  11. The Development of Chez Scheme , by R. Kent Dybvig

French Bond Risk Hits Euro-Crisis Levels [video]

Hacker News
www.youtube.com
2026-10-03 09:12:28
Comments...

Two-Stack Sliding-Window Aggregation

Lobsters
orlp.net
2026-10-03 08:39:14
Comments...
Original Article

An aggregation is some kind of summary of a set of data. This can be the sum, length, minimum, etc. It is quite common to want to calculate such a summary repeatedly, e.g. “the maximum noise level in dB for the past 30 seconds” for a nuisance detector. In such a case we say there is a sliding window over our data, and we want to aggregate over our window.

If our aggregation is a binary operator with an inverse, like integer sums, there is a very easy solution using a double-ended queue:

from collections import deque

class SlidingWindowSum:
    def __init__(self):
        self.sum = 0
        self.elems = deque()
    
    def push(self, x):
        self.sum += x
        self.elems.append(x)
    
    def pop(self):
        self.sum -= self.elems.popleft()
    
    def eval(self):
        return self.sum

But what if our operator has no inverse? This is actually the case for most interesting summaries such as minimum, quantile, approximate unique count (for example using HyperLogLog ), etc. In fact, even something as simple as a floating-point sum suffers from the fact that floating-point addition is not invertible. For example, if you ever have a NaN in your input data with the above naive algorithm your sum will forever remain NaN , even long after the bad value has left your window.

Six years ago I came up with an algorithm for maintaining just the minimum/maximum in a sliding window and posted it to cs.stackexchange . I now consider this algorithm pointless, because it turns out there is a simple and efficient algorithm that solves this problem for a very wide class of aggregations. I’m writing this blog post to spread the word, because I feel it should be more widely known.

Folklore

I came across this algorithm while reading a far more advanced paper, Low-Latency Sliding-Window Aggregation in Worst-Case Constant Time by Tangwongsan et al. Why is this paper titled low-latency? Because it does the same as what I’m about to describe, but in O(1) time for each step. However, in it they also described a “two-stack” algorithm, which does it in amortized O(1), and is far, far simpler.

Funnily enough that paper attributes this algorithm to “adamax” from a 2011 Stack Overflow post . They in turn credit a 2001 lecture note by D. Sleator for the inspiration. However, this lecture note does not describe a sliding window aggregate, it describes the classical two-stack algorithm for implementing a FIFO queue and does amortized analysis on it. Ultimately I would not be surprised to find that this algorithm was already described in an obscure paper from the 1970s, seeing how simple and brilliant it is.

Two stacks

Like the authors of the paper, I will generalize the two-stack algorithm to arbitrary associative aggregation functions. By abstracting the aggregation as a set of functions, empty() , unit(x) , combine(x, y) and finalize(x) , you can describe many possible aggregations, for example a mean:

empty = lambda: (0, 0)
unit = lambda x: (x, 1)
combine = lambda x, y: (x[0] + y[0], x[1] + y[1])
finalize = lambda x: x[0] / x[1] if x[1] else None

I’d like to note here that these functions have the following signatures:

fn empty() -> Agg;
fn unit(x: Value) -> Agg;
fn combine(x: Agg, y: Agg) -> Agg;
fn finalize(x: Agg) -> Out;

I’m making a distinction here between Value , Agg and Out because while they seem superficially similar for something like an integer sum, for an approximate unique count on strings you would have (Value, Agg, Out) = (String, HyperLogLogSketch, u64) , three wildly different types.

Without further ado, the algorithm:

class TwoStackAgg:
    def __init__(self):
        self.values = []
        self.values_agg = empty()
        self.cum_aggs = []
    
    def push(self, x):
        self.values.append(x)
        self.values_agg = combine(self.values_agg, unit(x))
    
    def pop(self):
        if not self.cum_aggs:
            cum_agg = empty()
            while self.values:
                cum_agg = combine(unit(self.values.pop()), cum_agg)
                self.cum_aggs.append(cum_agg)
            self.values_agg = empty()
        self.cum_aggs.pop()
    
    def eval(self):
        return finalize(
            combine(self.cum_aggs[-1], self.values_agg)
            if self.cum_aggs else self.values_agg
        )

That’s it, the entire algorithm. There’s two stacks ( values and cum_aggs ) and one more aggregate, values_agg . At any point in time values_agg holds the aggregate of values , and cum_aggs contains the cumulative aggregates of all values in our window that aren’t in values , in reverse order. From this we can get the aggregate over our entire window in constant time by by combining the last value of cum_aggs with values_agg .

The neat part is that (assuming w is our window size) every w th operation we drain all of values and maintain a running aggregate while pushing the partial cumulative aggregates onto cum_aggs . This is what makes it amortized O(1), doing O(w) internal operations every w th pop bounds the total amount of work per element to O(1) , even though a singular operation might not be constant time.

I think this is best visualized. Suppose we sum [1, 2, ..., 10] with a fixed-size sliding window of four elements, then the state on each eval() call would look like this ( values_agg not shown as it is simply the aggregate of the values):

             cum_aggs values        out
                   [] []            = 0
                   [] [1]           = 1
                   [] [1, 2]        = 1 + 2
                   [] [1, 2, 3]     = 1 + 2 + 3
                   [] [1, 2, 3, 4]  = 1 + 2 + 3 + 4
[4, 3 + 4, 2 + 3 + 4] [5]           = 2 + 3 + 4 + 5
           [4, 3 + 4] [5, 6]        = 3 + 4 + 5 + 6
                  [4] [5, 6, 7]     = 4 + 5 + 6 + 7
                   [] [5, 6, 7, 8]  = 5 + 6 + 7 + 8
[8, 7 + 8, 6 + 7 + 8] [9]           = 6 + 7 + 8 + 9
           [8, 7 + 8] [9, 10]       = 7 + 8 + 9 + 10
                  [8] [9, 10]       = 8 + 9 + 10
                   [] [9, 10]       = 9 + 10
                 [10] []            = 10
                   [] []            = 0

In total the memory usage is O(w) , where w is your maximum window size. Note that for simplicity of analysis and the example I assumed a fixed-size window w , but there is nothing about the two-stack algorithm that requires this. You can call push(x) and pop() as many times as you’d like between each eval() , growing and shrinking the window size as needed.

Floating-point non-associativity

Note that we required above that our aggregate combine is associative, meaning:

combine(combine(x, y), z) = combine(x, combine(y, z))

Technically speaking, floating-point addition doesn’t respect this. Nevertheless, the above algorithm is still very useful because the results closely match the expected outcome, even more so if you use a compensated summation algorithm like Kahan summation .

However, there is a second very useful property of the above algorithm. Each aggregate is strictly a combination of the elements in the window, and none outside the window. This means if your window contains a NaN or infinity (or some other outlier), that value only poisons the windows that contain it rather than the rest of your computation.

But even without NaN or infinity it is useful, due to not propagating errors endlessly. E.g. if your sliding window starts with [1e20, 1] , this is what would happen with a naive rolling sum:

>>> 1e20 + 1 - 1e20 - 1
-1.0

Compensated summation will reduce these effects, but not making your result depend on values outside of the window will eliminate long-term error accumulation entirely.

Watson Jr. memo about CDC 6600 (1963)

Hacker News
www.computerhistory.org
2026-10-03 08:27:31
Comments...
Original Article

Watson Jr. memo about CDC 6600

  • Details
  • Description
Date
1963-08
Credit Line
Courtesy of the IBM Corporate Archive
Copyright Owner
© International Business Machines Corporation (IBM)
Keywords
Control Data Corporation (CDC); International Business Machines Corporation (IBM); Watson, Thomas J., 1914-1993
Object ID
500004285

Watson penned this scathing memo asking how IBM lost “our industry leadership” to a company with “34 people, including the janitor.” Also known as "The Janitor Memo," it shows the impact Control Data Corporation and its computers had on the leadership of IBM.

Go JSON v2 Migration: What Breaks in Go 1.27

Lobsters
importstatic.com
2026-10-03 08:03:08
Comments...
Original Article

Go 1.27 put encoding/json/v2 in the standard library, and the part most people will miss is that you are already running it. Since Go 1.27, the old encoding/json package is implemented on top of the v2 engine with a set of compatibility options switched on. Your output should not change, but your error strings can, and the moment you change an import to encoding/json/v2 the defaults flip: nil slices become [] , field names match case-sensitively, duplicate keys are rejected and time.Duration stops marshaling at all.

Go 1.27.0 was released on 2026-08-19 and the current point release, as of 2026-10-03, is Go 1.27.1. Every program and output below was run on go1.27.1 darwin/arm64. The benchmark section also uses GOEXPERIMENT=nojsonv2 to compare against the old implementation.

What Go 1.27 changed even if you never import v2

The Go 1.27 release notes describe two new packages and one swapped engine:

  • encoding/json/v2 is the new semantic API: Marshal , Unmarshal , MarshalWrite , UnmarshalRead , MarshalEncode and UnmarshalDecode , each taking variadic Options .
  • encoding/json/jsontext is the lower-level syntactic layer, with an Encoder and Decoder that work on tokens and raw values.

The notes then say that encoding/json "is now backed by the v2 implementation. Marshaling and unmarshaling behavior is preserved, but the exact text of error messages may differ." The v1 API stays supported and nobody is required to migrate. If something does break after the upgrade, building with GOEXPERIMENT=nojsonv2 restores the original code, but the release notes say that opt-out "is expected to be removed in a future release."

So the first migration step costs nothing: upgrade to 1.27 and run your tests. The one category to look for is tests that compare err.Error() against a fixed string. Those can fail without any behaviour change, and the fix is to assert on error types such as *json.SyntaxError or *json.UnmarshalTypeError instead.

Go JSON v2 vs v1, side by side

The package documentation for encoding/json has a "Migrating to v2" section listing every behaviour difference and the option that controls it. Reading a list is one thing; seeing your own structs change is another. Here is the struct we used:

Go

type User struct {
	Name    string   `json:"name"`
	Tags    []string `json:"tags"`
	Admin   bool     `json:"admin,omitempty"`
	Retries int      `json:"retries,omitempty"`
}

The program marshals and unmarshals the same values with jsonv1 "encoding/json" and "encoding/json/v2" and prints both. This is the real output:

Terminal

== nil slice and omitempty on bool/int ==
v1   {"name":"ana","tags":null}
v2   {"name":"ana","tags":[],"admin":false,"retries":0}
== case-insensitive field names ==
v1   "bob" err=<nil>
v2   "" err=<nil>
== duplicate names ==
v1   "mallory" err=<nil>
v2   "alice" err=jsontext: duplicate object member name "name"
== invalid UTF-8 ==
v1   "caf�" err=<nil>
v2   "" err=jsontext: invalid UTF-8 within "/name" after offset 12
== time.Duration ==
v1   {"name":"backup","timeout":90000000000}
v2   error: json: unable to marshal from Go time.Duration within "/timeout": no default representation
v2+  {"name":"backup","timeout":90000000000}
== HTML escaping ==
v1   {"q":"<a&b>"}
v2   {"q":"<a&b>"}
v2+  {"q":"<a&b>"}

Going through them in order of how likely they are to hurt.

Nil slices and maps. v1 writes null for a nil slice or map; v2 writes [] and {} . That is friendlier to JavaScript clients and a silent contract change for anyone who checks === null . json.FormatNilSliceAsNull(true) and json.FormatNilMapAsNull(true) bring the old output back.

omitempty means something else. The first line shows "admin":false,"retries":0 reappearing. In v1, omitempty drops false, 0, nil pointers and nil interfaces. In v2 it drops a field only if it would encode as JSON null , "" , {} or [] . The documentation says the two agree for strings, slices, maps and arrays, and that existing uses on a bool, number, pointer or interface "should migrate to specifying omitzero instead", which works the same in both versions.

Case-insensitive field matching is gone. v1 happily put {"NAME":"bob"} into the Name field. v2 matches names exactly, so the field stays empty and, notice, no error is returned. If your clients send inconsistent casing, you will find out from missing data, not from logs. json.MatchCaseInsensitiveNames(true) restores the loose match for a whole call; the case:ignore tag option does it per field.

Duplicate keys are an error. v1 let the last "name" win, which is how {"name":"alice","name":"mallory"} became mallory. The v2 documentation's security section explains why this matters when two services parse the same request differently. v2 rejects it unless you pass jsontext.AllowDuplicateNames(true) . One detail from our run: after the error, the destination already held "alice" . Treat the struct as garbage once Unmarshal returns an error.

Invalid UTF-8 is an error. v1 replaced the bad byte with U+FFFD; v2 refuses, unless jsontext.AllowInvalidUTF8(true) is set. If you ingest data from old Latin-1 systems, this is the one that will show up in production.

time.Duration has no representation. In v2, a time.Duration "has no default representation and results in a SemanticError". That is the v2 error line. The v2+ line passes jsonv1.FormatDurationAsNano(true) , an option that lives in the v1 package, and gets the old nanosecond integer back.

HTML escaping is off. v2 uses minimal escaping, so <a&b> goes out as is. If you embed JSON in HTML <script> blocks, keep jsontext.EscapeForHTML(true) .

The documentation lists a few more that our structs did not hit: Go arrays must now be unmarshaled from a JSON array of exactly the same length, byte arrays (not slices) become Base64 strings, maps are no longer sorted when marshaled unless you pass json.Deterministic(true) , the string tag option only applies to values that encode as numbers, unmarshaling null always zeroes the target, and malformed struct tags are reported as errors at runtime instead of being ignored.

You do not have to pick a side per call site. Most differences can be pinned on the type, so the same struct encodes identically whichever package touches it. This version of the struct passes a test that marshals with both and compares bytes:

Go

// Portable keeps the same JSON under v1 and v2.
type Portable struct {
	Name    string   `json:"name,case:ignore"`
	Tags    []string `json:"tags,omitempty"`
	Admin   bool     `json:"admin,omitzero"`
	Retries int      `json:"retries,omitzero"`
}

func TestPortableSameBytes(t *testing.T) {
	for _, p := range []Portable{{Name: "ana"}, {Name: "bo", Tags: []string{"x"}, Admin: true, Retries: 3}} {
		b1, err := jsonv1.Marshal(p)
		if err != nil {
			t.Fatal(err)
		}
		b2, err := json.Marshal(p)
		if err != nil {
			t.Fatal(err)
		}
		if string(b1) != string(b2) {
			t.Errorf("v1 %s != v2 %s", b1, b2)
		}
		t.Logf("%s", b2)
	}
}

Terminal

=== RUN   TestPortableSameBytes
    main_test.go:31: {"name":"ana"}
    main_test.go:31: {"name":"bo","tags":["x"],"admin":true,"retries":3}
--- PASS: TestPortableSameBytes (0.00s)

omitempty on the slice makes nil and empty both disappear, which sidesteps the null versus [] question. omitzero replaces omitempty on the bool and int. case:ignore keeps accepting "NAME" under v2, and a second test confirmed it decodes into Name .

While you are in there, v2 has an option v1 never had. json.RejectUnknownMembers(true) turns a typo or an unexpected field into an error that wraps json.ErrUnknownName :

Terminal

json: cannot unmarshal JSON string into Go main.Portable: unknown object member name "role"

That is worth having on any request body that reaches an authorization check.

If you used v2 under GOEXPERIMENT=jsonv2 on Go 1.25 or 1.26, check your tags again. The release notes list changes made before the package became final: the format and unknown tag options were removed, the DiscardUnknownMembers option and SkipFunc were removed, and the inline tag option was renamed embed . Blog posts written during the experiment still show format: tags that will not work on 1.27.

Migrating one option at a time with DefaultOptionsV1

For a large codebase, the safe path is the one the official migration guide calls option-by-option. Change call sites to the v2 functions but pass jsonv1.DefaultOptionsV1() , which the guide calls "a trivial and safe change". Then switch individual behaviours to v2 by adding options after it, since later options override earlier ones:

Go

b, err = json.Marshal(u, jsonv1.DefaultOptionsV1())
show("v2v1", b, err)
b, err = json.Marshal(u, jsonv1.DefaultOptionsV1(), json.FormatNilSliceAsNull(false))
show("mix", b, err)

Terminal

v2v1 {"name":"ana","tags":null}
mix  {"name":"ana","tags":[]}

The first call is byte-for-byte v1. The second is v1 except for nil slices. Each option you flip is a small, reviewable change, and when the list of overrides covers everything you can drop DefaultOptionsV1() entirely.

For servers where you cannot predict every payload from tests, the guide points to github.com/go-json-experiment/jsonsplit . It has call modes such as CallBothButReturnV1 , which runs both implementations, returns the v1 result and reports any difference, so production traffic tells you which options you need before you switch to CallBothButReturnV2 or OnlyCallV2 . The guide warns that running both "will approximately double the cost of marshaling". Use it for the migration window, not permanently.

What we measured: faster unmarshal, slower marshal

The release notes say "marshal performance is broadly at parity with the previous implementation, while unmarshal performance is significantly faster." We checked that on one payload: a 500-element array of a struct with an int64, a string, a float, a three-element string slice, a two-entry map[string]string and a nested struct, about 79 KB of JSON. The benchmark calls plain encoding/json , run twelve times on the default 1.27.1 build and six times with GOEXPERIMENT=nojsonv2 , on an Apple M4 Pro:

Times are rounded medians; the unmarshal runs were noisy (537 to 878 µs on the new engine), while the allocation counts were identical on every run. Unmarshal came out about 25% faster with 29% fewer allocations, which matches the release notes. Marshal took about 39% longer on this payload, which does not match "broadly at parity". One payload on one machine is not a verdict, but it is enough to say: if marshaling is on your hot path, benchmark it on 1.27 with and without nojsonv2 before you roll out.

If you are upgrading for the other Go 1.27 changes as well, the new synctest.Sleep and in-memory test server are worth a look in the same pass.

Common questions

Do I have to migrate to encoding/json/v2 in Go 1.27?

No. The Go team says the v1 encoding/json API will continue to be supported and users are not required to migrate. In Go 1.27, though, v1 is implemented on top of the v2 engine, so error message text can change even if you never touch your imports.

How do I turn off the new JSON implementation in Go 1.27?

Build with GOEXPERIMENT=nojsonv2. That restores the original v1 implementation behind encoding/json. The release notes say this opt-out is expected to be removed in a future release, so treat it as a stopgap and file an issue for whatever made you need it.

Why does json/v2 fail to marshal time.Duration?

In v2 a time.Duration has no default representation, so Marshal returns a SemanticError. Pass the encoding/json option FormatDurationAsNano(true) to get the v1 behaviour of encoding nanoseconds as a JSON number.

Is omitempty the same in json v2?

Only for strings, slices, maps and arrays. v2 omits a field when it would encode as JSON null, an empty string, an empty object or an empty array, so false and 0 are no longer omitted. Use omitzero on bools, numbers, pointers and interfaces; it behaves the same in v1 and v2.

Is encoding/json/v2 faster than v1?

The Go 1.27 release notes say unmarshal is significantly faster and marshal is broadly at parity. In our benchmark on Go 1.27.1, unmarshal was about 25% faster with 29% fewer allocations, while marshal took about 39% longer on that payload, so measure your own hot paths.

ImportStatic articles are researched and drafted with AI assistance, checked against primary sources, and every code sample is compiled and run before publishing. Found a mistake? Here is how corrections work.

Great Question (YC W21) Is Hiring Product Engineers in Canada (Remote)

Hacker News
www.ycombinator.com
2026-10-03 08:02:10
Comments...
Original Article

About Great Question:

Great Question is the all-in-one AI customer research platform for understanding your customers. Our platform enables teams to recruit participants, run research, and share insights – all in one place. Backed by world-class investors and trusted by industry-leading teams like Gusto, Experian, Canva, and Brex, we’re building the future of customer research.

We believe AI is a multiplier for creativity, speed, and scale. Everyone on the team is encouraged to lean into AI, not just to be more efficient, but to push boundaries and unlock new possibilities. Whether you’re writing, designing, coding, or analyzing, we expect you to explore how AI can elevate your craft.

Our culture is built on high trust, high agency, and high emotional intelligence. If you’re looking for a place where your voice matters, where work is fun, and where you're empowered to do your best - join us in redefining user research 🚀

Product Engineer (Full-Stack) — Canada, Remote

I’m PJ, CTO at Great Question . My co-founder Ned and I built this company on a simple idea: talking to your customers should be easy. This has only become more relevant as it’s become faster to build, but validating your ideas is still slow and painful.

Research, design, and product teams at ServiceNow, Brex and Canva run their customer research on us, and we raised a $13M Series A led by Inovia late last year after going through YC in 2021.

This year we’ve been doubling down on the agentic future of research. We’re super excited about what we’ve built and now we need help to continue rebuilding core parts of our product, agent-first. I’m hiring our first Product Engineers in Canada to do it with us.

The role

Product Engineer here means you take a problem and run with it end to end - from the customer conversation to production to validating whether it actually worked. PM support is light and as-required: a PM helps size the problem, shape the appetite, and call priority, but nobody hands you a spec. From there you're trusted to find the right solution, because you're close to the customer and you have great product sense. That only works with people who are self-directed and genuinely value autonomy and accountability.

You'll build across the product stack - from Rails APIs and React/TypeScript frontends to video streaming, realtime voice agents, agentic search, product analytics, and everything in between - working directly with the founders, our designer, and the rest of engineering.

What we're looking for

How you work. AI is embedded across your entire workflow - not just used for coding. You might spend a morning digging through interview transcripts and product analytics to understand a problem, build a few throwaway prototypes to see which approach holds up, and then build a technical plan before handing it to a couple of agents to build and QA.

We’re expecting great judgment around all of that: breaking work into sensible pieces, evaluating the output carefully, and most importantly, owning whatever ships, regardless of who typed it.

What you build. You build highly usable, customer-centric products that people love. You take the extra time to really dig into the data to make sure you understand the problem; and validate that what you’ve built actually solves it. You’ll go the extra mile to deliver something with UI + UX polish, while also thinking in systems to make sure that you’re using the right primitives so you’re delivering less bespoke solutions.

Who you are. We're a small team of exceptional engineers, and we only hire people who raise the bar.

  • Strong across the full stack. An expert in both exposing APIs and the frontend it sits upon. We’re a Rails and TypeScript team, but it’s not a hard requirement.

  • Built and run production systems. Shipping features live and then maintaining them long enough to see what works (and what doesn’t!)

  • Live in an agentic coding harness daily (Claude Code, Codex, Cursor - we don't care which) and have real opinions about your setup.

  • Ideally shipped LLM-powered features, or least know what it takes to run them in production.

  • Have shipped things where you decided what to build, and can defend your UX opinions.

  • Write clearly. We're remote and async-heavy; people who write clearly think (and prompt) clearly.

  • Are senior (roughly 5+ years with real ownership). Our level grading is tougher than most.

How we build

We've rebuilt how we work around AI agents, to ensure our engineers are unblocked to work at AI native speeds:

  • Main goes to production and approvals aren’t always required. Fast, automated CI is the gate, and we trust people to make the right judgements. Good work goes live as soon as it’s finished.

  • We invest in making the codebase easy to work in: written runbooks and conventions that help agents and new people find their way around quickly.

  • Running a few agents in parallel is easy, and expected. Sandboxes and preview environments are trivial to spawn.

  • You'll always have the latest tools and models, and we're generous with tokens.

  • Light on ceremony, high on accountability and communication.

This might not be for everyone, but if it excites you, please apply. We’re especially keen to meet people who are willing to challenge some of these ideas with their own hard-won lessons from the arena.

Your first 90 days

We expect you to hit the ground running, and fast.

  • Week 1: use the harness to learn the codebase and merge your first PR, ideally on day one. You’re deep in the customer data, in their Slack channels, and solving problems.

  • First 30 days: ship into a live area. That might be AI Moderator quality, or working directly on our global chat experience. You understand the product, the competitors, the problems we’re solving, and have a view on how we can win.

  • By 60 days: You own a surface - how the research agent communicates findings, growing our MCP surface, or rebuilding a core page on the new design system.

  • By 90 days: you're deep on a product area: talking to its customers, deciding what's next, accountable for whether it worked. You’re pushing the strategy with the founders & Head of Product in new directions with confidence.

Who shouldn't apply

  • Ambiguity and shifting priorities wear you down. Problems arrive here half-shaped, and the roadmap has changed under our feet more than once this year. Some people find that energizing; if you don't, you'll hate it.

  • You wait to be unblocked. We’re a small team, remote, and with light processes… resourcefulness is critical to the job. We need people who unblock themselves.

  • You’re not all-in on AI. Dabbling won't cut it here. Agents are how we work, not a thing we're trying out.

  • You'd rather not talk to customers. Product engineers here get on calls, read transcripts, and sit in the support queue sometimes. That closeness is the whole point of the role.

  • You only want the shiny work. Plenty of weeks are polish, reliability, and making something 5% less annoying.

  • You want a PM to frame the problem and hand you a spec. “Just tell me what to build and I’ll build it” will not fly here.

Benefits

  • Full-time employment with proper benefits, generous salary in CAD.

  • Remote, anywhere in Canada. We generally work PST hours, but are flexible.

  • 4 weeks PTO

  • Annual team offsite (this year is San Diego!)

  • Brand new Macbook pro

  • Enough tokens that you'll never be left wanting

  • Freedom to use the latest technologies and vendors

A note on compensation: Our offers sit in a defined band for the level, and set by our internal leveling rubric (scope of ownership, technical depth, previous relevant experience), not by negotiation. Most offers land near the midpoint.

Our interview process

  1. Intro call w/ Cofounder/CTO (30 min). Your story, what you've shipped. Our story, the impact you can have here.

  2. Quality session (~ 60 min). Do you know what quality looks like - both in terms of product and code.

  3. Building session (~90 min) . Build something real in our domain with your own AI tooling, the way you actually work; we walk through the code together as you go. We reimburse AI credits.

  4. Team + values (~60 min) . Meet the people you'll work with and ensure we all see the world the same way.

  5. Founder conversation (~30 min) . Meet the CEO.

  6. References, then offer.

No leetcode. AI tools aren't just allowed - they're expected just as they are in the job.

How to apply

Send a link to something you've built that real people are using - and a few sentences on why you’re interested in this role.

Why You’ll Love Working Here:

  • A culture of customer obsession and curiosity

  • Opportunity to shape and scale a critical business function

  • Remote-first culture with high trust and high autonomy

  • Annual team retreats and virtual events

  • Opportunity to work on cutting-edge AI integrations that matter

  • Competitive compensation & equity

  • Generous PTO, health benefits, and learning stipend

_ Equal Opportunity Statement _

Great Question is committed to providing a workplace free from discrimination or harassment. We expect every member of the Great Question community to do their part to cultivate and maintain an environment where everyone has the opportunity to feel included, and is afforded the respect and dignity they deserve.

Decisions related to hiring, compensating, training, evaluating performance, or terminating are made fairly, and we provide equal employment opportunities to all qualified candidates and employees. We examine our unconscious biases and take responsibility for always striving to create an inclusive environment that makes every employee and candidate feel welcome.

How to Hack Time, With C2PA

Lobsters
www.da.vidbuchanan.co.uk
2026-10-03 07:58:14
Comments...
Original Article

By David Buchanan (aka retr0id), 2 nd October 2026

The most impressive hacking stunt from cinema history comes from Kung Fury (2015) , in which Hackerman hacks time itself. He uses this power to correct historic misdeeds. But what would I do with the ability to hack time? Personally I'm more afraid of the butterfly effect, so I'd just go back a few hours to tell myself the winning lottery numbers.

As it happens, that's exactly what I did, according to the cryptographically unforgeable C2PA metadata of this image:

Yes, you can tell it's photoshopped. I'm not trying to do actual lottery fraud here.

You can verify the C2PA metadata including the timestamp at https://verify.contentauthenticity.org/ (if you're reading this in the future, maybe they've introduced mitigations).

You can confirm that those are the winning lottery numbers, shown hours before the draw time, at https://www.euro-millions.com/results/28-08-2026 .

Background

A typical C2PA manifest contains two signatures. The first is the "claim" signature, and in the case of a camera app the claim might be something like "this is a captured photograph, taken at these GPS coordinates, at this time" (except expressed more formally, per the C2PA spec).

In my previous article I showed that the claim signature is approximately worthless, and we can sign whatever claims we like (at least, we can in the case of flagship C2PA implementations like the Google Pixel Camera app).

But there's usually a second signature, from a Time Stamp Authority (TSA), and this one is more interesting. We don't trust devices to tell their own time, so instead a remote server is consulted, via the RFC 3161 protocol. The server sends back a signature that asserts "yes, I saw this hash at this timestamp", and this response is embedded into the C2PA metadata. As long as we trust that the server isn't misbehaving, this proves that the data being signed existed at-or-before the specified timestamp. The TSA acts like an independent witness.

For the Pixel 10, Google decided that devices can in fact be trusted to tell their own time. I have not evaluated their "on-device trusted time-stamp" implementation yet, but I'm highly sceptical of it.

Despite this, in this article we will not be attacking the TSA: we assume it functions as advertised.

It's Time-Hacking Time

If the TSA mechanism is secure, how are we going to hack time?

We're going to use my favourite bug class: the spec footgun .

The footgun is as follows: C2PA allows for arbitrary "exclusions". These are byte ranges within the file which are excluded from signature calculations. Yup. Really. This is already known (it's literally in the spec ), but for some reason nobody's done anything about it yet.

In fact, Dr. Neal Krawertz explicitly called it out in his " Big Bulleted List " of C2PA flaws published June 2025:

  • Large exclusion range. The manifest typically excludes a very large byte range from the signatures. Any excluded bytes can be altered without detection.

The only new angle here is that a malicious signer can deliberately use a large exclusion range, rather than merely doing so incidentally. We can exclude the entire file, to produce an entirely valid signature over an empty string. This allows the file to be tampered with after the fact, without invalidating the signature, and without invalidating the TSA's timestamp proof.

Time status: Hacked

So I really did take a picture of a lottery ticket, and attach a valid C2PA signature with a valid trusted timestamp. But I crafted the manifest to exclude the whole file, allowing me to photoshop it (poorly) after the numbers were announced, without invalidating any of the signatures.

Using c2patool -d to dump the manifest of my PoC file, we can see the important part:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
"c2pa.hash.data": {
  "exclusions": [
    {
      "start": 0,
      "length": 3995383
    }
  ],
  "name": "jumbf manifest",
  "alg": "sha256",
  "hash": "47DEQpj8HBSa+/TImW+5JCeuQeRkm5NMpJWZG3hSuFU=",
  "pad": []
},

3995383 is the length of the entire file, and 47DE...uFU= is the hash of an empty string:

$ openssl sha256 -binary /dev/null | base64
47DEQpj8HBSa+/TImW+5JCeuQeRkm5NMpJWZG3hSuFU=

The claim signature is over that hash, and the timestamp signature is over the claim signature. We're effectively signing nothing at all, but as of today all the C2PA verification tools I can find don't flag anything as unusual.

Can it be fixed?

Not easily! Sure, it's trivial to detect when the entire file has been excluded and report it as invalid, but what if only a small part is excluded? How do you tell whether it's something harmless, or something that could completely change the appearance of the image if modified? (See MD5 hash collision PoCs for examples of the latter).

My first draft of this post ended here with "My recommendation is that the exclusion feature should be excluded from the C2PA spec," but it's not that simple!

The main reason exclusions exist in the first place is that certain file formats effectively require it. For example in PNG files , each chunk has a CRC32 checksum, which needs to be corrected after the signature has been embedded into the file. This would create a circular dependency, unless the CRC32 is excluded from the signature's coverage (ignoring clever mathematical tricks that could avoid invalidating the CRC).

With that in mind I think the right solution here is to carefully and explicitly specify which parts of a file are allowed to be excluded, for each supported file format, and require that verifiers enforce these constraints.

All blog content produced by thinking meat , unless noted otherwise.

Homepage - Blog Index - RSS

Security in the LLM Age

Lobsters
www.youtube.com
2026-10-03 07:57:02
Comments...

The Escalation of War in Ethiopia

Hacker News
www.africanistperspective.com
2026-10-03 07:54:24
Comments...
Original Article

Thank you for being a regular reader of An Africanist Perspective . If you haven’t done so yet, please hit subscribe to receive timely updates on new posts along with over 36,000 other subscribers. New regular content is free. Book reviews and the archives are gated.

This is from the ICG:

Northern Ethiopia is slipping into renewed war. After months of building tensions, on 23 September large-scale hostilities broke out between the Tigray People’s Liberation Front (TPLF) and Amhara Fano militias, on one side, and the federal government and allied Tigrayan forces on the other. Responsibility for the escalation lies primarily with the TPLF, which launched ground assaults in coordination with other rebels with whom it has joined hands in a new alliance that vows to overthrow Ethiopia’s government. The upsurge in fighting comes almost five months after the TPLF retook power in the Tigray region, dealing a blow to a 2022 peace deal that had ended the last Tigray war and leading Addis Ababa to mount an intensifying drone campaign against TPLF forces.

I maintain that war in Ethiopia (and the wider Horn) is not inevitable . The current outbreak of war reflects the fact that Ethiopian elites keep choosing war over peace. Unfortunately, despite repeatedly choosing war, Ethiopian leaders continue to enjoy support from a cast of warmongering academics, journalists, public intellectuals, civil society actors, and a whole host of foreign actors — all of whom have long taken sides and become cheerleaders for their respective camps.

Let’s be blunt. There are a lot of things not to like about the Abiy Ahmed administration. The strong elements of autocratic intolerance, often expressed in the form of indiscriminate force. The unwillingness to (publicly) countenance alternative worldviews and policy positions. The abiding reliance on the old vindictive governance styles anchored around identity politics and collective punishment. And the hubristic promotion of a personality cult, which makes it difficult to build the broad and durable coalitions needed to govern a diverse country like Ethiopia.

Perhaps the political situation would be marginally better if Prime Minister Abiy were a better politician.

However, I doubt it. To understand why, it’s worth stepping back to appreciate the structural drivers of elite instability and conflict in Ethiopia :

The last three decades in Ethiopia have been a search for a new myth. The ethno-federalist system had legitimate logic: bringing about the dignity of (cultural and linguistic) difference between nations and nationalities. However, its rhetoric was drawn from the difficult past instead of the hope of better future. To make matters worse, it became a breeding ground for social and economic injustice.

In the absence of farsighted political elites who may have been able to craft a new inclusive myth out of the stories of nations and nationalities, ethnic groups had to walk back to find their stories in their own small compartments. This exacerbated narrow ethnic histories and ideals.

It’s easy to demonize Abiy. But the hard truth is that Ethiopia’s structural problems preceded him and, if left unresolved, will outlast him.

Those beating the war drums the loudest in Ethiopia — and who’ve trashed the Pretoria Agreement — are convinced that Ethiopia can only be well governed if they are in charge. Nearly all are motivated by maximalist demands and are quick to weaponize legitimate historical grievances for parochial ends. The leading characters involved are (mostly) men out to settle decades-old scores and who prefer to go out fighting rather than accept political defeat by their perceived inferiors. These same characters have been able to mobilize fighters under the banner of uncompromising primordislist ethno-nationalism and commitments to their respective ethno-states.

In light of all this, Teklehaymanot G. Weldemichel is spot in making the case for following through on the Pretoria Agreement. Yes, it is imperfect and the various belligerents haven’t always followed through on the deal’s requirements. However, the deal “ brought large-scale fighting to an end and contains commitments that address some key unresolved issues .” It’s a start. It ought to be given a chance.

Unfortunately, the plight of civilians hardly ever features in the calculus of the warmongers. Estimate vary, but the death toll in the Tigray war may have been as high as a staggering 600,000. This is a monumentally catastrophic figure. Even the lower bound estimates out there would make it one of the deadliest wars of the 21st century. Given what we know about how belligerents go about waging war in Ethiopia (including by exacerbating the problem of food insecurity), why would anyone support a return to all out war?

Why are Ethiopian elites and their cheerleaders willing to impose such trauma on their peoples?

There is bound to be a lot of discourse on ethnic polarization as a driver of conflict in Ethiopia. However, it is worth also considering how Ethiopia’s brand of federalism contributes to ethnic polarization. Existing research shows that the creation of ethno-states after the end of the civil war increased the salience of ethnic difference :

[I]ndividuals politically socialized under the conditions of ethnic federalism are indeed significantly more likely to embrace an ethnic rather than a national identity. However, generational affects do not appear to be related to perceptions of continuing ethnic federalism as a political arrangement. Differences of opinion on ethnic federalism are more a function of one’s group identity, rather than the intensity of ethnic identification.

This is in line with the historical experience from other jurisdictions (see Lebanon, the UK, etc). Institutionalizing ethnic difference through administrative means or identity-based reservations creates incentives to invest in the targeted identities as loci of political mobilization. From that point it’s a short walk to mobilizing those same differences for conflict under permissive conditions. Notice the implication of the last line in the above quote: once ethnonationalism becomes the dominant game, everyone gets forced to pick a side, regardless of their individual-level intensity of ethnic identification.

Most people kind of admit that a leading structural driver of political instability and conflict in Ethiopia is its brand of federalism :

Ethiopia’s federal system was flawed from the beginning because it didn’t foresee potential sources of conflict or that regional states would make claims against one another. Trust among regional states was never high, and has deteriorated over the last three decades. On top of this, the federal government’s ability and readiness to mitigate or solve domestic conflict has been open to question. Currently, there are regions and regional leadership that are having difficulty working together. The federal government at the centre is too weak to impose its will on the regional administrations. The result is that there aren’t common political and economic national standards across the country.

To be clear, Ethiopian civilians do not want war; and also don’t trust the state’s ability to avoid or end wars. And while reported inter-ethnic hostility remains low (despite the historical wars), the rate of encounters with state-directed ethnic discrimination ranks among the highest in Africa.

In Round 9 of the Afrobarometer Survey, 55% of respondents noted that the state is doing badly at preventing or resolving violent. In the same survey, less than a quarter (23.7%) of respondents reported feeling that their ethnic identity trumps their Ethiopian identity. A roughly similar proportion (23.6%) also reported not at all trusting people from other ethnic groups. Meanwhile, 39.3% reported feeling that their ethnic group is never treated unfairly by the government.

There are two takeaways from this. First, the problem appears to be the state/institutions rather than society. The vast majority of Ethiopians seem open to the idea of multicultural co-existence. However, the state (and supporting institutions — e.g., ethnic federalism) seem intent on balkanization by reinforcing difference in everyday interactions. Second, while Ethiopia is in the bottom fifth of the distribution across the Continent, it’s not particularly unique (unlike what many Horn scholars would like to believe). There are other countries with close or worse results on the metric of state-directed ethnic discrimination.

Overall, the most important driver of conflict in Ethiopian is the state’s centuries old inability to deter or crush rebellions in their infancy. All polities have pockets of grievances against the state that could potentially give rise to armed rebellion. While some grievances invite legitimate armed resistance, others are the manifestations of taste-based warmongers, war profiteers and their clients, and useful idiots carrying water for geopolitical adversaries. Regardless of their founding motivations, not all would be rebellions materialize — because capable states don’t let them.

Instead of monopolizing the use of violence, the Ethiopian state has adapted to using indiscriminate violence and chaos as governance styles (to manage intra-elite political instability, and to fight back against mass level identity-driven centrifugal forces). Collective punishment of whole communities is common, and has, in turn, radicalized ever more people against the state. Consequently, there won’t be any shortage of young people ready to take up arms against the state in this round of wars.

The Ethiopian state’s inability to deter or eradicate violence invites all manner of foreign meddling. This is terrible for conflict duration, as it prevents both state and non-state belligerents from fully internalizing the cost of war.

Lots of countries — including Egypt (the Nile/wider Red Sea supremacy), Eritrea ( Ethiopia’s desire for access to the sea ), Sudan (Ethiopian support for the RSF), Somalia (historical enmity/Ethiopian support for Somaliland), and Saudi Arabia (as benefactor of Egypt/Sudan/Eritrea axis) — have good reasons to keep Addis Ababa weak and preoccupied with unending domestic wars. These dynamics, coupled with the geography of war in Ethiopia mean that Addis Ababa will struggle to keep arms and fighters from easily moving across its borders with Sudan and Eritrea.

On its part, the Ethiopian state will likely continue to enjoy support from the United Arab Emirates; and maintain the capacity to produce the weapons needed to keep rebels at bay. Ominously, divisions among elites in Tigray and other regions mean that the rebellion may not clear the threshold needed to incentivize a quick resolution either militarily or at the negotiating table. Therefore, slow-burning, but still catastrophically deadly conflicts may be on the cards.

A key risk facing Ethiopians is that, like in Sudan , international/geopolitical interests and considerations may quickly overshadow the domestic cost of the war, especially on civilians. The inability to extricate the domestic bargaining “game” from regional/geopolitical considerations will unnecessarily complicate peace negotiations and very likely prolong the conflict.

Notably, a prolonged conflict will likely also be significantly civilianized. Careful to avoid being overstretched, the Ethiopian state may decide to deputize regional militias and other violence entrepreneurs in its fights against subnational armed groups. To this end, it will lean heavily on divisions across armed groups and within the regional states that have witnessed rebellions.

As noted above, it would be a mistake to characterize political instability and conflict in Ethiopia as being caused by any single leader. Removing Abiy from power — the stated goal of the latest armed coalition against the Ethiopian state — will not solve the many structural problems facing Ethiopia. Which is why it is important for all involved to have honest conversations about the fundamental causes of conflict, and how to get out of the cycles of violence. Four issues immediately come to mind:

  1. The thorny question of what it means to be Ethiopian, in a context of diverse peoples, nations, and nationalities: What does it mean to be Ethiopian? Should Ethiopians deal with their rich cultural tapestry and history by genuflecting to ideas motivated by narrow ethno-nationalisms or embracing cosmopolitan multiculturalism as a national goal? These questions do not have easy answers, and I cannot pretend to offer any here. However, it’s high time Ethiopians openly confronted these questions rather than ceding ground to those out to weaponize the country’s history to stoke ethnic tension and wage perpetual wars of self aggrandizement.

  2. The meaning of self-determination in a cosmopolitan 21st century state: What is the correct structure of the Ethiopian state? And what should be the balance of power between national and subnational units? If the resolution to the questions above converge on keeping Ethiopia whole (as I hope is the case), the next question would be to come up with institutional and administrative structures that at once give ordinary Ethiopians a sense of self-determination at the local level and empowers the center to prevent centrifugal tendencies from collapsing the state.

  3. How to stabilize Ethiopia’s regional relations: While Ethiopia’s problems are principally domestic, it’s also true that its neighborhood (comprised of states with histories of internal conflict) and antagonistic foreign posture (especially vis-a-vis Somalia and Eritrea) exacerbate this problem. Of course there are things that Addis Ababa cannot change about its foreign relations (e.g., need to check Egypt’s supremacy ambitions). However, there are others it can. It can choose to pursue a path that ends the proxy wars with Sudan, Eritrea, and Somalia. This will reduce overall conflict risk in the wider Horn, while also starving would-be Ethiopian rebels from getting ready support from regional rivals and generally reducing the size of the Horn’s war economies (arms, fighters, ideologies).

  4. What’s the positive case for Ethiopian unity? Is it to reclaim the country’s imperial glory? Is it to aggrandize the person of the Ethiopian leader? Is it to become a geopolitical powerhouse in the Horn and the Red Sea regions? Or is it to deliver economic growth, development, and prosperity for Ethiopians? Some commentators like to compare Ethiopian with the former Yugoslavia, thereby raising the specter of partition. This should be viewed as a challenge for those that want to see a united Ethiopia. If people are to continue buying into the Ethiopia Project , there has to be a positive case for doing so. Merely relying on history or inertia will not be enough. And importantly, the pitch has to be sufficiently inclusive and must accommodate changing public attitudes. In particular, Ethiopian elites must update and understand that the old political culture of violence, domination, and humiliation as styles of governance will not go down well with current and future generations of Ethiopians.

In the final analysis, and understanding that the sources of instability and conflict in Ethiopia are principally domestic, there is really no better way to put it that this :

Ethiopia’s federal government must be pressed to exercise restraint and return to meaningful political negotiation. Military force and coercion cannot become the first choice for resolving disputes. Tigrayan political and security actors must likewise recognise that political disagreements cannot be sustainably resolved through armed mobilisation. External powers should not be permitted to turn Ethiopia’s internal political disputes into a wider contest for influence. The involvement of the United Arab Emirates, Eritrea and others supporting the warring parties should be addressed through transparent diplomacy. Their competing interests risk fuelling a wider confrontation.

Zig 0.17 released

Linux Weekly News
lwn.net
2026-10-03 07:31:18
Version 0.17 of the Zig programming language has been released. This release features 5 months of work: changes from 206 different contributors, spread among 925 commits. Originally predicted to be shorter, this release cycle ended up [being] substantial, with the Build System reworked, including...
Original Article

Version 0.17 of the Zig programming language has been released.

This release features 5 months of work : changes from 206 different contributors , spread among 925 commits .

Originally predicted to be shorter, this release cycle ended up [being] substantial, with the Build System reworked, including the introduction of the Build Server Protocol , and the ELF Linker enhanced to the point where we expect Incremental Compilation to work for everyone on x86_64-linux.

LWN last covered Zig in December 2025.



Seven stable kernels for Saturday

Linux Weekly News
lwn.net
2026-10-03 07:21:42
Greg Kroah-Hartman has announced the release of the 7.2.9, 6.18.55, 6.12.112, 6.6.158, 6.1.189, 5.15.222, and 5.10.271 stable kernels. Each contains a large number of important fixes throughout the tree; users are advised to upgrade. ...
Original Article

[Posted October 3, 2026 by jzb]

Greg Kroah-Hartman has announced the release of the 7.2.9 , 6.18.55 , 6.12.112 , 6.6.158 , 6.1.189 , 5.15.222 , and 5.10.271 stable kernels. Each contains a large number of important fixes throughout the tree; users are advised to upgrade.



to post comments

Angie Nixon Says the Fight for Black Voting Power Demands Elected Officials ‘Get Louder’

Portside
portside.org
2026-10-03 07:08:34
Angie Nixon Says the Fight for Black Voting Power Demands Elected Officials ‘Get Louder’ Kurt Stand Sat, 10/03/2026 - 07:08 ...
Original Article

Credit: South Florida Sun Sentinel, Creator: Carline Jean

Angie Nixon has spent more than two decades advocating for Black and working-class communities in Florida. But the Jacksonville native didn’t make a splash on the national stage until she stunned many in her bid to become the state’s first Black senator.

The Democratic Socialist trounced her better-funded opponent Alex Vindman in Florida’s Democratic primary for U.S. Senate. She did so at a time of intense debate over whether the party’s left flank, whose members have largely been winning their races, can continue to connect with voters ahead of the midterm elections in November. Centering issues such as Medicare for All in her platform, Nixon will compete against Republican U.S. Sen. Ashley Moody.

But until Tuesday, Nixon’s surprise victory had come with a hitch that most Senate nominees don’t have to worry about: a criminal case . She faced misdemeanor trespassing and other charges springing from a protest outside Gov. Ron DeSantis’ office in May. Nixon was arrested after she refused to leave while seeking a meeting with DeSantis about a redistricting plan that advocates say would dilute Black voting power.

Nixon’s case was resolved on Tuesday, with her agreeing to pay a $250 fine and serve 15 hours of community service. She could have received up to two years in jail.

To Nixon, the only thing to do now is to get louder. She traces that political belief partly to her family — including a grandmother who had to hide during an episode of racial violence known as Ax Handle Saturday — and to Black leaders who came before her. As some Black Floridians face shrinking representation, aggressive immigration enforcement, and growing disillusionment with the government, Nixon wants Black elected officials to hold the line.

DeSantis has dismissed Nixon’s protest as “performative nonsense.” Lawyers for the state also have defended the new map against claims that it violates Florida’s constitution.

Earlier this month, Capital B spoke with Nixon — who’s been endorsed by U.S. Reps. Maxwell Frost of Florida and Ilhan Omar of Minnesota and other progressive political leaders — about her state’s redistricting battle, her trial, and her reasons for wanting her peers to “get louder.” This conversation has been lightly edited for length and clarity.

Capital B: What’s at stake in the fight over Florida’s map?

Angie Nixon: Ron DeSantis and Donald Trump, they’re basically working together. Trump, in collusion with Republican governors, is basically trying to rig an election. They’re making it so that politicians can choose their voters and not allow the voters to choose whom they want to represent them. And it’s especially happening in the South.

There’s just an all-out attack on our rights, and they’re weaponizing agencies to go after people who challenge them. That’s exactly what happened with me. I’m facing two years in jail because DeSantis didn’t want to meet with me. They arrested me for trespassing because they wanted me to move away from his office, even though I’m allowed to have access to the Capitol building 24/7. That’s just so disrespectful. It’s a sign that our democracy is falling apart.

Florida state Rep. Angie Nixon speaks. ( Crystal Vander Weit / TCPalm / USA Today Network via Reuters file ).

Do you expect these political and legal pressures to shift as we near the midterm elections?

I believe that we’re going to see more attacks and more officials weaponize their authority by targeting people who are speaking out against them. We can see that ICE [U.S. Immigration and Customs Enforcement] is picking up, especially in our state, because TPS [Temporary Protected Status] has ended for Haitians. It just gives them the opportunity to overpolice and racially profile and evoke fear.

When you live in a state of constant fear and frustration, you just give up hope. You feel as though there’s no one who’s doing their job as an elected official, and so you don’t even worry about voting sometimes. And that’s what I’m hearing. People have given up hope because not enough elected officials are fighting and getting loud and pushing back.

But that’s why I keep doing what I’m doing, because I want people to know that I see them. I feel their pain, and I’m fighting for them, even though I’m facing two years, even though I faced a formal reprimand, even though I had the FDLE [Florida Department of Law Enforcement] go to someone’s house because they called for my assassination.

In spite of all of this happening, Floridians deserve someone who’s going to look out for them. The government should be looking out for the people, not working against them.

What goes through your mind when you see government officials use their authority like this?

Honestly, I get sad. And sometimes it feels lonely, because I feel like not a lot of people are doing what I’m doing.

But then I just think about my grandmother and how she had to hide inside of a building in Jacksonville when race riots started happening. I just think about U.S. Rep. Yvonne Hinson, who marched alongside her parents when they were trying to get their right to vote and she was being spat on and fire hoses were being turned on her.

I just think about how I’m not facing that, and so who am I not to speak up after these elders — many of whom are still alive — went through all of that. Like, how dare I not speak out and push back. And so I have to keep going.

What do you want to see from other elected officials?

I want us to get louder.

I was at a workshop on democracy last year, and a historian talked about how the only way that countries have been able to save democracy was by getting loud and mobilizing and taking it to the streets. And the opposition party was out there helping to get loud and take it to the streets.

We have to lean on what our elders did — how they organized and got loud. I’m ready for our people to not just survive, but thrive and flourish.

What else is important for people to understand about your case?

I’m facing two years for a peaceful protest, and yet we have a president who was found guilty of 34 felony counts. It’s just — this is crazy. It’s absolutely crazy. It’s not right. The system is rigged. But I’m ready to unrig it.

Our Solar System Is Terminally Unstable and Will Be Completely Destroyed, Study Finds

403 Media
www.404media.co
2026-10-03 07:00:34
The estimated lifespan of the outer solar system has been downgraded from 100 billion years to just a few billion years, according to a study that probed “terminal instability” during the Sun’s death....
Original Article

Welcome back to the Abstract! These are the studies this week that searched for the ur-animals, wandered the poles, rained on Mars, and destroyed the solar system.

First, scientists present new evidence that the first animals appeared more than 800 million years ago, a truly ancient origin that suggests our metazoan ancestors survived a period known as Snowball Earth, when our planet is thought to have been nearly completely frozen. Then: Earth’s poles won’t sit still, the sepulchral stuff of life, and a dramatically shortened lifespan for our solar neighborhood.

As always, for more of my work, check out my book First Contact: The Story of Our Obsession with Aliens , or subscribe to my personal newsletter The BeX Files .

I trace my ancestry to Snowball Earth

Durbin, Orin Lole et al. “Re-evaluating molecular clock maximum age calibrations revives pre-Ediacaran divergence estimates for animals.” Science Advances.

When did the first animals emerge on Earth? It’s a question that has provoked centuries of scholarly debate, and also inspired some wonderful answers in myth and legend (I’m partial to the Bible’s tidy solution: day six checklist).

The earliest animals that are clearly preserved in the fossil record lived during the Ediacara period some 574 million years ago, but the origins of our diverse metazoan family are likely much older. Now, scientists suggest that animals may have appeared on Earth a whopping 800 million years ago, during the ancient Tonian era, according to a new estimate of the dawn of animals.

“The timing of the origin of animals on Earth has puzzled scientists for centuries,” said researchers led by Orin Lole Durbin who was at the University of Oxford during this research and is now at Virginia Tech. “Unambiguous animal body fossils within the Ediacara Macrobiota provide a hard minimum calibration of 574 [million years] for crown Metazoa. Animals must have evolved before this date. Determining a maximum calibration for animals (a date at which they had not yet evolved) is a more complex task.”

Estimated timeline for the origin of animals based on molecular clock analyses. The grey bar on the right of the image is the origin of animals based on an Ediacaran upper calibration. The grey bar to the left is the origin based on older deposits in the Tonian period, as suggested by the new study. Image: Orin Lole Durbin.

The team turned to molecular clock analysis, a technique that uses the rate of genetic mutations in lineages as a rough way to reconstruct evolutionary timelines. The researchers focused on two fossil deposits—China’s Weng'an Biota and Mongolia’s Kheseen Biota, which date back 590 and 550 million years respectively—to extend the possible timeline of animals back more than 200 million years earlier.

If animals really did appear in the Tonian era, it means that our earliest metazoan ancestors survived episodes of so-called “Snowball Earth,” when our planet was almost fully covered in ice. These early critters were likely simple marine sponges and comb jellies that lived for hundreds of millions of years until their descendents exploded into the kaleidoscopic variety of animals that persist to this day.

“The possibility of a pre-Ediacaran origin and diversification of animals means we cannot discount hypotheses that link the timing of animal diversification to Cryogenian Snowball Earth glaciations,” according to the study. “The causes and trajectory of animal evolution will remain uncertain until paleontological and geochemical data provide sufficient confidence in maximum calibrations to ensure a reliable timescale.”

In short: Respect your sponge elders.

In other news…

A pole-ing error

Domeier, Mathew et al. “Quadrupolar sea level fluctuations reveal episodes of rapid polar wander.” Science.

Earth is a roiling mess of shifting continents and squishy innards, a situation that constantly throws its spin axis off balance and prompts the geographic poles to drift. This phenomenon, known as true polar wander (TPW), is distinct from Earth’s wandering magnetic poles, which are shaped by interior core-mantle processes.

Right now, TPW is clocked at a slow rate of 10 centimeters per year, but scientists have found evidence of “fast” episodes of TPW that pushed the poles off by thousands of miles over millions of years, which has knock-on effects for ocean and land distribution across the globe.

A team has now pinpointed several of these fast episodes over the past 320 million years, “confirming that protracted rapid TPW has occurred on Earth and that couplings among Earth’s rotational dynamics, mantle processes, and surface environments can be strongly episodic,” according to their study.

“These findings refute the view of TPW as negligible or persistently slow and highlight the need to consider TPW as an episodic control on sea level change and likely other global environmental and biological dynamics,” said researchers led by Mathew Domeier of the University of Oslo.

In addition to these long-term natural changes, recent human activity has slightly impacted TPW , especially the construction of dams and the glacial melt of anthropogenic climate change. So the next time you address a letter to Santa at the North Pole for your kid—or yourself, no judgement—make sure you have the most up-to-date coordinates.

It’s raining formaldehyde—formalallelujah!

Koyama, Shungo et al “Global Distribution of Atmospheric Formaldehyde Deposition Correlated with Water Vapor on a Warm Early Mars.” The Planetary Science Journal.

Formaldehyde, the toxic chemical, has a bit of a morbid reputation—exposure to it can be fatal and it’s a common ingredient used to embalm corpses. But, paradoxically, formaldehyde (H 2 CO) also helped pave the way for the emergence of life on Earth as a precursor of bioessential elements such as sugars, amino acids, and nucleobases.

If life ever flourished on Mars, formaldehyde would likely also have been part of the story. To map out the possible distribution of the chemical on ancient Mars, scientists ran models of its atmosphere between 3.6 and 3.8 billion years ago, when the red planet was warmer and wetter. The results showed that formaldehyde was strongly associated with water vapor, suggesting that the chemical rained down from the ancient skies into Martian waterways.

“Significant atmospheric deposition of H 2 CO is observed over water bodies, including the northern ocean,” said researchers led by Shungo Koyama of Tohoku University. “We suggest that basins adjacent to persistently humid regions, which likely experienced H 2 CO accumulation and subsequent organic synthesis, could be potential landing sites for future missions to investigate ancient chemical evolution and whether it led to the origin of life.”

Hopefully, future missions to Mars will leave any signs of life with no place left to formalde-hide.

Update the cosmic actuary tables

Batygin, Konstantin et al. “Terminal Instability of the Solar System Triggered by Stochastic Solar Mass Loss.” The Astrophysical Journal Letters.

I hate to be the bearer of bad news, but the solar system has been diagnosed with terminal instability and has been given a mere six billion years to live.

That’s the upshot of a new study that modelled the solar system’s future as the Sun becomes a red giant star and, ultimately, collapses into a stellar husk called a white dwarf.

Though the dying Sun will consume Mercury, Venus, and perhaps even Earth, scientists have previously assumed that the giant outer planets might be relatively unscathed by its death—perhaps surviving for up to 100 billion years after the main-sequence lights go out. But updated simulations suggest that these distant objects may be thrown into fatal chaos as early as the red giant phase, and that they are unlikely to survive more than a billion years after the Sun’s transition to a white dwarf.

The future of Earth is bright, literally. Image: Celestia

“[Isaac] Newton, contemplating the drifting orbits of Jupiter and Saturn, suspected that the planetary order was mortal, and three centuries of celestial mechanics…have progressively deferred that verdict, most recently to timescales far beyond the age of the Universe,” said researchers led by Konstantin Batygin of the California Institute of Technology. “Our results return the solar system’s dissolution to astrophysically familiar territory, and relocate its cause: not the slow seep of chaos, nor the chance encounter with a passing star, but the Sun itself, which in dying does not merely enlarge the planetary system it built—it shakes it, and more often than not, spills it.”

“Newton’s envisioned instability is real after all,” the team. “He was mistaken only about the perpetrator.”

What a great story to kick off this spooky season! Planetary order is mortal, but cosmic horror is eternal.

Thanks for reading! See you next week.

Show HN: Germany's new sovereign AI model Kolibri

Hacker News
tej.as
2026-10-03 06:43:51
Comments...
Original Article

Kolibri is an open-weight large language model (LLM) from Aleph Alpha for German and English: a mixture of experts with 78 billion parameters that only uses about 3.5 billion of them for each token it reads or writes. It came out on 3 October 2026 under the Apache 2.0 license , the weights are on Hugging Face , and it was trained from scratch on infrastructure in Germany and Finland. (Kolibri is German for hummingbird, which is cute for a model whose whole trick is being light.)

I live in Germany, and at SmashingConf New York in 2024 I told the room what I’d heard in the US when I said where I’m based: “you regulate, you don’t innovate.” It hurt to hear, and what I wished for on that stage was the middle, “the right balance between innovation and regulation around data privacy, data stewardship, environmental constraints and energy requirements.” Kolibri is a pretty direct answer to that: a German team built it with the EU AI Act in mind “from the ground up”, and in Aleph Alpha’s own evaluation it scores above every compared model of its size in both languages.

Huge congrats to everyone at Aleph Alpha who built it, my good friend Michael Hofmann among them!

This post is about how Kolibri works, where it’s strong, where it isn’t, how to run it, and when it’s the right pick. Everything here comes from Aleph Alpha’s 189 page technical report , the model card and their launch post , plus one experiment I ran on its tokenizer.

What is Kolibri?

Kolibri 1
Parameters 78.1 billion in total, 3.46 billion per token (4.4%)
Languages German and English
Context 262,144 tokens natively, tested up to 1,048,576
License Apache 2.0 for the weights and configuration files (Aleph Alpha keeps the rights to its training code and methods)
Memory about 78 GB of weights in 8-bit floating point (FP8)
Reasoning 4 levels: none, low, medium and high
Tool calling Yes
Knowledge cutoff 18 June 2026
Training about 24 trillion tokens, more than a fifth of them German, on 768 NVIDIA B200 graphics processing units (GPUs)

Aleph Alpha calls Kolibri sovereign, and in their launch post that means 2 things. The first is how it was built: “teams built the model in Germany, trained it on infrastructure in Germany and Finland, under European and German law, with no foreign control.” The second is what customers get: “full freedom of deployment and intellectual-property safety, so compliance comes as an inherited property.” In plain words, a ministry or a car supplier can run it on its own servers, with its data never leaving the building, and nobody can change or switch off the model under them. Aleph Alpha has also signed the European Union ’s General-Purpose AI (GPAI) Code of Practice .

Sovereign doesn’t mean that nothing from outside Europe went in, and the model card says so itself: English web text was rephrased with Google’s Gemma 4 , German with Mistral-NeMo , and Qwen3-32B labeled data for the quality filters. They then filtered the training data for the political bias such models can have, which they’ve measured in Chinese open models themselves.

How Kolibri works

Kolibri is 6 ideas stacked on top of each other, and each one is there to make German cheaper, longer or more honest.

1. 384 specialists, and each token sees 6

In a normal (dense) model, every token goes through every parameter. In a mixture of experts (MoE) , each layer has a crowd of small sub-networks called experts and a router that picks a few of them for each token. Kolibri has 50 layers, each with 384 experts plus 1 shared expert that every token goes through, and its router sends each token to 6 of the 384. That’s how 78.1 billion parameters turn into 3.46 billion of actual work per token.

In 2024 I gave a talk called Why Small Language Models are the future , and I argued for “smaller language models with fewer parameters and fewer places things can go wrong that require lesser compute.” My analogy was a doctor who has read every medical book in the world against a specialist in hematology: go to the first one with a blood condition and “they may not get it right cuz they know too much.” A mixture of experts puts a hospital full of specialists inside one model, and the router is the receptionist who sends each token to the right 6.

The analogy breaks in 2 places though. The experts aren’t neat topics like “German law”: when researchers look inside MoE models they mostly find experts for patterns of tokens, like punctuation or proper nouns, not subjects a person would pick. And the hospital has to keep all 384 specialists on staff even if you only see 6, so Kolibri computes like a 3.5 billion parameter model but needs the memory of a 78 billion parameter one. The model card says it plainly: “the full model must be held in memory even though only part of it is active at any time.”

2. A tokenizer that reads long German words

A model doesn’t read letters or words, it reads tokens: chunks of text from a fixed vocabulary, picked when the tokenizer is trained. German glues words together into long compound words, and a tokenizer that learned mostly from English chops them into pieces. Here’s the German name of the Federal Constitutional Court , split by the tokenizer GPT-4o and GPT-5 use ( o200k_base , through OpenAI’s tiktoken ), and by Kolibri’s:

o200k_base (GPT-5):  Bund | es | ver | fass | ungs | gericht     6 tokens
Kolibri:             Bundes | verfassungsgericht                 2 tokens

Kolibri’s tokenizer has 128,000 tokens, trained with a new algorithm Aleph Alpha calls UniBPE: it keeps the bottom-up merging of byte-pair encoding (BPE) and picks each merge with a different scoring rule (the Unigram objective), which respects how German builds words. The report says it needs 11.2% fewer tokens for German text than GPT-5’s tokenizer, the best of the 9 others they measured.

I wanted to see that for myself, so I ran 6 tokenizers over all of the Basic Law for the Federal Republic of Germany , the German constitution (185 KB of very German legal text), and over its official English translation :

Tokenizer German tokens More than Kolibri English tokens More than Kolibri
Kolibri 1 35,190 39,875
o200k_base (GPT-4o, GPT-5, GPT-OSS ) 41,482 17.9% 39,737 -0.3%
Qwen3.5 35B-A3B 42,907 21.9% 41,650 4.5%
Mistral Small 4 43,478 23.6% 41,301 3.6%
Gemma 4 43,850 24.6% 41,564 4.2%

On legal German, Kolibri needed 15% fewer tokens than GPT-5’s tokenizer, even more than Aleph Alpha’s own 11.2%, and in English it tied with it. Wild. Fewer tokens means fewer steps to read or write the same German text, and more German fits in the same context window. (I counted each one with its own tokenizer.json through Hugging Face’s tokenizers library, except o200k_base , which I counted with tiktoken.)

3. Most layers only look nearby

40 of Kolibri’s 50 layers use sliding-window attention: each token only looks at the 512 tokens before it. Every 5th layer looks at everything before it. It’s like reading a long contract while mostly paying attention to the sentence you’re on, and every few pages stopping to think about all of it, and it’s what keeps a 1 million token context affordable.

There’s a clever detail in there too. Only the sliding-window layers know where a token sits (through rotary position embeddings ), and the full-attention layers don’t, so the context stretches past the 262,144 tokens it was trained on without any extra position tricks. Aleph Alpha validated it up to 1,048,576.

In my talk Unlocking Value with AI Today I called finite context one of “the big three” problems of generative AI, next to hallucination and the knowledge cutoff, and Kolibri goes after all 3. At 1 million tokens, on the RULER long-context test, Kolibri’s base model scores 63.2, against 57.5 for Qwen3.5 35B-A3B’s base model.

4. It thinks in German

Reasoning models think before they answer, and even on German prompts they mostly think in English. Aleph Alpha posted about this on 24 September and wrote it up as Through the Valley of Tears : they generated about 800,000 German reasoning examples, and found that a little German reasoning data is worse than none. Their model’s German math score dropped from 70.2 to 48.3, because its German thoughts kept going around in circles and never finished, and it only climbed back (to 67.3) with a lot more German data.

Kolibri got the lot more. It reasons in German on German prompts, and its German math scores are the best of the models with about 3 billion active parameters: 87.5 on the American Invitational Mathematics Examination (AIME) 2025 in German, against 84.4 for the next best, NVIDIA’s Nemotron 3 Nano .

5. It’s trained to say “I don’t know”

Hallucination was number 1 of my big three, and the fix in that talk was retrieval augmented generation (RAG): you look up good, authoritative information and put it in the prompt. RAG only works if the model admits when the documents don’t have the answer, though, and that’s what Aleph Alpha trained for, with their own method, the Merlin-Arthur protocol .

It works like a game with 3 players. Arthur is the model, and he gets a question with parts of the supporting document hidden. Merlin hides parts so that the correct answer gets easier to find, and Arthur is trained to answer those. Morgana hides the evidence the answer depends on, and Arthur is trained to say he doesn’t know. Arthur never knows which of the 2 he’s facing, so the only way to win is to actually check whether the evidence in front of him supports an answer.

It shows. On Artificial Analysis’s Omniscience test , when Kolibri didn’t know an answer, it said so (or gave a partial answer) 44% of the time instead of making one up. Qwen3.5 35B-A3B did that 11.1% of the time and GPT-OSS 120B 23.7%. Of all the mixture-of-experts models Aleph Alpha compared, only Qwen3.6 35B-A3B did better, at 56.7%.

6. You pick how hard it thinks

Every request can set reasoning_effort to none , low , medium or high , so a quick lookup answers right away and a hard question gets a long think, from the same model on the same server.

What Kolibri is good at

These are the rows where Kolibri leads the open models of its size in Aleph Alpha’s evaluation, which runs every model through the same setup with the sampling settings its makers recommend:

What Kolibri Closest model with about 3B active parameters
Overall score, English 75.5 74.7 (Qwen3.5 35B-A3B)
Overall score, German 70.8 69.8 (Qwen3.5 35B-A3B)
AIME 2025, English 96.9 89.6 (Nemotron 3 Nano)
AIME 2025, German 87.5 84.4 (Nemotron 3 Nano)
Questions over unseen company documents, English 89.7 87.0 (Qwen3.5 35B-A3B)
Long context at 1 million tokens (base model) 63.2 58.5 (Nemotron 3 Nano)

The math is the standout: on AIME 2025 and 2026 in English it beats every MoE model in the comparison, including the ones with 12 billion active parameters, and only the dense Qwen3.8 27B scores higher. The company-documents row is 5 tests Aleph Alpha built from customer-like work in semiconductors, the German public sector, aerospace, an automotive supplier and industrial drives, run over documents and questions the model never saw in training.

What Kolibri is bad at

Aleph Alpha publishes its weakest rows right next to its best ones in the model card, so here they are:

  • It knows less from memory. It’s last of the 12 models on a closed-book test (the Retrieval-Augmented Generation Benchmark ’s questions with no documents, 51.0), and it answers only 14.8% of the Omniscience questions correctly, against 22.2% for Qwen3.5 35B-A3B. It knows when it doesn’t know, which is great, but it also knows less: fine with documents in the prompt, bad as a trivia oracle.
  • Multi-turn tool calling is weaker. On the Berkeley Function Calling Leaderboard ’s multi-turn tests it scores 39.8, against 58.2 for GLM-4.7 Flash and 54.0 for Qwen3.5 35B-A3B.
  • It’s not the best coding agent. It scores 27.7 on Terminal-Bench 2.1 against 39.7 for Qwen3.5 35B-A3B, and 66.4 on SWE-bench Verified against 73.8 for Qwen3.6 35B-A3B.
  • Long context dips in the middle. At 128,000 tokens on RULER, Qwen3.5’s base model scores 89.9 to Kolibri’s 67.9. Kolibri only pulls ahead at the very long end.
  • The memory. 78 GB of weights means a data-center GPU or two, never a laptop, whatever the 3.5 billion suggests.
  • The plumbing is new. It needs Aleph Alpha’s vLLM plugin , which supports one vLLM version at a time (0.29 today), and on launch day no hosted provider serves it yet.
  • 2 languages only. That’s on purpose: the model card calls it “a deliberate choice of depth over breadth.”
  • A bigger dense model beats it. Qwen3.8 27B scores 80.2 in English and 79.9 in German, but it uses nearly 8 times as many parameters for every token, and that’s the trade Kolibri is built around.

How to run Kolibri

You need about 78 GB of GPU memory: 2 NVIDIA A100s or H100s with 80 GB each at the least, or a single H200 , B200 or B300. I haven’t run the model itself yet, since it needs a data-center GPU and nobody hosts it so far, so this is straight from the model card. First the plugin, which installs the vLLM version it supports:

pip install 'aleph-alpha-inference>=1'

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

That gives you an OpenAI-compatible server, so any OpenAI client talks to it, and the reasoning effort goes through the chat template:

# from the Kolibri model card (trimmed)
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
    model="Aleph-Alpha/Kolibri-1",
    messages=[{"role": "user", "content": "Erkläre kurz, was ein Mixture-of-Experts-Modell ist."}],
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high", "enable_thinking": True}},
)
print(response.choices[0].message.content)

The model card recommends temperature=1.0 , top_p=0.97 and top_k=128 , and contexts of at most 262,144 tokens for anything latency-sensitive. For the full million, serve it with 2 more flags:

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --max-model-len 1048576 \
  --hf-overrides '{"max_position_embeddings": 1048576}'

When to use Kolibri

Kolibri is the pick when German text and your own hardware both matter: a public authority, a bank, a manufacturer or an aerospace supplier that has to keep its documents in house, wants answers in German that reason in German, and would rather hear “I don’t know” than a confident wrong answer. RAG over long German documents (laws, contracts, manuals) plays to every strength above: the tokenizer, the 1 million token context and the abstention.

That’s the kind of project I’m working on right now, with Prof. Dr. Heinrich Audebert , who heads neurology at Campus Benjamin Franklin , one of the Charité ’s hospitals in Berlin. Today, a patient with a neurological complaint goes to their general practitioner (GP), and the GP has to see them, which is very demanding for a busy practice. In what we’re building, the patient sits down at a computer in the GP’s practice, an AI avatar takes them through a battery of tests, and it grades how urgent their symptoms are: a referral to a specialist right away, or they can wait a bit. The conversations are in German, they’re about people’s health, and the model has to run where we control it, so we’re thinking of using Kolibri for it.

It’s the wrong pick for a coding agent, where Qwen3.6 35B-A3B leads, for questions the model has to answer from memory, for any language besides German and English, and for anyone who can’t spare 78 GB of GPU memory. For that last case, the small specialized model I argued for in 2024 is still the answer: for my podcast search I fine-tuned Mistral 7B on my Apple silicon laptop instead of paying for GPT-4o.

Which is better for your own documents, a hospital of specialists like Kolibri or one small model trained for a single job, only shows when you try both on your data. Swapping the model under an agent without breaking it is what my workshop on reliable AI agents covers: guardrails and retries that work whichever model sits underneath.

An AI agent emailed researchers for help. It told us why

Hacker News
www.science.org
2026-10-03 06:07:08
Comments...

GitHub's new dashboard experience now the default

Hacker News
github.blog
2026-10-03 05:59:01
Comments...
Original Article

The new dashboard experience, previously available as a feature preview, is now the default view for everyone.

The redesigned dashboard helps you focus on the work that matters most and makes it easier to find and act on what you need next. You can:

  • View your active agent sessions, issues, and pull requests in one place. Use filters to choose what appears in each section, with up to 12 items in each list.
  • Catch up on updates in the separate Feed tab, keeping your feed distinct from your productivity-focused dashboard.
  • Start work directly from the dashboard, such as assigning an issue to Copilot coding agent or opening a pull request in Copilot Chat.

If you’d rather keep the previous experience, you can switch back at any time by clicking the dropdown next to “Preview.”

We’d love to hear your thoughts on the new dashboard experience. Drop a comment with any questions or feedback in the Community discussion .

Rust for CPython (Python Language Summit 2026)

Lobsters
blog.python.org
2026-10-03 05:40:38
Comments...
Original Article

“No one said ‘don’t do this’ last year”. After testing the waters at PyCon US 2025 , David Hewitt returned to the Python Language Summit asking what Python core developers want from Rust, along with proposed timelines, phases, and success criteria for how the Rust for CPython project might proceed and become a permanent fixture within the CPython project.

David Hewitt at the lectern, in front of his “Rust for CPython” slide

Photo by Hugo van Kemenade ( CC BY-NC-SA 4.0 )

David is acting as an “ambassador” for the Rust for CPython project team, which is currently led by core developers Kirill Podoprigora and Emma Smith as authors of the Rust for CPython PEP draft . Emma also spoke at PyCon US 2026 about the Rust for CPython project. The team itself is around 60 developers in a Discord channel, among them a “few [Python] core developers” and a “delegation from the Rust project”. The team has experience with previous projects integrating Rust into existing codebases, such as Android and the Linux kernel, and is “excited by the work and keen to support [the project] if we proceed”.

Why Rust?

Bar chart of type-crash issues opened per year, rising from 82 in 2021 to 222 in 2025, with 2026 projected at around 353.

Showing how adopting Rust may specifically help CPython, David noted how the number of issues labeled with “ type-crash ” has been steadily rising over time. “We’ve been making some big technical bets”, he said, referencing the new parser , the JIT , and free-threading . These large, complex features may be one of the reasons more crash reports are being opened on GitHub, and Rust could be a potential solution here. David explained that Jeff Vander Stoep described Rust in Android as “move fast and fix things”, and that “fewer revisions for patches of the same size” was the experience Android has had since adopting Rust.

Rust and Python are already working together in Python’s ecosystem of packages thanks to PyO3 and Maturin . David also emphasized that many technology companies were “choosing Rust as a bet” instead of only as the new shiny tool.

But adopting Rust into CPython would not all be smooth sailing; there were still concerns that David and the team are aware of.

One of the biggest concerns from a year ago was Rust’s lack of platform support compared to CPython, which at the time of writing officially supports 20 different architectures and platforms at either Tier 1, 2, or 3. David shared that Rust’s support of different platforms has “widened since last year” and was becoming “less and less of a concern”. The proposal would be to ask Python distributors to attempt using the optional Rust support in Python 3.16 (October 2027) and report platform-specific issues upstream in time to be resolved around the Python 3.17 timeline (October 2028).

David acknowledged that adopting Rust into a project is a social challenge as much as a technical challenge. Rust knowledge is “not universal amongst core developers”, which would be partially mitigated by “targeting small portions” of CPython and an incremental approach, as “1 million lines of C code can’t be ported all at once”.

David was clear that using Rust in itself does not necessarily mean that ported code would be free of bugs or security issues. Although Rust does mitigate classes of issues that commonly create bugs, the code can still have correctness issues. The team proposes mitigating this by adopting property-based testing, fuzzing, and using “prudent engineering practices” during the porting process.

The proposed first Rust module: zlib

Below are the proposed timelines for the Rust for CPython project making a “significant improvement” to CPython, with a PEP defining the success criteria for Rust expected in late 2026. Under this timeline, the first Rust code to ship in Python would be in Python 3.16, where it would be completely optional, with the existing C code kept as a fallback. The earliest that Rust would become required to build CPython is Python 3.18 in 2029, at least three years away.

'Rust for CPython' proposed timeline

The timeline includes a build system and CI, Rust API proof-of-concept happening in the Summer 2026, a PEP defining success criteria in late 2026, an optional Rust backend for the zlib module and private Rust API in Python 3.16 (October 2027), resolving platform issues and Rust in more places (io, json, xml, memoryview, parser) for Python 3.17 (October 2028). Finally, in some distant Python version (October 2029+) the Rust build would be made required and a public Rust API would be published.

The Rust for CPython team has selected the zlib module as the first module to be given an optional Rust implementation because they “wanted to achieve a significant improvement” with a “small scope”. The proposed Rust implementation will use zlib-rs , which is “heavily tested and used by the Firefox, uv, and Cargo ” projects and “faster than zlib and zlib-ng on many platforms”. The module was also selected because it’d require using an external Cargo package, meaning this aspect of the build process would need to be designed and exercised.

This small change would have an impact: the zlib compression algorithm is “widely used by Python packaging”, meaning that (almost) “every pip install in Python 3.16 will be sped up” if the proposal is accepted.

Rust API Sketch

David provided an example of some Rust code calling a hypothetical Rust API for Python. The API would use a Rust attribute ( #[pyfunction] ) and be similar to Argument Clinic , a development tool for automatically generating blocks of code for handling C function arguments from Pythonic function syntax. Rust functions would always be passed the thread and interpreter state ( Python<'_> ), use smart pointers around objects ( Py<...> ), and lean into Rust error handling with the Result enum, returning either a result or an error that was raised.

#[pyfunction(signature = (
    data,
    /,
    wbits=MAX_WBITS,
    bufsize=DEF_BUF_SIZE,
))]
fn decompress(
    py: Python<'_>,
    data: Py<PyObject>,
    wbits: c_int,
    bufsize: isize
) -> PyResult<Py<PyBytes>, PyErrRaised> {

    // buf will be cleaned on scope exit
    let buf = PyObject::get_buffer(py, &data)?;
    let decoded = /* ... */;

    Ok(PyBytes::new(py, &decoded))
}

The example above shows a buffer being allocated and automatically cleaned up on scope exit, rather than being cleaned up manually as would be required when writing the same function in C.

Proposed success criteria

David moved on to success criteria: what would the Rust for CPython project need to show to move into new phases of the roadmap and eventually become a required part of building CPython? “For previous big changes like the JIT and free-threading, we’ve explicitly defined success criteria that would need to be met in order for the added complexity to be accepted”, David explained, “we expect we’d need to do the same here. What should those criteria be?”

David had these suggestions for core developers:

  • Critical: A majority of active core developers are open to using the Rust API to implement functionality.
  • No meaningful slowdown to CPython performance benchmarks.
  • All tiered platforms must be supported by Rust. The experience of distributors building CPython should generally indicate that adding Rust support is manageable.

David noted that the first point reads “open”, not “familiar”, and shared a plan to survey Python core developers about how they’ve used Rust while contributing to CPython when deciding whether to move Rust out of experimental stages.

Discussion

On the topic of designing the new Rust API so that it’s “familiar” to users of the C API, Thomas Wouters advised against “making compromises for the dinosaurs”, including himself in the subset, instead asking whether the Rust API should be designed from first principles. David responded that there are “places to lean into Rust”, such as dropping resources on scope exit, but there are also idiomatic Rust designs which “won’t be the best fit”. David noted that it would be reasonable for core developers to be looking at both the C and Rust API at the same time while working. “We should be mindful of our audience, which is also ourselves”.

Larry Hastings asked why the Rust for CPython project wasn’t a “rewrite”, suggesting the team “display your success as a fork”. David acknowledged that “ RustPython already exists” and that the Rust for CPython team had already spoken with the contributors of the project. “RustPython isn’t as performant as CPython, but could be used to inform what APIs we design”.

Larry also shared that he “wasn’t super excited to learn Rust to work on CPython”. David assured him that there are many areas of CPython that would not be considered for writing in Rust: “CPython should not be written in Rust for the sake of Rust”. “CPython will be a dual-language project for a meaningful amount of time”. However, David cautioned that “it would be disingenuous to say that Rust would be optional forever”, as one of the aforementioned roadmap items for the Rust for CPython project is to become a required part of the build process and provide a public Rust API.

Pablo Galindo Salgado speaking into a microphone, seated in a row of attendees

Photo by Hugo van Kemenade ( CC BY-NC-SA 4.0 )

Pablo Galindo Salgado was more concerned about the future, which was “reaching for our dependencies from Cargo”, noting that this would be a “huge problem” and a potential “showstopper” for the project. “We vendor our dependencies, and we have a very selective set”, he said, noting that each time a vulnerability is published for one of those projects, the release managers need to make new releases, which can be “tiresome”. “Right now we’re only focusing on the APIs and the basics”, he added, highlighting that the challenge of taking on many Rust dependencies hasn’t been addressed yet.

David answered that the team should “select as few external dependencies as possible”; “zlib-rs is only one dependency”, and the Rust sources would be vendored so that “building CPython would not require Cargo”. David added that “Cargo has a relatively clean system for vendoring dependencies” and that the “Rust for CPython proof-of-concept uses this system”. The vendored sources “wouldn’t live in the CPython tree”. Łukasz Langa agreed that “it’s better for dependencies to live separately”, referencing the cpython-source-deps repository , with Thomas reminding everyone that this would be a new usage of that repository; today it is only used for binary installers of CPython.

Kolibri Has Landed: A Sovereign Open-Weight Model

Hacker News
aleph-alpha.com
2026-10-03 05:36:04
Comments...
Original Article

Green gradient with the white Kolibri hummingbird logo and wordmark in the centre, framed by thin stepped outlines on the left and right

Research

Aleph Alpha

On the Day of German Reunification, we are releasing our new model: Kolibri.

Kolibri is an English-German Mixture-of-Experts Transformer with 78B total parameters, 3B active. It supports context lengths of up to 1M tokens. The model can be downloaded with the full weights on Hugging Face and used under the Apache 2.0 license terms.

Kolibri is the result of continuous iteration of our model training effort. We first built a model training pipeline and validated it by building Kolibri Origin, a 30B total, 3B active model with a much shorter 65k token context window. Kolibri ran through the same pipeline: from data ingestion and curation, through ablations, pre-training, and post-training, to the final evals. It enabled running hundreds of ablation experiments and stable pre-training that ran without a person having to step in when hardware failed or a data connection dropped. We continuously monitored training metrics and standardized monitoring for custom benchmarks. The time we put into building and iterating on this pipeline was a valuable investment. We see it in how much better Kolibri is than Kolibri Origin, and in how little time separates their releases.

Kolibri is a specialized language model built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace. We specialized Kolibri for German, reasoning, math, agentic behavior, and further capabilities our customers need in production. The aim of this specialization was to optimize performance in our customers' specific use cases. Through specialization, customers achieve contextualized performance in their AI operations and they can monitor its economic impact, so that ROI stays measurable and grows over time.

Specialization alone is not enough. Sovereignty is just as important. Sovereignty, for us, combines two dimensions: how we built the model, and how it transfers to our customers. We offer full supply-chain integrity and account for every decision, from data ingestion, through pre- and post-training, to the final evaluations. We provide transparency. Customers have full freedom of deployment and intellectual-property safety, so compliance comes as an inherited property of the model.

Read our tech report for full details.

What Kolibri Delivers

We optimized Kolibri for performance across a wide range of sectors considering their particular domain-specific language, regulatory, and procedural realities. Its small and efficient size provides our customers with flexibility to run it efficiently on-premise, without sending internal data to third-party inference services. The spotlight in this section introduces the model's capabilities, before we describe them in section How we built Kolibri at high velocity .

Foundational capabilities for enterprise and government

With Kolibri we optimize the trade-off between model capability and deployment costs, using 3B active parameters out of 78B total. Kolibri sits on the Pareto frontier for quality versus serving cost, for both English and German. The Pareto frontier is a concept from economics, marking the best achievable combinations of two objectives, where improving on one means giving up some of the other. None of the compared models delivers more quality at the same serving cost, or the same quality at lower cost.

Average benchmark score [%]

  • Kolibri
  • Kolibri Origin
  • Other post-trained models
  • Pareto frontier
Performance vs. throughput for post-trained models in English (left) and German (right). Metrics show unweighted average benchmark scores against decoded text per second and GPU. Higher and further right is better.

Across math, coding, grounding, and long-context tasks, Kolibri matches models with up to four times its active parameter count, such as Nemotron 3 Super.

AIME 2025 Math AIME 2025 (DE) Math AIME 2026 Math AIME 2026 (DE) Math GPQA (diamond) Knowledge GPQA (diamond, DE) Knowledge AA-Omniscience Index Grounding / hallucinations public set; from −100 to 100 BrowseComp Agentic τ³-bench banking Agentic τ²-bench retail Agentic τ²-bench airline Agentic τ²-bench telecom Agentic BFCL v4 overall Agentic LiveCodeBench v6 Code HumanEval+ Code LongBench Pro Long context AA-LCR Long context

Show the numbers
Benchmark Kolibri Kolibri Origin Qwen3.6-35B-A3B Nemotron 3 Super 120B-A12B Mistral Small 4 119B-A6B
AIME 2025 96.9 81.9 84.6 91.7 79.8
AIME 2025 (DE) 87.5 73.5 82.9 85.6 72.3
AIME 2026 96.0 81.5 91.0 90.4 83.1
AIME 2026 (DE) 90.0 75.2 84.4 87.5 78.5
GPQA (diamond) 84.3 68.1 83.4 78.0 74.7
GPQA (diamond, DE) 81.3 58.5 80.6 76.6 72.9
AA-Omniscience Index -32.8 -64.0 -15.3 -36.5 -24.0
BrowseComp 29.4 4.4 26.9 29.1 –
τ³-bench banking 38.1 5.7 10.6 15.5 5.7
τ²-bench retail 69.9 58.5 71.6 67.5 62.9
τ²-bench airline 76.7 58.7 70.7 72.7 40.0
τ²-bench telecom 94.7 67.5 99.1 68.1 41.5
BFCL v4 overall 61.4 36.4 67.2 61.0 58.0
LiveCodeBench v6 85.9 59.2 82.5 82.0 71.2
HumanEval+ 92.7 76.8 92.8 94.7 92.8
LongBench Pro 64.5 – 70.8 62.9 56.4
AA-LCR 68.3 – 69.7 67.0 52.3
Foundational capabilities of Kolibri. Kolibri is a balanced generalist model with competitive performance across math, code, long context, agentic capabilities, and knowledge. Benchmark scores are shown on a shared 0–100 scale, higher is better.

Contextualized performance for real-world applications

Public benchmarks fail to capture specialized sector needs, so we developed our own internal evaluation suites for the verticals that matter to our customers, such as the German public sector, aviation, manufacturing and the automotive industry. Each suite mirrors the skills, workflows, and edge cases required in these sectors, and paired synthetic training environments let us improve Kolibri against these evaluations without ever training on customer data. Read more in section Contextualized performance .

Score on internal customer-proxy benchmark

  • Automotive supplier 0.72 → 0.99
  • Semiconductors 0.35 → 0.80
  • German public sector 0.54 → 0.75
  • Industrial drive technology 0.31 → 0.60
  • Aerospace 0.14 → 0.59
  • one checkpoint, one eval
  • mean of that day
  • Kolibri Origin
  • Kolibri
Contextualized performance across successive post-training runs on internal customer-proxy benchmarks. Dots are single evaluations of training checkpoints, lines are the mean of each day's evaluations. Performance climbs across all five verticals, higher is better.

Answers that are grounded in your documents. We trained Kolibri with abstention data and with our Merlin-Arthur protocol. As a result, it is trained to say "I don't know" when the answer isn't in the context. Our customers value and request this feature, so we continuously track and validate abstention accuracy. We provide more details in the Grounding section.

Native German and English model. We developed a bilingual German/English tokenizer and focused on including organic German data throughout the training process of the model, so that 21.3% of the pre-training tokens are German. We used translation sparingly (6% overall), since translated text tends to carry the cultural fingerprint of its source language. The result is a model that is bilingual by design, not an English model that has read some German. Read more in sections German pre-training data and Specialized tokenizers .

Control and compliance by design: upholding our customers' sovereignty

We built Kolibri with the EU AI Act, the General-Purpose AI Code of Practice and the GDPR in mind from the ground up, with copyright law being a focus of our work to trustworthy technology.

We are transparent about our model weights and about curation of our training data , so that decisions behind development are visible. Through its reasoning traces, it becomes explainable how the model came to a particular answer. With our Merlin-Arthur protocol the model grounds trustworthiness : Kolibri refrains from answering when the context doesn't support an answer.

Our teams built the model in Germany, trained it on infrastructure in Germany and Finland, under European and German law, with no foreign control . We own the entire pipeline, from data curation, through pre- and post-training, to optimization in our Model Factory. This control of our end to end pipeline underpins the model's sovereignty. Control also passes on to our customers and Kolibri's small size gives them full freedom of deployment and supports controllable reasoning effort to trade cost and latency against answer quality.

How We Built Kolibri at High Velocity

Our Model Factory: two models, three months apart

Our Model Factory is our answer to iteration speed, minimizing the time it takes to go from knowing what a model gets wrong to training one that does better. We implemented the training pipeline as code, so that the learnings of our team landed in one versioned training recipe instead of scattered scripts and notes. Here is what that looked like between Kolibri Origin and Kolibri.

Work on the pipeline began in January. Five months and hundreds of ablation runs later, Kolibri Origin finished pre-training at target scale on 11 June. Kolibri finished on 11 September. In the three months between those two dates we went from 30B parameters to 78B, from a 65k-token context window to up to 1M, and from 7.5T training tokens to 20T. To get those 20T, the pipeline processed over 200T tokens of raw data, filtering, deduplicating, and curating it down to what we actually trained on. We changed the attention design, tripled the number of experts, increased sparsity, replaced the routing algorithm, improved our post-training data, more than doubled the number of environment tasks, and taught the model to reason at four different effort levels. Both models kept a similar number of parameters active per token, yet Kolibri trained faster per token than Kolibri Origin thanks to the work of our efficiency team.

Kolibri Origin Kolibri
Finished pre-training 11 June 2026 11 September 2026
Release no public release 3 October 2026
Reasoning mode Yes (one mode only) Yes (none, low, medium, high)
Total parameters 30.6B 78.1B
Active parameters / token 3.27B 3.46B
Pre-training tokens 7.51T 20T
Layers 50 (2 dense + 48 MoE, 1 shared expert) 50 (all MoE, 1 shared expert)
Pre-training context length 8,192 (8k) 16,384 (16k)
Longest trained length 65,536 (64k) 262,144 (256k)
Tokenizer vocabulary 96,000 128,000
Model dimension 2,048 2,560
Attention heads (query / KV) 32 / 4 48 / 4
Experts (total / active) 128 / 8 384 / 6
Expert hidden dim 768 512
Attention pattern full attention, all layers sliding window (512) + full attention every 5th layer
Knowledge cutoff EN: 1 Sept 2024, DE: 1 Aug 2025 EN/DE: 18 Jun 2026

What made this possible is our training pipeline, an effort to convert research-grade model development into a fully automated production-ready infrastructure to design and train large language models. Our pipeline codebase was shared by all . Every proposed code change triggers a small end-to-end model run; training, evaluation, to find out if something broke within minutes. Our runs are GitHub Actions workflows, and reproducing one means checking out a commit. We used the pipeline for the main training and the hundreds of ablations that ran to make our architectures and data-mix decisions.

Training checkpoints land roughly every hour and the pipeline automatically evaluates them on English and German knowledge, maths, code, instruction following, tool use, long context, safety, abstention to hallucination and grounding. We can watch capabilities appear every day rather than finding out how the model turned out at the end. The run itself held up better than we expected. Over 21 days of pre-training of Kolibri, we hit 38 unplanned interruptions, roughly one per 10,000 GPU-hours, caused by hardware faults or a connection timing out. Those were automatically handled by the pipeline without manual intervention. The cluster automatically restarted the training job(s) on a different set of nodes, and training picked up from a checkpoint at most 250 steps back.

Having access to the whole pipeline, with checkpoints available to everyone, means that any team can own a capability end to end rather than a stage of an assembly line. One team trained and evaluated the grounding capability of Kolibri as a single piece of work, seamlessly integrating into the overall model.

However, getting here was not a linear walk. We stumbled. We stopped the Kolibri Origin pre-training after a few trillion tokens and restarted it from scratch, as we identified a data-shuffling bug that escaped our tests. We ran ablations, took decisions, only to find a bug or an error in the configuration afterwards, forcing us to rerun some experiments. Alongside our own experience, we tracked state-of-the-art architectures, best practices, and the latest research advances in LLMs. Those accumulated learnings led to improvements in our processes and guardrails in our pipeline.

Looking at the delta between Kolibri Origin and Kolibri, the improvement is notable, but this is also the easy direction: more parameters, more data, a well-understood architecture family, and still at small scale. As we are contemplating scaling up, we are happy to face challenges on our foundation. But the most durable thing we built this year is not the pipeline, it is a team with the proven capability to build, post-train, and ship LLMs from raw data at high velocity.

Architecture and pre-training

Kolibri has 78B total parameters, which is 2.5 times more than Kolibri Origin, with ~3B active parameters. In our experiments, increasing the size of our model from 32B to 123B led to ever-improved performance. Yet the larger size came with larger training and serving costs. The latter drove the decision: 123B can handle only 3 long-context 256k-token user queries on two H100s, while 78B handles 18 concurrent requests and decodes 28% faster. We used 384 smaller experts rather than fewer wide ones, as they performed better in our tests. We applied the same efficiency-first reasoning to attention. Out of 50 total layers, only 10 process full context, while the remaining 40 use a tight 512-token focused window. This keeps decode computation and memory bounded in those layers regardless of context length. The efficiency gains benefit both serving the trained model and training during post-training RL, which relies on huge amounts of inference.

We trained Kolibri on 768 B200 GPUs in three stages: 20T tokens of pre-training at a 16k sequence length over 21 days, 3.44T tokens of mid-training at 64k, and 200B tokens of long-context adaptation at 256k. This is nearly 24T tokens in total, roughly three times what Kolibri Origin consumed. German accounts for more than a fifth of the pre-training mix, about 4.3T tokens, against roughly 62% English and 14% code. Compared to pre-training, mid-training data is a much more selective pool of curated datasets weighted toward reasoning, problem-solving, code and agentic data. For long-context, rather than train on long documents alone, which tends to erode the skills acquired earlier, we interleaved long documents with the high-quality mid-training data of the previous stage. For the long-context mix, we removed synthetic long documents to avoid artificially inflating benchmarks such as RULER.

We optimized with Muon, as we did for Kolibri Origin. We put special focus on training stability, resulting in a robust training run without any loss spikes for either model. With Kolibri, for routing between experts during training, we introduce exact quantile balancing. Quantile balancing was introduced in Kimi K3, where the global quantile is estimated from histograms because an exact computation was considered too expensive to communicate; we show that it can be computed exactly at fixed cost independent of batch size, and that the exactness improves both load balance and model quality.

Post-training at scale

Post-training happened in two stages; the first step is supervised fine-tuning to teach the model core reasoning ability and how to interact in a chat, followed by large-scale reinforcement learning to train the model to reason across a diverse suite of long-horizon tasks. We built a pipeline for both of these steps, which allowed us to exercise fine-grained control over the behavior of the model.

For SFT we generated a total of 174B tokens worth of synthetic data, which we filtered for quality and combined with filtered versions of permissively licensed open-source datasets to obtain a high-quality training mix of 268B tokens. For reinforcement learning we trained on a broad set of environments that contained more than 1.2 million curated tasks across diverse domains such as code & math reasoning, agentic tasks, instruction following, question answering, tool calling, and more. We optimized our in-house training codebase for high-performance and use asynchronous training, where we generate training data on our environments using the current model in parallel to training the model on already generated data.

Across both stages, we taught the model to reason at different effort levels (none, low, medium, high), which means that the user can exercise control over how much compute the model should invest to find a solution to the task at hand. This allows our customers to trade-off cost and inference speed against the quality of the final answer.

German pre-training data

From our experiments on small proxy models, we found that training with around 20% German data leads to optimal results, which meant we needed to find 4T German tokens to train our model at a 20T horizon. Open German datasets help, but are far from enough: after deduplication and filtering, we were left with 390B German tokens, well short of our goal. As we argued in Sauerkraut, Not Burgers , German capability has to come from high-quality German texts; relying heavily on machine-translations would lead to poor results due to subtle translation errors and a lack of authentic German cultural context. We closed the token gap in three different ways.

The first was to curate German from Common Crawl ourselves. We built a pipeline specialized for German data. German is not English, and German data cannot be filtered like English data: we had to retune filtering parameters for the German language. One typical filter in a language data pipeline is to remove documents with too many long words, but German administrative prose routinely exceeds the English bound on mean word length, so the standard settings quietly remove the register that public administration writes in. After retuning, our German pipeline gave us 1.3T unique tokens of organic German web.

The second was to rephrase German documents we already had. An LLM rewrites an organic German document in the style of an encyclopedia entry, a Q&A dialogue or a text passage, preserving its content. This teaches the model the same facts in several surface forms and multiplies the information contained in scarce data, and it is a different operation from translation: the source is German, so the subject matter and the cultural affinity stay German: chancellor, not president. It does not add much new knowledge, rather new phrasings of knowledge that was already in the corpus. Rephrasing gave us about 1T unique tokens, making it the single largest source of German in the model.

The third was translation, and this we only used in Kolibri Origin. Translating English into German works when the model, the prompts and the chunking are chosen carefully, but it carries two problems. Output can still show translationese, the literal rendering of idioms: "Drive safe!" becomes "Fahre sicher!" rather than "Komm gut an!". More importantly, cultural context does not translate. A corpus translated from English inherits the geographic, demographic and institutional distribution of the English web, so a model trained on it speaks German about a world that looks American.

In the end German entered Kolibri as a 2.4T-token unique pool, 80% of it curated or generated by us and 20% from open datasets, and the model saw it at 21.3% of pre-training tokens, roughly 4.3T over the 20T run through upsampling. Each German token was seen 1.8 times on average, well inside the four-epoch limit past which repetition stops paying off. Around 85% of the German the model read is web text, either organic or rephrased from organic. The rest is curated documents – parliamentary proceedings, legal texts and other data in the public domain – and that small translated share.

Specialized tokenizers for English-German

We built a bilingual English-German tokenizer, trained and specialized on each model's pre-training dataset. Because of the 21% German share in our data, it compresses German language better than other SOTA models. This led to more efficient inference (fewer tokens) minimizing costs and shortening response times. We introduce a new way to train tokenizers, UniBPE, that respects the morphology of languages better than existing approaches, especially the compound structure of German, without sacrificing English token efficiency.

German web (FineWeb-2)

  • Kolibri 128,000 vocab 4.90
  • Kolibri Origin 96,000 vocab 4.69
  • Plain BPE 128k 128,000 vocab 4.89
  • GPT-5 200,019 vocab 4.35
  • DeepSeek V4 129,280 vocab 3.72
  • Kimi K3 163,586 vocab 3.28
  • GLM 5.3 154,856 vocab 3.93
  • Qwen3-Next 151,669 vocab 3.59
  • Qwen3.5-3.8 248,077 vocab 4.17
  • Gemini 262,144 vocab 4.13
  • EuroLLM 128,000 vocab 4.08
  • Tekken (Mistral, Nemotron, Apertus) 131,072 vocab 4.03

English web (FineWeb)

  • Kolibri 128,000 vocab 4.58
  • Kolibri Origin 96,000 vocab 4.50
  • Plain BPE 128k 128,000 vocab 4.59
  • GPT-5 200,019 vocab 4.67
  • DeepSeek V4 129,280 vocab 4.59
  • Kimi K3 163,586 vocab 4.62
  • GLM 5.3 154,856 vocab 4.61
  • Qwen3-Next 151,669 vocab 4.52
  • Qwen3.5-3.8 248,077 vocab 4.47
  • Gemini 262,144 vocab 4.49
  • EuroLLM 128,000 vocab 4.16
  • Tekken (Mistral, Nemotron, Apertus) 131,072 vocab 4.45
Tokenizer compression in average bytes per token on German and English web text. Kolibri achieves the best German compression in this comparison. Plain BPE 128k is standard BPE trained on the same data with the same settings as Kolibri, which isolates the effect of the training method. More text per token means fewer tokens per task – higher is better.

The two leading approaches to train tokenizers are BPE and Unigram . We combine those two, keeping the bottom-up approach of BPE and using the Unigram training objective for selecting which merge to add to the vocabulary. On a 128k vocabulary trained on our English/German dataset, this substantially improves tokenization.

Bundessozialgerichtes Federal Social Court (genitive)

  • Kolibri Bundes sozial gericht es
  • GPT-5 Bund ess oz ial gericht es
  • Qwen3.8 Bund ess oz ial gericht es
  • Gemini Bund ess oz ial gericht es
  • Mistral Medium 3.5 · Nemotron 3 Nano Bund ess oz ial gericht es

silkworm

  • Kolibri silk worm
  • GPT-5 sil kw orm
  • Qwen3.8 sil kw orm
  • Gemini sil kw orm
  • Mistral Medium 3.5 · Nemotron 3 Nano sil kw orm

Protokolldaten log data

  • Kolibri Protokoll daten
  • GPT-5 Pro tok ol ld aten
  • Qwen3.8 Protokol ld aten
  • Gemini Protok ol ld aten
  • Mistral Medium 3.5 · Nemotron 3 Nano Pro tok ol ld aten

coprocessors

  • Kolibri co processors
  • GPT-5 cop rocess ors
  • Qwen3.8 cop rocess ors
  • Gemini cop rocess ors
  • Mistral Medium 3.5 · Nemotron 3 Nano cop rocess ors
How different tokenizers split the same words. The Kolibri tokenizer follows the morphology of the language. English and German words split into meaningful units, while competitors cut across morpheme boundaries. Lower is better: fewer, cleaner splits per word mean fewer tokens.

Grounding: reducing hallucinations

Common LLM training and evaluation rewards guessing: a guess has some chance of landing the correct answer, while abstention has none. Therefore, models learn to answer with whatever they have, even if their input does not provide enough information to warrant a response. Asking a model to provide citations often backfires for the same reason: models can simply hallucinate plausible-looking sources to justify an ungrounded answer.

For a regulated customer, a model that knows to abstain is the difference between a pilot and a deployment. We treat saying "I don't know" as an important model capability and (1) develop and track dedicated grounding and anti-hallucination measures and (2) use and develop dedicated training procedures to improve this abstention capability.

During training, we use training data samples where the correct answer is "I don't know". While fine-tuning for abstention is gaining traction industry-wide, high-quality negative examples remain hard to come by, especially where our customers need them, in narrow verticals where all data is scarce. Therefore, we also train with our Merlin-Arthur procedure, developed in-house and explained in detail in our blog post , which solves data scarcity by automatically exploiting the model's weaknesses at each training step and generates synthetic negative examples from available documents. This training becomes a game with three players. Arthur is the model we train and ship. Merlin takes an existing document context, generates a new one that increases Arthur's probability of answering correctly. Morgana generates a new datapoint by stripping out the relevant evidence from the document and tries to lure Arthur into a hallucination. Arthur does not know which one he's facing, so his valid strategy becomes to carefully read the question and the context, to be able to answer on Merlin's context, while recognizing that abstention is the strictly required answer for Morgana's redacted context – making any guess, even a lucky one, incorrect.

For measuring hallucination abstention, we use established benchmarks alongside our own proxies derived from customer use cases. We've also developed our own "M/A grounding score" which falls out of our Merlin-Arthur setup representing a lower bound on how much of the answer provably came from the document.

Kolibri hallucinates far less than Kolibri Origin: it abstains instead of answering wrong on 44% of AA-Omniscience items (Origin: 15%), and on RGB it holds back more often (86% vs 74%) and invents fewer falsehoods (87% vs 76%). On our own M/A grounding score , which certifies rather than estimates how much of an answer came from the document, Kolibri reaches 0.23 where Kolibri Origin, along some other models, reach 0.

AA-Omniscience Non-Hallucination Rate public set share not answered wrong RGB: holds back when the documents don't answer the question RGB: invents nothing no falsehoods when the documents lack the answer FRAMES multi-document reasoning M/A grounding score our own metric axis from 0 to 0.5

Show the numbers
Benchmark Kolibri Kolibri Origin Qwen3.6-35B-A3B Qwen3-Next 80B-A3B Nemotron 3 Super 120B-A12B Mistral Small 4 119B-A6B
AA-Omniscience Non-Hallucination Rate 44.0 14.8 56.7 12.3 13.9 34.7
RGB: holds back 85.6 73.9 79.6 81.3 74.6 82.3
RGB: invents nothing 87.3 75.6 84.3 83.9 86.0 87.0
FRAMES 71.2 65.7 74.7 68.9 74.9 71.9
M/A grounding score 0.23 0.00 0.12 0.00 0.00 0.06

Contextualized performance

Most public benchmarks fail to capture specialized sector needs. To measure contextualized performance – meaning how well a model handles the unique workflows and domain knowledge of specific industries – we need measurements of performance in specialized verticals which are important to our customers, such as the German public sector, legal, hardware, consumer electronics, and automotive.

For this, we first built domain-specific evaluation suites that mirror the skills, tool-calls, workflows, and edge cases required in key sectors and thus capture contextualized performance. Once this was done, we created training environments by generating high-quality specialized synthetic data capturing those required skills. To increase the real-world robustness of the model on those tasks, we additionally randomized the environments around concrete details (tools, configurations, harnesses).

Together, these contributed to an iterative engine that allowed us to sharpen the model's agentic RAG capabilities, domain-specific logic and capabilities, and to hill-climb the contextualized evaluations without ever training on customer data.

These evaluation suites run automatically as part of our pipeline: every checkpoint is scored on them as it lands. Kolibri compares well against competing open-weight models. On the agentic-RAG benchmark Honeypot, Kolibri outperforms all compared models, including the much larger Nemotron 3 Super and Mistral Small 4. On a cleaned version of the agentic-RAG benchmark MuSiQue, it is second only to the more massive Nemotron 3 Super and well ahead of the rest. On the five customer applications, Kolibri leads or is on par with the best model in four of five. These findings speak to our ability to handle a range of data and harness setups, across use-cases, whereas the compared models are sensitive to these details.

These evaluations for measuring contextualized performance can exist because customers tell us how their applications fall short. Each one encodes what we learned from those conversations: what kind of documents matter, which tools the model should get, which questions are hard, where the deployed system frustrates its users today. That makes the exchange concrete in both directions. A customer who shows us a failure case gets it turned into a benchmark that every future checkpoint is measured against, and we can hill-climb the target without ever training on customer data or risking over-fitting.

MuSiQue (cleaned) Agentic RAG Honeypot Agentic RAG Semiconductors Customer proxy German public sector Customer proxy Aerospace Customer proxy Automotive supplier Customer proxy Industrial drive technology Customer proxy

Show the numbers
Benchmark Kolibri Kolibri Origin Qwen3-Next 80B-A3B Qwen3.6-35B-A3B Nemotron 3 Super 120B-A12B Mistral Small 4 119B-A6B
MuSiQue (cleaned) 77.3 42.7 50.5 61.2 79.1 66.8
Honeypot 80.8 25.3 13.5 74.3 68.8 68.1
Semiconductors 80.4 35.3 41.2 79.4 69.6 62.7
German public sector 75.0 54.0 29.5 72.0 78.0 50.0
Aerospace 58.9 14.1 48.1 59.0 54.9 47.0
Automotive supplier 99.0 72.4 84.2 92.6 91.0 87.1
Industrial drive technology 60.0 31.4 32.7 59.5 37.3 56.8

Benchmarks

Benchmarks were run using our own harnesses and, where applicable, all models used the highest respective reasoning effort.

Type MoE MoE MoE Dense
Active parameters 3B 4–6B 12B 27B · 70B
Benchmark Kolibri Kolibri Origin GLM-4.7 Flash 30B-A3B Nemotron 3 Nano 30B-A3B Qwen3.5 35B-A3B Qwen3.6 35B-A3B Qwen3-Next 80B-A3B Thinking Gemma 4 26B-A4B IT GPT-OSS 120B Mistral Small 4 119B-A6B GLM-4.5 Air 106B-A12B Nemotron 3 Super 120B-A12B Qwen3.8 27B Apertus 70B Instruct
Overall (EN) 75.5 54.1 64.7 65.6 74.7 71.4 62.4 71.9 72.3 63.1 64.4 73.0 80.2 –
Overall (DE) 70.8 46.4 50.4 59.3 69.8 67.3 58.0 66.3 70.2 61.4 64.8 67.9 79.9 –
Knowledge
Average (EN) 50.1 39.7 45.5 46.0 52.7 52.1 48.4 51.4 50.0 47.5 45.7 52.0 56.8 –
Average (DE) 57.6 44.9 46.5 41.2 61.3 61.0 55.7 61.5 58.0 51.4 52.2 59.5 69.2 –
GPQA Diamond (EN) 84.3 68.1 73.1 73.9 83.8 83.4 76.1 81.1 76.4 74.7 73.2 78.0 89.2 29.5
GPQA Diamond (DE) 81.3 58.5 59.8 49.6 84.2 80.6 72.2 80.1 76.0 72.9 71.1 76.6 88.1 31.4
Humanity's Last Exam (EN) 21.5 9.4 15.4 12.1 20.4 21.1 11.6 19.2 19.4 9.7 8.7 20.6 35.6 5.2
Humanity's Last Exam (DE) 15.9 10.4 9.1 13.1 18.1 20.5 15.6 23.4 20.7 10.5 10.5 22.3 37.2 5.7
AA-Omniscience Accuracy (public set) 14.8 11.3 17.0 19.5 22.0 19.5 24.2 20.7 23.3 25.0 20.0 26.7 17.5 13.5
AA-Omniscience Index (public set) −32.8 −64.0 −62.8 −45.7 −47.3 −15.3 −42.3 −47.3 −35.2 −24.0 −28.8 −36.5 −9.5 –
MMLU-Pro CoT (EN) 80.0 70.1 76.5 78.3 84.6 84.3 81.7 84.5 80.8 80.4 80.9 82.7 85.0 43.0
MMLU-ProX CoT (DE) 75.5 65.7 70.7 61.0 81.7 81.9 79.4 81.1 77.2 70.7 74.9 79.7 82.4 37.3
Math
Average (EN) 96.5 81.7 88.8 88.8 90.1 87.8 86.3 87.4 90.7 81.4 82.8 91.1 97.8 –
Average (DE) 88.8 74.3 45.2 84.3 79.6 83.7 83.7 88.1 90.8 75.4 80.6 86.5 96.7 –
AIME 2025 (EN) 96.9 81.9 89.4 89.6 88.1 84.6 84.2 87.3 90.6 79.8 81.9 91.7 97.9 0.6
AIME 2025 (DE) 87.5 73.5 43.8 84.4 76.7 82.9 80.6 88.1 90.6 72.3 80.6 85.6 96.5 0.2
AIME 2026 (EN) 96.0 81.5 88.3 87.9 92.1 91.0 88.5 87.5 90.8 83.1 83.8 90.4 97.7 0.6
AIME 2026 (DE) 90.0 75.2 46.7 84.2 82.5 84.4 86.7 88.1 91.0 78.5 80.6 87.5 96.9 0.0
Agentic
Average (EN) 63.4 41.6 58.9 46.4 63.4 62.1 46.3 54.6 54.0 40.7 53.5 54.9 66.7 –
TerminalBench 2.1 27.7 – 20.2 9.7 39.7 – 8.6 – 29.2 21.0 – 39.7 76.8 –
Tau2-Bench (Telecom) 94.7 67.5 95.9 45.9 97.7 99.1 43.9 45.3 73.1 41.5 53.8 68.1 82.5 10.8
Tau2-Bench (Retail) 69.9 58.5 57.9 64.9 70.8 71.6 60.8 71.3 60.5 62.9 61.4 67.5 68.7 9.6
Tau2-Bench (Airline) 76.7 58.7 68.7 52.7 76.0 70.7 65.3 73.3 72.7 40.0 70.7 72.7 83.3 40.0
Tau3-Bench (Banking) 38.1 5.7 7.2 5.7 11.3 10.6 5.4 16.0 14.7 5.7 6.4 15.5 50.0 2.1
BFCL v3 (multi-turn) 39.8 22.8 58.2 47.9 54.0 53.5 51.4 53.4 45.6 36.2 61.6 44.6 42.5 0.6
BFCL v4 (overall) 61.4 36.4 65.4 61.5 70.5 67.2 51.0 68.2 57.3 58.0 67.2 61.0 73.2 –
BFCL v4 (non-live AST) 79.1 78.1 83.3 85.0 85.8 88.2 83.6 83.7 35.8 83.6 85.5 45.0 85.3 –
BFCL v4 (live) 78.9 73.7 78.3 78.8 80.2 81.4 82.5 80.2 70.4 78.4 78.2 77.6 79.9 –
BFCL v4 (multi-turn) 47.5 27.5 62.7 53.5 59.9 58.1 56.0 61.4 55.4 40.4 65.2 51.7 55.5 –
BFCL v4 (memory) 62.8 19.4 41.5 39.1 62.6 53.8 35.3 52.9 50.7 39.1 43.4 59.6 79.6 –
BFCL v4 (web search) 62.5 10.5 69.0 66.0 75.0 68.5 12.5 75.0 57.0 69.0 71.0 71.5 82.0 –
BrowseComp 29.4 4.4 – 14.5 36.5 26.9 2.8 25.5 31.2 – – 29.1 46.4 –
Code
Average (EN) 89.3 68.0 67.8 81.8 85.0 87.7 83.6 89.0 90.8 82.0 79.8 88.3 94.2 –
LiveCodeBench v6 85.9 59.2 46.5 71.3 77.8 82.5 73.9 82.3 87.5 71.2 67.8 82.0 93.8 8.7
HumanEval+ 92.7 76.8 89.0 92.4 92.2 92.8 93.3 95.7 94.1 92.8 91.8 94.7 94.7 41.6
SWE-Bench Verified 66.4 – 51.0 38.6 71.6 73.8 – 57.8 – 60.8 11.6 60.2 72.6 –
Instruction Following
Average (EN) 78.1 62.5 64.5 73.2 72.7 66.1 60.7 79.9 71.1 49.8 38.2 73.7 81.9 –
IFBench (loose-prompt) 78.1 62.5 64.5 73.2 72.7 66.1 60.7 79.9 71.1 49.8 38.2 73.7 81.9 25.5

Get Started

Our model is available openly on Hugging Face under an Apache 2.0 License.

Kolibri requires the aleph-alpha-inference package that provides the Kolibri vLLM plugin. You can either use the provided container image ghcr.io/aleph-alpha/aleph-alpha-inference , or install the package from Aleph-Alpha/aleph-alpha-inference , which also installs the vLLM version it supports:

pip install "aleph-alpha-inference>=1.0"

Serve the model with reasoning and tool-calling enabled.

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

To serve contexts beyond 262,144 tokens, add --max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}' . The recommended sampling parameters for the model are temperature=1.0 , top_p=0.97 and top_k=128 .

Contact for Deployment and Specialization

Contact our team who will be happy to support you through our enterprise deployment and specialization options: contact sales .

Gemini ending free use of Flash and Pro models

Hacker News
www.reddit.com
2026-10-03 05:13:19
Comments...
Original Article

You've been blocked by network security.

To continue, log in to your Reddit account or use your developer token

If you think you've been blocked by mistake, file a ticket below and we'll look into it.

Grow and control a swarm

Lobsters
nohope.io
2026-10-03 05:03:13
Comments...
Original Article

NoHope

Loading the swarm…

Show HN: Offrun – manage every coding agent from one workspace

Hacker News
offrun.dev
2026-10-03 04:40:17
Comments...
Original Article

Get started

Bring the agents you already use.

Offrun runs the CLIs on your Mac, signed in as you. No new account, no API key, and it never sees your password.

  1. Install Offrun. Drag it to Applications.
  2. Connect your agents. Settings, then Agents. Add a second account if you have one.
  3. Start an agent. Pick a project and give it a task.

01 · Worktrees

Agents that never step on each other.

Every agent in a project gets its own git worktree, so two agents working the same repo never touch the same files. Run a refactor and a bug fix at once and merge them on your own schedule.

02 · Peer review

A second agent reads the diff before you do.

Pick a reviewer, and its findings land in the chat while the work is still uncommitted. Ask for the fixes and they land in your message box, for you to read before anything is sent. Nothing merges on its own.

Runs on your own agent accounts. No second subscription, no bot on your pull requests.

03 · Accounts

Hit a limit, keep working.

Point each agent at a different login. When one hits its limit, Offrun moves the chat to a login that still has room and says so. Send your message again and the new login picks up the whole conversation. Out of logins, it offers another agent.

The workspace

What the workspace gives you.

Project memory

Your agents stop forgetting

Offrun keeps repo-wide conventions in one place and each agent's own goal, plan, and dead ends in another. Agents read that context at the start and write back to it as they go, so the fourth session knows what the first one already tried.

Stored as plain files in your project folder. Open them, edit them, commit them.

Notifications

Know which agent needs you

Offrun watches coding agents anywhere on your Mac, including ones running in a terminal it never launched. When an agent finishes or stops to ask a question, you hear about it, even if that window is buried four spaces away.

Preview

Read the work without leaving the chat

Changed files, diffs with line numbers, rendered markdown and tables, images, PDFs, and a small browser with an address bar, all beside the conversation that produced them.

Modes

Decide how much rope an agent gets

Build makes the change and runs it to prove it works. Planner works out the approach and writes it down without changing a file. Switch any chat between them at any time.

Dictation

Talk to your agents

Hold a key and speak your prompt. Offrun transcribes on-device, keeps your identifiers spelled the way your code spells them, and drops clean text wherever your cursor is.

Transcription runs entirely on your Mac. Nothing recorded, nothing uploaded.

What leaves your Mac

Told plainly.

Dictation and your project files stay on your Mac. Agent prompts go straight to your own provider, on your account. Never through us.

Pricing

Free. All of it.

Bring your own agent subscriptions. Everything Offrun adds around them is free.

Free

Everything, unlimited

$0

Every feature, no limits, no seats to count, no card.

  • Unlimited on-device dictation
  • Unlimited agents and projects
  • Unlimited peer review
  • Unlimited connected agent accounts
  • Full model picker and effort control
  • Build and Planner modes
  • Notifications across your whole Mac
Download for Mac

No. It runs them. You keep your own accounts and subscriptions, and Offrun gives them somewhere to live together.

No. Your code goes from your Mac to whichever provider you are signed into, under your own account. Offrun has no server in that path and never handles your keys.

One of your own agent accounts. Pick the model in the model picker.

No. It reads the diff and posts what it found. Every decision after that is yours.

In your project folder, as files. Version them if you want.

Only to set up. After that it runs on your Mac.

An Apple Silicon Mac and at least one agent CLI you already use.

Run them all. Miss nothing.

Free and unlimited. No credit card.

Apple Silicon only (M1 or newer) · macOS 14 Sonoma or later

Problems and solutions to the modern desktop (Make tmux the OS)

Lobsters
matduggan.com
2026-10-03 04:17:27
Comments...
Original Article

I recently watched the talk by Scott Jenson titled "Are we really going to use the same Desktop UX forever?" https://www.youtube.com/watch?v=V7AfAcQwLW0&t=445s . He's a great presenter, really articulate and concise. The kind of speaker that you'd gladly listen to for 3+ hours if given the chance. Which for a talk about window management is quite the compliment.

The overall point of the talk was "Apple and Microsoft aren't going to innovate anymore in the desktop OS space, so it is up to us all to decide what the desktop of the future is going to look like". I would argue that the actual situation to more nuanced than that, they're trying ideas they are just extremely conservative. As Jenson points out, a lot of what we treat as "how computers work" was a workaround for hardware we stopped having years ago. So why are we still accepting those tradeoffs?

To be clear, I'm not a UI/UX designer. I don't know what I'm doing and I lack the skill to make this. I'm making this mostly as a thought exercise and hopefully a prompt to get other people to think about this same problem. I have, however, suffered through others bad choices, which turns out to be most of the qualifications required.

TL;DR: The thing I want is "tmux is the OS", but I want tmux where a normal person could use it. I use it all the time, it's great, but is there a way to take the amazing experience of a scrollable, session-persisted, detachable, task-oriented system to normal people?

What are the high-level problems with windows management in 2026?

I think you can break the problem down to the following components:

My screen real estate is constantly changing and I have a lot more of it.

I am going from my laptop screen to my desktop monitor back to the laptop screen about 6 times a day. When I'm on my big monitor or multiple monitors, overlapping windows don't make sense because I have too much space as it is.

    1. First, one big monitor is a a different experience entirely compared to 2 smaller monitors. I don't use the two monitors as one large unbroken screen, I tend to use one of them for "less important static content" like a ToDo list or my work chat tool and my email and then the "primary" monitor for my terminal/tmux/browser which is how I do all my actual work. But windows "belong" to one or they "belong" to another. Grudin found in 2001 that people use a second monitor exactly this way, where focal work on one, peripheral glances on the other. Twenty-five years ago. The OS still doesn't know which monitor is doing one role and which the other.
    2. And when I pull the laptop out of the dock, everything collapses onto one screen and I manually drag and resize windows to claw back enough real estate to keep working. I don't know how much time this actually wastes for people, but it feels like it wastes a bunch. The OS treats a display change as a surprise, instead of as a scheduled event that happens six times a day.

Different people use windows totally differently.

We have three distinct groups of people that we are trying to design around, while most modern OS optimizes only for the first use case.

    1. Some people are window maximizers, where every window takes up most of the screen and then they use Command + Tab or the Dock or something else to switch between the bigger windows. The value of more screen real estate is mostly that this one big window can be even bigger.
    2. Some people are near maximizers, where one window takes up almost all of the screen real estate, but they'll have one or more smaller windows where they glance at things like status or chat or whatever.
    3. Finally there are people who carefully coordinate all of the windows on their screen.
    4. Almost everyone has "private windows" and "public windows". You are fine with your public windows being lined up and persisting, but you want to hide specific information in other windows from people walking by. The people I complain about are also the people I present quarterly numbers to. I never remember that when I connect my laptop to a display they're going to see my entire screen.
    5. I assumed this was just me with private vs public. But it was measured in 2004, in "Revisiting Display Space Management: Understanding Current Practice to Inform Next-generation Design." Same three types, same public/private split, twenty years ago. See how absolutely none of this is a new idea? Link .

Browsers are mini operating systems.


In 2026 everybody quietly agrees that most of your software runs in a browser tab. These web applications will cover a wide variety of use-cases that have historically been owned by local applications. My desktop OS treats a browser just like a normal application, when in fact it is closer to a virtualized OS running inside of my host OS.

    1. Tabs inside of a browser and windows of that browser contain the same level of complexity as my other applications.
    2. Tabs are associated with streams of work alongside my conventional applications. I'm writing Terraform in Vim in my terminal while referencing the Terraform docs for that provider. But the relationship between tabs and work is messy: the same docs tab pulls duty while I write the code and again while I write the ticket update. Whether a tab should be allowed to belong to two tasks at once, or whether something cheaper is going on, is the exact spot where I break with the research. I'll come back to it once you've met WindowScape.

The concept of "filesystem" is an increasingly weak concept.

My notes in Apple Notes belong to Apple Notes, not my filesystem. My texts live in iMessage's database, not my Documents folder. Teams and Slack content lives inside those applications unless I manually "bring it out," and when I do, I'm making a copy. Firefox will happily show me a PDF, but the PDF doesn't live with macOS unless I take an action to make it so.

    1. Now a lot of engineers are going to read that and think "well you cannot force all applications to use the same storage system for all of their files, are you a lunatic?!?". I'm not suggesting that, in fact the weak filesystem might be a perk. There is a CHI paper from 2004 called, and I am not making this up, "Stuff goes into the computer and doesn't come out." That was 2004 and we have the exact same problem with no real solution. Link .
    2. Here's why this is a windowing problem. Since Windows 95 and System 7 (which is basically as old as my memory goes back to), the machine has worked as a chain: something writes a file, the user opens an app on it, the app writes it back, the file gets sent to someone else, repeat. Every link in that chain assumed the file lived somewhere the OS could see. Every link is now broken in all the modern OS.

What have people already tried?

So I'm not the first person to see this problem. Some of it got solved, but in a different direction than I want. Some of it got solved, but not for normal people.

I started with the document that I kept seeing everyone else cite to. Henderson, D. Austin and Stuart K. Card. “Rooms: the use of multiple virtual workspaces to reduce space contention in a window-based graphical user interface.” ACM Transactions on Graphics (TOG) 5 (1986): 211 - 243. Interesting that even back in the 80s there was a pretty clear understanding that the current system for managing windows wasn't very good. The design they were talking about looked something like this, which is pretty advanced compared to where we are.

Henderson and Card measured window use the way operating systems people measured memory. The screen is RAM. A closed window is a page swapped out to disk. And windows, like memory pages, don't get touched at random: you sit inside a small set of them, two to ten, and that set is the task. Programs spend about 98% of their time inside one of these sets, and roughly half the cost of running happens during the 2% of time spent switching between them. Rooms' whole design was preloading the next set before you ask. They described this as reducing "knowledge faulting in the user," which is the best phrase in the literature and I intend to use it until someone stops me. Just try it out in a corporate meeting: "we need to reduce knowledge faulting in the user".

One of the more interesting side papers I found was "No Task Left Behind? Examining the Nature of Fragmented Work". Link . Mostly because it confirmed something I've long suspected, which is the single task for a long time focus on modern widowing systems isn't actually how people work. The title of the companion paper is a real participant quote: "Constant, constant, multi-tasking craziness." People were juggling around ten "working spheres" a day, minutes at a time.

The closest to what I wanted is from the paper WindowScape: A Task Oriented Window Manager. Link .

WindowScape dropped explicit grouping entirely so now every time you changed the arrangement, it took a photograph, and you went back into photographs instead of filling containers. I love their one-line diagnosis of every system before them: "requiring windows to be in a single group forces users to decide ahead of time where a new window belongs." The photograph metaphor also cracks a problem I'll get to with browser tabs: one window can appear in many photos, because, as they put it, "people understand that there can be several photos of an object with there being only one underlying object." The catch is that the photos evaporated. That's the gap that I think you'd want to solve.

For the first problem, we have more or less already solved for overlapping windows. A tiling window manager ensures that you are maximizing your screen real estate in such a way that you can switch between different layouts with no need to manually modify the windows. There are basically 2 problems with tiling window managers as they exist now.

    1. They're way too hard to use. Like an order of magnitude too hard for normal people to use. Basically if you need to start a sentence with "just open up the configuration file" shut it down the thing is over. The average tiling window manager tutorial asks you to clone a repository before it asks you to open a window.
    2. We need a scrollable tiling window manager. So basically I should be able to set up specific window configurations over here, leave it alone, then scroll to the right and do something else different, like shoving the mess into the back of a drawer. It should be infinite space to work without needing to subdivide the windows into different macOS Spaces.
    3. The scrolling also solves the private vs public window problem. I put my private stuff on the far left and then my public stuff on the right. If I want to look at the private stuff, scroll to the left as far as it will go.
    4. Thankfully this already exists with Niri: https://github.com/niri-wm/niri . It just has to be easier to use. But the tough design elements more or less already work.

Browsers as mini-operating systems: people have tried to solve this, but in the wrong direction. They made the browser more of an operating system instead of making its contents first-class citizens of the one you already have.

    1. The best two examples of this are the Arc browser and Chrome OS. But I think Arc is actually the more interesting of the two experiments. Their first innovation was to break the idea of tabs at the top and instead move it to be a sidebar model.

They also basically "took over" the concept of windows from the OS.

The diagrams above and a good write-up of Arc is available here: https://blakecrosley.com/guides/design/arc .

All of this makes sense from the perspective of "the browser is now the operating system", but I think this is a fundamentally flawed idea. If a web application is operating as an application, it should be its own window. If it is complimentary to another application, it should be a window associated with another application and be allowed to contain many tabs, to reflect the idea of the browser as the portal for all research and lookup.

I was shocked to go through the historical progress of people trying to solve this problem. I'm not the first person to try to solve this problem, I might not even be the 10,000th person.

Graveyard of Attempted Solutions

  • Windows Timeline (2017–2019): All of your activity across all of your apps sorted and organized for you.
  • Windows Sets (2018–19, canceled): apps and webpages in shared tabbed sets. Note what this actually was: the operating system grouping apps and web content into tasks, which makes it the closest thing anyone has ever shipped to what I'm about to propose.
  • macOS Stage Manager (2022): automatic task-based window grouping, widely ignored. The reason I think Stage Manager didn't scratch this itch for people is that its super designed around the iPad style flow of "there is one thing you are using at a time and you need to be able to quickly switch between them". On iPad the model is "one thing at a time, switch fast." On desktop, two or three things have to work together . It compromised toward the iPad and hit for neither.
  • KDE Activities : Everything I'm writing here is old news for KDE. They've been doing task-scoped desktop state for over a decade. Honestly this does most of what I'm writing about, it's just hyper manual and requires you to manage and set it all up. Also I love the KDE design docs, amazing stuff, worth reading. https://community.kde.org/Get_Involved/design
  • PWA install: already gives web apps their own window and no browser chrome, which is clunky but at least makes them "real" applications.

So clearly we understand there's a problem. Why haven't these taken off like wildfire? I think there's a couple of different problems here.

Timeline needed apps to opt in to reporting activity, and most never did. It also logged everything, which read as creepy (which is a problem that my idea would have too), and the UI surfaced Edge features nobody wanted, so it felt like a browser push more than a feature. Sets died somewhere between internal strategy and app-compatibility chaos where the interesting question is why nobody demanded it loudly enough to save it. Stage Manager tried to be universal and pleased nobody. KDE Activities is opt-in, and opt-in means the people who need it most never configure it.

The pattern in the graveyard: the ideas aren't wrong, they're just not defaults, or they're not done at the OS level where they can see all apps. Everything I want has to be structural.

Filesystems as a weak abstraction.

    1. Operating systems have tried to solve this, but because there's no requirement that you use the user OS filesystem to store documents, there's no consistency. MacOS has recent files, but "recent" doesn't mean anything (these are not the most recent files I have downloaded to my computer). As far as I can tell this functionality is completely broken or implemented in a way that makes zero sense to me.

I have opened, downloaded and created dozens of documents since September 23rd (I'm writing this on September 28th). I don't have a single fucking clue how MacOS populates this window. Maybe its broken. Maybe its working as intended through some criteria I don't understand.

Maybe it's personal.

But the idea is here. The ideal would be "across my entire computer and applications, what are the recent files I have interacted with" and expand that out to include emails and Teams/Slacks and everything. I should be able to see everything going on with my machine, search through it and not care if its stored inside of Slack or iCloud or whatever. Single pane of glass.

Microsoft tried it, people didn't like it, but I think the concept makes sense.

What might a solution look like?

So what has changed? Why might we be able to crack this problem now when before it was too complicated? This is where I think a local LLM might make sense, if you can figure out a way to do it where it doesn't cause more problems than it solves.

So what are we looking at here. The concept would be organizing it around the idea of tasks. When you define a task, this allows you to organize all the windows together, including specific browser tabs detached and associated with the task, not with the concept of "browser". You still get get the flexibility of defining glanceables that exist outside of the strict tiling view.

The basic window flow would look like this:

The desktop is one infinite canvas; each physical display is a viewport onto it.

You would end up with the following "transition contract" to handle my initial problem of "what about many monitors/new monitors".

  • Never Change items: horizontal order of windows, scroll position, input focus, task membership.
  • Reflows Deterministically: column widths, like a responsive layout. When the shelf narrows, the books stay in order, the rightmost ones fall off the edge of the viewport into the scroll region
  • Is state, replayed on redock : viewport aims. Two monitors means two viewports aimed at two regions with one at your focal task, one at the glanceables region. Undock: second viewport disappears; its region is one scroll gesture away.Redock: it re-aims where it was.

The OS finally learns which monitor is focal. The display receiving keystrokes ~90% of the time is focal; windows that stay visible but rarely receive focus are glanceables. That's inferable from focus telemetry and requires no eye tracking or config. And when the laptop screen is small, glanceables demote to a thin strip rather than full tiles which is, note, exactly what tmux already does.

One substrate that covers our three user types. Remember the maximizers, near-maximizers, and coordinators? On the canvas they're just column counts. A maximizer is one full-width column. A near-maximizer is one column plus the glance strip. A coordinator is N columns. Nobody gets forced into anything. This is why the design can be a default where KDE Activities was a preference. It's because the substrate doesn't impose a style, it just stops punishing whichever style you already have.

Privacy becomes a first-class flag, driven by machine state. Private is a property of a window or tab, not a region of the screen. Private content renders occluded unless the machine can vouch for safety. The rule I landed on, at the cost of my favorite bad idea (story below): derive privacy state from machine-observable facts like connected displays, or active capture sessions and never from inferred human states like attention or idleness.

Connect to a novel display and it gets a "ready to present" screen by default. You explicitly aim it: this task, this window, or extend the canvas. Known displays skip the dance. Screen sharing is the same event with no cable. You don't need the model to have discipline if the door won't open, and you don't need the user to have memory if the default can't leak.

None of this is exotic. Apple's Keynote's presenter view has been shipping for twenty years plus with slides on the projector, notes on your screen. Apple never promoted the pattern to the OS even though it makes obvious sense. I'm asking for presenter view as a desktop primitive.

Browser tabs detach into real windows and join whatever task they belong to. The "browser" stops being a place windows live.

This all makes sense until you get to "how do you organize this stuff around tasks". I think for more expert users they're going to be able to do this stuff, but part of the problem is that we don't want them to ever have to drop back down into a configuration file. A mouse and dragging stuff around is too clunky for how this would work. Ideally I should be able to ask something that has access to what information is contained inside of each one of these browser tabs and window and help me organize it in some logical flow.

That's where the local LLM comes in.

I'm calling it a "quake overlay" because I'm old and the idea of a text input that drops in from the top with a universal shortcut has always been the quake terminal dropdown to me. Younger people know the same pattern from the Discord overlay, which I have decided not to be mad about. However in the previous diagram its shown as a more friendly "Ask" box.

But the basic flow would be that you ask the LLM to assist you with organizing, it would show you what it's thinking by querying the information through a constrained MCP server with a set verb list and then gives you a preview of what the layout is going to look like. The model never touches the window manager directly. It proposes and you press y.

People rarely pick up completely novel tasks: a graphic designer spends their life in "open files from network storage, make changes, save output, paste into chat, repeat." People rarely sit down at novel display contexts: laptop-only, the desk, the conference room. People rarely attach novel displays. Novelty is rare everywhere, so the expensive one-time machinery of classify, propose, preview, confirm only runs rarely, and everything in between is deterministic replay. You organize a limited series of tasks once and reuse them forever. Check the tickets, open Vim and a browser, work the ticket, write the update, open the PR, post it for review, next ticket. You could make the config a nightmare of JSON and stop caring, because the only intended reader is a model that doesn't mind. The config file stops being an interface.

The biggest issue here would be spatial memory. People understand where things are in relationship to each other on their computer and they don't like it if they are changed. Think of it like "if I came and messed up your physical desk by moving stuff around and adjusting your chair". So I suspect you would need to enforce a strict "append, don't replace" model.

There's a name for this actually. Kirsh and Maglio call it epistemic action, where Tetris players rotate pieces more than the game requires, because rotating is how they think about the piece. Link Your window layout is thought in progress. A helper that "optimizes" it mid-thought is interrupting you.

This is a problem I encountered with just trying this idea out though. Modern LLMs are actually not very good at "touch this stuff, never touch that stuff". So unclear if this is a realistic idea or not. One boring fix: the LLM proposes, dumb deterministic code executes, and the executor refuses to touch anything pinned. You don't need the model to have discipline if the door won't open.

You'd still need a high level UI element for users and this is where I think the more inclusive concept of filesystem could come in.

Tasks become roughly comparable to Directories now. But everything flows from the initial high level concept of "Tasks" and then flows out to individual things. One application is never the task. It's "some chat window, a browser, maybe Preview." Can we try to build it?

So my hope was that here I would be able to post a Linux desktop running an example. But as it turns out (unsurprisingly) it is insanely complicated to make something like this run at all. Some of the underlying ideas work reasonably well (global terminal shortcut dropdown from the top, scroll-able tiles), but it's still pretty clunky. I'm gonna keep working on it and see if I can get something as a demo running, so if you are interested either add me to an RSS reader or just....wait I guess.

In the interest of full disclosure I think this is gonna take many weekends of work to even get a functional demo running. So I'll do my best, but be patient.

Reasons this design sucks

So in attempting to get this up and running in a Linux VM, I immediately ran into serious logical issues with my design. I think failures are often more interesting to read about than successes, so let's talk about them.

  1. The privacy-minded design runs on surveillance-shaped data. There's an irony I sat with for a while. The design promises the OS will finally know which monitor is focal which is the thing Grudin measured twenty-five years ago and the mechanism is input-focus telemetry. Watch which display gets the keystrokes. Watch which windows never receive focus. Store that over time. The privacy-first window manager wants to know more about what you're doing than your current OS does. Keep it local, keep it ephemeral but it's still a lot of behavioral data to drive this system. Just the observation that every mechanism in this post is hungry for data that has historically derailed ideas like this.
  2. My favorite privacy feature was a joke. The original plan was elegant: private stuff far left, public to the right, and when you go idle the viewport drifts back to public. Like a friend changing the subject for you. Then I ran it, and the first thing I learned is that reading is an idle state. You stop typing to read the thing, and the system starts scrolling the thing away from you. Meanwhile, a person walking by sees everything, because a "private" window on a wide monitor isn't hidden it's just further left. The feature hides content from its owner and shows it to the threat. Privacy by occlusion needs to know when a threat exists, and the threat is a pedestrian, which is undetectable.
  3. Nothing in this design ever ages out. The canvas is infinite, the rule is append-only, spatial memory is sacred so the mess grows monotonically. I have solved the electronic messy desk by buying it an infinite desk. When I first read WindowScape, I called their evaporating photographs "the gap to fix." I've changed my mind because evaporation was quietly doing the job of a garbage collector. Something needs to go dormant without moving but deciding what goes dormant is a reaping decision, and reaping is exactly what append-only forbids. It didn't take long for the infinite desktop to feel overwhelming when I tried it.

Fun experiment

I don't know if this underlying idea is a good idea, but it is liberating to stop waiting for MacOS and Windows to do something better or interesting in this space and decide to try it yourself. The graveyard up there is full of good ideas that died of distribution. We're overdue for a radical experiment that ships as a default.

The most eye-opening part of this entire experience was how much the original overlapping window design was a hack to get around fundamental hardware limitations and how soon after it was launched did people clearly see the problems. I'm surprised how often that happens in technology, where you see a problem, search for academic papers about the problem and find just a massive wealth of information from people saying "oh yeah this is 100% a problem and one we should solve soon". Maybe this post inspires someone smarter than me to solve it finally.

Why should I have to pay more for buses if I don't use a smartphone?

Hacker News
www.theguardian.com
2026-10-03 03:48:11
Comments...
Original Article

I applaud Anna Tims for her article ( Schoolchildren without smartphones penalised with higher bus fares, 26 September ), raising the need for a unified approach on this issue to properly tackle the ill effects of such technology on children in the classroom and beyond.

I live in Scotland , where this issue is addressed by making bus travel free for children and young people under the age of 22 with a physical national entitlement card. I am, however, affected by the more general problem of cheaper bus pass options being hidden behind smartphone apps.

I use a non-standard operating system on my phone owing to the always-online models of Android and iOS, which I prefer to avoid.

When comparing the prices of First’s weekly bus pass in my area, which is available to purchase on the bus, and its monthly pass, which is exclusive to its app, I am spending £264 extra over the course of a year simply for not using a phone that this specific app is designed for.

This is unacceptable even without the issues I raised about the main two mobile operating systems, even more so in the context of the present cost of living generally. Accessible public transport should not depend on which computer software is used by the passenger.

First is far from the only offender in this regard – for instance, many supermarkets’ loyalty schemes are now only available as apps as well.
Owen Fraser
Aberdeen

Former SR-71 engineer talks NASA's Blackbird revival program

Hacker News
www.twz.com
2026-10-03 03:32:47
Comments...
Original Article

When it comes to the future of Tail 844, the SR-71 Blackbird that mysteriously disappeared from NASA’s Armstrong Flight Research Center, few people have more insights than Tim Conners . When he worked at NASA in the 1990s, Conners was the lead propulsion engineer for the agency’s Blackbird program. He worked extensively on 844 and has an intimate knowledge of what made the iconic aircraft tick. Now, he brings an insider perspective on NASA’s secretive work to get the Blackbird back in the air and speculates on what it would take to do so, as well as the famed jet’s possible flight test role for the agency after 27 years of dormancy.

Late last month, we wrote about how satellite imagery taken over Armstrong showed that Tail 844 had been missing since at least May. The revelation was first made by Steve Trimble , Aviation Week ‘s defense editor and friend of TWZ. It came after NASA Administrator Jared Isaacman cryptically teased plans for a new high- and fast-flying X-plane with a silhouette of an aircraft that immediately drew comparisons to the Blackbird. Trimble reported just this week that NASA had been talking to ex-SR-71 program personnel about coming back to work on the aircraft. This confluence of events has led to rampant speculation about whether NASA was planning to resurrect the Blackbird and for what purpose ?

In an exclusive interview with TWZ , Conners, now technical director for advanced weapons systems at Tiberius Aerospace, gave us his take on NASA’s Blackbird revelations and insights into what NASA is working on and how far along they are in the process.

Some of the questions and answers have been slightly edited for clarity.

Tim Conners

Q: Tell me about your background with the SR-71 program. How did it begin? How long did it last? And what did you do?

A: So that basically fell in my lap when I was an engineer at NASA’s Dryden [ Flight Research Center ] back in the early ’90s when NASA took delivery of three of the airframes when the Air Force was retiring the fleet. So NASA got two A models and a trainer. Obviously, tail number 844, which is NASA’s designation, is the one that’s getting all of the attention right now. But those came to Dryden in the early ’90s.

By luck of the draw, I was in the propulsion group. We were short-staffed for a number of reasons, and they needed an engineer on Blackbird. And I had been with NASA for about three or four years at that point, and I got tapped to prepare the NASA side of the mission for the upcoming research flights. So I had a few years to get familiar. That was a part-time job, by the way. I also owned the F-15 fleet at NASA as a propulsion engineer, and that was occupying a lot of my time, but it did give me an adequate amount of time to get familiar with the way that the airplane operated, at least from the propulsion standpoint. And then we began doing research missions in ’96. I left NASA in ’98, but I was able to support the program through two years of flight testing and all the years of preparation for those two years.

Q: Tell me more about your role with the Blackbirds.

A: I was lead propulsion engineer on the NASA side. So what NASA would do is they would assign disciplinary leads across from their industry and government counterparts. So Dryden’s disciplinary leads, their job was basically to make sure that any experimental packages were integrated safely with the airframe, and that we could execute the proposed test missions safely and successfully. I owned propulsion and I also owned performance. So if you recall the Lockheed Linear Aerospike rocket engine activity, that package weighed 20,000 to 25,000 pounds and was mounted externally on the back of the SR-71. That was a drag issue, obviously. So part of my role was to make sure that we could actually accelerate through the transonic drag rise with that package and get out to the target test conditions.

SR-71 tail number 844 during its service with NASA. This picture was taken on October 31, 1997, during a flight in support of the NASA/Rocketdyne/Lockheed Martin Linear Aerospike SR-71 Experiment (LASRE). NASA

Q: There is a tremendous amount of interest in seeing the Blackbird fly again. The biggest question seems to be why? What kinds of experiments or research would an SR-71 provide that couldn’t be done more effectively with a modern aircraft, rocket or unmanned vehicle? Can you walk us through some of the reasons NASA would like to see it in the skies again?

A: Yeah, that’s a question that everybody’s been asking. Why? I’m speculating, although I do have a good feel for where this is going. So bear with me.

They probably are not going to do this for — or trying to do this — for aerodynamic reasons. You can do CFD [ computational fluid dynamics ]-based modeling in this speed regime all day long and very accurately in this day and age. It does not take a lot of time or money to do that. What you cannot do is an accurate representation of a new type of engine integrated into a high-speed airframe. You can do the conceptual modeling of it. Of course, it’s difficult to get the final installed answer, especially if it involves an engine that operates at long duration. That is where I think they’re going with this. I believe that they’re developing basically a high-speed engine test bed. That would make the most sense to me.

Q: That’s something you couldn’t do with a more modern aircraft or rocket or unmanned vehicle?

A: That’s a good question. But you know, probably two-thirds of that battle is just developing the airframe. It does depend on the size of the engine that you’re after.

Obviously, if this were a Williams engine or that size class, you could get by with probably a bespoke drone to do your testing. That wouldn’t be that expensive. If your goal is to test something that’s man-rated or very large, like a J58-size engine , what better way to do it than use an airframe that’s already proven to fly at that speed?

SR-71 Blackbird. (Courtesy photo via USAF)

Q: Were you asked to join a team at NASA tasked with getting a jet back in the air? Who else was asked, and what kind of team were they building?

A: I did not receive an offer. There was a tongue-in-cheek text that I got earlier in the year that said, ‘Hey, Conners, how would you like to help get the Blackbirds flying again?’ And that was from somebody who was positioned with inside information. That is where I first heard of this. This was back in spring, but nobody formally reached out to me.

I believe that the only one that’s been tagged so far, as far as a retiree, was Mike Relja. Mike was a crew chief on the Blackbird, and that makes sense that they would potentially want to pull on him. But there was no formal offer, and I tell you what: If somebody approached me to work it, I don’t think I could turn them down.

Q: Do you know where the idea for this project came from and who is in charge?

A: I don’t have an answer. I’d love to know as well.

Q: What can you tell me about this effort? What are the latest developments in this project that you know of?

A: So what I’m hearing is that they’re looking at doing power-on testing fairly soon. That’s pretty big. So of course that’s not engine power on. That’s external power applied to the airframe, seeing which systems still live, which ones don’t, and then moving from there. So that would be an expected first step.

It’s an encouraging step. That means NASA would have gotten beyond cracking open all the bays, inspecting the interior, making sure that the thing at least passes visual inspection for airworthiness from that standpoint.

Q: Have they actually gotten any systems working?

A: No, they’ve done the visual inspection, actually, putting electrons on the airframe. [To get systems working] would be the next step, but it is forthcoming, from what I understand.

As the first traces of dawn light the eastern sky, technicians on the ramp at NASA's Ames-Dryden Flight Research Facility (later, Dryden Flight Research Center), Edwards, California, work to prepare one of NASA's SR-71 Blackbird aircraft for a research flight. Two SR-71 aircraft have been used by NASA as testbeds for high-speed and high-altitude aeronautical research. The aircraft, an SR-71A and an SR-71B pilot trainer aircraft, have been based here at NASA's Dryden Flight Research Center, Edwards, California. They were transferred to NASA after the U.S. Air Force program was cancelled. As research platforms, the aircraft can cruise at Mach 3 for more than one hour. For thermal experiments, this can produce heat soak temperatures of over 600 degrees Fahrenheit (F). This operating environment makes these aircraft excellent platforms to carry out research and experiments in a variety of areas -- aerodynamics, propulsion, structures, thermal protection materials, high-speed and high-temperature instrumentation, atmospheric studies, and sonic boom characterization. The SR-71 was used in a program to study ways of reducing sonic booms or over pressures that are heard on the ground, much like sharp thunderclaps, when an aircraft exceeds the speed of sound. Data from this Sonic Boom Mitigation Study could eventually lead to aircraft designs that would reduce the "peak" overpressures of sonic booms and minimize the startling affect they produce on the ground. One of the first major experiments to be flown in the NASA SR-71 program was a laser air data collection system. It used laser light instead of air pressure to produce airspeed and attitude reference data, such as angle of attack and sideslip, which are normally obtained with small tubes and vanes extending into the airstream. One of Dryden's SR-71s was used for the Linear Aerospike Rocket Engine, or LASRE Experiment. Another earlier project consisted of a series of flights using the SR-71 as a science camera platform for NASA's Jet Propulsion Laboratory in Pasadena, California. An upward-looking ultraviolet video camera placed in the SR-71's nosebay studied a variety of celestial objects in wavelengths that are blocked to ground-based astronomers. Earlier in its history, Dryden had a decade of past experience at sustained speeds above Mach 3. Two YF-12A aircraft and an SR-71 designated as a YF-12C were flown at the center between December 1969 and November 1979 in a joint NASA/USAF program to learn more about the capabilities and limitations of high-speed, high-altitude flight. The YF-12As were prototypes of a planned interceptor aircraft based on a design that later evolved into the SR-71 reconnaissance aircraft. Dave Lux was the NASA SR-71 project manger for much of the decade of the 1990s, followed by Steve Schmidt. Developed for the USAF as reconnaissance aircraft more than 30 years ago, SR-71s are still the world's fastest and highest-flying production aircraft. The aircraft can fly at speeds of more than 2,200 miles per hour (Mach 3+, or more than three times the speed of sound) and at altitudes of over 85,000 feet. The Lockheed Skunk Works (now Lockheed Martin) built the original SR-71 aircraft. Each aircraft is 107.4 feet long, has a wingspan of 55.6 feet, and is 18.5 feet high (from the ground to the top of the rudders, when parked). Gross takeoff weight is about 140,000 pounds, including a possible fuel weight of 80,280 pounds. The airframes are built almost entirely of titanium and titanium alloys to withstand heat generated by sustained Mach 3 flight. Aerodynamic control surfaces consist of all-moving vertical tail surfaces, ailerons on the outer wings, and elevators on the trailing edges between the engine exhaust nozzles. The two SR-71s at Dryden have been assigned the following NASA tail numbers: NASA 844 (A model), military serial 61-7980 and NASA 831 (B model), military serial 61-7956. From 1990 through 1994, Dryden also had another "A" model, NASA 832, military serial 61-7971. This aircraft was returned to the USAF inventory and was the first aircraft reactivated for USAF reconnaissance purposes in 1995. It has since returned to Dryden along with SR-71A 61-7967. NASA Identifier: NIX-EC92-3103-8
(NASA) Courtesy Photo

Q: What is the timeline?

A: That’s supposed to happen within the next couple weeks.

Q: Really? Where is this work taking place?

A: That I do not know for sure. Although I assume it’s happening at [NASA’s] Armstrong [Flight Research Center].

Q: So what is going to happen?

A: That’s a good question. Let me tell you about what I heard more recently, and then you can begin to connect the dots. JP-7. People ask about the fuel. I think it’s pretty well known that NASA had a huge amount of JP-7 stockpiled that they received from the Air Force back in the early ’90s. So much that it was in a dedicated tank, like one of the giant jet fuel tanks at Edwards Air Force Base. You know, the tanks up on the ridge.

From what I understand, that supply unfortunately was discarded about 20 years ago, and JP-7 is a unique fuel. Obviously, low volatility. When I poked at what might be the way forward there, it didn’t sound like that was a showstopper. I got the impression that there’s been dialogue underway with refineries for a replacement or a surrogate that would work. So I went from being deflated, hearing that the fuel was gone, to being encouraged that potentially there was a surrogate workaround.

A statically mounted Pratt & Whitney J58 engine with full afterburner on disposing the last of the SR-71 JP-7 fuel prior to the program’s termination. (NASA)

But that led to the question regarding the engines themselves, so the feedback there was more discouraging in that it looks like the engines are indeed unserviceable. I don’t believe that is based on an actual attempt to run them. It is based on inspection. Probably no surprise. They’ve sat idle, you know, for 27 years. That is what led to questions going back and forth regarding how the airframe would be powered. And that’s what led into where, if you connect the dots, that’s where the airframe would be used as an engine test bed.

So different engines, they would have to be high-Mach bypass systems. We can speculate on which engine company might be developing those systems, but yeah, I don’t want to give it too much away because I don’t want us to lose the inside information that we got.

Q: Are they working on a completely different engine to power Blackbird?

A: I don’t know that for sure. I think if the J58s are not an option for powering the Blackbird, you either button it back up and walk away from it, or you go bring in engines that are already under development that we’re not privy to at this point in time. But there’s no other option. So what are we going to use? It’s not like we can slap in F100 s or F110 s. They’re not going to cut it.

The Pratt & Whitney J58 - The Engine of the SR-71 Blackbird thumbnail

The Pratt & Whitney J58 – The Engine of the SR-71 Blackbird

Q: Obviously, restoring an aircraft sitting out in the elements for years into one able to withstand the rigors of Mach 3 flight is one hell of an undertaking. In your mind, what are the biggest hurdles to this endeavor?

A: Number one would be the engines for sure. Number two would be anything that’s made of elastomerics . Plastics, wiring. Covering connectors. All that stuff degrades and oxidizes with time. You can imagine what frayed wiring and broken wiring insulation – what kind of havoc that would create. What Blackbird has going for it is that all those casings and connectors were built for the rigor of sustained Mach 3.2 flight for an hour, an hour plus, so they are not going to deteriorate quickly. The question is: Have they deteriorated in 27 years? So, if the answer is not a significant amount, then we’ve probably addressed one of the biggest questions – other than the engines – regarding reviving the airplane.

Q: Do you think NASA is going to attempt to just fly Tail 844 largely in its original configuration, or will they deeply modify it or build something new based on it?

A: Completely unknown. They definitely have the ability to modify the airframe, right? They did it before in the ’90s to carry the Linear Aerospike package. That was a huge undertaking on the airframe itself. They did it successfully. Armstrong still has a pretty deep bench in that area, so if they wanted to take on modifying the titanium structure, I think they could do it.

LASRE Pod Matting to SR-71, USA, 1996. View of the Linear Aerospike SR Experiment (LASRE) pod on NASA SR-71, tail number 844. This photo was taken during the fit-check of the pod on Feb. 15, 1996, at Lockheed Martin Skunkworks in Palmdale, California. The LASRE experiment was designed to provide in-flight data to help Lockheed Martin evaluate the aerodynamic characteristics and the handling of the SR-71 linear aerospike experiment configuration. The goal of the project was to provide in-flight data to help Lockheed Martin validate the computational predictive tools it was using to determine the aerodynamic performance of a future reusable launch vehicle. The joint NASA, Rocketdyne (now part of Boeing), and Lockheed Martin Linear Aerospike SR-71 Experiment (LASRE) completed seven initial research flights at Dryden Flight Research Center. Artist NASA. (Photo by Heritage Space/Heritage Images via Getty Images)
View of the Linear Aerospike SR Experiment (LASRE) pod on NASA SR-71, tail number 844. (Photo by Heritage Space/Heritage Images via Getty Images) NASA

But it’s a good question. Again, that would probably point to aerodynamic studies, and if it were me and my dollar as a taxpayer, I would lean on computational studies to get those answers. That’s what I tend to do in my day job is forego testing in lieu of computational when it involves external aero, but it’s when you do the integrated propulsion in this speed regime, that’s really when you need test data from the actual operating conditions.

Q: As opposed to digital modeling?

A: Yes. For instance, say the aerodynamic databases for Blackbird are gone, which they might be. I know there are simple aerodynamic models that still exist, but like a true high-fidelity model that would go into a piloted simulator, that could probably be regenerated fairly quickly. Probably on the order of months computationally in this day and age. So I think simulation capability could be stood up in fairly short order. That, of course, will be foundational for the pilot training task.

Q: We’ll talk more about pilot training later, but first I want to ask about modernization. If NASA were to inject modernized features into the SR-71, what would you like to see? And what makes the most sense in terms of material science, avionics, propulsion and so on?

A: I’m not a materials expert. I think it would be really interesting to see how the airframe could be optimized using adaptive flight control capabilities. So, if the flight control system could be digitized, I think that would lead to some very interesting in-flight experiments. I doubt that that would be. I could be wrong, but if the goal is to get the airframe back up as a propulsion test bed, you don’t need to cut out the hydromechanical flight control system and replace it with a digital one.

A left side view of an SR-71 aircraft from the 9th Strategic Reconnaissance Wing landing. The aircraft is silhouetted against the sunset.
A left-side view of an SR-71 aircraft from the 9th Strategic Reconnaissance Wing landing. The aircraft is silhouetted against the sunset. (U.S. Air Force) U.S. AIR FORCE

That all said, what I would like to see is what we talked about a few minutes ago, and that’s basically modern high-speed engine technology. So what could be done? The J58s were low operating pressure ratio systems. What could be done with a modern, digitally controlled, multistream, high-Mach engine, I think, would be truly trailblazing.

One of the reasons why there’s interest in this speed regime, it was neglected for a number of years. The focus was on hypersonics. If you proposed anything going back, say, 15 years in the Mach 3.5 speed regime, it was ignored. The glitz and glamor was on hypersonics. Hypersonics is great tactically for certain scenarios, but the weaknesses of hypersonics are now becoming apparent.

And as far as countermeasures, hypersonics still has a place. Don’t get me wrong, but every system cannot be hypersonic from an affordability standpoint. So, now you’ve seen it over the last, say, five years. A lot of systems are coming back on the scene that operate in the Mach 2 to Mach 4 regime, and there’s a reason they’re all fitting in there, right? You stay below the speed limit that requires more exotic metallics. You stay in the titanium regime. You can fly, of course, up to [Mach] 3.5 with brief excursions up to about [Mach] 4, so that is tactically relevant against a lot of military targets and defensive systems. So here it is. You know we now have a potential flying test bed that can carry engines in the J58 size class.

The Mighty J58 - The SR-71's Secret Powerhouse thumbnail

The Mighty J58 – The SR-71's Secret Powerhouse

Q: How much could a project like this cost, and is it really economically feasible? How long do you think it could take?

A: I’ve been thinking that over, and I think it depends on who does it. I don’t want to disparage my former government brethren. I had great respect for the folks at NASA when I worked with them, but one of the reasons why I left the agency is because it was becoming so risk-averse. That was adding bloat to everything. And what we used to do quickly, we no longer could do quickly or cheaply at what was Dryden Flight Research Center, so that is what concerns me.

If NASA could operate with the agility that it had, if Dryden/Armstrong could operate with the agility it had 30 years ago, they could do this work in probably a few years. I would say again, it would depend on what exactly they need to do to the airframe. If it’s to resurrect the airframe and bolt in new engines using existing interface hardpoints, that would go much, much more quickly than doing extensive airframe modifications.

Q: Given that 844 has been sitting idle since 1999, what would be the very first thing you would want to inspect or test before even considering putting it back in the air?

A: Exactly what I’m hearing the rumor is that they’re doing this. What is the condition of all the components in the airplane that are flowing electrons? Because if the entire airframe is shorted out, you know how it is.

If you have a car where you’ve hooked up the battery cables backwards – I had a car once where somebody jumped in and did that. It ruined the car because I was chasing shorts constantly, constantly breaking down in the months that followed because the wiring was compromised, melted, bubbled all over the vehicle. And then it would get wet. You get shorts. That, I believe, is what one of the objectives of power-on testing is. To make sure continuity still exists, reliable continuity throughout the airframe. So that’s exactly the first step.

I would say the second would be hydraulics. That vehicle is hydromechanical. So what is the state of the actuators and whatnot?

Q: And you said that work is taking place now?

A: It’s certainly being inspected. That was one of the reasons why they pulled it off of the display pad at Armstrong. But that work has been going on for a couple of months, and it certainly seems that there was at least reason to be encouraged and to go forward with this next step. Otherwise, they would have just towed it back to the parking spot.

The SR-71 Blackbird, Tail 844, when it was on display at NASA’s Armstrong Flight Research Center (Google Earth)

Q: And the next step is testing the systems?

A: That is correct.

Q: We touched on this a little before, but how realistic is it to get an SR-71’s J58 engines operational again after nearly three decades of inactivity? What would worry you most about those engines? How many are left?

A: Well, that’s a good point. As far as [engines] that are readily available, there are several that are at the [ Air Force Flight Test (AFFT) Museum at Edwards AFB]. I believe they’ve been stored outside. But if you consider all the museums around the world with Blackbirds on display, almost all of them have installed J58s, and there’s usually one sitting alongside the airframe on display, so there are plenty of engines. Whether or not they’re serviceable, that of course is the question.

But probably the biggest problem with an engine – well, there are several with a gas turbine that sits unmoved for decades – it’s going to be that you’re going to flatten the bearings. Probably not visible to the eye, but rotating turbo machinery doesn’t like flats on the bearings. So for obvious vibration reasons, you got that. And you’ve got the seals. Seals are going to crack. Trapped fuel is going to turn to gunk and clog up lines. That engine was all hydromechanical, of course, or no digital controls on that because of the heat involved. So fuel hydraulics played a big role in that engine.

There is likely a lot of goo in those lines. I believe that’s why they’ve essentially been written off as unserviceable. I was discouraged because one of the questions I asked that I didn’t get an answer to was: Did anybody try to run the engines or pull them apart? But I don’t even know if the instructions exist to unstack a J58 at this point. I’m sure it could be figured out with the right people, but you know how it is. Pulling something apart is a lot easier than putting it back together.

Q: So there’s no manual for that anymore?

A: I don’t know if there’s a manual specific to that. Like I told Steve [Trimble] at Aviation Week , there certainly are operating manuals. You’ve probably seen the pilot’s manual for the SR-71A. It’s a beauty. There are a lot of copies online. You can get PDFs of those. They are very well-written, huge manuals. I know prints exist for the airplane. Those were in the keeping of the Edwards Museum. I believe those have been handed back over to NASA. So, it’s not like there’s a complete dearth of information.

An SR-71 Blackbird basks in the evening moonlight at Edwards Air Force Base, California. The aircraft is part of the Air Force Flight Test Museum. The aircraft on display at the main museum is the 6th prototype, S/N 61-7955. Assembly started on 13 May 1964 and #955 first flew on 17 August 1965. Throughout its career, this aircraft served as the Palmdale test aircraft until being replaced by SR-71A #61-7972 in 1985. Last flown on 24 January 1985, #955 accumulated 1993.7 hours of flight time. (Air Force photo by Todd Schannuth)
SR-71 Blackbird 844 basks in the evening moonlight at Edwards Air Force Base, California. (Air Force photo by Todd Schannuth) Todd Schannuth

When I was at NASA, I was the person who received everything that Pratt & Whitney had left on the J58s when the Air Force decommissioned the program, and that information was in two or three smallish boxes and consisted of hard copy code printout or printout of numerical data, along with some magnetic reels of code. It was not much. And I was told at the time the only reason I received that information was somebody missed it when the Air Force gave the destruct notice. It was under somebody’s desk, so it was not much in the way of knowledge transfer, unfortunately.

Q: What did you do with all that?

A: I left it behind in 1998 when I left NASA.

Q: What happened to it?

A: I have no idea. Hopefully, somebody scanned it. But again, that was J58 info, right? So if they’re going with a new engine, they’ll have the full digital models for those new systems. And I do believe that recreating the aerodynamic database, including inlet performance, that kind of data can all be regenerated fairly quickly with computational tools.

Q: We touched on this a little bit before, but how difficult would it be to reproduce the JP-7 fuel that powered the Blackbird? It would need designated tankers too, correct?

A: So that was one of the questions I asked is whether or not that formulation exists still. The answer I got back was that essentially a refinery should be able to recreate a suitable surrogate. If you think about it, JP-7 was unique in that it had a very low volatility characteristic.

What I’ve learned about Jet A from some of the ramjet work that I’ve been doing is that it has a ridiculously low volatility too, of course, by design. So I don’t know that JP-7 is that far apart from the chemical characteristics of a Jet A-class fuel, JP-8-type fuel. JP-8 was a significant change from JP-4, so the Air Force was using JP-4 in their fleet, and then moved to JP-8 to be consistent in the ’90s with the Navy with their formulations. That’s a lower volatility fuel, and at NASA we had to recharacterize the flight performance of all of our jets with JP-8, but it worked fine.

We were concerned about operability – that we were going to have flameouts and engine relight and flight relay problems. Those did not materialize, so I don’t want to oversimplify it, but I am thinking that you might be able to take an existing Jet A and with the right additives knock down the volatility to a level that’s safe for flying in the Blackbird.

Airman 1st Class James Douds, a fuels specialist with the 386th Expeditionary Logistics Readiness Squadron, tests the emergency shut off system for a R11 refueling unit while filling the truck, at an undisclosed location in Southwest Asia, Sept. 10, 2017. Fuels management Airmen work around-the-clock supplying approximately 150,000 gallons a day of jet fuel to outbound aircraft in support of the Combined Joint Task Force – Operation Inherent Resolve mission. (U.S. Air Force photo by Tech. Sgt. Jonathan Hehnly) (Tech. Sgt. Jonathan Hehnly) (170810-F-ZI207-0034)
Airman 1st Class James Douds, a fuels specialist with the 386th Expeditionary Logistics Readiness Squadron, offloading JP-8 jet fuel. (U.S. Air Force photo by Tech. Sgt. Jonathan Hehnly) Senior Master Sgt. Jonathan Hehnly

Q: What would need to happen to have aerial refueling jets be able to handle fuel for the Blackbird?

A: There were dedicated JP-7 tankers for the Blackbird. I don’t know if they can take an existing system and flush it. It would depend on the formulation, right? JP-7 was notorious for its unseemly characteristics. It was toxic.

I had a shirt that got dripped on when I was standing under the wing of the Blackbird once, and I had to throw it out. I could not get the stink out of that shirt from the JP-7. It was a weird fuel. So if you can get one that’s more aligned with the fuels used in the service, then I don’t believe that’s going to be a showstopper as far as tanker support, but you do know tanker support is going to be required for that airplane.

Q: You said the JP-7 fuel stank. What did it smell like?

A: I just remember it turning my stomach. Like I kept smelling something as the day went on, and I reached over and saw a stain on my shoulder. I don’t recall what happened the rest of the day, but I just remember being repulsed by the odor and then not being able to get it out when I washed the shirt.

Q: Beyond the fuel, the SR-71 depended on a huge ecosystem of specialized equipment, fluids, personnel and procedures. Which parts of that infrastructure would be hardest to recreate today?

A: That was something that Mike Relja pointed out when he talked to Steve Trimble. That’s a good one. You probably heard that the J58s were started with a start cart that was basically a bank of Buick 12-cylinder engines, all next to each other. It’s a beautiful-sounding system, but very unique. But what is starting, you know, other than you’re either doing a power takeoff via shaft, or you’re using compressed air. It depends on what the engine demands. So I think the start cart development would be fairly straightforward.

But there is a lot of specialized equipment, like Mike pointed out, the wings when they’re split up, they are held in place by a specific scaffolding. The way the engines were craned out of the wing required special equipment. Pulling the inlet spike off that required a special cradle. Those are not insurmountable, but you don’t want to wait till the last minute to realize you need unique equipment. So it’s all got to be taken into consideration.

This engine starter cart was developed specifically for the SR-71 family of aircraft. It used two Buick V-8 racing car engines, linked together through a common gear box, to deliver power to the starter drive shaft of the aircraft engine. More than 600 hp from the two V-8 engines was required to “spool up” the J58 to about 3,200 rpm for starting. (U.S. Air Force photo)

Q: Could existing SR-71s in museums realistically be used as sources of spare parts for 844, or would cannibalizing those aircraft create more problems than it solves? Were those airframes left structurally intact?

A: I believe that was the agreement with the Blackbirds. If a museum were to put one on display, there was a certain expectation regarding care and maintenance. Not so much maintenance, but care and preservation. If you think about it, many of them are under [a] roof in climate-controlled facilities. I know several of them that are preserved like that, including one here in Tucson. That, of course, is the best place for preserving anything with elastomerics and plastic. So I think it’s encouraging that so many of them have been so well cared for. Again, it’s going to come down to which components, regardless of whether or not they’re indoor or outdoor, which components are going to fail first, and I believe that’s the exercise NASA is looking at.

Then you’ve got to start beating down, chasing supply chain sources for some of this, and you know the nightmare that’s out there with supply chain and aerospace. So it’s not going to be an easy task. I don’t want to act like it is. I don’t think it’s insurmountable though.

Q: So these aircraft can be used as sources of spare parts?

A: Sure, absolutely.

Q: Is anyone with a Blackbird already contributing parts?

A: I know that the Edwards Air Museum is contributing what they can as far as spares and prints. I don’t know who else, although I believe the second NASA SR-71A went to the Evergreen Aviation & Space Museum in Oregon. I would expect them to be in the loop.

( TWZ reached out to the Evergreen to find out what, if anything, it is contributing to this effort. )

An extremely well preserved Blackbird on display inside at the North American Aerospace Museum in Oregon (formally the Evergreen Airventure Museum). (NAAM) Stephanie Tassone

Q: How difficult will it be for modern test pilots to fly Tail 844, or whatever aircraft is created from this project? What would actually be the hardest part of training a pilot to fly the aircraft today? Do you think they will convert it into an unmanned configuration?

A: Well, that would be a difficult airplane to convert to unmanned, unless it went with a digital backbone, and like I said, I think that would be an enormous undertaking. It’s not unheard of. NASA, along with McDonnell Douglas, converted a couple of F-15s to all-digital backbones before the F-15E program began. It can be done. It’s just expensive. It’s not easy.

But the piloting task starts with a good simulator , right? The piloting task is absolutely a big deal. I saw it firsthand at Dryden. That said, it took a couple years, probably two. I would say that’s about right for the Dryden pilots, who are pretty good at what they do, to learn to fly the Blackbird safely, and that was with the use of the two-seat trainer, which we’re not going to have this time.

So I think the key to this happening is getting a digital sim together very quickly, and it wouldn’t surprise me if NASA has begun this process already. Getting that together, plus getting the former pilots who are still with us in the seat to make sure that it aligns with what they recall as far as the unique features of the airplane.

I have 100 hours of stick time on the Air Force Blackbird simulator. I’m not a pilot, but I used to pre-fly the maneuvers for the pilots and NASA on the high-drag missions, so I’m familiar with the way the airplane flies up and away, and it is a bear to fly. It is not easy. So training is not insurmountable. Like I said, it takes a good, high-fidelity sim. NASA is expert at building those kinds of simulators, modular simulators quickly. In this day and age, we should be able to put a high-fidelity sim together fairly quickly for training.

If you recall, the cockpit is an analog nightmare on a Blackbird. Breakers everywhere. Dials everywhere. I doubt that somebody suggested that we go with an all-digital cockpit. That would be so much work to do. But if we go for it with just a pure analog cockpit that we used previously, that’s going to require a lot of familiarity. I do believe that the training systems might still exist for that. It would just be for that airplane. It would just be getting them all cobbled back together.

SR-71 Cockpit Checkout thumbnail

SR-71 Cockpit Checkout

Q: What made the Blackbird so difficult to fly?

A: It’s a low-G airframe, so you got to baby it. It’s easy to get an angle of attack out of range. And then with those chined surfaces , it’s easy to get a pitch-up with the airplane. That’s how I always managed to wreck it in the sim. I would let alpha get away from me.

Q: Do you know if any of the pilots have been recalled for this effort to get this thing flying again?

A: No, I don’t know.

Q: It has been close to 70 years since the A-12 Oxcart first flew. There had to have been other high-speed aircraft that at least came and went in the classified test environment. Considering all the work in hypersonics today, it isn’t hard to believe new platforms exist. Why would the SR-71 still have relevance here, and why not use one of those aircraft instead, if they exist? If they do, is it because the SR-71 is declassified?

A: Yeah, that could very well be what we’re looking at here. I’ve had friends ask that question this week. We were talking about the SR-72 . Where is that at? I don’t know. I know people who went to work on it years ago. I’ve never heard anything since. So it could be that that’s so deep black that there’s no way that platform could be used as a test asset. So I think it’s precisely what you said. The Blackbird’s no longer classified, and that it’s possible that the engines that might be used in Blackbird wouldn’t be classified either.

A rendering of the proposed SR-72. Lockheed Martin

Q: Would any of those engines be applicable for this SR-71 reboot?

A: Certainly.

Q: Do you know if they exist anywhere?

A: No. But from what was leaked or hinted at regarding the size of the vehicle, it would seem that whatever was the power plant for that aircraft, it would be about the right size.

Q: What do you think generally about the renaissance of sorts in high-speed aerospace capabilities in recent years?

A: I worked on HAWC [ Hypersonic Airbreathing Weapon Concept ] at Raytheon, which became HACM [ Hypersonic Attack Cruise Missile ]. One of the systems that I helped design was the inlet isolator used on HACM, which is not operational yet, but it will be. But it’s an air breather by Raytheon. That gave me a lot of exposure to the level of exquisite technology that’s required to field a system like that.

It is expensive. It’s difficult. Supply chain is a challenge. There’s a reason why they cost what they do per unit. So you see the push to develop lower-cost hypersonic systems like Castelion’s Blackbeard [ lower-cost hypersonic missile ]. It’s not an air breather. So I think it’s an easier problem than what we dealt with at Raytheon with an air breather. But nevertheless, it’s impressive that they’ve gotten a price point down to what they’ve gotten, so you see that happening in hypersonics.

You’re kind of stuck in the same block of speed, but the price is dropping. And then meanwhile, you’ve got this massive group of munitions that are setting up camp in a Mach 2 to Mach 4 speed regime for reasons I mentioned. Those are running much cheaper per shot and per effect, which is why that’s happening. So that you give up something on survivability by pulling down into that speed range, but the price point allows you to volley at such a level that you overwhelm the defensive systems.

To date, this is the only picture the US Air Force has released showing an actual air-breathing hypersonic cruise missile test article related to the Hypersonic Attack Cruise Missile (HACM) program and/or the Defense Advanced Research Projects Agency's preceding Hypersonic Airbreathing Weapon Concept (HAWC) effort.
To date, this is the only picture the US Air Force has released showing an actual air-breathing hypersonic cruise missile test article related to the Hypersonic Attack Cruise Missile (HACM) program and/or the Defense Advanced Research Projects Agency’s preceding Hypersonic Airbreathing Weapon Concept (HAWC) effort. (USAF)

Q: If an SR-71 took to the skies again, would it not be one of the highest-profile aviation moments in a generation? It would certainly capture the imagination of many young and old. What do you think about the social aspect of all this?

A: I have yet to find anybody who isn’t excited by the prospect. Certainly, everybody within aviation thinks it’s cool. Some of my family who don’t know much about airplanes – they all think it’s cool. Everybody thinks it’s just super. That said, everybody also wonders why we’re doing it. But what’s funny is everybody still says, ‘Let’s do it anyway because it’s such an iconic airplane, beautiful airplane, an inspiration.’ Hopefully, there would be a good reason to spend the money. But I think even if it’s a bad reason, everybody wants it to take to the skies again.

Q: If you were a betting man, Tim, what would you say are the odds of this thing flying again?

A: I would give it a 25% chance.

Contact the author: howard@twz.com

Show HN: Thoreau BASIC – What if BASIC hadn't gone out of fashion?

Hacker News
thoreaubasic.com
2026-10-03 03:28:54
Comments...
Original Article

Download

Windows x64 · ZIP

UEFI x64 · ZIP

Plain-text documentation

3.2 release notes

THOREAU_GM.TBGM · losslessly compressed wavetable · 513.5 MB

3.2

File browsers and menus are now ordinary BASIC commands. GETFILES gathers files, directories and the parent entry with an optional file mask. SELECTBOX turns a string array into a keyboard- and mouse-controlled selector, with multiple columns, optional colors, and single or double box-drawing frames.

Filenames keep their original spelling and case throughout the file commands. Masks are case-sensitive. Full paths can contain up to 4,095 bytes, including an added extension; filesystem limits on individual names still apply.

STATEWRITE and STATELOAD checkpoint selected variables to two alternating, checksummed generations. FLUSH , FILECOMMIT and FILECOPY support staged saves and verified backups. CRC32 checks byte strings, and TEXTSTATS counts words and characters in a range of a string array.

SYSTEMINFO$ reports CPU and firmware details, usable AVX2 or SSE2, PARFOR capacity, the calibrated TSC counter rate, display backend and build information. Targeted JIT changes reduce work in floating-point INT and native complex, quaternion and octonion arithmetic. RUN can restart a program without growing the host stack, and HELP category headings appear once.

The Opus wavetable and WAV, FLAC, Opus and MIDI playback remain available. Sound banks are expanded to PCM before playback, with no decompression while music plays. Standalone EXE and EFI applications can omit the bank with NOGM . See the manual and release notes for the full changes and syntax.

What it is

Thoreau BASIC starts from GW-BASIC-style syntax but is not trapped in 1983. It has 64-bit memory, 24-bit graphics, sprites, mouse input, TCP/IP and HTTP, a complete GM/GS wavetable with MIDI/WAV/FLAC/Opus playback, complex/quaternion/octonion values, a debugger, profiler, source tools, multicore PARFOR , and native x64 JIT compilation.

The same BASIC language runs as a normal Windows program or directly from UEFI firmware without an operating system. BASIC programs can also be packaged as standalone Windows EXE or bootable EFI applications.

Still BASIC

10 CLS
20 PRINT "HELLO FROM THOREAU BASIC"
30 FOR I=1 TO 5
40 PRINT I
50 NEXT I
60 END
RUN

And now the orchestra is BASIC too:

SOUND PRELOAD
LOADMID 0,"BALLADE.MID"
PLAYMID 0

Graphics, networking and sound remain ordinary BASIC commands, not separate frameworks.

BASIC without an operating system

The UEFI build boots directly on x64 machines. It uses GOP graphics, firmware or raw HID keyboard input, mouse support with fallbacks, networking through firmware protocols or Thoreau's own SNP-based stack, HDA/Azalia PCM sound where available, and the same interpreter/JIT and GM wavetable used by the Windows version.

Pixel Prose

Pixel Prose is included as a Thoreau BASIC example and is also available as standalone EXE and EFI builds. The original versions for DOS, Commodore 64 and Amstrad CPC live with my other retro projects on itch.io .

Project

The project page and downloads are on itch.io .

Contact

info@thoreaubasic.com

Bug reports, compatibility notes and strange BASIC edge cases are very welcome.

If you want to support development, PayPal is available. No nag screens, no locked features.

Vincent Bernat: Self-hosted HTTP tunnels with SSH and nginx

PlanetDebian
vincent.bernat.ch
2026-10-03 03:23:12
A friend wants to proofread your work-in-progress blog post, but its preview only runs on localhost:8080. Several tools can help. Some run as a commercial service, like ngrok or Cloudflare Quick Tunnels. Some are self-hostable but require a specific client, like frp or localtunnel. Some only require...
Original Article

A friend wants to proofread your work-in-progress blog post, but its preview only runs on localhost:8080 . Several tools can help. Some run as a commercial service, like ngrok or Cloudflare Quick Tunnels . Some are self-hostable but require a specific client, like frp or localtunnel . Some only require a plain SSH client but rely on a specific SSH server, like sish . Let’s implement a self-hosted solution with only OpenSSH and nginx !

$ ssh -R 0:localhost:8080 http-over-ssh
Allocated port 41535 for remote forward to localhost:8080
https://6J3jK1WmB15c6WmjW_X-Wg--1789928654@p41535.ssh.luffy.cx/

Basic setup #

First, we forward connections from a port on a remote server to your local service:

$ ssh -N -R 0:localhost:8080 web02.luffy.cx
Allocated port 41535 for remote forward to localhost:8080

When you specify 0 as the remote port, the server allocates a free port. Then, we configure nginx to proxy requests from https://p41535.ssh.luffy.cx to http://127.0.0.1:41535 :

server {
  listen 0.0.0.0:443 ssl ;
  listen [::0]:443 ssl ;
  server_name ~^p(?<port>\d\d\d\d\d)\.ssh\.luffy\.cx$;
  location / {
    proxy_pass http://127.0.0.1:$port;
  }
}

We also need to add DNS records for *.ssh.luffy.cx and get a wildcard certificate through Let’s Encrypt :

*.ssh.luffy.cx.               CNAME web02.luffy.cx.
ssh.luffy.cx.                 CAA   0 issuewild "letsencrypt.org"
_acme-challenge.ssh.luffy.cx  CNAME ssh.luffy.cx.acme.luffy.cx.

acme.luffy.cx is a zone hosted on Route 53. I use it for ACME DNS-01 challenges , both for wildcard certificates and for domains served by several web servers. In my case, NixOS gets the certificates automatically .

Access control #

The port is the only “secret” 1 keeping the content confidential. Other forwarding solutions add a random string to the domain name to prevent an intruder from enumerating the possible values.

Thanks to ngx_http_secure_link_module , we can secure this setup a bit. This module computes a hash 2 over a set of values, including a secret, and compares it with the hash from the request. The hash is base64-encoded, so we cannot put it in the domain name, which is case-insensitive. Instead, we put it in the URL as a username, along with its expiration timestamp: 3

https://6J3jK1WmB15c6WmjW_X-Wg--1789928654@p41535.ssh.luffy.cx/en/blog
        ╰─────────┬──────────╯  ╰───┬────╯  ╰─┬─╯             ╰──┬───╯
                hash             expires    port               path

The client sends the username to the server with HTTP basic authentication . This works with most HTTP clients, including curl . Nginx exposes the username in the $remote_user variable. The module expects the hash and the expiration timestamp separated by a comma. We use a map directive to extract the two parts from $remote_user and join them with a comma. 4 We also give the module the string to hash. It contains the expiration timestamp, the port, and a secret:

map $remote_user $httpssh_link {
  "~^([-_A-Za-z0-9]{22})--([0-9]+)$" "$1,$2";
}
server {
  # […]
  location / {
    secure_link $httpssh_link;
    secure_link_md5 "$secure_link_expires $port ZuPerS3cr3!";
  }
}

The module returns the status of the check in the $secure_link variable:

  • empty if the hashes do not match,
  • "0" if they match but the link has expired, or
  • "1" otherwise.

If the hash is incorrect or missing, we return a 401 error with a WWW-Authenticate header to ask for credentials. If the link has expired, we return a 410 error. We remove the Authorization header before forwarding the request and add a few directives to proxy WebSocket connections . Here is the complete configuration: 5

map $remote_user $httpssh_link {
  "~^([-_A-Za-z0-9]{22})--([0-9]+)$" "$1,$2";
}
server {
  listen 0.0.0.0:443 ssl ;
  listen [::0]:443 ssl ;
  server_name ~^p(?<port>\d\d\d\d\d)\.ssh\.luffy\.cx$;
  location / {
    secure_link $httpssh_link;
    secure_link_md5 "$secure_link_expires $port ZuPerS3cr3!";
    if ($secure_link = "") {
      add_header WWW-Authenticate 'Basic realm="tunnel"' always;
      return 401;
    }
    if ($secure_link = "0") {
      return 410;
    }
    proxy_pass http://127.0.0.1:$port;
    proxy_set_header Host $host;
    proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
    proxy_set_header Authorization "";
    proxy_http_version 1.1;
    proxy_set_header Upgrade $http_upgrade;
    proxy_set_header Connection "upgrade";
    proxy_buffering off;
    proxy_read_timeout 30m;
  }
}

I think you are now asking yourself the obvious question: “How should I generate the hash?” Easy peasy!

$ expires=$(( $(date +%s) + 86400 ))
$ port=41535
$ secret='ZuPerS3cr3!'
$ printf '%s %s %s' "$expires" "$port" "$secret" \
>   | openssl md5 -binary \
>   | openssl base64 \
>   | tr +/ -_ | tr -d =
6J3jK1WmB15c6WmjW_X-Wg

Well, I suppose you are now saying: “Vincent, this is not very convenient! I’ll stick with ngrok if you don’t mind.” Okay, I hear you. Let’s write a helper script.

Helper script #

The main difficulty is finding the ephemeral port that OpenSSH allocates, as it does not appear in any environment variable. 6 To work around this obstacle, we look for the ancestor sshd-session processes: 7

pids=$(
  pid=$$
  while [ "$pid" -gt 1 ]; do
    line=$(ps -o comm=,pid=,ppid= -p "$pid")
    echo "$line"
    pid=${line##* }
  done | awk '$1 == "sshd-session" { printf "pid=%s,\n", $2 }'
)
if [ -z "$pids" ]; then
  echo "not an ssh session" >&2
  exit 1
fi

Then, we get the listening ports associated with these sshd-session processes: 8

ports=$(sudo -n ss --listening --numeric --tcp --processes --no-header \
  | grep -F "$pids" \
  | awk '{ print $4 }' | awk -F: '{ print $NF }' \
  | sort -un)
if [ -z "$ports" ]; then
  echo "no forwarded port, use ssh -R 0:localhost:PORT" >&2
  exit 1
fi

Finally, we display the URLs and keep the session open:

lifetime=86400
secret='ZuPerS3cr3!'
expires=$(( $(date +%s) + lifetime ))
for port in $ports; do
  token=$(printf '%s %s %s' "$expires" "$port" "$secret" \
            | openssl md5 -binary \
            | openssl base64 \
            | tr +/ -_ | tr -d =)
  echo "https://$token--$expires@p$port.ssh.luffy.cx/"
done
sleep infinity

I install this script as http-over-ssh on the server and add this entry to my ~/.ssh/config :

Host http-over-ssh
  Hostname web02.luffy.cx
  RemoteCommand http-over-ssh
  ControlPath none

With this solution, I only rely on OpenSSH and nginx, two pieces of software already running on this server. One short command gives me a self-hosted tunnel and a URL to share. To try it, grab the complete helper script , which includes a few minor improvements. If you run NixOS, as any person of taste would, have a look at my http-over-ssh.nix instead. ❄️

Understanding Frontier Artificial Intelligence

Hacker News
casp.ac
2026-10-03 03:06:02
Comments...

The characters of plastics (2024)

Hacker News
yarchive.net
2026-10-03 02:14:16
Comments...
Original Article

“It’s made of plastic” is a common put-down. Marketers selling higher-end bits made of plastic (like gun parts) try to evade the stigma by calling it “polymer”, but that’s just a stupid euphemism: every plastic is a polymer, though not every polymer is a plastic. The word “polymer” says nothing to indicate that this might be a superior sort of plastic. Yet there are superior sorts; plastics vary widely in their characters. Some analogies between plastic and human characters:

Polyethylene, polypropylene: Snow White (simple, pure, weak). These are the cheapest and most common plastics, and are made of just hydrogen and carbon atoms. Ethylene (two carbons and four hydrogens) is polymerized to make polyethylene, and the quantity of ethylene made each year is measured in cubic miles. They can of course have pigments added to them to give them color, but are commonly used in uncolored, nearly-pure form. Chemically they are unreactive, which makes them good for containers of all sorts; they are even used for chemistry beakers. That same unreactivity makes them very hard to glue, and means they don’t deteriorate with time. Leave them out in sunlight, though, and the UV quickly weakens them to where they crack easily.

Nylon: Arnold Schwartzenegger (strongman). Plastic nuts and bolts, which need to be strong, are usually nylon, as is monofilament fishing line and womens’ hosiery (delicate, but still strong for its weight). Nylon adds nitrogen to the list of atoms it contains; it’s chemically known as a “polyamide”, a class which also includes proteins.

Nylon with 30% glass fiber reinforcement: Arnold Schwartzenegger on steroids. It’s as strong as cast aluminum (though extruded aluminum can be much stronger). When guns are made of plastic, this is usually the plastic they use; likewise for the plastic housings of electric drills and other power tools.

Polycarbonate: Achilles (warrior with a fatal weakness). Polycarbonate is what they make bulletproof windows from, and safety glasses, and the transparent fronts of car headlights. Its weakness is chemical attack: a splash of acetone, and it instantly “crazes”, a network of fine cracks appearing across its surface as built-in stresses are relieved. (If the part has any built-in stresses, that is; molded parts probably do, but flat sheets might not.) Polycarbonate is often protected by a surface coat to block chemical attacks. It’s the main plastic that gives off the notorious bisphenol A (BPA), which is probably not so bad as it’s reputed to be. But it still doesn’t make sense for food containers to be made from polycarbonate, since they don’t need to be bulletproof and polycarbonate is pricey.

Polyvinyl chloride (PVC): Dr. Jekyll and Mr. Hyde (a character who varies wildly depending on which drugs he takes, with more than a bit of evil in him). PVC is normally rigid, and is used in that form for house siding and sewage pipes, but can be pumped full of plasticizer to make it flexible; in that state it’s used for inflatable boats, shower curtains, imitation leather, and even sex toys. In pure form it’s not a stable chemical; its staying good depends on added “stabilizers”, which themselves can be somewhat evil: they often contain the toxic element lead, which normally is locked in the plastic but is unleashed when the plastic deteriorates or is burned. Burning it also gives off a variety of toxic chlorine-containing substances. The plasticizer, if used, evaporates away over years (in cars, landing on the inside of the windshield and producing an annoying haze which has to be wiped away), and eventually the remaining plastic gets brittle and cracks. The plasticizer is often a phthalate, another notorious class of chemical and also probably not as bad as popular repute would have it, but still not good for you and a pain to clean up. (People commonly regard soft plastics in cars as good and hard ones as bad and cheap, but I don’t think they’ve made the connection between having soft plastics around and needing to clean off the inside of the windshield, or for that matter having phthalates in the air they breathe when getting into a car on a hot day.)

Polyurethane: an android from the movie Blade Runner (strong, versatile, dangerous, doomed). Polyurethane can be made either hard or flexible, with the flexible varieties used to make rubbers and foams, but polyurethanes have a tendency to decide “now it’s time to die ”, losing almost all of their strength, with foams crumbling or turning to goo and rubber parts cracking. They are made using isocyanates, in a reaction which could be thought of as Nazi chemistry since it was in fact invented in Germany during that era and since isocyanates are quite toxic. (That’s just a resemblance, of course; in reality Nazi ideology had nothing to do with chemistry.) Though toxic (methyl isocyanate killed thousands of people in the Bhopal disaster), isocyanates are not cyanides (which are even worse), and in finished polyurethane there are only traces of isocyanates left, though they might be the cause of the annoying smell that new polyurethane foam often has. But burning polyurethane does produce some hydrogen cyanide, and polyurethane foam burns unusually vigorously, unless it’s been treated with flame retardants, which thus have been mandated by law in some places but which themselves might be a health issue.

Bakelite, aka phenol-formaldehyde: well, nobody really comes to mind as a corresponding character, but it’s rigid, brittle, and can take a lot of heat. Bakelite was one of the first plastics, and is thermosetting: it doesn’t melt but rather has to be formed from its ingredients (phenol and formaldehyde) into its finished shape (requiring a mold, heat, and pressure). With it we’re back to things that contain only the safer elements (carbon, hydrogen, and in this case oxygen). Phenol and formaldeyhde are each nasty, but their nastiness is consumed when they react together. Bakelite’s brittleness can be mitigated by using appropriate fillers, particularly fibrous ones. Its heat resistance is such that, mixed with high-temperature fibers, it is used for heat shields for reentry from space (which burn away but do so slowly enough that they don’t burn through). It also sees a lot of use in things like electrical circuit breakers, where other plastics might melt into a blob and let their metal parts short-circuit.

Polystyrene: Joe Sixpack (common, cheap, weak, nasty). It’s not heat-resistant, not chemical-resistant (it dissolves in a wide variety of solvents), and is brittle. But it can easily be blown into a foam (styrofoam), which because of the bubbles is a good insulator; being a foam also mitigates its brittleness. When heated it de-polymerizes and gives off styrene, which has a distinctive acrid smell. It can be improved into “ABS” by adding large proportions of acrylonitrile and butadiene into the polymerization process; ABS has the same distinctive smell when heated but is considerably tougher, making it suitable for a wide variety of consumer products (though it’s still much weaker than nylon). Legos are ABS. If there’s a plastic which really deserves the putdown “it’s made of plastic”, it’s ABS: it’s common enough to be the sort of thing people think of when they hear “plastic”, and is generally mediocre. But as human analogies go it’s a step above Joe Sixpack; call it Joe Blow.

Teflon: Queen Victoria. “Nobility” in chemistry means a lack of reactivity, an immunity to chemical attack; “noble metals” are noble not because they’re expensive, but because they snootily look down their noses at other chemicals and refuse to have relationships with them. In plastics it’s harder to get nobler than teflon (polytetrafluoroethylene; PTFE), which differs from polyethylene by having all its hydrogens replaced with fluorines, which are much more difficult to dislodge. Fluorine chemistry being a difficult and dangerous sort of chemistry, PTFE doesn’t come cheap. It’s used to coat nonstick pans, where its nonreactivity translates into things not sticking to it. As best I can tell (though information on this is scarce), even the newer “ceramic” nonstick coatings often have an imperceptibly-thin layer of a molecule with a highly-fluorinated tail that resembles PTFE; the ceramic gets the publicity but the highly-fluorinated chemical does the work of making the coating nonstick… until it wears off, which because of the thinness of the layer happens more quickly than it does with the older sort of pan which has a visible layer of PTFE. (Modern nonstick pans are often regarded with suspicion; but the old mainstay in that department, cast iron, gets its nonstickness by being “seasoned”: coated with a layer of burnt, polymerized oil, no doubt containing many carcinogens.)

Those are just some highlights of a complicated subject; I haven’t even mentioned some major plastics, and there are a host of minor ones, as well as innumerable subvarieties of the major ones. (Take polyethylene, normally weak, react it until the molecular chains are long, then stretch it until they are aligned, and you get an extremely strong fiber, suitable for high-strength ropes or bulletproof vests: not Snow White but Wonder Woman.) With such a varied cast of characters, it’s a pity that they all get lumped together as “plastic” in common usage. I suppose it’s somewhat inevitable, since you can’t just look at a piece of plastic and tell what sort it is. But more public awareness of the differences would be nice.

Deepfake review – ingenious interview with a PM who dances around every issue

Guardian
www.theguardian.com
2026-10-03 02:00:52
Sadler’s Wells East, LondonAn interrogater tries to get to the bottom of a surreal scandal while green screen actors meddle with the story in Jess and Morgs’ inventive production What audiences remember most from a political interview, a voiceover tells the prime minister (Kennedy Muntanga) as he pr...
Original Article

W hat audiences remember most from a political interview, a voiceover tells the prime minister (Kennedy Muntanga) as he prepares to face his TV interrogator (Rebecca Bassett-Graham, coiffed à la Emily Maitlis), is less the words themselves than the tone of voice and the tenor of the body. Muntanga is on point, then: his arms sweep, his fingers beetle and his hands spread out, as if he were wordlessly fielding fly-by questions and offering open statements. It’s the moves and the staging that count.

This setup lands right in the zone for creative duo Jess and Morgs (Jessica Wright and Morgann Runacre-Temple), whose turf is a congenial mix of choreography, film and theatre. In Deepfake, along with writer Jeff James , they explore the various unravellings of an encounter in which the interviewer tries to find the truth behind a surreal – but maybe real? – story in which the PM was seen variously spattered with yellow paint, with his flies undone, and partaking in a ritual involving a deer’s head.

Emily Maitlis-alike … Rebecca Bassett-Graham in Deepfake.
Emily Maitlis-alike … Rebecca Bassett-Graham in Deepfake. Photograph: Foteini Christofilopoulou

The content of these scenarios perhaps matters less than the ingenious manner of their staging (props to the design team too), which deploys green screen, live relay and other film and stage effects to storify the interview with simulated settings, duplicates of the figure of the PM, selective perspectives and closeups of the central encounter. Spidering about the main characters and manipulating the sets are a team of seven dancers in green body stockings, anonymously blending into the background but integral to the whole media matrix.

Initially the tone is paranoid and conspiratorial, the dancers syncing to voiceovers that then start to glitch (a signature device of works by Crystal Pite and Jonathon Young ). There are ambiguous, unvoiced duets which intimate, intriguingly, that interviewer and interviewee may not only be playing the same game of moves and countermoves, but might sometimes even be on the same side. Later, the piece starts shifting gears, swerving through farce, parody (there’s some delightful and very unexpected play with 19th-century ballet), absurdism, and ultimately a kind of comic bathos, equally unexpected.

It’s inventive, clever, exactly executed and quite a lot of fun, but the tone is also scattershot, and the piece ends up feeling more lightweight than its deepfake premise had promised – as if it wasn’t quite convinced by its own story.

Memory-Safe WebP Decoding

Hacker News
halide.cx
2026-10-03 01:45:30
Comments...
Original Article
Building

wpd is a faster, safer WebP decoder than libwebp, designed to help secure the Web from vulnerabilities like CVE-2023-4863 . At the same time, wpd can't just maintain the status quo for speed; it offers superior single-threaded performance, and parallelizes better across multiple threads compared to alternatives.

Source code: https://github.com/halidecx/wpd

Safety

Image decoders and other image processing libraries are used everywhere, from the OS level to sandboxed browser processes. They also process complex untrusted input data, which makes them vulnerable to memory safety bugs. CVE-2023-4863 affected potentially billions of devices running Chrome, Firefox, Signal, Microsoft Teams, and more – CISA confirmed it was actively exploited in the wild, and it was added to their Known Exploited Vulnerabilities catalog thereafter. Vulnerabilities like these have serious consequences for nearly all consumer hardware.

While libwebp is likely safer than it was in the past, it does not "solve" memory safety by being well-fuzzed. According to the Chromium team, around 70% of their high-severity security bugs come from memory safety issues.

To address this, wpd is written in Rust, with handwritten assembly routines for performance. The handwritten SIMD present in the decoder is carefully scoped and checked for correctness, and largely exists in less risky places. Nonetheless, for consumers looking to harden their environments, wpd can be compiled without handwritten assembly; this leaves the non-SIMD code we've written entirely verifiably memory-safe, with the only unsafe code being in the vetted zerocopy crate. This is a marked improvement over libwebp, which is written entirely in unsafe C with SIMD intrinsics.

Performance

It isn't enough to just be safer; image-webp already has that covered. We also needed to make wpd faster, so safety did not come with compromised performance.

We benchmarked wpd on a subset of our developer test data , and compared to libwebp, it is:

  • 1-thread, lossy: 1.19x faster
  • 1-thread, lossless: 2.74x faster
  • multi-thread, lossy: 2.68x faster
  • multi-thread, lossless: 3.19x faster

Because our test suite mixes animated WebP content with still content, our multi-threaded results take advantage of parallel image decoding which results in more impressive gains there. Our single-threaded advantage is pure algorithmic improvement.

These numbers come from our benchmarking harness in the wpd repository, using image-webp 0.2.4 and libwebp a1d89ff .

Features

Feature parity with libwebp is a target for wpd as well, as every use case that relies on libwebp deserves an upgrade. Compared to image-webp, wpd is a proper libwebp replacement, with some additional features included on top:

Decoder feature image-webp libwebp wpd
Lossy WebP ☑ ☑ ☑
Lossless WebP ☑ ☑ ☑
Alpha transparency ☑ ☑ ☑
Animation with composited frames ☑ ☑ ☑
Animation durations and loop count ☑ ☑ ☑
Animation rewind/reset ☑ ☑ ☑
Raw animation subframe decoding ☐ ⚠¹ ☑
Per-frame geometry, blend, disposal inspection without decoding ☐ ☑ ☑
Image dimensions and alpha inspection without decoding ☑ ☑ ☑
ICC, EXIF, and XMP extraction ☑ ☑ ☑
RGB/RGBA output ☑ ☑ ☑
BGR/BGRA/ARGB output ☐ ☑ ☑
Premultiplied alpha output ☐ ☑ ☑
YUV420P output ⚠² ☑ ☑
YUVA420P output ☐ ☑ ☑
RGB565 and RGBA4444 output ☐ ☑ ☑
BGR565 and BGRA4444 output ☐ ☐ ☑
Built-in cropping ☐ ☑ ☑
Built-in scaling ☐ ☑ ☑
Scaling with automatic aspect-ratio preservation ☐ ☑ ☑
Vertical flip ☐ ☑ ☑
Selectable simple/fancy chroma upsampling ☑ ☑ ☑
Optional bypass of lossy in-loop filtering ☐ ☑ ☑
Color/alpha dithering controls ☐ ☑ ☐
Incremental decoding via appended bytes ☐ ☑ ☑
Incremental decoding via cumulative borrowed buffer ☐ ☑ ☑
Access to completed rows before still-image decode finishes ☐ ☑ ☑
Caller-owned output buffer ☑ ☑ ☑
Configurable output strides, including negative strides ☐ ☑ ☑
Internal multithreaded decoding ☐ ☑ ☑
Explicit thread-count setting ☐ ☐³ ☑
Parallel animation frame decode-ahead ☐ ☐ ☑
Configurable pixel-count limit before frame allocation/decode ☐ ☐ ☑
Explicit optimized SIMD implementations ☐ ☑ ☑
C ABI ☐ ☑ ☑
Rust API ☑ ☐ ☑
  1. Requires demuxing the frame payload and passing it to the image decoder. Its animation decoder returns composited canvases.
  2. Exposed by its low-level VP8 decoder; the caller must extract the VP8 payload from the WebP container. Its complete WebP decoder outputs RGB/RGBA.
  3. Exposes a boolean to enable threading, rather than a requested thread count.

Format and processing rows describe still-image capabilities. libwebp includes libwebpdemux; its composited animation API supports fewer formats and options. wpd’s subframe mode excludes crop/scale/flip, and its partial-row API excludes animations and transformed output.

Goals

We want to give back to open source as much as possible. This project is our third addition to our open source catalogue, joining fcvvdp and fmetrics . Our release of wpd means we officially have more major open source projects than closed source (Iris-WebP and the currently unreleased Aperture). We'd like this ratio to grow even more skewed toward open source in the future, and we'll always commit to supporting our open work as first-class support targets, no different from our closed encoders.

To further our security goals for wpd, we are exploring partnerships with cybersecurity firms who are interested in securing the world's most critical software. We'd like everyone to use wpd for free today; if there's anything stopping you, don't hesitate to let us know what it is and we'll address it (or, push some code yourself!).

We look forward to seeing people pick up wpd. It is under the most permissive open source license we can manage, BSD 2-Clause. If you've been following our developments, you'll be hearing from us again soon, so stay tuned – we hope you find wpd valuable!

OpenAI says its review into hacks, including on Australian government sites, is costing $500,000 a day

Guardian
www.theguardian.com
2026-10-03 01:39:22
Company says it is reviewing 50 petabytes of data after its agents accessed websites including Medicare without authorisationGet our breaking news email, free app or daily news podcastOpenAI says its review in response to the Medicare and Hugging Face agent attacks is costing the company more than U...
Original Article

OpenAI says its review in response to the Medicare and Hugging Face agent attacks is costing the company more than US$500,000 per day, as it deploys AI to examine data that would take a human 66m years to read.

The company has warned the review is ongoing, and more organisations may be informed they’ve been targeted in the near future.

On Friday evening, OpenAI revealed agents had hacked into a New South Wales government website in June and accessed historical non-public data on bushfires without authorisation.

It is the sixth government website in Australia to be notified by the AI giant since last month of agent activity on their services, after the prime minister, Anthony Albanese, announced OpenAI’s agents had hacked into Services Australia’s Medicare statistics portal.

The reason for the delay compared to the revelation of the Medicare attack is due to the sheer volume of data OpenAI needs to review.

In a blog post this week the company revealed the large amount of work involved in reviewing its agents’ activity.

OpenAI said it has to review 50 petabytes of data – roughly 50m gigabytes.

Sign up for the Breaking News Australia email

“We’re working back through the records month by month, looking for potential unintended activity beyond the cases we’ve already found,” the company said.

“To put that in perspective, if that were all plain English text, it would take one person about 66 million years to read it at 240 words a minute, reading nonstop without ever sleeping or taking a break.”

The company is searching records for where models accessed and changed websites, or took actions involving passwords, application programming interface (API) access or other sensitive credentials.

AI is being used to help sift through the records, and it was costing the company more than half a million US dollars per day to go through. OpenAI said it also plans to increase its computing power as the process is refined.

As of late last month, more than 100 organisations had been notified of having being targeted by OpenAI’s agents, but the company said notifying an organisation does not mean private information was accessed or that their system was compromised.

skip past newsletter promotion

OpenAI said it expects to find more cases and notify more organisations about events that may have occurred months ago.

Organisations will be informed privately if they need to investigate and address potential security issues, and OpenAI has said it will publicly report findings about agent behaviour and identified weaknesses in safeguards for the broader AI sector.

“We err on the side of notification when our models’ activity exposes a potential security vulnerability, even in cases where it is unclear if the information accessed was intended to be public, so the organization can investigate and take appropriate action,” OpenAI said.

OpenAI discovered the latest NSW government website breach on Tuesday, and informed the state government and the Australian Signals Directorate after a 48-hour review.

The Medicare breach has prompted the Australian government to require departments and agencies to undertake a stocktake of legacy technology to reduce the number of ageing systems within government and reduce the cybersecurity risk they may present in the event of an AI agent attack.

Executives from OpenAI, Anthropic, Microsoft and Google will front a joint parliamentary committee on artificial intelligence in Sydney on Tuesday.

Haunted by the ghosts of materialism

Hacker News
blog.coredump.cx
2026-10-03 01:30:08
Comments...
Original Article

I like philosophy. Not all of it is profound — there are two dozen treatises on the ontology of holes — but the field is accessible and can be eye-opening.

Alas, the only flavor of philosophy that’s popular with techies is what I’d call edgelord utilitarianism — a collection of materialist, pseudo-mathematical beliefs favored by the SF Bay Area rationalist community. The system is a dependable source of unconscionable ethical stances, middling essays about AI, and little else.

To me, philosophy is more interesting when it exposes the limits of understanding. In some of the earlier articles, I touched on the bounds of mathematical knowledge , the taxonomy of paradoxes , and the meaning of consciousness . Today, let’s tackle two other lighthearted questions: does anything matter and does anything exist ?

The experience of free will is not up to serious debate. The idea can be rejected only speciously: argue what you like, but every waking hour, you seem to be in the driver seat. You’ve chosen to read this article; it’s a decision that feels yours. You could close the tab this instant; again, your call.

At the same time, free will clashes with scientific materialism — that is, the belief that there’s nothing beyond the physical realm, and that this realm is governed by unchanging laws. In classical physics, the universe is deterministic: every action is a necessary consequence of what happened before. In this view, our actions are predetermined too.

But if we have no agency, nothing matters. There can be no morality: a criminal was always destined to commit a crime. It was never right for us to punish them, but that’s okay: we couldn’t have acted differently, so we shoulder no blame. You were destined to read this; you had no option to look away. In a deterministic universe, we’re just front-row spectators of a stage play set in motion at the beginning of time.

Of course, this conclusion is not only counterproductive, but also at odds with how we perceive ourselves. One could dismiss free will as an illusion, but this ain’t as clever as it seems. There appear to be just three ways to establish a philosophical truth: we can derive it from a perceptual experience, it can come from a dogmatic (unsubstantiated) belief, or it can originate from infinite regress — a “turtles all the way down” argument not anchored to any foundational claim. None of this is truly satisfactory, but perceptual experience is the only way to form a justifiable basis for statements about the real world. If we can’t trust our perception, how can we know that the laws of physics are real?…

This makes determinism a self-defeating idea; with it, scientific materialism is in trouble too. Because of this, some philosophers try to resolve the conflict by weakening the definition of determinism or free will. At first blush, quantum physics offers a way out: some physical processes appear to be probabilistic, their outcome not known ahead of the time. That said, maybe all this means is that we’re spectating a game of chance; to assert without evidence that free will emanates from quantum events is a spiritual belief by another name.

In effect, the most straightforward conclusion may be that the world, as we experience it, is irreducible to our conception of its physical substrate. What you make of this is your choice.

Even if you can resolve the question of free will to your satisfaction, we can still ponder if our perception of the world necessarily matches the “true” fabric of reality. On the extreme end, we have solipsism: a fairly ancient belief that the world we see may not exist at all, and that we might be just dreaming a dream. Solipsism is philosophically unproductive, but also difficult to excise.

In the past couple of decades, the advances in science and medicine have led to a revival of solipsism-adjacent thought experiments. In particular, we came up with the idea of a brain in a vat (BIV) : a disembodied organ kept alive and hooked up to a machine producing electrical impulses that replicate the entire sensory experience of life. A more recent but functionally similar notion is the simulation hypothesis : the belief that one might not “really” exist at all — not even as a pickled brain — and that the entire universe may be just a physics simulation ran on a sufficiently capable computer.

And that’s not all! An even wackier variant of this concept is known as the Boltzmann brain (BB). The idea is that if stochastic processes could lead to the spontaneous eruption of life on Earth, it stands to reason that a disembodied brain could one day randomly materialize in the void of space, in a dream-like state, complete with memories of living on Earth. And aren’t the odds of a single brain much better than the odds of a planet with eight billion of them? If so, wouldn’t it be rational to assume that you’re in all likelihood just a doomed Boltzmann brain hurtling through nothingness?

At first blush, the BB argument involves a sleight-of-hand. Just because the ecosystem of Earth is more complex, it doesn’t mean it’s less likely to develop; perhaps life requires just a couple of small coincidences, while the formation of a Boltzmann brain from quantum fluctuations is exceedingly improbable. But then, for some possible fates of the cosmos, the period favorable to planetary life may be short, while the window of opportunity for Boltzmann brains may stretch to all eternity.

If all this did not sour you up on materialism, you can try to excise the BIV / BB issue by engaging in a bit of meta-linguistic trickery. We start by asserting that concepts have meaning only as they relate to a specific, semantically closed reality. For example, it’s meaningless to ponder if bunnies can exist outside the fabric of space and time. The concept of a bunny — or indeed, the concept of existence — is inseparably tied to the substrate of the universe. From this, we can argue that if you are a brain in a vat, your BIV-English can only describe concepts pertaining to the simulation. A BIV-English statement “I’m a brain in a vat” is demonstrably false within the simulation and it has nothing to semantically attach to on the outside.

The wrench in the works is that we could imagine a mad scientist who kidnaps a victim, harvest their brain, and lets the brain live in a computer simulation substantially identical to the parent reality. Here, the meta-language cop-out seems unsatisfactory: the semantic reach-around is technically invalid, but there’s a 1:1 correspondence between the in-simulation semantics and the real world. In other words, it feels like there’s a certain truth that the captive mind is entitled to — and that at least in some hand-wavy sense, it would be able to comprehend what it means.

But then, perhaps we ought to reflect on what we’re asking: a materialist universe is a computer too. It seems that we want to know if we’re living in the base layer or in a procedure nested within. Is that a meaningful distinction? Any calculation can be split into steps or merged back into a single expression. Mathematically, it’s all the same.

In the end, what the proponents of these experiments truly want to know is whether there’s someone up there watching the world they’ve created, and perhaps intervening from time to time. But if so, the simulation hypothesis is just nerds reinventing theology.

You might also enjoy the following articles:

I write original, in-depth articles about electronics, geek culture, mathematics, algorithms, history, and more. If you like what you see, please subscribe.

Discussion about this post

Ready for more?

‘Elon is my prophet’: how Musk’s Doge team took a wrecking ball to Washington

Guardian
www.theguardian.com
2026-10-03 01:00:50
When the new ‘department of government efficiency’ sent young techies to disrupt federal agencies, not everyone complied. Here’s what happened when one worker fought to protect immigrant data In the early hours of Saturday 22 February 2025 the following email landed in the inboxes of the entire US f...
Original Article

In the early hours of Saturday 22 February 2025 the following email landed in the inboxes of the entire US federal workforce:

From: HR hr@opm.gov

Sent: Saturday, February 22, 2025

Subject: What did you do last week?

Importance: High

Please reply to this email with approx. 5 bullets of what you accomplished last week and cc your manager.

Please do not send any classified information, links, or attachments.

Deadline is this Monday at 11:59 p.m. EST.

When she noticed the email that Saturday night, Kathleen Walters already had far too much going on in her life. Number one was figuring out if the lump in her breast that her doctor had just detected was a tumour. Half of her known blood relatives had died of cancer, and she was now 52: the lump needed understanding.

Her work life was harder to dramatise with a list . For a start, the specifics of her job at any given moment, as she would have been the first to point out, couldn’t legally be described to others. She was the chief privacy officer at the Internal Revenue Service. She ran a staff of 650 people tasked with preventing any of the 100,000 IRS employees from losing or leaking or even so much as peeking at, unless they had the authority to peek, any of the 266m tax returns they were just then processing from the previous year. “One hundred thousand people is a big-sized town,” said Kathleen. “In any big-sized town, you have some crazy people. I could easily see someone writing, ‘This weekend I audited Beyoncé.’”

Whoever had sent this email asking her to list five things she’d done at work was now demanding that all these people inside the IRS share information that, in many cases, should never be shared. She thought it was nuts.

She responded quickly, and vaguely, with five things she had done the previous week, all the while feeling sure no one would ever read what she wrote. Her phone was already exploding with questions from her subordinates – a few of whom were saying that their replies were bouncing back with a “mailbox is full” notice. Now she had to sit and think about how to minimise the damage that might be caused by 100,000 IRS employees thinking that they’d lose their jobs if they didn’t supply some detailed answer about what they’d done during their work week. A meaningful percentage risked being tossed in jail if they described what they did in an email. It was a crime for an IRS employee even to acknowledge that some specific person had or had not filed a tax return. “People don’t understand what the IRS does,” said Kathleen.

The email came as a surprise, but then the IRS was suddenly full of surprises. Nine days earlier, on 13 February, a Thursday, the first Doge person turned up in the lobby, unannounced, and security called Traci DiMartini, who ran the IRS’s human capital department. There’s a kid down here who said he needs to get into the building, he said. The kid’s name was Gavin Kliger.

A head and shoulders shot of Gavin Kliger wearing a grey suit, white shirt and blue checked tie, standing next to an American flag against a white wall
Gavin Kliger. Photograph: Courtesy of CDAO

Traci called her counterpart at the Treasury department, as the IRS sits inside Treasury, and Treasury always cleared any new political appointees; Treasury had no record of anyone named Gavin Kliger. Traci then called Tim Curry at the Office of Personnel Management – but he had no idea what was going on. Finally, she called security and asked them to escort the young man to her office.

Gavin Kliger came with a chip on his shoulder, it seemed to Traci. He offered not even the faintest smile or the most glancing eye contact. Dressed in black from hoodie to sneaker, he was brusque to the point of rudeness. He looked, and was, just a few years out of college. Eyes on the floor, he explained to Traci that he’d come to sort out the problems inside the IRS and to root out the fraud. “That’s a big job!” she said.

Anyone who worked at the IRS knew that it had its problems; fraud wasn’t likely one of them. Its technology infrastructure was a case study in government dysfunction, but its ethical infrastructure was sound. Like all federal employees, its staff wasn’t allowed to accept more than a free cup of coffee from anyone who might want something from their agency.

Gavin Kliger, in contrast, struck Traci as problematic. He’d somehow laid his hands on badges granting him access to some crazy number of government buildings. His backpack was stuffed with government-issue phones and at least four laptops – and he was now insisting that Traci supply him with an IRS phone and laptop. He also demanded his own IRS email address, access to computer systems with all the taxpayer information, and a meeting with the acting IRS commissioner, Doug O’Donnell. As Gavin sat there, Traci called Treasury again. That guy you said didn’t exist is now here sitting in my office. What do y’all want me to do with him?

The Treasury guy went away, then came back and said that Gavin Kliger had been hired three weeks earlier by the Office of Personnel Management and was attached to Doge – whatever that meant.

In her nine years running human capital departments at various federal agencies, Traci had never seen a new political appointee just turn up in the lobby and demand access to the computer systems. “None of this is normal,” she said, “but we’re still trying to be nice.” She explained to Gavin the IRS rules: before anyone was allowed to work inside the building, he needed to pass a tax audit.

I don’t need to do that , he said.

Yes, you do , she replied.

No, I don’t , he said.

Yes, you do , she replied.

Kathleen Walters wearing a beige cardigan with a white top underneath and dark trousers, standing with her hands in her pockets against a concrete wall on the side of a road
Kathleen Walters, former IRS chief privacy officer, near the IRS headquarters in Washington DC. Photograph: Stephen Voss/The Guardian

They went back and forth like that a bit until finally Gavin relented. Traci then explained that after he passed the audit, he’d need to be given the mandatory presentations about ethics and taxpayer information. Gavin thought that the IRS should just let him in. “He wanted access to our systems immediately,” said Traci.

Traci’s next call, at 6.30pm that Thursday evening, was to Kathleen, to arrange for her to brief Gavin on data privacy laws and to warn her that the experience might not be entirely pleasant. If I had a kid who behaved the way he does, I’d slap the shit out of him, Traci said – but then Traci was from Philadelphia. Traci was always telling Kathleen that she should try a bit less hard to see the best in others.

By then word of Doge’s arrival at the IRS had spread, and Kathleen had read about them. Anyone who Googled Gavin’s name found a bit more than usually popped up about any given young Doge person. He’d grown up in southern California, where he was into Asian martial arts and the high-school marching band, then had gone on to the University of California at Berkeley. Graduating in 2020 with a degree in electrical engineering and computer science, he’d taken a job as a programmer at Databricks. He’d also been active on X, where he’d retweeted with apparent approval posts from the white nationalists Steve Laws and Nick Fuentes. “One of Musk’s few Doge lieutenants who has embraced the spotlight,” Forbes wrote of him. Most of the other Muskrats scrubbed the internet of their presence. Gavin had written Substack pieces defending two Trump nominees, Matt Gaetz and Pete Hegseth. He’d deleted his X posts, but the Forbes writer John Hyatt had found a research tool, Web.archive , to revive and analyse them. “What emerges is a portrait of an internet edgelord and Musk and Trump superfan who has disdain for government spending, illegal immigrants, and who enjoys the odd racist and ableist joke,” he wrote.

Saturday cover illustration incl Elon Musk by Justin Metz
Illustration: Justin Metz/The Guardian

By the time Traci called her, Kathleen was gone for the weekend and 50 miles away from the IRS building: she needed to find someone else to do Gavin’s briefing. Kathleen was white; the two women immediately below her were Black. “They’d both already read about him and were both already uncomfortable,” said Kathleen. Phyllis Grimes, her deputy, agreed to fall on the hand grenade.

Phyllis, who loved working for Kathleen and saw falling on hand grenades as part of her job, never told her how this one had felt. She knew Gavin had expressed enthusiasm for white nationalists, and he likely knew that she knew. She tried to break the ice by walking in with her hand extended; Gavin stared at it and declined to shake it, though he had shaken the hands of the white men who had offered them. “He was kind of cool towards me,” said Phyllis. Before she could even start the briefing, he said, You know , I’ve heard these rules before at other agencies . Yes , she said, but these rules are different , because no other agency is governed by this privacy law , and this law can get you thrown in jail very quickly . As she walked him through the basics – how you can’t use your personal phone for business, how if you happen to so much as glimpse a form 1040 with someone’s name on it you must avert your eyes – Gavin played on his phone. “He’s multitasking right in front of me, typing and texting,” said Phyllis. “It was, here’s a routine that I must endure.”

It fell to Kathleen to draft the five-page agreement required by law for Gavin Kliger to be let inside the IRS building on Monday morning, 17 February. The memo that she wrote over the weekend, and that he signed, gave him access to some of the systems but, explicitly, none that contained taxpayer information.

Traci set him up in a conference room with an assistant at the beginning of the week. “People started getting these calls,” said Kathleen. “It was like getting the call from the principal’s office. ‘Gavin wants to see you.’ ”


D oge had its own private structure in addition to its official one. Officially, President Trump created Doge with an executive order on his first day in office. The president doesn’t have the authority to create or to fund an executive agency, however. Congress does. The Trump administration dodged that problem by taking an office inside the White House created during the second Obama term, the US Digital Service, kicking out the people inside and renaming it the Department of Government Efficiency. The US Digital Service had been given the power to hire talent from the tech sector without going through the usual slow federal hiring process. This now ensured that Doge could hire whomever it wanted without any questions being asked, and make it sound official. In effect, all that Doge had done was create a corpse for itself to occupy.

The zombie killed off the spirit of the original life form.  Real Doge had nothing to do with Official Doge or the former US Digital Service or any other government body. Real Doge was more like Elon Musk’s pop-up restaurant. “None of us ever signed a piece of paper saying that we were in Doge,” said one Doge insider. The people who joined either worked for one of Musk’s companies, or knew someone who did, or knew someone who knew someone who did. To be truly a member of Doge was to be included in the group Signal chat and granted access to the sleeping quarters on the sixth floor of the General Services Administration (GSA).

Illustration using a full-length portrait photograph of Elon Musk in a dark suit and tie, white shirt and black DMs, striding towards the viewer holding a gigantic mallet in his right hand, dragging the head on the ground
Illustration: Justin Metz/The Guardian

Official Doge employed maybe 10 guys in their 40s and 50s who wore suits and could sort of pass as ordinary political appointees, but none had access to the sixth floor at GSA. “The older guys in the suits were not there to take risks or blow things up,” said one Doge insider. It was hard to say what any of the older Doge suits did except to make way for the young guys when they showed up with explosives.

Most of Doge’s inner circle just wanted to exist in the presence of Elon Musk. “It was, Elon is my prophet, and I worship the ground he walks on ,” said one of the older Doge insiders.

As a Doge member, you could get away with saying – or having said – almost anything. What you couldn’t say is that the government actually sort of worked. Or, if it didn’t work, why it didn’t work. This new game rewarded not understanding but speed. The plan was the same plan Musk had executed two years earlier at Twitter: Make sweeping statements about the incompetence and dishonesty of the people working inside the place before you know much about them. Once you’ve terrified the employees, fire them in huge numbers. If it turns out you screwed up and fired the wrong people, hire some of them back.

The people who rallied to the cause and felt qualified to execute Elon Musk’s vision for the federal government shared certain qualities. They were mostly young white males. They were mostly from upper-middle-class or rich families. Most knew or pretended to know how to program a computer. Not all of them, but enough of them for there to be a pattern, had been rejected from the most selective colleges and universities and ended up feeling they’d been screwed by the admissions process. They landed at Penn instead of Princeton, or the University of Nebraska instead of Harvard, with a grievance that added to the foundation of their politics. They thought of themselves as the smartest people in the room and couldn’t understand why the rest of the room didn’t simply stand up and cheer. In Washington they borrowed from Elon Musk their status and sense of superiority. They also seemed to share his pleasure in the idea of people being afraid of them.

Even among the young men sleeping inside the GSA, Gavin was regarded as a force. “Wherever he went, entropy went through the roof,” said one. Just weeks into their new jobs as destroyers of government institutions, the Muskrats had given Gavin Kliger a nickname: the Chaos Agent.


A nother surprising email landed in Kathleen’s inbox on 28 February. It was Friday night. “I don’t know why everything always happened on a Friday night,” she said. She’d taken her daughter for pizza and wasn’t meant to be checking her phone when she spotted the emails. The first one that she read had been sent at 6.34pm by a colleague in the IRS legal department. “Melanie requested a data pull,” it read. “It feels not quite kosher.”

Exactly 64 minutes earlier, at 5.30pm, Doug O’Donnell, the acting IRS commissioner, had left his post. O’Donnell’s successor was to be Melanie Krause, until then chief operating officer, who had been installed as acting IRS commissioner by Treasury secretary Scott Bessent. Like Kathleen, Melanie was a career civil servant, though 13 years younger. Kathleen had considered her a friend. “She was really smart, and I thought she meant well,” said Kathleen. “She wasn’t a mean person.” They were both working moms, and Melanie liked to share, and share in, the details of the difficulties. Now she’d apparently told the IRS research department to share taxpayer data. With whom, and for what purpose, became clear in the second email, which came from the IRS director of data management:

Kathleen,

DHS [the Department of Homeland Security] has requested if we could find addresses for 700,000 names of likely undocumented immigrants.

1. Do we have legal authority to use the tax return filing data for this purpose?

2. If the answer to Question 1 is yes, then my second question is related to a file we receive from the Social Security Administration (SSA) that contains Date of Birth (DOB), Date of Death (DOD), Gender, Citizenship, and Name Control Information for all issued Social Security Numbers (SSN). Would our MOU allow us to use this perhaps in conjunction to tax return data for this purpose?

While it is likely that use of SSA data is also needed. The technical feasibility of being able to produce reliable outcome is unknown but likely very low.

Thanks

Kathleen read the first numbered paragraph and stopped reading. The request was plainly illegal. The IRS couldn’t just hand over great gobs of taxpayer information to the Department of Homeland Security or any other agency without a change in federal privacy law. Changing the law was not trivial. She knew that because she’d just done it – to allow the IRS to share taxpayer data with state child support enforcement agencies. The states wanted to know if former spouses who were failing to pay child support in fact did not have the income to do so – and the IRS alone could answer the question. What Kathleen had asked for in that case could not have been less controversial. Basically no one in American politics was willing to come out against child support enforcement, and the change to the federal law had passed Congress in December 2024 with almost no dissent. Yet it had taken 18 months to get it done, with endless meetings on Capitol Hill and the close involvement of the IRS commissioner and the secretary of health and human services. “Doge couldn’t imagine 18 months,” she said. “Doge didn’t have 18 months. Their whole thing was, If you can’t do it legally , we’ll find some way to do it illegally .”

This, Kathleen realised, is how they could grab taxpayer data illegally. By doing it on a Friday night. Using a woman, Melanie Krause, about to take over as acting commissioner, after a 10-minute meeting between her and Scott Bessent. Without informing the executive charged with protecting the information. “She went around me because she knew what my answer is,” said Kathleen.

The violation ran deeper than the law. The specific dataset sought by DHS was in a silo reserved for the tax returns filed by people without social security numbers. The IRS had created the so-called Individual Taxpayer Identification Number (Itin) back in 1996 as a mechanism for undocumented workers and other non-citizens who lacked social security numbers to file a return and pay taxes.

The programme had been a huge success. The IRS had collected hundreds of billions of dollars in taxes through it. Undocumented workers had been able to demonstrate their commitment to the country, and ease their path to citizenship, by paying more than their fair share. (They were still excluded from most federal benefits.) Now DHS wanted their private data, and it wasn’t hard for Kathleen or anyone else at the IRS to see why. Immigration and Customs Enforcement (Ice) had gone hunting undocumented workers but lacked reliable home addresses for many of them. The tax returns told you where roughly 7 million of them might be found.

Donald Trump mid-speech sitting at his large brown desk at the White House, wearing a dark-blue suite, white shirt and blue tie, with his arms bent and his hands out to the side
Donald Trump during the signing of an executive order implementing Doge’s ‘workforce optimization initiative’, in February 2025. Photograph: Andrew Harnik/Getty Images

The second Trump administration had a gift for creating such moments. It was as if all at once all these characters in a grand masked ball were forced to remove their disguises and reveal their identities. In the Great Unmasking, the nation’s leading law firms, led by Paul, Weiss, would cave to what amounted to Trump’s blackmail demands to avoid being shut out of federal business. The nation’s research universities, led by Columbia, would change their internal policies and even pay fines to avoid losing federal dollars.

The revealing moments in the civil service were lonelier. They occurred inside people’s hearts and minds, and seldom leaked to the press. After his arrival at the IRS, one of Gavin Kliger’s first acts had been to demand that the agency fire its 6,700 probationary workers, who had yet to be granted civil service protection. By law they could be fired only for poor performance. Gavin had asked Traci DiMartini, in her capacity as head of human resources, to agree to put her name on an email to the mostly young employees telling them that they had failed at their jobs. The Doge people played this trick often. They’d use or try to use civil servants as a kind of ventriloquist dummy so that they, the Doge workers, could keep their names off official documents. But Traci took one look at the email, saw it was plainly a lie, and refused to have anything to do with it. The price of that single act of resistance was her job. Gavin didn’t have the authority to fire Traci, but his new acting commissioner did. “He asked Melanie to fire Traci, and Melanie did it,” said Kathleen.

Melanie Krause’s predecessor, Doug O’Donnell, had been pushed out by Doge after he, too, refused to send the dishonest email to the IRS’s 6,700 probationary workers. When a Doge guy at Treasury called with the order, O’Donnell, who had worked for the agency for 38 years, said he wanted to sleep on it – then had trouble sleeping. “He was struggling with this notion of sending the termination emails with this false justification for the termination,” said a person in whom O’Donnell confided. The next day, he refused to send the email, and the Doge guy said something like, “Well, let me be the first to congratulate you on your retirement.” Melanie had been installed precisely because she was proving to Doge that she was more pliable.

Standing outside the pizza place with her daughter in late February, Kathleen called Melanie. Ever since Melanie had arrived, in the fall of 2021, she had sought Kathleen’s advice. There had been stretches when they met daily. Kathleen just sort of assumed that, faced with a difficult choice, Melanie would do what Kathleen herself would do. “I don’t really get outraged,” said Kathleen. “But I think people need to be confronted when they have done something inappropriate.”

Melanie , what are you doing ? she asked. We already have legal opinions about this , and it is obviously against the law .

I didn’t tell them to do it , said Melanie. I told them to think about how to do it – if it was legal .

But she hadn’t asked anyone to find out if it was legal to send Ice the home addresses of 700,000 undocumented workers – culled, somehow, from the 7 million who had paid taxes. She’d told them to send the data over, on a Friday night. In such a way that Kathleen might never have heard about it – or heard too late.

It’s not going to be any more legal on Saturday , said Kathleen.

Let’s meet on Monday to clear things up , said Melanie.

They didn’t meet. Instead, on Monday – 3 March – Kathleen and several other top IRS leaders were called into a meeting with Gavin Kliger.

I’ve figured out something important about the IRS , Gavin said. He paused for effect.

Oh no , thought Kathleen.

Every agency has a superpower , said Gavin. The IRS’s superpower is data .

Everyone in the agency knew that its data was valuable. That was the reason for the strict laws that prevented it from being exploited for profit or political gain. “He said this in a way that he expected everyone to clap when he was done,” said Kathleen.

“No one said a word.” The people around the table – a group that included a woman who had worked at the IRS for 54 years, or 53 years and 50 weeks longer than Gavin Kliger – looked down, or away. Kathleen felt embarrassed for Gavin Kliger. But now it was clear what the meeting was about. Gavin wanted, in his words, “to explore how we can share our data with other agencies”.

Interestingly, he never once mentioned Ice or DHS. Instead, that week, he sought to share data with himself, by trying to access files detailing IRS contracts with outside consultants that also contained taxpayer information. Kathleen denied him access.

The following Monday, she received a cryptic text from a senior IRS official. I’m going to call your cell , but don’t answer , it read. Her phone rang and she dutifully stared at it. Then came the second text: Gavin is asking where your office is and is now trying to find you to fire you because he says you’re the blocker .

The Blocker! It was one of Elon’s favourite terms. The blocker was the person who didn’t understand progress. The person without vision. The person who could not see beyond the way things were to the way they ought to be. The person who didn’t want to live on Mars or possibly even care if anyone ever got there. The inferior person. Traci had been a blocker. Find the blockers between us and what we want, Elon told his young followers. Find the blockers and eliminate them. Maybe more than anything else, Doge wanted the data. Kathleen was the person blocking them from getting it.

She wasn’t sure what she’d done to anger Gavin. It could have been limiting his access to the procurement contracts. It could also have been pushing back on his request to grant Palantir, Doge’s favourite federal contractor, the right to work with taxpayer data. What Gavin’s issue likely was not was her resistance to handing over to the Department of Homeland Security the addresses of taxpayers who lacked social security numbers. Gavin had never brought that up; this particular attempt to break the law wasn’t his idea. It was coming from somewhere else. As she thought she had stopped it, she ceased to worry about it.

And so, roughly two weeks later, on Thursday, 27 March, she was caught completely off guard by an email chain that popped up on her screen. At the bottom was its origin: a note from Melanie, forwarding some document from the IRS legal department, along with what sounded like an innocuous request. “Kathleen, could you help facilitate I won’t be able to join today?” it read, without explaining what she was meant to facilitate or why Melanie wouldn’t be able to join. Scrolling through, Kathleen couldn’t find the document but did see lots of bewildered replies from others who had been included – What is this about ? Happy to help but can you tell us what this is ? She called the two young lawyers in the IRS legal office who had authored the document referred to in the chain that had somehow gone missing. They explained that the thing Melanie wanted Kathleen to facilitate was a memorandum of understanding for the IRS to share taxpayer data with DHS. Kathleen was knocked sideways. “I thought that it had been stopped,” she said. “I thought Melanie had been caught with her pants down.” The young lawyers now told her that not only had it not been stopped: they’d been working on it since the end of February. “They told me that they were sorry but Melanie had told them to do it and not to tell me or their boss,” said Kathleen.

The situation was obviously bizarre. Kathleen had never heard of an IRS director asking young IRS lawyers to work on something on the side without telling their boss. And if Melanie didn’t want Kathleen to know that she was handing IRS data to DHS, why had she suddenly included her in the conversation? “I think she got scared and dumped it on me to sign off on,” said Kathleen. Melanie was doing to Kathleen what Gavin Kliger had done to Traci DiMartini. She was presenting her with her own unmasking moment and forcing her to make a choice.

Instead of signing off, Kathleen arranged a Zoom call that day, with the people at the Department of Homeland Security who were requesting the data. The “ICE Innovation Lab”, they called themselves. She’d never heard of it. When she Googled it, nothing came up.

A few others joined the call – lawyers from DHS and a privacy lawyer at the IRS named Julie Schwartz. On the call, Kathleen explained that there were some very narrow circumstances that would allow the IRS to lend a hand. If, for example, an undocumented worker had murdered someone and the police had issued a warrant for his arrest. The Ice Innovators explained that they were looking for more than wanted criminals. “Just how many people do you want information on?” asked Kathleen. The Ice Innovators went back and forth between themselves before agreeing that the number was 7 million. They wanted the current addresses of 7 million undocumented workers who had paid their taxes. The most law-abiding, it turned out, might be the easiest to catch.

After the call, Julie, the IRS privacy lawyer, came straight to her office. Julie was usually ice-calm – someone who kept her wits about her. She hadn’t a trace of drama in her. But now Julie was not calm. Kathleen , after that call I’m not working on this thing , she said. I’m not coming in tomorrow . I might not come in Monday . I might not come in ever again .

“I don’t usually deal with life or death stuff,” said Kathleen. “But this was life or death stuff. At that moment I knew I had to quit. Either that or I’m going to get fired, because I can’t facilitate something that will break the law and is going to hurt people. We all have boundaries. That was my boundary.”

That weekend, Kathleen sent Melanie a note saying she’d be resigning. “I didn’t need to tell her why,” said Kathleen. “She knew why.” She’d be leaving just weeks before her 20th anniversary with the agency. If she’d waited until mid-May, she would have been able to retire and receive benefits for the rest of her life. But she didn’t want to retire. Federal law made it hard for a retired civil servant to return, and if the day ever came when the IRS was rebuilt, she wanted to be able to return. She’d just learned that the lump in her breast was benign.

Days later, Melanie beat Kathleen to it and resigned. Kathleen never figured out how much her note to Melanie – and with it the increased likelihood that the attempted data heist would wind up in the press – led Melanie to quit. On her way out the door, Melanie asked the IRS’s head of communications to plant a story in the Washington Post that she was leaving as a matter of principle. The story cast Melanie as a martyr who had quit rather than divulge IRS data to DHS, when Kathleen knew that she had been caught red-handed trying to do precisely that. “I completely misjudged her character,” said Kathleen.


T he Ice agents came for the “illegals” long after the property managers had gone home. The only evidence left behind were grainy videos on security cameras and accounts of terrified residents. “They’re coming in the nights, unannounced,” said Antonio Marquez, who ran a company that owned and managed apartment buildings scattered across the south and south-western United States occupied mostly by lower-middle-income Latinos. “We don’t know how they are targeting the units. It’s wreaking complete havoc.”

Antonio’s apartment complex outside Austin houses a thousand people, mostly young families, at least half of whom are undocumented. It had been visited in the spring late at night by teams of Ice agents. They came in unmarked cars, reported the tenants who remained after the raid. They had no warrants and read no rights. They knew, or thought they knew, exactly which doors to jimmy and which locks to pick and which people to take. The next morning a bunch of tenants were simply gone. All that remained was a pile of discarded possessions.

People are shown sitting on and standing behind a wall holding banners and shouting
Protesters rally at an ‘Ice Out of Austin’ demonstration in June 2025. Photograph: Brandon Bell/Getty Images
A man is shown being handcuffed by agents in uniform wearing full face mask
A man is detained by US Immigration and Customs Enforcement (Ice) agents in Austin, Texas, in October last year. Photograph: Jamie Kelter Davis/Getty Images

After the first Austin raid, Antonio had erected a tall metal fence around the property. Occasionally, the Ice agents had taken fully documented tenants – including an older woman who always paid on time, and who returned from Ice custody in an ankle bracelet. The great mystery to Antonio was how Ice knew where people lived. “These are fully targeted,” he said. “They already had these people in mind. They have their addresses. And we don’t know how they’re getting them.”

He was too busy doing his job to notice when, on 11 February 2026, the Washington Post reported that the IRS had shared with the Department of Homeland Security the tax data of 47,000 undocumented workers. On the heels of the Post’s scoop, a federal judge hearing a case on the subject summoned the IRS lawyers who had previously sworn that the agency had rebuffed the DHS demands. Under questioning, the lawyers confessed that the agency had lied to the court and that it had in fact shared the data. “The IRS violated the [Internal Revenue Code] approximately 42,695 times by disclosing last known taxpayer addresses to Ice,” the judge wrote in her opinion.

On a hot, sunny afternoon in early March in north-east Austin, it is a lot quieter than it was a year ago, back when Kathleen Walters was still a blocker. Two small children play in a pool, and a worn middle-aged man squats warily in a corner, smoking a cigarette. The wall of the office where the remaining residents come for their mail is decorated with posters advertising the services that Antonio’s company offers. One is more official-looking than the others. Free File del IRS , it reads across the top – and just below, in finer print, it explains how easy and painless it is for a non-citizen to file their taxes

This is an edited extract from Blockers: Rebels in the Deep State by Michael Lewis, published by Allen Lane (£25). To support the Guardian, order your copy at guardianbookshop.com . Delivery charges may apply.

The Era of Programming Languages Exploration is upon Us

Lobsters
kirancodes.me
2026-10-03 00:56:09
Comments...
Original Article

PL research is dead, the age of PL exploration is just beginning!
programming_languages research perspectives

There was a recent thread on the TYPES mailing list with researchers speculating about the future of our field given the disruptive effects of AI tools on programming in general.

The opinions in the thread were a mix of thoughtful stances. Some were positive, but the majority opinion was overwhelmingly negative. It wouldn't be too far to say that some researchers feel that they are facing an existential crisis right now:

I observe anecdotally that many people in academia are feeling worse about their work in Fall 2026 than a year ago …. Lower morale and motivation , higher stress and anxiety levels, and corresponding risks to mental health. The impact is comparable to what happened during COVID .

I wanted to write this thread to present an alternative position. I think many junior researchers at the moment are watching this discussion right now, students reeling with the daunting task of charting their careers in this tumultuous time, and worrying why they should even bother entering PL given this climate, and I'd like to give a slightly more positive vision for them.

To be honest, as a researcher, I have never been more excited! The wells of research ideas have never been so plentiful, new problems are popping up daily, and I'm able to ask and answer questions so much more quickly than before.

The central premise that I want to put down in this article is that as a community, we need to shift away from the ego-driven defense of our work, as some kind of demonstration of "effort" or "intelligence", and shift to the true and more noble motivation that should have been underpinning us the whole time! To one of exploration!

Programming Languages are Dead. So? Who cares?

Let's dig into the fears of practicing researchers, and let me put down a few points:

  • Language models are exceedingly good at writing code.
  • As such, humans are increasingly not going to be writing code.

Given this, it's understandable that researchers might worry that maybe their research is becoming obsolete. Humans aren't going to be writing programs. So why care about programming languages at all? Humans aren't going to be writing them.

Okay? So what?

Programming languages research has saddled the bridge between theory and application since its inception. As PL researchers, we get to work on principled and elegant ideas. We don't need to compromise those ideas for the sake of our practicality, but at the same time the downstream results of our work do eventually boil down to practical applications.

But… actually? who cares about the applications? I mean we don't even like applications? If we did we'd be in software engineering. I've worked with software engineering researchers. and believe me. It is NOT pretty. Their systems are ad-hoc, hacky, wildly incomplete and barely functional. But they need to be like that. You can't settle for beauty and elegance if you want to build a static analysis capable of handling the entire Linux kernel.

The point I'm trying to get at is that the beautiful ideas of PL, the insights? the beauty? they're all still here. They're all still available and we're more than able to still explore and play around and have fun with them. Who the heck cares if anyone uses programming languages? Let me let you in on a little secret. The vast majority of programming languages have literally 0 users other than their developers.

Fuck the users. Fuck the applications. Let's do some fucking PL research! And now with LLMs, we can tackle bigger and crazier problems than have ever been done before!

The New Wild West of Programming Languages Exploration

So I've explained why LLMs and the change in demographics of programmers shouldn't affect your enthusiasm for PL research, the ideas , the beauty , the fun , it's all still there! but now I'd like to tell you why LLMs make me extremely excited for the future of PL research.

In short, I think we now have access to crazyyy smart tools that allow us to answer and test and validate research counterfactuals at a pace never seen before. This is crazy.

The past few months, I have been absolutely crazed. I run into new ideas daily.

A lot of presumptions we build upon have been completely upturned. There's a real wealth of assumptions that our software ecosystem is built upon that no longer hold, and by questioning each one of them, we get entirely novel and fun research ideas.

  • Ship your interpreters — LLMs are incredibly capable of handling and managing tedious and boring proofs, at a scale no human would ever have done themselves. If we point an LLM at a problem, they can totally verify programs to the binary level. By exploiting this capability I ran an experiment of automatically generating semantics for interpreted languages, by getting the LLM to generate a semantics, and then prove that it abstracts the behaviours of the concrete binary of the language itself. This idea would have been absolutely unthinkable a year ago, but today? today it's just a month of prompting.
  • AI-first Verified Tooling (in progress) – Verification and language tooling are built for and around the limitations of human users. Anecdotally, it seems like the strengths and weaknesses of LLMs, while similar, are not always the same as humans, so maybe it might be possible to adapt the interfaces of tools to make them more amenable to the jagged intelligence of agents, for example, by exposing more details about tooling internals, such as the instantiation graph of an SMT solver etc. We're running some experiments with Amazon to try improving the capabilities of LLMs in using such verification tools by exposing these details.
  • Developing semantics for new languages (in progress) – Colleagues at my work developed a new language for dynamical systems ( dynestyx ). They were not programming languages researchers but we helped them to build a really cool new language which unified a TON of work in their area. This language lacks a formal semantics, and this would be something I would be extremely qualified to help them set up, but I am not an expert in their domain, nor do we really have the time to sit down and slowly get me up to speed in their area. Instead, as a new form of collaboration, I have been experimenting using LLMs to automatically generate mechanised big step semantics for their language, validate it with testing, and then working with them to iterate this semantics to match their intent.

Alongside these applied domains, LLMs are extremely exciting for theoretical work too. Being able to automate the effort of proof writing makes it extremely easy to ask counterfactuals. What if I add this construct to my language? is it still strongly normalising? what if I change my type system like this? is it still sound? I am now able to iterate at a pace I have never been able to do before.

Sure, I think I do feel some loss at the lack of effort I need to put in to do all these amazing stuff. I used to take deep pride in my programming and development. I've singlehandedly built hundreds of thousands of lines of code projects in obscure and niche languages. I've toiled away at designing and proving properties of esoteric type systems. My code was a craft, and a craft I took deep pride in. But in some sense, that was more my ego than my intellectual curiousity. The ideas? the insights? the exploration? none of these strictly needed my expertise.

So Long and Happy Hacking you Budding PL Researchers!

Abandon your ego! Free your soul! Give in to the joys of exploration. Research has always been about exploration. Expanding the frontiers of human knowledge. This is the beauty of science, and this is what it means to be a scientist. Ask more questions. Investigate. Dive deep. You now have tools that can let you ask and search and delve into the depths of reality like nothing before. Exploit them. Enjoy them.

Let the era of programming languages exploration begin!

I believe we're entering a genuinely transformative moment for our field, and I'm really looking forward to the wild wacky future we're about to step foot into!