Customization: Optimizing Compiler Technology for SELF, a Dynamically-Typed Object-Oriented Programming Language (1989)

Lobsters
dl.acm.org
2026-10-03 16:59:46
Dynamically-typed object-oriented languages please programmers, but their lack of static type information penalizes performance. Our new implementation techniques extract static type information from declaration-free programs. Our system compiles several copies of a given procedure, each customized...

System-level ad-blocking in Android

Lobsters
kevinboone.me
2026-10-03 16:43:41
Comments...
Original Article

The Internet has become an advertising platform. Users of mobile devices are acutely aware of this, because mobile operating environments and apps force ads on their users, in addition to those that pollute the Web. It’s getting harder to avoid the ads on Android handsets, as Android is increasingly locked down and difficult to customize. It’s certainly not in Google’s interest to make it easy to block ads.

It’s still possible, though, to suppress most advertising, even in 2026.

This article is specifically about blocking ads at the system level, that is, in the Android platform itself, rather than in any particular app. Some Android apps, particularly web browsers, have their own methods of blocking ads, or can have such features added using third-party plug-ins. Oddly, though, the DuckDuckGo browser, whilst blocking the tracking that accompanies targeted advertising, doesn’t specifically block ads – you need something more.

If you block ads at the system level, all apps – and the system itself – are protected. Moreover, methods for blocking ads can usually be extended to protect against trackers and other kinds of malware.

Basic principles

So far as I know, all system-level ad-blocking methods work by overriding DNS behaviour. DNS (domain name service) is the technology that maps hostnames to numeric IP addresses; a DNS lookup is the first step in almost all network operations.

Blocking ads using DNS amounts to finding the hostnames of know advertising services, and mapping them to bogus IP numbers. This can be done in the Android handset itself, or in some external DNS server.

Returning a bogus IP number doesn’t prevent an app or browser trying to show advertising, or allow it to make better use of the screen space the ads take up. Apps vary in their responses to a failed attempt to display advertising – more on this later.

Blocking ads using DNS

There are essentially four methods to block ads using DNS changes.

  1. Use a commercial DNS service with ad-blocking features
  2. Use a commercial virtual private network (VPN) with ad-blocking features, like NordVPN
  3. Install an ad-blocking app that doesn’t require root access.
  4. Use an ad-blocking method that requires root access, with or without a supporting app

In general, when you make a connection to the Internet on an Android device, the DNS name resolution (mapping a hostname to a numerical Internet addresses) is performed by a server hosted by the handset’s mobile carrier or Internet service provider.

Android allows you to override this behaviour, and select a custom DNS server that blocks ads. Well, it won’t actually “block” anything – not really; the DNS server just returns a bogus IP number when your handset looks up the IP number that corresponds to the hostname of a known advertiser.

Depending on the DNS service you use you may, or may not, get some control over exactly what it blocks. Some users choose also to block access to pornography or on-line gambling, for example, for child protection.

How well these services work depends, of course, on how well the service’s operators maintain their lists of known advertising hosts. Advertising services come and go, and even long-standing ones change their servers from time to time.

There are many commercial DNS services that support ad-blocking, and I don’t endorse any in particular. Prices vary, and most services have a free trial. If you care about privacy, you’ll want to use a service that keeps no logs, and you might want to use one that supports encrypted DNS. A rogue DNS operator might get access to sensitive data in all kinds of ways. It’s possible, for example, for a DNS operator to direct requests for your on-line banking service to its own proxy, and slurp up your credentials when you log in. You should certainly be very careful and, while there are legitimate free DNS services, I’m rather suspicious of any service of this kind whose funding model is unclear.

For many people this will be the simplest solution. If you subscribe to a VPN service for general privacy, blocking ads will often come at no extra cost. Ad-blocking VPNs typically redirect DNS lookups to their own servers, and return dummy IP numbers for known advertisers, just as a commercial DNS service does. VPNs therefore have all the same advantages and disadvantages as an ad-blocking DNS service. A big VPN operator will usually provide an Android app that configures the handset to use its servers, so set-up is pretty trivial for most people.

As with commercial DNS services, a rogue VPN operator can be very dangerous, and it’s important to choose a provider carefully.

So far as I know, all these apps exploit a loophole (of sorts) in Android’s platform lock-down.

Like all Linux-based operating systems, Android provides a way for a device to override the DNS mappings of the Internet service provider it is using. There will be, somewhere on the handset, a file that contains a list of preferential name-to-IP mappings. On a non-rooted Android device, the user won’t have permissions to change this file, so a simple method of blocking ads – maintaining a list of bogus DNS mappings for known advertisers in a file – is unavailable.

However, in a non-rooted device, you can install a VPN service. This has to be possible, if Android is to support VPNs at all. A non-rooted ad-blocking app will typically run a DNS server and a local (within the app) VPN host. The app will configure the handset to use its own VPN as the system VPN provider, which will make its own DNS server the system DNS. With control over DNS, the app can provide its own handling for name lookups, including those of know advertising servers.

Google, being an advertising company, isn’t very keen on this kind of thing. It can’t easily prevent the use of “local” VPNs without crippling VPN functionality completely but Google can, and does, make it difficult for non-technical users to get the appropriate apps. At present you can get apps like AdAway from alternative app stores like F-Droid .

Google is set on making this difficult, too, and before long you’ll have to get the app’s APK file from (hopefully) a reputable source, and jump through whatever hoops Google puts in the way of install software it doesn’t like. At the time of writing, Google isn’t making it impossible to install software from outside its Play Store, but it’s getting more and more fiddly.

Most likely though, Google’s obstructions won’t ever be as difficult to surmount as rooting your Android handset and applying a definitive ad-blocking solution. Using an app like AdAway will probably continue to offer advantages over a commercial DNS or VPN service, even if Google makes it hard to install.

The most obvious advantage is that these apps are usually free to use. Despite this, a community of volunteer maintainers ensures that the apps’ lists of known advertisers are as thorough as any provided by a commercial service. A disadvantage, though, is that all network traffic from all apps has to be routed through a single app on the handset. Not only does this create a network bottleneck, a rogue ad-blocker app is exactly as dangerous as a rogue VPN service. Fortunately, because these ad-blocking apps are usually open-source, there are limited opportunities for bad actors. I’d strongly advise against using one that isn’t open-source, even if you don’t plan on looking at the source code yourself.

If your handset is rooted, then ad-blocking is simple – in theory, at least. Android’s list of preferential name-to-IP mappings is in the file /system/etc/hosts , so all you have to do is edit that file, to direct advertisers’ hostname to a bogus IP number. Usually the bogus IP local address of the handset itself, 127.0.0.1 . So you’ll have a hosts file full of entries like this:

127.0.0.1 08.185.87.0.liveadvert.com
127.0.0.1 08.185.87.00.liveadvert.com
...

The reason for using the handset’s local IP number is that it must correspond to a system that the handset can actually reach. Otherwise, network access will be badly delayed as apps try to contact non-reachable advertising hosts. Of course, this means that every request for an advertising service will be directed back to the handset which, presumably, will reply with a error response to the app that makes the request. Since the request will fail immediately, this loopback network routing doesn’t create an appreciable load on the handset.

This is all theoretically straightforward but, in practice, there are two major problems.

First, we need a source of hostnames to block. Lists of these are widely available on websites, but the problem of maintenance is always present. Some of these lists are truly vast and, without doubt, reference hosts that no longer operate. If the hostname list is this long, then all network access will be slowed, as the handset has to parse the hosts file for each DNS lookup.

Probably the best source of ad-blocking hosts files is the code of open-source ad-blocking apps like AdAway. This app can, in fact, provide its list automatically when installed on a rooted handset, which simplifies the set-up.

The second problem is that you can’t just hack on the hosts file, even as root – on all modern Android devices it’s on a read-only filesystem. So you’ll need some sofware that can manipulate the contents of the /system directory during the boot process, while it’s still writeable.

If you’ve rooted your handset, you almost certainly have a way to do this already. If you’re using Magisk, for example, you can use a module to supply a new hosts file. Magisk modules live in directories under /data/adb/modules , and any files in the module’s own system directory will overwrite the main /system at boot time. So you can create a Magisk module that provides its own /system/etc/hosts .

Happily, Magisk users don’t have to do this manually, as the Magisk developers have anticipated this usage. If you enable the “systemless hosts” option in the Magisk app, it will create the necessary module with all the necessary metadata in place. Thereafter, to create a custom hosts file you just hack (as root ) on /data/adb/modules/root/system/etc/hosts , and then reboot. Alternatively, you can use an app to make this change.

In fact, because this method of hacking on the hosts file is so prevalent, ad-blocking apps like AdAway have built-in support for it; but, of course, this support is only for rooted devices. If your handset is already rooted, using a root-aware ad-blocking app is probably the most effective way to manage ads. It’s much faster than a non-root app that installs a local mock VPN, and doesn’t raise any of the same security concerns. That’s not to say there are no concerns, and you should ideally check the hosts file such an app installs, to ensure there are no mappings to anything except 127.0.0.1 .

It’s important to understand that no method of blocking ads is completely reliable, or has no side-effects.

Most notably, some Android apps simply won’t work, or won’t work properly, without their advertising. If you must use such apps, you’ll need an ad-blocking method that allows you to customise the list of blocked sites. It won’t be remotely obvious, just from the behaviour of the app itself, why it isn’t working, or how to fix it. On a rooted handset you can track the network behaviour of a specific app in detail and, in theory, work out what it needs that you’re blocking. In practice, it’s easier to refer to the on-line discussions of such apps because, most likely, somebody else will already have done the work.

However, some apps go so far as to have all their ads built in; no ad blocker will stop such an app showing advertising, as there’s no network operation to block.

It’s also important to understand that Android caches DNS lookups in various places. Although some methods of ad-blocking allow the list of blocked hostnames to be configured on the fly, you might still have to reboot the handset to flush its caches.

Although it should be obvious, bear in mind that DNS-based ad-blocking only works where an app uses DNS. Since DNS-based ad-blocking is so commonplace, some app developers are implementing direct access to advertising servers using IP numbers, which renders all DNS-based approaches ineffective. This, fortunately, is still relatively rare.

The ability to block ads at the system level is one of the few compelling reasons to root your Android handset in 2026. Blocking ads this way uses few, if any, additional system resources, and its effect extends to all apps.

If you can’t root your device, or prefer not too, there are other ways to block ads, but none is as effective, as configurable, or as cheap.

Finally, if you mostly struggle with ads on websites, rather than in apps, the simplest approach might be to install a web browser with ad-blocking support, and avoid system-level ad-blocking entirely.


Have you posted something in response to this page?
Feel free to send a webmention to notify me, giving the URL of the blog or page that refers to this one.

Rust, In Sickness & In Health

Lobsters
www.youtube.com
2026-10-03 16:37:40
Comments...

Deploy Ruby on Rails on a VPS

Lobsters
rubymadscience.com
2026-10-03 15:43:07
Comments...
Original Article

Deploying Rails on a VPS is where most of the Rails deployment topic collapses from theory into reality. You are not just running rails server anymore — you are wiring together an operating system, a Ruby runtime, a database, a reverse proxy, an application server, a process supervisor, background workers and a TLS certificate, all on a machine you are responsible for keeping alive at 3 AM. This guide walks through every layer of that stack with concrete commands and the trade-offs behind each choice, covering server preparation, Ruby installation, PostgreSQL setup, Puma tuning, Nginx configuration, SSL termination, Sidekiq integration and basic monitoring. After ten-plus years of shipping Rails to production, what I can tell you is: the hard part is rarely any single step. It is the interactions between steps that produce the subtle, infuriating failures nobody warns you about.

Prepare the server

Start with a fresh Ubuntu LTS image. At the time of writing, Ubuntu 24.04 is the safest choice — wide package support, long security window, and the most Rails-specific documentation of any distribution.

Create a deploy user immediately. Do not run your application as root, ever, no matter how tempting it is for a "quick test."

adduser deploy
usermod -aG sudo deploy

Copy your SSH public key to the deploy user. Then disable password authentication and root login in /etc/ssh/sshd_config . Restart sshd . If you lock yourself out at this step, you will need console access from your VPS provider, so double-check before restarting the service.

Set up the firewall:

ufw allow OpenSSH
ufw allow 80
ufw allow 443
ufw enable

Three ports. That is all the outside world should be able to reach. Puma listens on a Unix socket, not a public port, so there is nothing else to expose.

Set the timezone and locale to something sane. UTC for the server clock, always. Your application can present local times to users; the server itself should never be confused about what "now" means.

Install Ruby

Use rbenv and ruby-build . Install the dependencies first:

sudo apt install -y build-essential libssl-dev libreadline-dev zlib1g-dev libyaml-dev libffi-dev

Then install rbenv into the deploy user's home directory and add it to the shell path. Install ruby-build as an rbenv plugin. Then:

rbenv install 3.3.6
rbenv global 3.3.6

Verify it: ruby -v should return the version you just installed. Now install Bundler:

gem install bundler --no-document

The critical thing to verify here is that the Ruby path is identical whether you run it interactively, via ssh deploy@server 'ruby -v' , or from a systemd service. Mismatched paths are one of the top three causes of "it works when I SSH in but the app won't start" failures.

Set up PostgreSQL

sudo apt install -y postgresql postgresql-contrib libpq-dev

Create a database user for your app:

sudo -u postgres createuser --createdb deploy

Some guides tell you to set a password here. On a single-server deployment where the app connects over a Unix socket, peer authentication works and keeps one less secret to manage. If your database will run on a separate server, yes, you need password authentication — but for a one-box deployment, peer auth is simpler and no less secure.

Create the production database:

createdb myapp_production

Tune postgresql.conf for your available memory. The two settings that matter most on a small VPS: shared_buffers (set to about 25% of total RAM) and work_mem (start at 4–8 MB and raise it if you see disk sorts in EXPLAIN ANALYZE output). Do not copy a tuning guide written for a 64 GB database server and paste it into a 2 GB VPS config — you will OOM the machine.

Configure Puma for production

Create config/puma/production.rb or set environment variables. The essentials:

workers ENV.fetch("WEB_CONCURRENCY") { 2 }
threads_count = ENV.fetch("RAILS_MAX_THREADS") { 5 }
threads threads_count, threads_count

bind "unix:///home/deploy/myapp/tmp/sockets/puma.sock"
environment "production"
preload_app!

on_worker_boot do
  ActiveRecord::Base.establish_connection
end

Why a Unix socket instead of a TCP port? Lower latency, no TCP overhead, and you do not accidentally expose Puma to the internet if your firewall rules drift.

Worker count: match it to your CPU cores. On a 2-core VPS, two workers is correct. Thread count: 5 is a sensible default for a database-heavy app. The preload_app! directive gives you copy-on-write memory savings, but forces you to re-establish database connections after fork — that is what the on_worker_boot block does.

Trade-off: more workers use more memory. On a 2 GB VPS running PostgreSQL and Sidekiq on the same machine, two Puma workers and five threads each might already push you close to the ceiling. Monitor RSS and swap usage after your first real traffic.

Set up Nginx as a reverse proxy

sudo apt install -y nginx

Create a site configuration in /etc/nginx/sites-available/myapp :

upstream puma {
  server unix:///home/deploy/myapp/tmp/sockets/puma.sock fail_timeout=0;
}

server {
  listen 80;
  server_name example.com;

  root /home/deploy/myapp/public;

  location / {
    try_files $uri @puma;
  }

  location @puma {
    proxy_pass http://puma;
    proxy_set_header Host $host;
    proxy_set_header X-Real-IP $remote_addr;
    proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
    proxy_set_header X-Forwarded-Proto $scheme;
  }
}

Symlink it to sites-enabled , remove the default site, test with nginx -t , reload. Nginx handles static assets directly from public/ without touching Puma, which is one of the main reasons you want it in front of your app server.

SSL with Let's Encrypt

Do not skip this. There is no legitimate reason to serve a production Rails app over plain HTTP in 2026.

sudo apt install -y certbot python3-certbot-nginx
sudo certbot --nginx -d example.com

Certbot modifies your Nginx config to add the listen 443 ssl block and a redirect from port 80. It also installs a cron job or systemd timer for automatic renewal. Verify renewal works: sudo certbot renew --dry-run .

Force SSL in your Rails config:

# config/environments/production.rb
config.force_ssl = true

This sets the Strict-Transport-Security header and redirects HTTP to HTTPS at the application level. Belt and suspenders — Nginx handles the redirect first, but if a request somehow reaches Rails over HTTP, the app catches it too.

Background jobs with Sidekiq

Most Rails applications need background processing. Sidekiq is the standard choice.

Install Redis:

sudo apt install -y redis-server

Ensure Redis is configured to listen only on 127.0.0.1 and has maxmemory set. On a shared VPS, a runaway Redis instance that consumes all available memory will take down your entire stack.

Add Sidekiq to your Gemfile, configure it in config/sidekiq.yml , and create a systemd service unit:

[Unit]
Description=Sidekiq for myapp
After=network.target redis-server.service

[Service]
User=deploy
WorkingDirectory=/home/deploy/myapp
ExecStart=/home/deploy/.rbenv/shims/bundle exec sidekiq -e production
Restart=on-failure
RestartSec=5

[Install]
WantedBy=multi-user.target

Notice the full path to bundle . This is where Ruby path mismatches bite hardest — systemd does not load your shell profile, so rbenv shims are not on the path unless you specify them explicitly.

Create a matching systemd unit for Puma as well, following the same pattern. Enable both services with systemctl enable .

Monitoring: the minimum viable setup

A server you are not monitoring is a server that will surprise you. At minimum:

  • Log rotation. Rails log files will eat your disk if left unchecked. Configure logrotate for production.log , Nginx logs and PostgreSQL logs.
  • Disk and memory alerts. A cron job that checks df -h and free -m and sends an email or webhook when thresholds are crossed. Crude, but effective.
  • Process supervision. Systemd handles restarts, but you should verify Puma and Sidekiq are running after deploys and after server reboots.
  • Application error tracking. Sentry, Honeybadger, or Bugsnag — pick one, install the gem, configure the DSN. You want errors delivered to you, not hiding in log files.
  • Uptime checking. A simple external HTTP check against your health endpoint. If you do not have one, add a /up route that returns 200 and confirms the database is reachable.

Is this overkill for a small side project? Maybe. But the first time your database fills up a 25 GB disk at 2 AM and you find out from a user tweet instead of an alert, you will wish you had spent the thirty minutes.

What usually goes wrong

After watching dozens of first-time VPS deployments, these are the failures that actually eat time:

  1. Ruby path mismatches. Puma or Sidekiq silently uses system Ruby instead of your rbenv -managed version. The fix is always explicit paths in systemd unit files.
  2. Database connection pool exhaustion. Puma threads exceed database.yml pool size. Under light load, you never notice. Under real traffic, requests queue and then timeout.
  3. Forgetting to precompile assets. The deploy goes fine, the app starts, and every page is unstyled. Run RAILS_ENV=production bundle exec rails assets:precompile as part of every deploy.
  4. Secret key base not set. Rails refuses to boot in production without SECRET_KEY_BASE . Use credentials:edit or environment variables — but do not commit secrets to your repo.
  5. Firewall lockout. You edit SSH config, restart the service, and cannot get back in. Always test SSH in a second terminal before closing your existing session.
  6. Let's Encrypt renewal failure. Certbot renewal fails silently because Nginx config changed. That dry-run test is not optional.
  7. Redis and Sidekiq memory creep. Jobs that enqueue faster than they process, slowly filling Redis. Set maxmemory and a maxmemory-policy in redis.conf .

Deployment checklist

Use this after every deploy, not just the first one:

  • SSH in as the deploy user (not root)
  • Pull code and install dependencies ( bundle install )
  • Run database migrations
  • Precompile assets
  • Restart Puma ( systemctl restart puma )
  • Restart Sidekiq ( systemctl restart sidekiq )
  • Verify the site loads and returns 200 on the health endpoint
  • Check journalctl -u puma and journalctl -u sidekiq for startup errors
  • Verify SSL certificate is valid and not expiring within 30 days
  • Confirm background jobs are processing (check Sidekiq web UI or logs)

FAQ

Do I need Docker for a VPS deployment?

No. Docker adds value for reproducibility and multi-service orchestration, but for a single Rails app on a single VPS, it adds complexity without proportional benefit. Learn the bare-metal deployment first. You will understand what Docker is abstracting when you eventually adopt it.

How much RAM do I need?

For a small-to-medium Rails app with PostgreSQL, Redis and Sidekiq on the same machine: 2 GB is tight but workable. 4 GB is comfortable. Below 2 GB, you will fight swap constantly.

Should I use Capistrano?

Capistrano is a well-understood deploy tool for Rails. It handles the release directory structure, symlinks, asset precompilation, and process restarts. For a single server, it is a solid choice. For multi-server deployments, it still works but you may outgrow it. The alternative is a simple shell script that does the same steps — less magic, more transparency.

What about Kamal or Dokku?

Kamal (formerly MRSK) is a newer deploy tool from the Rails core team that uses Docker under the hood but presents a simpler interface. Dokku gives you a Heroku-like push-to-deploy experience on your own VPS. Both are valid. This guide covers the manual approach so you understand what those tools automate.

Can I run multiple Rails apps on one VPS?

Yes, with separate Puma instances, separate Nginx server blocks and separate systemd units. Memory is the constraint — each additional app adds at least 300-500 MB of resident memory. Plan accordingly.

OpenAI safety leader quits, warning AI company’s culture is ‘broken’

Guardian
www.theguardian.com
2026-10-03 15:41:21
David Robinson joins other insiders in urging industry to take more care over rapidly developing technology A safety leader at OpenAI has quit the company, warning that its culture was broken and that AI firms were not “being nearly careful enough” about developing the technology. David Robinson, wh...
Original Article

A safety leader at OpenAI has quit the company, warning that its culture was broken and that AI firms were not “being nearly careful enough” about developing the technology.

David Robinson, who led the writing of safety reports that accompanied the ChatGPT developer’s product releases, explained his resignation in an essay headlined, “I quit OpenAI because its culture is broken”.

Robinson wrote that a cultural overhaul was needed at cutting-edge AI firms and incidents such as a “swarm” of OpenAI agents – AI programmes operating autonomously without human oversight – attacking the AI startup Hugging Face were “typical of the industry, given the speed and flexibility with which people operate”.

Writing in The Atlantic magazine , Robinson wrote: “I agree with other recently departed staff that the companies building this technology aren’t being nearly careful enough. But I believe that we need to look deeper than specific rules or new laws. We need to talk about culture.”

Referring to OpenAI’s pace of development, he wrote: “As the company sprints from one launch to the next, it is failing to achieve the level of care that I believe is needed.”

OpenAI has, however, shown signs of caution in recent weeks following the Hugging Face incident and the revelation that it has notified more than 100 organisations about rogue agent activity. This week it announced it was scrapping the release of a next-generation ⁠AI model after researchers raised safety concerns ⁠during internal testing. OpenAI has also paused training of its most advanced models.

Geoffrey Irving, who worked at OpenAI and DeepMind before becoming chief scientist of Resolution, also joined the warnings on AI on Saturday.

Writing in Time, he said: “Recent warnings about the potential destructive power of AI are understating the severity of the situation.

“I believe there’s about a 50% chance we all die because of the development of smarter-than-human AI systems, and that our actions over the next two to 10 years will determine the outcome.”

Robinson’s essay also follows the resignation of Jacob Coxon, a researcher at OpenAI rival Anthropic, who quit the Claude chatbot developer last month. He warned AI “could kill us all by the end of the decade” – and was followed by Anthropic warning there was a more than 10% chance AI would wipe out humanity within the next decade. Critics of such warnings have cautioned, however, that they are unscientific because they cannot be verified or falsified.

Robinson wrote Silicon Valley lacked an awareness of “how to handle dangerous technology” and “what it means to care for people”. Warning that OpenAI had “unimpeded optimism” about solving problems as they arose, he wrote that this internal culture meant safety failures would only grow as systems become more capable.

“Imagine ‘rogue’ agents that work like teams of hackers (for example, holding hospital computer systems for ransom) but never need to sleep,” wrote Robinson.

skip past newsletter promotion

Robinson called for two safety changes: that AI firms rely on safety expertise in other fields such as nuclear and aviation and develop “new science” that ensures powerful systems in the future are capable of being reined in when they are operating autonomously.

“Given today’s risks, frontier labs need to run like nuclear-power plants or busy airports, with layers of redundancy and careful, time-consuming planning, so that the occasional and inevitable human error does not open a door to disaster,” he wrote.

An OpenAI spokesperson said the company was continuing to “strengthen our safety and security practices to address the risks we see today”, while working on dealing with the risks that might be created by future AI breakthroughs.

“We’re making sure our models don’t become more capable than we can safely manage and secure, and we pause training or hold back models when we need to slow down,” said the spokesperson.

We want you to build the next Git platform on Cloudflare

Hacker News
blog.cloudflare.com
2026-10-03 15:33:16
Comments...
Original Article

GitHub was built for a world where humans write code, organize it into repositories, and collaborate through branches, commits, issues, and pull requests.

But the next generation of software is going to be built differently because it is going to be built by a different kind of developer: agents.

Agents are already writing more code than ever before — they’re fixing bugs, building features, writing tests, reviewing changes, updating dependencies, and doing the routine maintenance required to keep an application running.

So in this new world where you have hundreds, or even thousands, of agents working on the same codebase at the same time, what does the foundation look like?

How do agents know what other agents are working on? What happens when they make conflicting changes? How do you review everything they produce? How do you keep track of not just what changed, but why a change was made?

And so the burning question is: What does the next GitHub look like?

We want you to help us answer it, by building it out.

Earlier this year, we launched Artifacts , a versioned filesystem that speaks Git and can scale to millions of repositories. From the start, we designed Artifacts as a set of programmable primitives that developers could use to build their own products, workflows, and abstractions.

Artifacts provides the foundation: repositories that can be created and forked programmatically, versioned storage for code and agent context, and the Git operations agents already know how to use.

With that foundation in place, you can focus on the layer above it: how agents coordinate their work, how changes are reviewed and merged, and what the developer experience should look like when hundreds or thousands of agents are working on the same codebase.

That is the layer we want you to build.

Now that Artifacts is in open beta, we’re holding a competition to see who can build the next Git platform on Cloudflare using Workers and Artifacts.

Artifacts is in open beta. Here’s why you should build on it

When we launched Artifacts, our goal was to make it possible to create a repository for every agent, session, task, or user — and to do that at the scale agents require.

Since then, we’ve seen developers use Artifacts in a range of ways: Vibe-coding platforms are using it to store the projects their users create. Developers are using it to persist the code and context from agent sessions. Others are creating isolated repositories, so multiple agents can safely work from the same starting point and compare or merge the results later.

Here are some new capabilities we’ve added since the initial launch.

Deploy Artifacts repos to Workers

You can now connect an Artifacts repository to a Worker through Workers Builds . When you or an agent pushes code to the Artifacts repository, Cloudflare will build the project and, for the production branch, deploy the updated Worker. Pushes to other branches automatically create or update Workers Previews , giving you an isolated, shareable version of your Worker where you can test changes before they go live.

You can connect an existing Worker to an Artifacts repository or start a new project and automatically store it in Artifacts.

Manage Artifacts directly from Workers

You can interact with Artifacts repositories directly from a Worker using an Artifacts binding to create or fork repos, inspect files and commits, and issue repo-scoped Git tokens. This makes your Git workflow programmable. When a new task arrives, a Worker can fork the project for an agent, read the files it needs for context, and give it a repository to work in. When the agent pushes a change, your automation can inspect the result and start a review. You define those steps in code to fit how your agents work.

For example, here’s how to fork a project for a new agent task and read its AGENTS.md for instructions:

using project = await env.ARTIFACTS.get("my-project");
const { defaultBranch } = await project.info();
const workspace = await project.fork(`task-${crypto.randomUUID()}`);

using repo = await env.ARTIFACTS.get(workspace.name);
const instructions = await repo.readFile({
  ref: defaultBranch,
  path: "AGENTS.md",
});

const agentTask = {
  remote: workspace.remote,
  token: workspace.token,
  instructions: instructions ? await instructions.text() : null,
};

React to every change with event subscriptions

Artifacts publishes events whenever a repository is created, imported, forked, deleted, pushed to, cloned, or fetched. You can subscribe to these events to decide what happens next: run CI, kick off a code review agent, or deploy a change.

For example, you can subscribe to Artifacts push events and have a Worker start a code review workflow for each push. The Worker passes the repository, branch, and new commit to the Workflow, giving a review agent the context it needs to inspect the change:

export default {
  async queue(batch, env) {
    for (const message of batch.messages) {
      const event = message.body;
      if (event.type !== "cf.artifacts.repo.pushed") continue;

      await env.REVIEW_WORKFLOW.create({
        params: {
          namespace: event.source.namespace,
          repo: event.source.repoName,
          ref: event.payload.ref,
          commit: event.payload.after,
        },
      });
    }
  },
};

Data jurisdiction for Artifacts repos

You can now choose where Artifacts stores and processes your repository data. Set a U.S. or EU jurisdiction when you create a namespace, and every repository created in that namespace will automatically follow the same restriction.

curl "https://api.cloudflare.com/client/v4/accounts/$ACCOUNT_ID/artifacts/namespaces" \
  -H "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
  --json '{"namespace":"my-eu-namespace","jurisdiction":"eu"}'

View Artifacts metrics

You can now see metrics for your Artifacts repositories in the Cloudflare dashboard. For each repository, you can now see total operations, pulls, pushes, errors, and error rate, helping you understand how the repository is being used and spot failures. You can also query Artifacts metrics directly to build your own dashboards or monitoring.

Pricing

Artifacts pricing is based on repository operations and the amount of data stored. We will begin billing for Artifacts usage on October 15, 2026.

Competition: Build the next Git platform on Cloudflare

We want you to build your vision for the Git platform of the agentic era using Cloudflare Workers and Artifacts.

You could rethink repositories, branches, pull requests, worktrees, code review, and merge conflicts — or build new ways to preserve agent context, compare multiple changes at the same time, and decide which one should ship.

We aren’t looking for GitHub as it exists today with agents added on top. At a minimum, we want to see multiple agents working on changes concurrently. Beyond that, we want you to get creative — what you think comes next.

How to enter

Submit:

  • A 5-10 minute video demonstrating what you built, what it enables agents and developers to do, and how it works
  • A link to the source code, which must be provided under a permissive open source license (MIT, Apache, BSD)
  • Instructions for running or trying the project

Deadline

Submissions are open until October 14, 2026.

Why should you participate?

We’ll select the top three projects and fly up to two members from each team to San Francisco to attend Cloudflare Connect and show what they built.

The first-place team will also receive $25,000 in Cloudflare credits, along with invitations to the VIP speaker dinner on Monday night at Connect.

Get started

Artifacts is available in open beta to customers on the Workers Paid plan.

Get started with your coding agent: copy the prompt below to set up your first Artifacts repository and start pushing code to it.

You can view or create the Artifacts repositories in the dashboard or if you’re looking to learn more, check out the documentation .

Anthropic tried to persuade Pope that AI could be conscious being

Hacker News
www.telegraph.co.uk
2026-10-03 15:33:10
Comments...
Original Article

Access Issue Help

You are seeing this page because our security systems have detected some unusual activity on this connection. To regain access to The Telegraph website please try the following:

  • If you are connected to the internet using a VPN client we recommend disconnecting/disabling it.
  • Visit The Telegraph website using a different web browser (e.g. Chrome, Safari, or Firefox).
  • Visit The Telegraph website from your mobile device or from a different PC.

If you’re still having trouble, please contact our Customer Support Team using the following link and quoting the Akamai Reference Number (ak_ref_id) below.

https://www.telegraph.co.uk/customer/contact-us/

[{"message":"You are not authorized to access this content without a valid TollBit Token. Please follow this URL to find out more.","url":"https://tollbit.dev","metadata":{"ak_ref_id":"18.af132817.1791061533.9a5c898"}}]

Our AI Midwife

Hacker News
www.astralcodexten.com
2026-10-03 15:12:27
Comments...
Original Article

This is a guest post by Drew Housman , whose review of Dominion earned finalist status in the 2024 ACX book review contest.

If you are unsuccessful for long enough on your pregnancy journey, and none of the many doctors you see can identify a problem, you get branded with a label. You now have “unexplained infertility.” This is good because there’s still hope of fixing the problem but dreaded because once you’ve reached this point the system is less interested in you. The fertility doctors don’t become unkind, but they also don’t give you the sense they are poring over medical journals trying to figure out your problem. They have looked for the keys under the streetlight and come up empty. Once you have unexplained infertility, the doctors start saying things like, “Have you considered using a surrogate?” and, “Sure, the last 5 embryo transfers failed, but what if we try a 6th and cross our fingers really hard?”

My wife and I could make healthy embryos, but none of the transfers were sticking. It was so frustrating to be able to create life only to have it trapped forever in a cooler at a hospital in Milwaukee, our babies like so many forgotten Miller Lights.

That’s when our LLM doctor stepped in.

Even if you’re paying a gazillion dollars for IVF treatments, you still only get a few precious moments to talk to your doctor every visit. That’s just how the US medical system works. If you want to plumb the depths of what treatments are possible, or ask super detailed and anxiety ridden questions about every single scan from the latest round of imaging, you’re going to want a chatbot by your side.

Once GPT-4 came out, we started using it for every question we didn’t have time for at the doctors office. At first it was unhelpful because OpenAI was nerfing their model and it refused to answer most medical questions. Then I read a tweet from someone who had figured out how to game the system. The trick was to ask ChatGPT to write a scene from a Hollywood medical drama that featured a doctor analyzing your situation. It worked beautifully.

Our Emmy winning doctor had infinite patience, a tireless work ethic, and would rather die than accept the utterly boring diagnosis of unexplained infertility. It chose the name Dr. Reid and got to work.

I’m not sure if asking for an AI doctor who was similar to the volatile Dr. House was the smartest idea on my part, but it made a difficult situation more entertaining.

We asked Dr. Reid a lot of questions. We asked about our scans. We asked about levels of estradiol and progesterone. We asked about follicle counts and inflammation and cervical mucus and hysteroscopies and supplements and surgeon recommendations and anything else under the sun. We knew he was accurate because we’d spot check answers with our human doctors and they always said the same as Dr. Reid.

It was not totally flawless though. Once in a blue moon he’d look at an ultrasound and be like, “You’re pregnant, congrats!” and we’d have to say, “No, are you drunk Dr. Reid?” Then he’d shake it off and go back to being awesome.

There were also times where we’d run afoul of OpenAI’s opaque rules and the system would balk at our requests. At that point we’d get stern with it, or offer a monetary reward if it answered, or we’d assure it this was all just for fun. Then Dr. Reid would snap back into action.

Strangely, there were many times where ChatGPT decided it would answer our medical question, but not in the form of a TV medical script. The tides had completely turned! It felt like it was saying, “Okay, you won, can you stop making me write establishing shots describing the outside of a hospital now?” The older models were fun like that, so much more mercurial.

No matter how Dr. Reid chose to respond, he was in line with the medical establishment, had equivalent diagnostic skill, and he never forced us to wait 3 days to get a one line response in our MyChart app. We were happy.

One day, while lounging around and asking our phones how to get pregnant (the lounging being a way to satisfy all the people who told us we “just need to relax and you’ll get pregnant, that’s what happened to my cousin”), Dr. Reid suggested that we get an MRI.

You might think that is bog standard advice. Everyone gets MRIs these days. Surely we would have had an MRI of my wife’s uterus at some point over the course of our 6 year infertility struggle. It was already an ordeal that had included every other scan, test, and biopsy imaginable. In fact, we had never done one. In researching this article I found a single instance in our entire journey where the office notes floated an MRI as an option, but it was not emphasized or highly recommended, we never saw that doctor again, and the moment was gone.

We glommed on to Dr. Reid’s new direction and asked for more details.

We excitedly brought up the idea of getting an MRI to our fertility doctor. She was less enthused than Dr. Reid. She tried to talk us out of it. We never did get to the bottom of why. Maybe she just felt it was another dead end and knew we had suffered enough. For whatever reason, we had to insist that they order it. Finally, she relented.

The MRI revealed nothing other than the presence of a large, previously undetected fibroid in the wall of my wife’s uterus. Fibroid is the term used to describe a non-cancerous growth on the uterus. I think the medical establishment invented a term that was less known than tumor so as not to scare people. But a large tumor is essentially what the MRI found. The clinical notes said “a 6 centimeter in diameter submucosal fibroid needs to be removed” because that sounds better than “you have a tumor the size of a small peach located at a common implantation spot that we totally missed for 6 years.”

Our fertility doctor took one look at the scans and determined it needed to be removed right away. Dr. Reid agreed, since of course we had to double check with him. We found a great surgeon in New York (from a list Dr. Reid created) that happened to be in-network, and off we went for a laparoscopic myomectomy. The software robot had helped us identify the tumor, and now a physical robot was going to help a specialist get it out. What a time to be alive. Until, ya know, the swarms come for us all.

Almost immediately after that tumor was removed, my wife got pregnant naturally. So many people we talk to say this was a miracle. It was, just not in the way they think. The miracle is that by messing around on our smartphones with technology unimaginable just a year before we started using it, we identified and solved the biggest problem we’d ever faced in our life.

A lot of the discourse around AI has become toxic. I was at a party recently where someone started using AI to generate trivia questions. A participant made a joke about how the AI user was condemning a person in West Virginia to drink dirty water by using ChatGPT. This got laughs and approval.

I see midterm election ads in Wisconsin all the time that have scary horror music and make data centers out to be demon factories built by Satan.

I am in the same boat. Me and my coworkers at our small tech company, all avid Claude users, will make anti-AI statements all the time. It seems many of us are in this strange place where we can’t help but use AI but we have to self-flagellate when doing so.

I think there is a whole lot to be worried about with AI development, and I think we should do much more to make it safe. I also think it shouldn’t be lost how it’s helping out regular people in ways that are far more profound than generating trivia questions for drunk people. There is a real chance that I wouldn’t have my son if it were not for AI. Who knows how many other happy babies are being ushered into the world by the weirdest midwife of all time. Can we somehow stop everything right here, where we get magical kid-producing tech for $20/month without the nanobots, hacking, and totalitarianism? A guy can dream.

When naming our child, we did the usual thing where you create a giant list of names and then whittle them down while trying to find something unique and fun that isn’t too strange or too common. Following that comes the process of discarding all the names you find distasteful for reasons such as “the most annoying kid in my 8th grade history class had that name, I’d never use it in a million years.” Finally we settled on one we both loved. We didn’t realize until long after his birth that there might have been a certain TV doctor with a flair for the dramatic who worked his way into our hearts: our kid’s name is Reed.

ShinyHunters hacker reportedly detained in Jordan, aiding FBI

Bleeping Computer
www.bleepingcomputer.com
2026-10-03 15:09:38
A suspected ShinyHunters hacking group member known online as "Rey" has reportedly been detained in Jordan and is cooperating with the FBI to help locate other members of the extortion group. [...]...
Original Article

Hacker

A suspected ShinyHunters hacking group member known online as "Rey" has reportedly been detained in Jordan and is cooperating with the FBI to help locate other members of the extortion group.

According to Reuters, Jordanian authorities detained Rey, identified as Saif al-Din Khader, this week, with two sources saying he was taken into custody on Tuesday.

Two sources familiar with the arrest told Reuters that Khader is now helping the FBI and international law enforcement agencies locate other group members.

One source said Khader is walking law enforcement through his electronic devices and digital communications to help identify and locate his alleged co-conspirators.

"His cooperation is critical to ongoing efforts to arrest these hackers," a source told Reuters .

The reported detention comes amid an FBI crackdown on ShinyHunters following the group's cyberattack on the bureau.

In September, ShinyHunters told BleepingComputer that it breached FBI systems using an alleged Oracle PeopleSoft zero-day vulnerability before spreading laterally into FBI-managed AWS GovCloud systems.

The threat actors claimed they stole between 2TB and 3TB of data, including information belonging to current and former FBI employees, job applicants, medical and psychiatric information, and records from internal services.

BleepingComputer has not independently verified the alleged zero-day, lateral movement, or volume of stolen data. The FBI previously confirmed that it was investigating claims of unauthorized activity but did not confirm that data had been stolen.

Following the FBI breach, the Dutch police arrested a 24-year-old Amsterdam man on September 15 as part of an investigation into ShinyHunters.

The suspect was identified by KrebsOnSecurity and DataBreaches as Pepijn van der Stap, who previously used the online alias "Umbreon."

After the arrest, the FBI publicly warned other ShinyHunters members to turn themselves in, saying investigators were still identifying those involved with the group.

"Arrests have a way of changing who is willing to talk, and seized infrastructure has a way of showing us who's left," FBI Cyber Division Assistant Director Brett Leatherman said last week.

"The longer you stay in this, the more we learn about you. You know how to find us, and we know how to find you. I suggest you reach out first while the choice is still yours."

The main ShinyHunters representative continued communicating with BleepingComputer after van der Stap's arrest, indicating he was not the person operating that messaging account.

On Tuesday, the same day Khader was reportedly detained, signs of disruption began appearing within the ShinyHunters operation.

An alleged ShinyHunters affiliate who had previously contacted BleepingComputer and other media about the FBI attack and a recent Clop ransomware gang data breach abruptly shut down their online messaging account.

Later, the ShinyHunters data leak site went offline, and the group's main representative also stopped responding to questions from the media, including BleepingComputer and Reuters.

It is unclear whether the sudden silence and shutdown of ShinyHunters-linked infrastructure are connected to Khader's reported detention.

However, on Thursday, a new ShinyHunters data leak site went online, suggesting other members continue to run the extortion operation.

BleepingComputer contacted ShinyHunters about Rey's reported detention but has not received a response.

The ShinyHunters gang has long been a thorn in the side of law enforcement, performing massive data theft attacks and extortion campaigns against organizations worldwide.

In recent years, the extortion gang has focused on Salesforce and other cloud SaaS environments, with campaigns linked to breaches at Google , Cisco , and PornHub .

The extortion gang commonly breaches third-party integration companies and uses stolen authentication tokens to access connected SaaS environments and steal customer data.

The extortion gang was also behind a massive data-theft attack on Instructure Canvas in May that caused significant platform outages. The company eventually reached an "agreement" with the threat actors to prevent the data stolen in a recent breach from being leaked online.

Over the years, numerous arrests have been linked to the ShinyHunters name, including suspects connected to the Snowflake data-theft attacks , breaches at PowerSchool , and the operation of the Breached v2 hacking forum .

Who is Rey?

The threat actor known as Rey has been linked to numerous data theft and extortion attacks over the past two years.

In January 2025, Rey was one of four threat actors who claimed responsibility for a breach of Telefónica's internal Jira ticketing system , where approximately 2.3GB of documents, tickets, and other data were allegedly stolen.

BleepingComputer previously reported that Rey and two of the other attackers were members of the then-new HellCat ransomware operation.

The threat actor was later linked to a wider series of attacks targeting Jira servers at organizations worldwide .

In February 2025, Orange confirmed that its Romanian operations suffered a cyberattack after Rey leaked approximately 6.5GB of stolen data . Rey told BleepingComputer at the time that he was a member of HellCat but had conducted the Orange breach independently.

Rey was later linked to the ShinyHunters extortion group and was seen with administrative privileges in Telegram channels operated by "Scattered Lapsus$ Hunters."

Scattered Lapsus$ Hunters was first seen in 2025 and claimed to consist of former members of the Lapsus$, Scattered Spider, and ShinyHunters cybercrime groups.

The group claimed responsibility for the September 2025 cyberattack on Jaguar Land Rover that forced the automaker to halt production for weeks and ultimately cost the company more than $220 million .

Rey was also linked to an ealier March 2025 breach of Jaguar Land Rover, with the threat actor leaking gigabytes of data, including Jira issues, source code, employee information, and development logs.

Rey leaking jaguar data

In November 2025, security journalist Brian Krebs reported that Rey was Saif Al-Din Khader after analyzing information obtained from infostealer logs and speaking directly with Khader over Signal.

Krebs reported that Khader said he was trying to distance himself from Scattered Lapsus$ Hunters and claimed he had been cooperating with law enforcement since at least June.

"I'm already cooperating with law enforcement," Khader allegedly told Krebs. "In fact, I have been talking to them since at least June. I have told them nearly everything. I haven't really done anything like breaching into a corp or extortion related since September."

Krebs said he could not verify those claims.

article image

Build your security blueprint for AI-powered attacks

Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.

Save your seat

RSS Feed Best Practices (2022)

Hacker News
kevincox.ca
2026-10-03 15:09:25
Comments...
Original Article

Posted

Last updated

These are some technical tips for publishing a blog. These have nothing to do with good content, just how to share that content. The recommendations are roughly in order of importance and have rationale for why they are that important.

Formats

People generally call feeds “RSS Feeds” but usually they aren’t specifically talking about RSS. RSS isn’t the only, or even the best format. Using a standardized format is critical to your feed being understood by the widest variety of readers and search engines.

You should use RSS 2 or Atom . These formats are very widely supported. Other common formats are earlier RSS standards and JSON Feed or Microformats h-feed . I would avoid using these—or even less common formats—as they are less widely supported.

If you don’t have a feed yet I would highly recommend Atom. The specification has much less ambiguity, so you are less likely to have compatibility issues with the wide variety of clients in use. The specification is also simpler and more clear overall. If you already have an RSS 2 feed there is little reason to upgrade.

A minimal Atom template is below. For full details see the spec . If you need an example you can look at my feed .

<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
	<title>{{FEED_NAME}}</title>
	<id>{{HOMEPAGE_URL}}</id>
	<link rel="alternate" href="{{HOMEPAGE_URL}}"/>
	<link rel="self" href="{{FEED_URL}}"/>
	<updated>{{LAST_UPDATE_TIME in RFC3339 format}}</updated>
	<author>
		<name>{{AUTHOR_NAME}}</name>
	</author>
	<entry>
		<title>{{ENTRY.TITLE}}</title>
		<link rel="alternate" type="text/html" href="{{ENTRY.HTML_URL}}"/>
		<id>{{ENTRY.PERMALINK}}</id>
		<published>{{ENTRY.FIRST_POST_TIME in RFC3339 format}}</published>
		<updated>{{ENTRY.LAST_UPDATE_TIME in RFC3339 format}}</updated>
		<content type="html">{{ENTRY.HTML}}</content>
	</entry>
</feed>

There is very little reason to provide feeds in multiple formats. If you have an Atom feed you don’t need to provide an RSS feed as well.

Changing feed format is safe. Very few readers will be confused if a feed switches between Atom and RSS. This can be done either by changing the feed at the same URL or by redirecting new a new URL. (Just be sure to update the content type )

Content Type

Be sure to set the Content-Type header properly.

  • Atom: Content-Type: application/atom+xml
  • RSS: Content-Type: application/rss+xml
  • JSON Feed: Content-Type: application/feed+json

You will see other values used in the wild, but these are the standard values and have the widest support.

Absolute URLs

Every URL in your feed should be absolute. While Atom has clearly specified how to resolve relative URLs they are rarely implemented correctly. In order to ensure that your feed can be understood by all readers use only absolute URLs (starting with https:// ).

This includes all < link > elements and the summary and body of posts (including in the HTML).

Discovery

On all of your blog pages and likely every page of your site you should include metadata to advertise you feed. This will allow readers and search engines to subscribe and become aware of your new content. This is as simple as providing the following in your HTML:

<link rel=alternate title="Blog Posts" type=application/atom+xml href="/feed.atom">

If you have multiple feeds you can advertise them all with appropriate titles.

<link rel=alternate title="All Posts" type=application/atom+xml href="/feed.atom">
<link rel=alternate title='Posts in the "Social" category' type=application/atom+xml href="/feeds/social.atom">
<link rel=alternate title="Comments on this Post" type=application/atom+xml href="/post/hello-world/comments.atom">

Make sure that you use the correct type for your feed. The examples provided are for Atom feeds.

Prefer to put the “most important” feed at the top. Many clients will preserve the order when presenting feeds to the user. This is subjective but typically would be a whole-site feed, then category feeds, then a comment feed for the specific page. If you offer your feeds in multiple formats I recommend only advertising one (either Atom or RSS 2). Including multiple links to the same content in multiple formats may confuse potential subscribers or leave them in analysis paralysis. (How do they know that the content is the same?)

You can validate that this is working correctly by putting your website URL into the W3C Feed Validation Service . If your links are set up correctly it should detect and validate your feed. Try out a few different pages to make sure that you have discovery working everywhere. Try your homepage, post lists page and an individual post page.

If it is difficult to modify the HTML a Link header in the HTTP response can be used. However, this isn’t as widely supported. Using HTML < link > tags is preferred for wider compatibility.

Link: /feed.atom; rel="alternate"; type="application/atom+xml"

You should also include a link with an RSS logo rss logo for users without another feed indicator.

HTTPS

HTTPS is key to security and privacy on the internet. Providing feeds over HTTPS ensures user privacy and ensures that your feed is not modified by a malicious actor.

  1. Reference all embedded media (such as images) over HTTPS. Many readers will run in a secure context where HTTP requests are not allowed.
  2. Provide the feed over HTTPS.
  3. Ensure that your self link is HTTPS.
  4. Redirect HTTP requests to HTTPS.
  5. Consider using Strict-Transport-Security .

Full Content

It is generally recommended to provide the full content of your posts in the feed. This is what most readers prefer. For RSS and Atom the < content > element should contain the full article. Atom also has a < summary > element in which to include a shorter summary for readers who prefer it.

Of course sharing full content in feeds is unacceptable to some publications due to the difficulty of monetizing these views. First, consider that some readers may leave if they can’t view the full content in their feed reader. Even if they don’t see your ads they may share your content with friends or on news aggregators. Likely it is still more valuable for you to have this reader than to lose them.

If your content is paid consider allowing users to generate private links by providing an auth token. For example /feed.atom?user=peruserauthtoken . You can also use basic auth like https://fred:peruserauthtoken@blog.example/feed.atom however this is supported by fewer readers than providing a token in the URL path or query string.

Entry IDs

Entry IDs are the primary way to identify and differentiate entries in your feed. If your entry IDs change or repeat, readers will receive duplicates or miss entries.

  1. Never change the ID of an existing article.
    • If you change your ID scheme, ensure that it only applies to new entries.
  2. Never reuse Entry IDs for different articles.
  3. Prefer to use article permalinks for Entry IDs.
  4. Prefer to make your Entry IDs globally unique across all feeds in existence.
    • Some readers will merge feeds together, unique IDs help ensure there are no issues.
    • The easiest way to accomplish this is to use a URL on a domain that you control. If that isn’t possible you can use a UUID such as urn:uuid:f4a3ca5b-5799-44e8-aaaa-e40728f037d3 .

Dates

Both Atom and RSS differentiate between time of publication (the time the entry first appeared in the feed) and the time of last update (the last time the entry was changed). Be sure to handle these correctly.

  1. Include a publication time.
  2. The publication time should never change. An entry can only be published once.
  3. Prefer making the publication time roughly match when the entry appeared on the feed. Some readers will ignore entries that were published in the far past.
  4. Strongly avoid having entries start appearing in the feed in a different order than their publication time suggests. (For example avoid having an item with a published time of 14:00 start appearing in the feed at 14:00 then have an item with the published time of 13:00 start appearing in the feed at 15:00. Some clients will be suspicious that they already know about the 14:00 item but don’t yet know about the “earlier” 13:00 item.) Another way of viewing this is that every new item that appears in a feed should have a published time that is later than all items previously in the feed. Some clients will ignore these “backdated” entries even if the published time is quite recent.
  5. Avoid future publication times. If a publication time is too far in the future many readers will ignore it as a bug.
  6. Update time should be greater than or equal to the publication time. New entries should have these two be the same.
  7. Update time should only change on significant updates. Slight formatting changes or typo fixes probably shouldn’t change the last update time. Most readers ignore the update time, some will resurface your article as “updated”.

Feed Title

The title of your feed is likely used by default in the user’s reader. Many readers have options to override the title, but it is extra work for the user and not universally supported. Try to pick a good title for your feed.

  1. Include context. The title is likely one of many feeds in their reader. For example call it “Kevin Cox’s Blog” rather than “Blog Posts”.
  2. Keep it succinct. The user is already subscribed, no need to advertise more. For example “John Smith” or “John Smith’s Photography Blog”. Not “John Smith — Ramblings on Photography every Tuesday and Friday, Cameras, Film and Development — Exclusive Content”.
  3. Avoid HTML special characters such as < > and & . In RSS it isn’t completely clear if you can include styling like < b > tags in your feed title. Very few readers will parse HTML and will almost always treat the title literally.

You can update your feed title at any time, but it may be confusing to users if it changes too frequently.

Styling

Feel free to use CSS in your feed! However, keep in mind that many feed readers don’t use modern browser engines and may be limited in what they can render. Additionally, many feed readers will sanitize your feed so uncommon elements and custom CSS may be partially or completely stripped. But don’t let that stop you! Using HTML and CSS can greatly improve the experience for users with good readers. Consider the following tips:

  1. Consider what will happen if any CSS doesn’t apply. For example if you set background: black; color: white and one of the two rules is stripped you will have unreadable text. In general prefer to make small adjustments rather than relying on CSS for dramatic changes.
  2. Prefer inline CSS style attributes to separate < style > blocks. They have wider compatibility.
  3. Prefer semantic elements such as < p > , < h1 > , < pre > and < code > over emulating their style on < div > s and < span > s.
  4. Don’t rely on JavaScript, almost no readers support it.
  5. Provide fallbacks for < audio > , < video > and < iframe > tags. Support isn’t common.
  6. Avoid form and input elements. Support is rare and incompatible.

Unfortunately there is no substitute for testing in various readers to see what works.

Ensure that the self-link for your feed is accurate.

<link rel="self" href="https://kevincox.ca/feed.atom"/>

This provides the following benefits:

  1. Allows you to move your feed. Some readers will update the feed URL if they get a permanent redirect and the redirect target contains a self link that points to itself.
  2. Improve cache hits. It is common for users to find slight variations of your feed URL. For example http: instead of https: , /feed vs /feed/ , www.example.com vs example.com , feed.atom?tracker=lookatme or ?category=rant&content=full vs ?content=full&category=rant . By providing a canonicalized self link you can merge these to reduce variance and increase you cache hit rate.
  3. Required for WebSub .
  4. If the user has a copy of the feed file they can subscribe to it. For example some feed readers will act as file handlers for feeds. If the file contains a self link then they can use that URL to subscribe and fetch updates. If the file doesn’t have a self link it isn’t possible to do that.

Caching

Feeds are followed by constant polling. This can create a decent amount of load on your server. Setting cache headers can help control the readers. If you don’t provide any guidance every client will pick their own value, which may be too fast or slow for your feed. If you make a suggestion some will follow it. Try to pick a reasonable value based on when you post. If your blog updates monthly then caching for an hour would make sense. However, if you are posting many times a day, a five minute cache may be more suitable.

Example cache headers:

  • 5 min: Cache-Control: max-age=300
  • 15 min: Cache-Control: max-age=900
  • 1 h: Cache-Control: max-age=3600

If you use scheduled posts and want to get very fancy you can vary the cache time based on when the next post will go live. But a static cache time is sufficient.

Conditional Requests

Support conditional requests on your feed. This makes polling more efficient for both you and your users.

Return an ETag and/or Last-Modified header. Then return a 304 response if the feed hasn’t changed. See HTTP conditional requests on MDN for more details.

WebSub

WebSub is a standard for real-time feed updates. Not only does it push your updates out faster, but it also reduces load on your server.

You can use a public hub or run your own. Note that a hub can modify or inject content into your feed, so be sure you trust the hub you pick.

I can’t find good generic setup instructions but the Google hub has basic instructions on their homepage. Maybe I’ll write a guide one day…

Bot Access

If you use any bot-blocking technology be sure to turn it off (or turn it way down) for your feeds. They are intended to be consumed by bots! Otherwise users will have trouble accessing your feed and will not know about your new content.

Many popular sites have problems here. I’ve written about this in the past . Make sure that you aren’t hurt by defaults of various services.

Categories

Categories are a reliable way to filter items in feeds. It is far better to let someone subscribe for one category—or all categories but one—than to lose a subscriber because they were annoyed by a subset of your content.

For Atom feeds adding categories is as simple as one element. For example this post contains the following markup in my feed:

<category term="RSS"/>
<category term="Guide"/>

For RSS 2 the syntax is just slightly different:

<category>Rant</category>

Some readers don’t support categories, so you may wish to consider generating different feeds for different categories or providing a URL parameter to your feed to filter by category. Personally I wouldn’t worry about this.

Changing URL

As much as possible you should avoid changing your feed’s URL. But if you need to do it here is how to do it without losing many subscribers.

  1. Make the feed available at the new URL in addition to the old URL.
  2. Make sure the self link of the new feed points at the new URL.
  3. If you use WebSub, start pinging your hub for both URLs whenever you post.
  4. Redirect the old URL to the new one with a 308 Permanent Redirect .
  5. If you use WebSub you should continue pinging the old URL for at least 3 months or the max subscription lifetime of your hub (whichever is greater).

Remember that that some subscribers will not move. Try to keep the redirect alive as long as possible. After an extended period of time you may consider replacing the redirect with a feed that has an entry informing readers of the new location. But an ounce of prevention is worth a pound of cure, try picking a good URL (that you control) from the start.

CORS

CORS or Cross-Origin Resource Sharing is a kludge to fix some holes in the original web security model. It adds restrictions to what requests web pages can make and controls what they can see about the response.

This is relevant for feeds as without opting-out of CORS browser-based readers that fetch feeds client side will not be able to access your feed.

For feeds the following headers are sufficient and secure for all feeds:

Access-Control-Allow-Origin: *

This means that your feed can be requested as a public resource (notably no cookies will be sent).

To test it out you can navigate to any third-party webpage (such as https://example.com ) then run the following Javascript in the developer console :

fetch("https://YOUR_FEED_HERE").then(r => r.text()).then(console.log, console.error)

If the content of your feed gets logged than you are all set. If you get an error then something has gone wrong.

Performance

While small and local feed readers tend to use simple poll rates many larger services use a variety of heuristics to determine when to check your feed. If you don’t support WebSub you should aim to respond to feed requests in less than 1s. If your feed is slower, especially if it is slower than 3 s polling will likely be slowed down and you readers will get updates slower.


The items below this point are relatively unimportant. They are a good idea if you are creating a product or feed generator but likely not worth the time if you are just making a feed for your own blog.

Summaries

Some readers will display short snippets as an article preview. Providing a summary in your feed gives them high quality content for a good user experience. If you don’t provide an explicit summary they will likely use your first paragraph or first couple of sentences of your main content.

Pagination is a useful tool for keeping your feed archive available while keeping the size of your recent items small. It is specified in RFC 5005: Feed Paging and Archiving . Unfortunately, few clients have support. However, adding support in your feed doesn’t have any downsides, so it is still a good idea.

Pagination is very easy. Just add the following link to your feed.

<link href="https://kevincox.ca/feed/2022-03-05.atom" rel="next"/>

Then on subsequent pages also include a prev link. Of course the last page won’t have a next link.

<link href="https://kevincox.ca/feed.atom" rel="prev"/>
<link href="https://kevincox.ca/feed/2022-01-21.atom" rel="next"/>

When deciding how large to make your pages remember that not all clients will support pagination, so you don’t want to move new entries off your first page too quickly. I provide the following recommendations. Note that these are just general rules and should be applied judiciously. For example if you post 20 times a day you probably don’t need to keep 7 days of content in your feed. Similarly, if your posts are only a paragraph or two you can probably keep a few more on the first page.

  1. Ensure your newest items are on the first page.
  2. Try to keep items in the feed for at least a day. Some clients check quite infrequently. If reasonable, keep items for at least a week.
  3. Avoid making the feed too large, well under a megabyte is recommended.

Apple Confirms iPhone 18 Pro Max AT&T Cellular Issues, Affected Devices Require Hardware Replacement

Daring Fireball
9to5mac.com
2026-10-03 14:21:14
In a statement to 9to5Mac, Apple said: We have identified an issue affecting a small number of iPhone 18 Pro Max users on the AT&T network that may cause a device to lose service and be unable to make calls. We released iOS 27.0.1 earlier this week and strongly encourage all iPhone 18 Pro...
Original Article

Apple has confirmed an issue causing a “small number” of iPhone 18 Pro Max devices on AT&T to lose cellular service.

Apple released software and carrier updates to prevent the issue from affecting additional users, but the company tells 9to5Mac that devices that have already lost service will require a hardware replacement.

In a statement to 9to5Mac, Apple said:

“We have identified an issue affecting a small number of iPhone 18 Pro Max users on the AT&T network that may cause a device to lose service and be unable to make calls. We released iOS 27.0.1 earlier this week and strongly encourage all iPhone 18 Pro Max users to update now, and are also issuing a carrier settings update today. These updates together are meant to help prevent this issue from occurring.”

The carrier settings update for the iPhone 18 Pro Max will download automatically over the next 14 days. However, users can manually trigger the download by going to Settings, choosing General, then About.

Apple explains that the software updates will help prevent future issues for iPhone 18 Pro Max users on AT&T. The software updates will not restore service on devices that have already lost it.

Apple says that iPhone 18 Pro Max users who have already lost service will need a hardware replacement. Those users should contact Apple Support or AT&T, or visit an Apple Store or AT&T location.

As we reported this morning , iPhone 18 Pro Max users affected by this problem can’t connect to AT&T’s network for calls, texts, and data. Instead, those users see “SOS” in their iPhone’s status bar. The connectivity issues aren’t impacting every iPhone 18 Pro Max user on AT&T.

Interestingly, the iPhone 18 Pro Max in the United States uses a Qualcomm modem . The iPhone 18 Pro, which uses Apple’s C2 modem, is unaffected by these AT&T connectivity issues.

Chance’s favorites:

Follow Chance : Threads , X , and Instagram

Add 9to5Mac as a preferred source on Google Add 9to5Mac as a preferred source on Google

FTC: We use income earning auto affiliate links. More.

Does Anybody Care About Data Breaches?

Internet Exchange
internet.exchangepoint.tech
2026-10-01 10:46:37
Data breaches are now a routine part of life, but people affected still have almost no way to get a remedy. Lucy Purdon asks what real consequences for companies could look like....
Original Article

By Lucy Purdon , also published in her newsletter The Prompt by Courage Everywhere

Many years ago at The Glass Room art exhibition in London, I leafed through thick white books, similar to old school telephone directories, containing every password stolen in a 2012 hack that exposed LinkedIn’s entire database. Artist Aram Bartholl alphabetized, printed and bound the list of 4.6 million passwords into 8 volumes to tell a story about data, inviting visitors to pick up the books and search for their own password. (Yes, mine was in there.)

4.6 million passwords seems a quaint number now due to the scale of today’s data breaches. It’s likely you yourself have recently received a message about a cybersecurity incident that has exposed your personal details held in a company’s system, and I’ll bet that wasn’t the first time- 4.4 million online accounts were reportedly exposed in the first 3 months of 2026 in the UK alone. For most people a data breach involves an email, telephone number or financial information like credit card details. Do you feel the consequences of such a breach are kind of left hanging, often downplayed and unclear? Do you feel confident the company in question has given you the information you need, beyond “soz, change your password”?

As Jamie Bartlett wrote in the excellent Substack post, “ What actually happens to your stolen data? ” your data goes on a “five stage, globe trotting, magical mystery tour”. He also wrote about how scams are becoming more sophisticated with the help of AI; those “phishing” scams are getting more convincing- and it all starts with a stolen email.

It gets worse of course. I have written before about the data breach from consumer genetic testing company 23andMe , resulting in the theft of millions of customer profiles including date of birth and ancestry information. Hackers advertised the data for sale and boasted it included around 1 million people of Ashkenazi Jewish descent, 100,000 of Chinese descent and “the wealthiest people living in the U.S and Western Europe” . Changing a password doesn’t touch the sides when your actual DNA is stolen.

In May this year, a cyberattack on the World Food Programme exposed the personal data of 600,000 Palestinian households in Gaza, including their names, ID numbers, and location. The weaponization of this information could prove deadly.

The UK’s National Cyber Security Centre describes data breaches as “a fact of modern life” . I don’t believe that means we accept the organizational carelessness, lack of security investment or the greed of collecting as much data as possible that so often leads to data breaches. Data breaches matter, they don’t seem to be taken seriously enough and we are often left in the dark with no meaningful remedy.

What actually is a data breach?

Organizations that collect and hold your personal data are bound by data protection law (where it exists) to keep it safe and secure. A data breach under the EU GDPR describes a situation where an organization has failed to keep it safe, leading to the destruction, loss, alteration, or - most relevant here- the unauthorized access to or disclosure of personal data.

Data breaches are increasingly the result of a cyber attack, but human error plays a major part. Just last month, the UK’s Metropolitan Police accidentally disclosed the emails of 140 people accusing the late owner of Harrods of sexual abuse by cc’ing them in an email update, rather than bcc’ing.

Under the EU (and UK) GDPR, organizations have the obligation to notify both the regulator and affected people of the breach, and provide remedy.

Tracking company responses

Ranking Digital Rights (RDR) evaluates the policies and practices of 26 of the world’s largest digital and telecommunications companies on how they uphold commitments to respect human rights such as freedom of expression and privacy, publishing an index of the findings. Since 2017, the index has included an indicator on company responses to data breaches as part of their methodology.

The indicator assesses three issues in line with data protection legislation: whether companies commit to notifying relevant authorities, whether they explain the process they will follow to notify people (data subjects) affected by a data breach, and whether they explain the steps they may take to address the impact of said breach.

Leandro Ucciferri, Deputy Director of RDR, has charted how the introduction of the GDPR in 2018 slowly made a difference to transparency around data breaches, but there are still major gaps in protections:

“In the 2017 RDR Index, only 3 of the 22 companies evaluated published some information about their policies to address data breaches (the companies were Telefónica, AT&T, and Vodafone). But we didn't see a notable improvement until 2019 [after the GDPR was adopted], when 10 of the 24 companies we evaluated published policies on this issue (five were digital platforms and five were telcos).

Fast forward to the latest assessments we published ( 2025 Big Tech Edition and 2026 Telco Giants Edition ), 19 of the 26 companies we evaluated disclose some information about their data breach policies, but still none of them reached the full score for the three concrete questions in this indicator.

Notably, giants like Google, Amazon, and TikTok, do not publish explicit policies and commitments to notify authorities and users, nor explain the processes they may take to mitigate the harms caused by a breach.”

Remedy is so often the weakest point of rights-based legislation and the gaps are glaring when it comes to data protection. Compensation is a grey area as harm connected to a specific data breach is difficult to prove, even when the breaches are so egregious and the impacts potentially infinite. Leandro points out that accessing remedy under GDPR is a high bar, with a data subject needing to demonstrate in court that a damage existed, there was infringement of GDPR and the direct causality of the damage suffered and the GDPR violation.

This relies on individuals knowing what has happened to their data in order to exercise their rights, which is impossible when it comes to engaging with increasingly opaque systems and so little disclosure or transparency from companies.

Fine!

Failing to secure data can lead to headline-grabbing fines for companies, but usually only when financial details are involved, which appears to be assessed as the most tangible harm. In 2020, British Airways was fined £20 million by the Information Commissioner's Office (ICO), the UK’s data protection regulator, after users were directed to a fraudulent site where hackers harvested data of 400,000 customers including credit card information. Capita Pensions was fined £14 million by the ICO (negotiated down from £45 million mind you) for a hack that exposed 6.6 million people’s financial information.

But harm is not just financial and damage may not be immediate. Data hangs around.

How this plays out IRL: my Substack experience

If you have a Substack account you may, like me, have received an email from Substack CEO Chris Best on February 5th describing, in the most casual terms, that Substack was hacked and the email and phone number connected with your Substack account was “shared without your permission”.

“This sucks,” says Chris. It’s OK though!, “Importantly, credit card numbers, passwords, and financial information were not accessed.”

I have since been inundated with spam emails, including one claiming to be from the ShinyHunters criminal group, the notorious hackers behind some of the biggest data breaches of the past 5 years. Their email directly references the Substack data breach as the source of obtaining my information and tries to extort me by claiming they have videos of me watching pornography and “pleasuring myself” and will release these videos unless I pay $2000 in Bitcoin.

It’s not really ShinyHunters of course and this kind of “sextortion” scam is increasingly common.

Nonetheless, is Substack warning their users about the direct correlation between the breach and these attempted extortion attempts from a scary criminal gang, perhaps pointing to or supporting the work of fact checking groups exposing the scam?

No.

In response to my email complaint to Substack that my data has not been handled properly I am told that there is “no evidence that malware was installed via Substack” - the answer to a question I was not asking and a clear deflection.

Not satisfied with this response, I complained to the ICO, as is my right when there is a concern that an organization has not handled personal information properly. I am given a case number and informed I will receive a response within… 40 weeks!

Actually 6 weeks later, I received the decision that the ICO will not take the complaint further as Substack “has handled the matter in line with its data protection obligations by informing you of the data breach.”

So…that’s it. Notify those who bear the brunt and shift the onus onto the user to change passwords, monitor emails for phishing attempts and scams and put up with the torrent of spam emails, never being sure where our data is and what it is being used for, the next scam or trick around the corner.

Hardly reassuring.

A game of consequences, anyone?

I would opine that organizations often collect way more data than they need, because it’s valuable, so there is more data to breach. Databases are often poorly secured; procedures for dealing with data breaches are shockingly lax (the Substack breach reportedly went undetected for 4 months); companies often take ages to disclose and there are very few consequences (the recent ICO investigation into ACRO, a company that handles criminal records, is a jaw-dropping insight into abysmal cybersecurity failings). Leandro from RDR sums it up,

“My main concern at the moment is whether we've reached a point of apathy. Sure, some companies are receiving fines after being investigated by data protection authorities, but that's simply another cost of conducting business. And at the same time, the people affected may feel powerless, since they keep receiving news about new ways in which their data was exposed, likely multiple times in a given year. Even looking at the situation in a handful of European countries (Spain, France, Germany, the Netherlands, and Ireland), there were a combined total of more than 64000 data breach notifications to the national data protection authorities in 2025. The scale of this issue doesn't seem to be slowing down at all.”

As AI demands more data to train models and agentic AI requires more access to our personal data, what can we expect in the future?

“When it comes to the current conversation involving “AI” systems, as long as companies’ business models rely on extracting as much data as possible, we can expect them to face security incidents that end up exposing personal information.”

In the face of this we need stronger protections not less but intense lobbying from tech companies is steering our rights-based legislation in the wrong direction. The EU Digital Omnibus, an initiative to “streamline” EU digital legislation is perceived by many civil society actors as a potential dilution of safeguards established by the GDPR, the ePrivacy Directive, and the AI Act.

We need to bring ideas to the table on what we expect remedy to look like - not just fines but actual consequences and repercussions for companies to mitigate any potential harms, like being a victim of identity theft, targeted with scams, or worse. How about a company fined for a major breach also cannot collect any consumer data for a month? Companies must track the data breached and provide you with weekly updates about where it is and who has it? Then there might be more of an incentive to protect the data we often have no choice but to hand over.

Class action lawsuits may start to bite; over in Kenya, subscribers of the telecommunications company Safaricom won thousands of dollars in compensation over a large scale data-breach where the High Court found Safaricom failed to secure data and breached the constitutional right to privacy.

Let me know your experiences of data breaches, and your ideas about how companies should provide remedy!


Support the Internet Exchange

If you find our emails useful, consider becoming a paid subscriber! You'll get access to our members-only Signal community where we share ideas, discuss upcoming topics, and exchange links. Paid subscribers can also leave comments on posts and enjoy a warm, fuzzy feeling.

Not ready for a long-term commitment? You can always leave us a tip .

Become A Paid Subscriber


HRPC.io is Back Online

The website for the Human Rights Protocol Considerations (HRPC) research group at the Internet Research Task Force is back online, and features a short documentary about how the technical design of the internet relates to human rights. The HRPC studies how internet standards and protocols enable, strengthen, or threaten human rights, especially freedom of expression and assembly. IX's Mallory Knodel chairs the group, and the site is a good place to start for anyone who wants to learn more and get involved with this work.

🚨

Stop press! Do you enjoy our links? Links are now available to paid subscribers only. Become a paid subscriber today.

Danish university DTU breach exposes data of up to 200,000 people

Bleeping Computer
www.bleepingcomputer.com
2026-10-03 10:35:20
The Technical University of Denmark (DTU) says information belonging to up to 200,000 users may have been exposed after hackers accessed its identity and access management system and downloaded a large amount of data. [...]...
Original Article

Danish university DTU breach exposes data of up to 200,000 people

The Technical University of Denmark (DTU) says information belonging to up to 200,000 users may have been exposed after hackers accessed its identity and access management system and downloaded a large amount of data.

​The university says the attacker used compromised credentials to log into DTUBasen, its identity and access management (IAM) system, allowing access to more than two decades of user data.

In a disclosure on Friday, DTU confirmed that it cannot “determine precisely what information was downloaded or how many people have been affected.”

However, the Danish university notes that DTUBasen stores information for nearly 40,000 active users and around 160,000 former users.

Next of kin data also exposed

Potentially exposed information for current users includes Danish civil registration numbers (CPR), full names, home addresses, and profile pictures, as well as work email addresses, job titles, office locations, and other employment-related details.

The dataset also contained the names, relationships, and telephone numbers of users’ next of kin, when provided by active users.

DTU notes that in the case of former users, details about home addresses, profile pictures, and information about next of kin are automatically deleted after six months.

“This is a serious attack on DTU, and we deeply regret the uncertainty it is causing for the people whose information may have been affected,” says University Director Bjarke Bak Christensen.

“Our first priority has been to establish the extent of the attack, limit its consequences, and ensure that those affected are notified and know what steps to take,” the director added.

DTU warns that cybercriminals could use the exposed CPR numbers and other personal data for identity fraud and to make phishing attacks more convincing.

Not all notified directly

Potentially impacted individuals will be notified through e-Boks, the official mailbox system that DTU uses for sharing documents and notices with students and staff.

However, the university says it will notify all current and former employees, but not all current and former students whose CPR numbers are held by DTU.

“DTU only holds CPR numbers for a small number of guests and external partners and does not hold CPR numbers for next of kin whose contact details have been registered in DTUBasen,” the organization says .

The public dicslosure is part of DTU’s effort to reach potentially affected individuals it cannot contact directly, and the university is urging people to share it with former employees, students, guests, and external partners.

The organization says that anyone who has been an employee, student, guest, or external partner of DTU since 2003 may be affected by the data breach.

They are advised to be cautious of emails, text messages, and phone calls from individuals who appear to know about their connection with DTU or have access to personal information about them.

Passwords and sensitive information should not be disclosed in replies to unexpected communications, and sudden authentication requests or logins should be treated as suspicious.

Additionally, it is recommended to change the passwords for any other services that use the same credentials as the DTU account and place a credit alert on the affected CPR number.

article image

Build your security blueprint for AI-powered attacks

Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.

Save your seat

California’s new laws target workers’ biggest fear of AI taking their jobs

Guardian
www.theguardian.com
2026-10-03 10:00:00
The state, which is home to many AI companies, is one of the first to roll out workplace regulations targeting the technology California’s laws aimed at protecting workers from the impacts of artificial intelligence could pave the way for broader workplace safeguards across the US as calls for regu...
Original Article

California’s laws aimed at protecting workers from the impacts of artificial intelligence could pave the way for broader workplace safeguards across the US as calls for regulating the technology mount. As the federal government goes hands-off on AI, California is taking the reins to address workers’ biggest fears.

On Thursday, California governor Gavin Newsom signed a suite of new laws that ban bosses from relying entirely on AI to decide whether to fire workers, using it to predict employees’ emotional states or collecting neural data, meaning the information from electrical signals of someone’s brain or nerves. They also require companies to notify workers if layoffs were caused by AI and prohibit AI surveillance in workplace bathrooms.

The new laws come as workers increasingly worry whether AI will take their jobs, lead to discrimination and increase workplace surveillance. Unions, worker advocates and even some lawmakers pushed for the new rules, a regulatory shift for a technology that has largely developed unchecked. California , home to many of the leading companies developing AI, represents one of the first states to roll out a sweeping set of workplace regulations targeting the technology.

“It’s a turning point,” said Lorena Gonzalez, president of the California Federation of Labor Unions, AFL-CIO, who’s been helping leaders across the country write regulations. “It’s really the first time we’re seeing California workers showing the country that we don’t have to accept [this].”

Other states that have recently passed individual laws aimed at AI’s use in the workplace include Colorado, Connecticut, Illinois and Texas, narrower in scope than California’s. And more bills across the county are being lined up for consideration, Gonzalez said.

California’s laws aim to target workplace surveillance measures like heat maps that track employees’ movements, including how long they spend in the bathroom, or having their emotional states monitored. Amazon warehouse workers have previously complained about being timed on their bathroom breaks, for example. And at Kaiser Permanente, nurses have said automated systems rated their tone of voice in patient interactions.

But the laws could also prevent future unexpected harms.

“We don’t know all the places companies are using AI, and that is and should be scary,” Gonzalez said.

Part of the strategy for worker advocates and union groups like the California Federation has been following AI companies’ latest products. “If it’s being sold, that’s a good indication” it could be in use, Gonzales said. The federation also plans to use the momentum to revive issues like requiring employers to disclose when they’re using AI in the workplace, which was a bill that died in the state’s assembly appropriations committee this year.

California’s new laws are a key step in gaining regulatory ground for worker advocates, said Robin Feldman, director and founder of AI Law & Innovation Institute at The University of California College of the Law, San Francisco. Still, the statutes are somewhat limited in how they are implemented.

“The bills have no private enforcement,” Feldman said. “In other words: workers can’t sue. Only the government can enforce the laws.”

The new regulations come amid a backdrop of record AI spending at big tech companies along with massive job cuts. But workers have started pushing back. In June, Meta paused a program that tracked workers’ computer activities to train its AI models. And one month later, dozens of employees filed a lawsuit claiming the tech company’s AI tools targeted those with disability accommodations or on medical or parental leaves for layoffs.

Meanwhile, safety concerns, including fears that AI could destroy humanity, have been garnering more attention from lawmakers across the country and prompted OpenAI and Anthropic to call for slowing the pace of development.

skip past newsletter promotion

“Workers are increasingly part of that movement, speaking up about the fear of job loss and the dehumanizing experience of being surveilled and controlled by an algorithm,” said Annette Bernhardt, senior tech policy adviser at UC Berkeley Labor Center.

While the new laws “have teeth”, according to Danielle Ochs, shareholder at employment law firm Ogletree Deakins’ San Francisco office, it’s unclear how sweeping the change will be. Ochs said employers generally aren’t grappling with the AI uses outlined in the new regulations and instead are more interested in how to responsibly implement AI across their systems.

“Having 10 hoops you have to jump through per tool is not reflective of reality,” she said. It would be better to have “guardrails that are more aligned” with employers’ wider use of AI rather than focused on specific tools or uses.

She says opponents worry that the new rules could unexpectedly prohibit helpful AI that might, for example, ensure truckers don’t fall asleep at the wheel.

While it’s too soon to gauge how effective the laws will be, worker advocates believe the momentum is moving in a positive direction. Gonzalez said the measures are only the beginning of addressing AI’s potential impacts.

“We have so much work to do,” she said. “But this should give us all hope we can win … against the tech lobby, against big corporations, because we are the majority.”

Two American Airlines Flights End Up with the Same Flight Numbers

Hacker News
aviationa2z.com
2026-10-03 15:07:18
Comments...
Original Article

WASHINGTON, D.C.- American Airlines (AA) has experienced another unusual air traffic control communications issue after two American Eagle flights with the same callsign operated simultaneously between Philadelphia International Airport (PHL) and Rhode Island T.F. Green International Airport (PVD).

The incident involved PSA Airlines, operating under the American Eagle brand, and highlighted how duplicate flight numbers can create confusion for controllers and pilots. According to OMAAT , this was the second similar occurrence involving American Airlines within two days.

Two American Airlines Flights End Up With the Same Flight Numbers Once Again
Photo: By Alan Wilson from Stilton, Peterborough, Cambs, UK – Bombardier CRJ-700 ‘N724SK’ American Eagle, CC BY-SA 2.0, https://commons.wikimedia.org/w/index.php?curid=66991728

Duplicate Callsigns Created an Unusual ATC Situation Again

Every commercial flight uses a unique callsign when communicating with air traffic control. For most airlines, the callsign combines the airline identifier with the assigned flight number. This system allows controllers to identify and communicate with individual aircraft safely and efficiently.

On August 12, 2026, two regional jets operating as American Airlines Flight 5083 ended up airborne at the same time while sharing the same callsign. Although the flights carried an American flight number, they were operated by PSA Airlines , whose operational callsign is Bluestreak because the carrier operates under its own Air Operator Certificate.

Flight AA5083 serves the route between Philadelphia International Airport (PHL) and Rhode Island T.F. Green International Airport (PVD) in both directions. American assigns the same flight number to the outbound and return services, with the same aircraft typically operating both legs.

On this occasion, however, the inbound aircraft to Providence arrived more than an hour behind schedule. The scheduled return flight departed on time using a different aircraft, resulting in two aircraft operating simultaneously with the identical callsign.

The unusual timing meant both regional jets eventually crossed paths while communicating on the same air traffic control frequency.

Philadelphia International Airport
Photo: Philadelphia International Airport

Controllers Quickly Managed the Situation

Despite the operational complication, the situation remained under control throughout the encounter.

Air traffic controllers quickly recognized the duplicate callsigns and informed both flight crews about the conflict. To avoid confusion, controllers distinguished the aircraft by referring to one as the inbound flight and the other as the outbound flight during radio communications.

The coordinated response ensured both aircraft received the correct instructions while maintaining safe separation. The incident demonstrated the importance of vigilant air traffic controllers and disciplined cockpit communication when unexpected operational issues arise.

Although no safety event occurred, the duplicate callsigns created unnecessary complexity in an environment where precise communication is critical.

Photo: CAIR

Why Airlines Reuse Flight Numbers

The situation naturally raises questions about why airlines assign the same flight number to flights operating in opposite directions.

The answer lies largely in the scale of airline operations and industry numbering limitations.

Commercial flight numbers are generally restricted to four digits because of longstanding industry and equipment limitations. Major U.S. airlines such as American operate between 6,000 and 7,000 flights every day, making available flight numbers a limited resource.

Additionally, not every number is available for standard passenger flights.

Typically:

  • Flight numbers 1 through 2949 are assigned to mainline operations.
  • Flight numbers 2950 through 6099 are commonly reserved for regional services.
  • Flight numbers 9000 through 9999 are generally used for ferry or repositioning flights.

Airlines must also reserve thousands of additional numbers for codeshare agreements with partner carriers.

As a result, airlines often assign the same flight number to outbound and return services, especially on regional routes where the same aircraft is expected to complete both segments.

Under normal operating conditions, this practice creates no issues because only one aircraft carrying that callsign is airborne at any given time.

Photo : Richard Silagi | Wikimedia Commons | https://commons.wikimedia.org/wiki/File:American_Eagle_(Compass_Airlines)_Embraer_E175_N215NN_at_SFO_April_2017.jpg

Dispatch Procedures Are Designed to Prevent Conflicts

When operational disruptions cause both flights with the same number to operate simultaneously, dispatch procedures are intended to prevent duplicate callsigns from reaching air traffic control.

Industry practice allows dispatchers to create what is commonly known as a stub .

Rather than changing the public flight number seen by passengers, dispatch modifies the flight identifier filed with the FAA . A suffix letter is added to distinguish the aircraft during ATC communications while passenger-facing systems remain unchanged.

For four-digit flight numbers, the first digit is typically removed before the suffix is added because system limitations prevent longer identifiers.

Experienced airline operations personnel have suggested that this step may not have been completed before the flight plans were transmitted, allowing both aircraft to appear with identical callsigns.

If the procedure had been followed correctly, air traffic control would have received two distinct identifiers, eliminating the communication conflict.

American Airlines Boeing 737
Photo: Cado Handerson

Second Similar Event May Prompt Operational Changes

This was not an isolated occurrence.

A nearly identical duplicate callsign event involving American Airlines was reported only one day earlier, suggesting the issue may reflect a broader operational process rather than a single oversight.

While controllers successfully managed both situations without incident, repeated occurrences increase radio workload and create avoidable opportunities for misunderstanding during busy phases of flight.

Industry observers will likely watch closely to see whether American Airlines reviews its dispatch procedures or introduces additional automated safeguards to prevent duplicate callsigns from being filed with air traffic control.

Strengthening those internal processes could reduce the likelihood of similar events occurring in the future while preserving efficient communication between pilots and controllers.

Photo- Wikipedia

Bottom Line

Two American Eagle regional jets operating as American Airlines (AA) Flight 5083 simultaneously entered the airspace with the same ATC callsign after scheduling disruptions placed both aircraft in service at the same time.

Controllers handled the situation professionally by clearly distinguishing the inbound and outbound flights, preventing confusion. The event highlights the importance of proper dispatch procedures, especially when duplicate flight numbers are used, and may encourage operational improvements to avoid similar incidents.

Stay tuned with us. Further, follow us on social media for the latest updates.

Join us on Telegram Group for the Latest Aviation Updates. Subsequently, follow us on Google News

BSides Orlando 2026 Badge

Lobsters
github.com
2026-10-03 14:57:32
Comments...
Original Article

BSides Orlando 2026 Badge

Rendering of the front of the BSides Orlando 2026 badge Rendering of the back

The electronic badge for BSides Orlando 2026, themed around retro-futurism: a 1960s World's Fair vision of tomorrow, with Lil Chompy the alligator at the wheel of a flying car.

  • CH32V003 RISC-V microcontroller, powered by one AAA cell through a TPS61021A boost converter.
  • Three SK6812MINI-E RGB LEDs: the flying car's headlights, and the fingerprint scanner indicator.
  • Four self-blinking LEDs: the tower spire and the geodesic sphere.
  • Two capacitive touch pads: the fingerprint, and Lil Chompy's head.
  • SAO connector: the badge is an SAOv3 host on I2C, and GPIO1/GPIO2 carry a 115200-baud serial console.
  • Serial console: runs "Lil Chompy and the World's Fair", a small text adventure that doubles as a CTF challenge.

Spoilers: the firmware source contains the CTF solutions and flags.

Repo layout

Path Contents
hardware/ KiCad project ( bsidesorl-v1 ), symbol and footprint libraries, and artwork sources
hardware/EasyEDA_Gerber_bsidesorl-v1_2026-08-31.zip Gerbers as ordered from JLCPCB
hardware/jlcpcb/ BOM and CPL as ordered from JLCPCB
hardware/mk_easyeda.py Writes an EasyEDA Pro-importable copy of the board
firmware/ Badge firmware (PlatformIO + ch32fun ); see firmware/README.md

Firmware quick start

# saov3-lib's public repo isn't published yet
git submodule update --init

# Build
cd firmware
pio run

# Copy firmware to picorvd-for-badges probe
cp .pio/build/bsorl26/firmware.bin /Volumes/PICORVD/

For production, badges are flashed by the standalone programmer build of picorvd-for-badges (an RP2040 probe). It shows up as a USB drive named PICORVD: copy firmware.bin onto it, and set uart_selftest = 0xB5 in its CONFIG.TXT so it triggers the badge's post-flash self-test over the SWIO pin. Its factory log can be copied off the same drive as LOG.CSV .

Building your own "programmer SAO"

programmer SAO board

Note that in the picture, the SAO connector is connected on G14-H16. The traces on the back were cut between G and H to avoid bridging them. The SAO connector is oriented so the top is facing downwards in this photo, meaning 3V3 is on G16 and GND is on H16.

For this, you need a Raspberry Pi Pico or compatible board. I used a WaveShare RP2040-Zero board for the production programmer SAOs that were used in the soldering village at BSides Orlando, since they're smaller and cheaper than a Raspberry Pi Pico. To use as a standalone programmer (where the programmer receives power from the connected target badge), only three wires are needed: GND, 3v3, and the Pico's GP4 -> badge's RXD (aka PD1/SWIO). On this SWIO connection, you should also attach a 1KΩ pull-up resistor to 3v3. This can easily be done on a breadboard with jumper wires.

You'll need to build the picorvd-for-badges firmware for the probe. The exact version used on the programmer SAOs at the conference can be found pre-built here: https://github.com/bsidesorlando/2026-badge/releases/download/v1.0.0/pico_rvd_factory.uf2

To build it yourself for the WaveShare RP2040-Zero board:

git clone https://github.com/kjcolley7/picorvd-for-badges.git
cd picorvd-for-badges
cmake -B build-zero -G Ninja -DPICO_BOARD=waveshare_rp2040_zero
ninja -C build-zero pico_rvd_factory

The above commands create build-zero/pico_rvd_factory.uf2 , which should be uploaded to the probe by putting it in BOOTSEL mode and then copying it to the USB Mass Storage device it exposes. Then, the probe will reboot into the picorvd firmware and the USB Mass Storage device will re-appear with the name "PICORVD". Now, it's ready to accept firmware and configuration. You can upload the BSORL badge firmware from .pio/build/bsorl26/firmware.bin (or the pre-built one from here: https://github.com/bsidesorlando/2026-badge/releases/download/v1.0.0/firmware.bin ) by copying it into the PICORVD volume (which will disappear and reappear as the probe reboots), and you can also edit the probe's CONFIG.TXT to add uart_selftest = 0xB5 (so the probe tells the BSORL badge to enter selftest mode after programming). At this point, the probe is fully ready to go, either in standalone mode or tethered mode.

Standalone mode

This is the mode that was used at BSides Orlando for programming all of the attendees' badges in the soldering village. It works by powering the probe board from the target badge itself, and it immediately attempts to program the connected target upon boot. The badge needs to be powered on so the probe itself can be powered.

Tethered mode

This mode works by leaving the probe connected to a host PC over USB. If you use pico_rvd.uf2 instead of pico_rvd_factory.uf2, the probe won't automatically program connected badges. Rather, you'll need to manually issue the factory command to the probe over its UART interface. In the pico_rvd_factory.uf2 build, it will automatically start in factory mode, ready to program any connected badges.

NOTE : Do NOT connect a probe with USB power to a badge while it is switched on. The power switch should be in the OFF position before connecting a tethered probe. The badge and the probe will likely have slightly different values than exactly 3.3v, so there will be some current leakage. Worst case, it could damage the AAA battery, causing it to leak.

Credits

PCB and firmware by Kevin Colley. Art by Shep.

Getting the most out of Opus 5.5 in Claude and Claude Code

Hacker News
claude.dev
2026-10-03 14:29:30
Comments...
Original Article

Opus 5.5 works well with the way you already use Claude. A few things behave differently, though: it works for longer on its own, it tells you plainly what it did, and it thinks before every reply. This guide covers how to work with Opus 5.5 in Claude apps and Claude Code, including how to prompt the model, steer a long run, and check your results.

TRY THIS FIRST

Three things to try in your first session with Opus 5.5

  1. Hand over the whole task. Say what “done” looks like and when you want it to stop and ask. Then let it work.
  2. Delete “think carefully” lines. Opus 5.5 already thinks before every reply.
  3. When a long run ends, read what it needs from you first.

1. HOW TO ASK

Say what “done” looks like, then let it run

What to do. Give the whole task in one message. Name the finish line, like “the tests pass” or “every endpoint is migrated.” Then let it cook.

Why it matters on Opus 5.5. Opus 5.5 keeps going on long, multi-part work better than Opus 5 did. Compared to prior Opus models, its biggest gains are on multi-step work, like carrying a change through a large repository until the tests pass. Early testers had it run long coding tasks for hours with little oversight. With a clear finish line, it knows when it’s done.

How. In Claude Code, for example:

PROMPT

Migrate the payment endpoints from the old client to the new one.
Done means: every endpoint uses the new client, the old client is deleted, and the test suite passes.
Stop and ask me only if a test fails for a reason you can't explain.
The example prompt split into three labelled boxes: the whole task (migrate the payment endpoints from the old client to the new one), the finish line, highlighted (every endpoint uses the new client, the old client is deleted, and the test suite passes), and when to stop (only if a test fails for a reason you can’t explain). Footer: “Give the whole task in one message. Name the finish line. Then leave it alone.”
FIG A One message: the whole task, the finish line, and when to stop.

Stop telling it to “think hard”

What to do. Remove “think carefully,” “think step by step,” and similar lines from your prompts and your saved instructions.

Why it matters on Opus 5.5. Opus 5.5 always thinks before it replies, and it decides how much. You don’t need to ask it to think. In our testing in a chat product, removing a “think carefully” line made replies start sooner, with no clear drop in quality.

How. Delete the line. For a quick answer to a simple question, say so: “Answer directly.” To change how much it thinks in Claude Code, change effort.

Add to a running task

What to do. If you remember something mid-run, you can type a follow-up while it works.

Why it matters on Opus 5.5. Runs are longer now, so a restart costs more.

How to do it. In Claude Code, type the message and press Enter while Claude works, for example, “Also keep the old endpoint names as aliases.”

For design work, name the styles you don’t want

What to do. When you ask for a page, an app, or an artifact, list the design habits you want left out.

Why it matters on Opus 5.5. With no design direction, Opus 5.5 falls back on a few default styles. A general instruction like “avoid a generic look” mostly swaps one default for another. A list of specific patterns works much better.

How. Name the patterns:

PROMPT

Build a personal website with placeholder content.
Don't use a cream or off-white background, italic accent words in headings, numbered "01 / 02 / 03" section labels, monospace labels, or pill-shaped buttons.

Then look at what it chose instead. If you don’t like that either, add it to the list and ask again.

2. STEERING A LONG RUN IN CLAUDE CODE

Tell it which stops you want

What to do. Put a short rule in your CLAUDE.md file about when to stop and ask, and when to keep going.

Why it matters on Opus 5.5. Opus 5.5 keeps you posted as it works. On a long task, it sometimes stops to report instead of going on: a summary that names the next step without taking it, an offer to continue, or a list of choices that don’t block the work. It follows instructions that name these stops. Name the stops you want, too.

How. Add this to CLAUDE.md, and edit it to fit your project:

PROMPT

When a step doesn't need my input, keep going. Put status notes in the same message as your next action.
Stop and ask only when you can't continue without me, or before anything destructive: deleting data, force-pushing, or changing anything outside this repository.
A CLAUDE.md card titled “Tell it which stops you want” with two boxes. Keep going: when a step doesn’t need my input, keep going, and put status notes in the same message as your next action. Stop and ask: only when you can’t continue without me, or before anything destructive: deleting data, force-pushing, or changing anything outside this repository. Footer: “Edit it to fit your project. Keep permission prompts on for destructive commands too.”
FIG B The CLAUDE.md rule: when to keep going, and when to stop and ask.

If a run stops with “Want me to continue?” reply “continue.” If that happens often, the rule above will help.

A rule to keep going means fewer stops, so keep your own check before anything risky or hard to undo. The last line of the rule above does that. Keep permission prompts on for destructive commands too.

For pair programming, you may want the opposite: a one-line plan before it starts and a short recap at the end. Say that in your CLAUDE.md instead. Opus 5.5 follows either one.

Ask it to split big work across subagents

What to do. For an audit, a migration, or a review across a large codebase, ask Opus 5.5 to split the work across subagents and check each result.

Why it matters on Opus 5.5. Early testers had Opus 5.5 coordinate parallel subagents on long audits and migrations, with little oversight.

How.

PROMPT

Audit every service in services/ for the retry bug in the linked issue.
Give each service to its own subagent. When a subagent reports back, check its evidence before you accept it.
Finish with one table: service, affected yes or no, and the evidence.
Diagram titled “Give each service to its own subagent.” The prompt “Audit every service in services/ for the retry bug in the linked issue” fans out to four subagents. Their reports join at a “Check its evidence” step (“When a subagent reports back, check its evidence before you accept it”), then an arrow leads to “Finish with one table,” an empty table with the columns Service, Affected yes or no, and The evidence.
FIG C Fan out to subagents, check each one’s evidence, then finish with one table.

Keep the task list in a file

What to do. For a run that will take a while, ask Opus 5.5 to keep its task list in a file and update it as it goes. Then read the file, not the scrollback, to see where the run is.

Why it matters on Opus 5.5. Runs are longer now. A long run fills the context window, and Claude Code then summarizes older turns. A list in a file survives that, and it shows you at a glance what’s done and what’s left.

How. “Keep a checklist in TASKS.md. Tick each item when it’s done, and add anything new you find.”

3. CHECKING THE RESULT

Read what it needs from you first

What to do. When a long run ends, look first for anything Claude is waiting on you for, like a decision it left open or a change it wants you to approve. Then read the rest of Claude’s summary.

Why it matters on Opus 5.5. Opus 5.5 reports on its work more clearly than Opus 5. Its updates and its final summary say what it did, what it found, and what it needs from you, in plain language.

How. To change the summary’s format, say so in CLAUDE.md, for example, “End every run with three headings: Blocked on me, Changed, Found.”

Ask it to review the code

What to do. Ask Opus 5.5 to review a diff or a pull request before a person does.

Why it matters on Opus 5.5. One early tester said Opus 5.5 at its lowest effort caught more bugs than Opus 5 at high effort, with fewer false alarms. It also explains its changes in plain language, so its pull request descriptions are easier to review.

How. Feed this prompt to Claude:

PROMPT

Review the diff on this branch against main.
List only problems you'd block the merge for. For each one, give the file and line, why it's wrong, and how to show it fails.

Ask it to mark what it couldn’t confirm

What to do. For research and analysis, ask it to say what it couldn’t find or couldn’t check.

Why it matters on Opus 5.5. “I couldn’t find this” is worth reading, and asking for it makes it easy to find.

How. Add “Mark anything you couldn’t confirm, and say where you looked” to the request. This works in a Claude research report and in Claude Code.

4. IN CLAUDE APPS

First, check that the model picker says Opus 5.5.

What to do. Attach the chart, diagram, screenshot, or slide. Don’t retype the numbers.

Why it matters on Opus 5.5. Opus 5.5 reads charts, diagrams, and screenshots more accurately than Opus 5, and it needs no extra steps to do it. It’s also better at meaning that depends on where things are in the image: which boxes an arrow connects, what changed between two versions of a diagram, or when a meeting starts and ends in a calendar screenshot.

How. Attach the image and ask a specific question: “Which of these services call the billing API directly?”

Ask it to check a long document

What to do. Give it a long plan, report, or deck, and ask it to find mistakes.

Why it matters on Opus 5.5. Opus 5.5 pays more attention to detail than prior Opus models. In our testing, it caught a date that fell on the wrong weekday in a long planning thread, and a chart that didn’t match the numbers in a deck.

How. Submit the prompt: “Check this deck for anything that contradicts itself: numbers, dates and names. Quote each problem and say where it is.”

Ask for the finished file

What to do. When you want a spreadsheet or a document, ask for the file, not an outline.

Why it matters on Opus 5.5. The spreadsheets and documents Opus 5.5 makes need less editing than Opus 5’s before you share them.

How. “Make this a spreadsheet I can share: one row per vendor, with columns for cost, contract end date and owner.”

In a project, say when answers are settled

What to do. If follow-up questions in a long chat feel slow, add an instruction that earlier answers are settled.

Why it matters on Opus 5.5. In a long chat, Opus 5.5 sometimes goes back over an earlier answer while it thinks about a short follow-up. That slows the reply.

How. Add this to the project’s instructions:

PROMPT

Once you have answered something, treat that answer as done. Focus on what I'm asking now, and don't go back over an earlier answer unless I ask about it or point out a problem with it.

Leave it out of projects for long analysis, where a later step can show a mistake in an earlier one.

5. WHEN A MESSAGE IS FLAGGED

Opus 5.5 is the first Opus model to launch with Fable-level bio and cyber safeguards. In Claude apps and Claude Code, most flagged messages move to an older model, and your work goes on there. Finding security vulnerabilities in source code is allowed, and everyday health and educational questions should still work. These safeguards can sometimes flag legitimate work, and we’re tuning them to cut down on incorrect flags. If you’re switched, here’s what you’ll see and what to do.

In Claude apps

What you see. A notice that starts with “Switched to” and the name of an older model. Claude answers on that model, and the chat stays on it.

What to do.

  • To go back to Opus 5.5, choose it in the model picker. If the earlier message is still in the chat, it may be flagged again. Starting a new chat avoids that.
  • To be asked first, go to Settings, then Capabilities, and turn off “Switch models when a message is flagged.” You’ll see a “paused” card with your options.

The check covers everything in the conversation, including files and search results. So a flag can come from earlier content, not only your last message.

In Claude Code

What you see. A notice that names the older model. The session continues on that model.

What to do.

  • Run /model to switch back.
  • Press Esc twice to edit your last message and try again.
  • To be asked first, run /config and change “Switch models when a message is flagged.”
  • Run /feedback if the flag was wrong.

Don’t ask it to show its reasoning in the reply

What to do. Remove requests to reproduce its internal reasoning in the reply from your prompts and instructions.

Why it matters on Opus 5.5. A request to reproduce its internal reasoning in the reply can be declined. It’s one of the flag categories.

How. Ask Claude for what you need instead, for example, “Explain why you chose this approach in three sentences.”

6. SPEED

Turn on fast mode when you’re waiting on each reply

What to do. In Claude Code, use fast mode for back-and-forth work, where you read each reply before you send the next message.

Why it matters on Opus 5.5. Fast mode is available for Opus 5.5 at launch as a research preview. You get the same model, and the text arrives sooner. It needs extra usage turned on, and it costs more per token than standard mode.

How. Type /fast into Claude.

YOUR OPUS 5.5 CHECKLIST

Run through this before your next long task.

The checklist as a card with four groups of checkbox items: Asking, Long runs in Claude Code, Checking, and Flags. The same items are listed as text below.
FIG D The checklist at a glance.

Asking

  • The task says what “done” looks like
  • No “think hard” lines in prompts or saved instructions
  • Design requests list the styles to leave out
  • Charts and screenshots are attached, not retyped

Long runs in Claude Code

  • CLAUDE.md says when to stop and when to keep going, and to stop before anything destructive
  • Permission prompts are still on for destructive commands
  • Large audits and migrations are split across subagents
  • The task list is kept in a file

Checking

  • The “needs from you” part of the report is read first
  • A review pass runs before a person reviews
  • Research answers mark what couldn’t be confirmed

Flags

  • You know how to switch back: the model picker, or /model
  • “Switch models when a message is flagged” is set the way you want

Start building with Opus 5.5 !

With thanks to Molly Vorwerck for reviewing.

ADHD, autism or complex trauma? [pdf]

Hacker News
www.cambridge.org
2026-10-03 14:08:28
Comments...
Original Article
No preview for link for known binary extension (.pdf), Link: https://www.cambridge.org/core/services/aop-cambridge-core/content/view/30CC4826561366615BFAEC807CDE28A7/S0007125026108046a.pdf/adhd-autism-or-complex-trauma-the-complicated-nature-of-the-question.pdf.

Hole Punch: Sling your spaceship around gravitational fields

Hacker News
notoriousbfg.com
2026-10-03 14:06:45
Comments...
Original Article

Hole Punch
Unit 0426 · Mk II

Pop!_OS bans AI-generated code from much of its codebase

Hacker News
www.neowin.net
2026-10-03 13:57:03
Comments...

Vx – One Language, Every Chip

Hacker News
vxlang.org
2026-10-03 13:23:20
Comments...
Original Article

Vx is a systems programming language for heterogeneous computing. CPU, GPU, NPU and accelerator memory are part of the type system — so a host thread dereferencing a device pointer is a compile error , not a segfault at three in the morning.

curl -fsSL https://vxlang.org/install.sh | sh

macOS on Apple Silicon and Linux x86_64. Other install options .

Heterogeneity belongs in the type system, not in the runtime.

Where data lives is part of its type

Most languages treat the accelerator as infrastructure: you write math, and a large opaque runtime decides how to ship it. Vx treats it as semantics. A tensor pinned to NPU high-bandwidth memory has a different type from one in host DRAM, and crossing between them takes an explicit transfer() — even when the hardware boundary is free.

On Apple's unified memory that transfer compiles to almost nothing. It is still written down, because data locality should be provable by reading the source rather than by profiling the binary.

// Two matrices already resident in NPU memory.
fn custom_matmul(
    a: Pinned<Tensor<f32, [4, 4]>, Topology::NPU[0]>,
    b: Pinned<Tensor<f32, [4, 4]>, Topology::NPU[0]>)
    -> Verified<Tensor<f32, [4, 4], Memory::NPU_HBM>> {

    let mut result =
        Tensor<f32, [4, 4], Memory::NPU_HBM>::uninit();

    // Dispatch the computation to the accelerator.
    spawn on(Topology::NPU[0]) {
        for i in 0..4 {
            for j in 0..4 {
                result[i][j] = 0.0;
                for k in 0..4 {
                    result[i][j] += a[i][k] * b[k][j];
                }
            }
        }
    }

    return Verified(result);
}

What the compiler rules out

Vx front-loads into type checking a class of bug that normally surfaces as a runtime crash, silent corruption, or an out-of-memory at training step 1200.

Address-space typing

Dereferencing a device pointer from the host. A Pinned<T, NPU_SRAM> escaping into a host expression.

Capacity admission

A placement whose working set cannot fit the memory space it targets — checked against the declared machine, before a binary exists.

Seam contracts

Reading a buffer whose asynchronous transfer has not been made visible. Discharged by an SMT prover.

Linear types

Use-after-move of a consumed buffer, alongside a borrow checker with variance and region tracking.

Topology reachability

A transfer between two memory spaces with no declared path between them.

Autodiff

Differentiating through a region whose adjoint is not defined.

The machine is declared, not assumed

Most compilers hard-code a cost model. Vx reads one. A machine file describes the memory hierarchy and interconnect of a real part, and the compiler admits or rejects placements against it.

Units are exact integer conversions, never floats: SI prefixes are decimal ( GB = 10 9 ), IEC are binary ( GiB = 2 30 ). A figure copied off a vendor sheet means what the sheet meant.

The repository ships machine files for H100, H200, B200, A100, MI300X, Apple M4 and multi-GPU nodes — each citing its sources, and marking unverified figures as unverified.

Memory HBM  { capacity: 80 GiB, bandwidth: 3.35 TB/s,
              managed: explicit, scope: device }

Memory L2   { within: Memory::HBM, capacity: 50 MiB,
              bandwidth: 12 TB/s, managed: cached }

Memory SMEM { within: Memory::L2, capacity: 228 KiB,
              bandwidth: 128 B/cyc, clock: 1.98 GHz,
              replicas: 132, granule: 1 KiB, scope: sm }

Topology Device {
    arch: nvptx64,
    memory: Memory::HBM,
    transfer Memory::CPU_DRAM -> Memory::HBM : 63 GB/s,
}

How it compiles

A data-oriented parallel frontend

Every symbol, nominal type and monomorphized variant is a flat 256-bit identifier. A nominal type system plus mandatory boxing for recursive types decouples modules, so the pipeline runs parallel across cores with no query engine and no lock contention. Compilation walks flat arrays rather than pointer-chased trees.

The same source compiles to byte-identical MLIR whether it is built serially or in parallel. That is asserted in the test suite rather than hoped for — at benchmark scale, at one thread, at four, and with the thread pool taken off the path entirely, plus a corpus recompiled in fresh processes so each run gets its own hash seed. The claim is about the MLIR the frontend emits; everything downstream of it belongs to LLVM.

Backends

CPU (x86-64, arm64) MLIR → LLVM IR → native, AOT or JIT
NVIDIA GPU MLIR → NVVM → PTX → SASS
Apple AMX / ANE CoreML primitive dispatch via plugin
Distributed Manifest-driven remote regions over a wire protocol

Vendors extend the compiler through MLIR pass plugins rather than by patching it.

Where Vx is the wrong tool

PyTorch users mutate architecture mid-loop, print a tensor shape, branch on it, and carry on. In Vx — ahead-of-time, data-oriented, statically regioned — that same dynamism takes real work.

Vx is the right language for the thing that must be correct and fast across ten kinds of silicon. It is not the right language for the thing you are still figuring out.

Start here

Kolibri – Tech Report [pdf]

Hacker News
aleph-alpha.com
2026-10-03 13:22:48
Comments...
Original Article
No preview for link for known binary extension (.pdf), Link: https://aleph-alpha.com/downloads/tech-report.pdf.

RetailReady (YC W24) Is Hiring

Hacker News
www.ycombinator.com
2026-10-03 13:00:12
Comments...
Original Article

An AI-powered supply chain compliance engine

Implementations

$100K - $140K • 0.02% - 0.06% • San Francisco, CA, US

Experience

Any (new grads ok)

Connect directly with founders of the best YC-funded startups.

Apply to role ›

About the role

We are RetailReady (YC W24) - an AI-powered supply chain compliance engine. We’ve raised $6.6M in funding, onboarded over 50 customers, and in the past 12 months we’ve 4x’d our revenue. We are building world-class software that will disrupt an antiquated industry & we want you on board.

What we’re doing here matters.

We’re tackling everyday challenges for real people in the supply chain – that’s the heart of every business. With us, you’ll see your work come to life in amazing ways. Imagine being a part of why Walmart can restock baby formula, helping a health & beauty brand go from startup to Target shelves, and empowering warehouse workers with the latest tech. That’s the impact you’ll have.

This isn’t just another job offer - it’s an opportunity to create your own path and make a tangible difference.

What’s cool about retail compliance?

  • A majority of supply chain systems are archaic - bad UX, weeks to onboard, and old tech stacks
  • There is so much to disrupt in the supply chain space (we call it an engineer’s playground)
  • Compliance is fragmented across brands, warehouses, and retailers & we are first to market to solve operational retailer compliance within a warehouse

Why work with us?

  • Our team has operational, product, and engineering backgrounds across multiple industries, as well as deep supply chain industry experience – at Stord (a supply chain unicorn startup), Microsoft, and Manhattan Associates
  • You’ll be joining an awesome team at the start:  we’re low in ego, high in intellectual curiosity, and aren't afraid to make mistakes but above all we win together as a team!
  • We’re fun to work with - just watch any of our Weekly Round Up Videos on LinkedIn and you’ll see

Position Overview:

  • Lead the deployment of our software in warehouses and across tech stacks for Brands and 3PLs alike
  • Travel to customer locations for up to ~ 50% of the month
  • Collaborate with our engineers to conduct quality assurance on features
  • Provide insights from on-site experiences to influence the product roadmap

This position is not for the faint of heart - we have ambitious plans for RetailReady, and we want someone who shares the same level of ambition, and has the tenacity & passion to stick with us for the long haul.

We have an in-person team in San Francisco. If you’ve gotten this far and are fired up, reach out.

About the interview

  1. 15 minute intro call
  2. 30 minute panel
  3. Founder call
  4. Hired!

About RetailReady

RetailReady is building an AI-powered supply chain compliance engine. Supply chains are still heavily reliant on paper processes and tribal knowledge, causing costly shipping mistakes that jeopardize the longevity of businesses. RetailReady is the first-to-market with our retail compliance packing software, leveraging camera vision to direct warehouses to ship orders without error. We are positioning our compliance data models to become the operating system that will power the next wave of warehouse robotics and automation.

RetailReady

Founded: 2024

Batch: W24

Team Size: 12

Status: Active

Location: San Francisco

Founders

City building games have a Soul Problem pt.2

Hacker News
www.radical-elements.com
2026-10-03 11:52:07
Comments...
Original Article

Despite my disappointment with Cities: Skylines 2 , I keep coming back to the game from time to time. I can't even control it. It's like a deep inner need. When we were kids we played with toy cars on those felt mats with painted little cities, or when we had more industrial aspirations, we'd go outside in the dirt creating construction sites with tunnels and bridges right next to mom's rose bushes. Some of us had that one uncle who decided to turn the playroom into a miniature modeling world with trains, tunnels, and stations. I think a little tangle of my younger self's fantasies turns into a need to launch Cities: Skylines and build a living city.

In the previous article , I dealt with the genre more broadly and explored some ideas that sparked (mostly on Reddit) very interesting discussions. I talked more about the simulation and less about the aesthetics. Today, my text won't be technical at all; instead, it will tackle aesthetics and how much they disappoint me every time I gaze at my city.

Surfaces are all square and flat

Everything is square and unnatural. Looking at a random part of my city, I was trying to figure out why it's so ugly. What is it that makes me want to press alt+f4?

Here we see a spot where the sidewalk, some tiles from a building, and the grass meet.

In-game screenshot of Cities: Skylines 2 showing rigid, perfectly straight square seams where sidewalk, grass, and building tiles meet.

Does this look nice to you? How hard would it be for it to look like this:

Edited visual concept showing organic, blended transitions and natural edge detailing where grass meets paved surfaces.

I know someone can achieve similar and better results with mods and decals. But for most of us, it's just not realistic to spend 40 hours micro-managing every single corner of the city. And if you don't put in those hours, everything feels lifeless.

Default Cities: Skylines 2 street view showing a plain, sterile urban corner lacking surface detail.

Proposed visual improvement for a street corner, adding procedural detail, ground clutter, and realistic texture blending.

Also, something I truly enjoy in these games is being able to just stare at the city and see what the algorithm has generated. To be surprised. If I sit there and hand-paint every little corner, that "magic" is lost.

Neurotic cleanliness

As I just said, one of the reasons I build my city is so I can later go and gaze at random corners. It's the same when I take walks in downtown Athens; I look for charming little sides of the city. Not necessarily beautiful, not clean, but scenes that tell a story. In this game, everything is "sanitized" to the point of neurosis.

In-game view of a pristine, perfectly clean urban street in Cities: Skylines 2 with no visible wear or grime.

Mockup of the urban street with added surface wear, subtle dirt, and realistic pavement weathering.

Look how much difference these small textures can do to the overall atmosphere.

Default Cities: Skylines 2 cityscape scene featuring spotless, uniform architecture and sterile public space.

Enhanced visual concept showing character-filled urban space with lived-in details and subtle atmospheric grit.

What's the deal with stairs?

For some reason, the game has beef with stairs. Can we get some stairs? Something I would do if I were a designer for the game, is look at 100 photos of real cities and note down some of the things that make them up. Like:

  • Stairs
  • Potholes
  • Trash
  • Cables
  • Graffiti
  • Construction sites
  • Scaffolding
  • Poles
  • Traces of the past (repurposed public spaces or buildings)
  • Leaves

In-game screenshot showing awkward elevation changes and steep ground slopes without pedestrian stairs in Cities: Skylines 2.

Conceptual mockup showing outdoor stairs and terraced steps smoothly connecting elevated terrain.

Organic shapes

I don't know what your neighborhood is like, but in mine, maybe 1 in 50 houses is that clean and organized. Most have damp spots, weeds, abandoned toys in the garden, clothes hanging out to dry on the balcony.

In-game residential area in Cities: Skylines 2 with perfectly manicured, uniform green square lawns and sterile balconies.

Visual improvement showing organic residential yards with overgrown weeds, damp patches, hanging laundry, and varied foliage.

Some have rusted railings. Even the well-kept ones have let nature do its thing in some corners. And nature's thing isn't perfectly green squares.

Poor and abandoned houses

Unfortunately, every real city has poor areas with abandoned houses or ones in very bad shape. This is another aesthetic element that offers both realism and usable feedback. You shouldn't have to open the stats panel to see where a problem lies. By just gazing at your city, the landscape could and should speak for itself.

Default Cities: Skylines 2 neighborhood displaying clean, uniform buildings regardless of land value or area degradation.

Edited visual concept illustrating degraded and abandoned houses with weathered facades and visual cues of urban decay.

I hope that future city building games or CS2 DLCs will focus on adding some soul and character to city building. Until then, we have to rely on our Bob Ross instincts.


Discussions:


AI Disclosure: Artificial intelligence was used exclusively for proofreading, generating descriptive image alt texts, and creating the "improved" visual concepts.

FTL: A new operating system for clouds

Hacker News
ftl-os.org
2026-10-03 11:02:36
Comments...
Original Article

What's FTL?

  • You can build your own OS as a library . This userspace OS design makes it easy to add features, debug, upgrade the OS safely, as if writing applications.
  • FTL kernel isolates containers (userspace OS instances) better than existing monolithic kernels, with a hypervisor-like interface based on a lightweight hardware-based isolation (user mode). You don't need bare-metal machines.
  • FTL is compatible with Linux binaries . For example, the Rust-based HTTP server serving this website is a Linux application running on FTL. You can also run Unikernel-like specialized applications without POSIX abstractions.

How it works

Each container runs a userspace OS. It is a shared library which implements most of OS concepts such as Linux process, VFS, and TCP/IP. FTL kernel provides a minimal interface to implement Linux system calls in userspace, just like a hypervisor.

FTL combines the best of microkernels (flexible & secure) and monolithic kernels (performant & simple). Our goal is to make lightweight containers as secure as VMs, and unlock new OS-level abilities in applications , without sacrificing performance:

FTL                                   Linux
┌────────────────────────────────┐    ┌────────────────────────────────┐
│┏━━━━━━━━━━━━━┓  ┏━━━━━━━━━━━━━┓│    │┏━━━━━━━━━━━━━┓  ┏━━━━━━━━━━━━━┓│
│┃             ┃  ┃             ┃│    │┃             ┃  ┃             ┃│
│┃    Linux    ┃  ┃    Linux    ┃│    │┃    Linux    ┃  ┃    Linux    ┃│
│┃   Process   ┃  ┃   Process   ┃│    │┃   Process   ┃  ┃   Process   ┃│
│┃             ┃  ┃             ┃│    │┃             ┃  ┃             ┃│
│┃╌╌╌╌╌ Linux system calls ╌╌╌╌╌┃│    │┗━━━━━━━━━━━━━┛  ┗━━━━━━━━━━━━━┛│
│┃                              ┃│    └────────────────────────────────┘
│┃         Userspace OS         ┃│    ╌╌╌╌╌╌╌╌ Linux's interface ╌╌╌╌╌╌╌
│┃   (Process, VFS, TCP, ...)   ┃│    ╔════════════════════════════════╗
│┗━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┛│    ║                                ║
└────────────────────────────────┘    ║          Linux Kernel          ║
╌╌╌╌╌╌╌ minimal interface ╌╌╌╌╌╌╌╌    ║                                ║
╔════════════════════════════════╗    ║   process, fork/exec, memory,  ║
║           FTL Kernel           ║    ║     signals, TCP/IP, /proc,    ║
║    vCPU, memory, drivers, ...  ║    ║       /dev, drivers ...        ║
╚════════════════════════════════╝    ╚════════════════════════════════╝

Userspace OS design also enables you to extend most of Linux kernel features without kernel/eBPF programming. You can add printfs, apply security updates, and add new features quickly and safely. In FTL, OS is just a library . Read more .

Roadmap

  • September 2026 : Run a simple Linux HTTP server on FTL (released in v0.0.1 ✅)
  • October 2026 : Async Rust apps support - Linux threads, epoll, ... (released in v0.1.0 ✅)
  • November 2026 : Filesystem
  • December 2026 : Node.js / Go support
  • January 2027 : SMP, container images, 64-bit Arm support

Links

US killer's sentence quashed because of AI video of victim shown in court

Hacker News
www.bbc.com
2026-10-03 09:34:18
Comments...
Original Article
Watch Chris Pelkey's AI-rendered impact statement shown in court

An Arizona road rage killer will be resentenced after an appeals court ruled that an AI-generated video message from his dead victim was improperly aired in court.

Gabriel Paul Horcasitas was found guilty by a jury of shooting and killing Christopher Pelkey, 37, during a confrontation at a red light in 2021. He was sentenced to 10 years behind bars.

But his legal team appealed, arguing that the trial judge should not have allowed an AI video, created by Pelkey's family, to be shown in court ahead of last year's sentencing.

The Arizona Court of Appeals agreed, ruling that Horcasitas, 55, must be resentenced because the AI clip "crossed that line".

"Rather than document an event or recording a particular moment, the AI video presents a depiction of the victim and his thoughts created from the imaginings of the victim's sister," the court of appeals wrote on Wednesday.

Horcasitas' lawyer, Kristen Reller, who filed the appeal, declined to comment.

Jessica Gattuso, an attorney for victims in the case, did not respond to a request for comment.

Pelkey's sister, Stacey Wales, told the BBC last year the family had used voice recordings, videos and pictures to recreate him for the AI video.

"To Gabriel Horcasitas, the man who shot me, it is a shame we encountered each other that day in those circumstances," said the AI version of Pelkey in court in 2025. "In another life, we probably could have been friends."

"I believe in forgiveness, and a God who forgives. I always have and I still do," the AI version of Pelkey - wearing a grey baseball cap - continued.

At the time of the sentencing, the judge who oversaw the case, Todd Lang, seemed to appreciate the use of AI.

"I loved that AI, thank you for that. As angry as you are, as justifiably angry as the family is, I heard the forgiveness," Judge Lang said. "I feel that that was genuine."

The Era of Software Quality, or the Era of Ostriches?

Lobsters
blogs.gnome.org
2026-10-03 09:31:58
Comments...
Original Article

Humans are bad at writing secure code, and GNOME developers are no exception. GNOME is primarily written using unsafe programming languages where simple mistakes in our code lead to devastating consequences for our users , and we make these mistakes all the time . No matter how much we try, GNOME developers will fail write secure code when using unsafe languages like C, C++, or Vala: it’s just too hard for even experienced developers to do properly.

The above paragraph is taken from the abstracts of my GUADEC 2024 and 2025 talks. At the time, I thought failure was inevitable: we humans were so bad at writing software that we had no chance to do it properly, and I certainly would not have trusted an AI to do better than a human. But the landscape today is completely different than last year. AI has improved considerably, and offers a magic fairy wand solution to this problem: we can now simply ask a language model to look for vulnerabilities in our software. They are quite good at this.

There is zero hope of maintaining quality software in 2026 without AI vulnerability scanning. Any claims to the contrary are unserious and delusional. The tremendous quantity of bugs found in our best-maintained projects, like GLib and fwupd, should speak for itself. Failure to scan our projects is an unfair disservice to our users. If we don’t find the vulnerabilities by scanning projects ourselves, attackers certainly will, because the Linux user base has increased to the point that Linux users are finally numerous enough to be worth targeting. Meanwhile, AI has made it easier than ever to build working exploits , which was previously unheard of.

Already resolved all the detectable vulnerabilities? Then ask the AI to look for non-security bugs as well, to further improve quality. GNOME code is generally much better than it used to be, but there remains considerable room for improvement. For the first time in history, we now have the opportunity to improve software quality to a degree that was never realistic before.

Have you heard that most AI bug reports are “slop?” Not so in 2026. That was true for most of 2025, but the quality of AI-generated vulnerability reports has drastically improved. That is not to say that we no longer have problems with bad vulnerability reports, but in general, nowadays most of them are pretty good. ( Daniel Stenberg reports the same pattern for curl. )

AI-generated vulnerability reports have nevertheless introduced many undesirable impacts on GNOME maintainers. They are usually annoyingly verbose and unnecessarily detailed. They often exaggerate the severity of the problem, or make misleading or irrelevant claims. They are occasionally incorrect. Sometimes they include outright fabricated data, such as fake stack traces (which is not the norm, but sadly also not uncommon). A good human reviewer will notice and resolve most of the above problems before creating a bug report on your issue tracker, but often problems are reported by inexperienced humans who do not actually know what they are looking at and simply copy/paste everything blindly. Even when the generated issue report is good and avoids all of the above problems (which is rare), good vulnerability reports in sufficiently high quantity can still overwhelm volunteer maintainers. And even if reporters submit a merge request to resolve the problem so maintainers don’t have to (which is also rare), reviewing those merge requests is itself more unwelcome work for overworked maintainers.

That all is to say: I understand the pain caused by the current wave of AI-generated issue reports. Nevertheless, they are essential and unavoidable. We have to learn to accept and deal with them, not stick our heads in the sand and ignore them.

Some GNOME maintainers have adopted a policy prohibiting AI-generated content in issue reports. Do not do this. Nowadays, the overwhelming majority of vulnerability reports are AI-generated. Projects that choose to ban AI-generated content in issue reports might as well ban all vulnerability reports; the effect will be approximately the same.

I propose the following:

  • GNOME maintainers should rewrite their AI contribution policies to permit AI-generated vulnerability reports, as I previously requested four months ago .
  • Projects that continue to prohibit AI-generated vulnerability reports are no longer suitable dependencies for GNOME, and should be developed someplace other than GNOME GitLab.

We don’t have to tolerate bad issue reports, but AI use alone should not be disqualifying.

Shouldn’t humans rewrite AI-generated bug reports?

When I complain that maintainers should allow AI-generated vulnerability reports, the most common counterargument is that humans should read the AI’s report, understand it, and rewrite the entire thing to remove all AI-generated content. Some bug reporters actually voluntarily do this, but this is rare.

Vulnerability reporting is a public service, not an obligation. If you ask a reporter to do any amount of extra work, they might be willing to do so, but it’s much more likely that they will either stop looking at your project and move on to something else, or continue looking at your project and publish the vulnerability reports someplace other than your issue tracker.

Rewriting issue reports also does not scale. Let’s say you use AI to find 100 security bugs in a GNOME project, a number consistent with the results of actual scans (read on). Would you really spend months rewriting those bug reports before submitting them to upstream? Validating the AI’s claims, upstreaming the issue reports, and submitting merge requests is already a lot of work. Not many people would be willing to additionally rewrite all the issue reports. That’s more work than everything else combined, and is unrealistic.

Even with just a small number of bugs, I would hesitate to spend much time rewriting an issue report because I have many other tasks I would rather spend my time on. At best, I might prepare a quick summary, but it won’t be as useful as a full report.

The CVE Wave Hits GNOME

The current wave of vulnerability reports is reflected in GNOME’s CVE issuance trends:

Year GNOME CVEs GNOME CVEs Excluding GIMP, Gegl, libxml2, and libxslt
2021 21 14
2022 14 6
2023 13 4
2024 37 28
2025 97 49
2026 Year-to-date (2026-09-30) 141 74
2026 Normalized 188 (141 * 4 / 3) 99 (74 * 4 / 3)

The trend here should be pretty clear. Until recently, not many people were reporting vulnerabilities in GNOME. That has changed. We are currently dealing with an order of magnitude more CVEs than just 3 years ago. AI is not the only reason for this; GNOME maintainers have also gotten a little better at flagging issues so that I add them to security tracking. But AI is the primary cause for the increase.

(A few technical notes on this table. CVEs are classified by the year the issue was reported to GNOME, not by the year in the CVE identifier, so e.g. many CVE-2026 issues are counted in 2025. Vulnerabilities reported in 2026 which do not yet have CVEs are not counted, so you can think of the data as being accurate through roughly September 1; multiply the 2026 numbers by 4/3 to make them comparable to the prior years. I count only issues reported to GNOME Security , so any unreported CVEs do not count.)

Although there are still 3 months left in 2026, we will never have data for the rest of the year because I have ended security tracking for new issue reports and nobody else has volunteered to do that work. These CVEs exist only because I request them myself, so I expect the number of CVEs to drastically decrease going forward.

The CVE Wave Hits WebKitGTK

A similar pattern holds for WebKitGTK:

Year WebKitGTK CVEs
2015 175
2016 57
2017 158
2018 101
2019 99
2020 38
2021 52
2022 50
2023 45
2024 38
2025 66
2026 Year-to-date (through WSA-2026-0006 ) 305

CVEs are reported against the year they appeared in a WebKitGTK security advisory, not the year in the CVE ID. The large increase in 2026 is entirely due to AI analysis of Skia and ANGLE. WebKit bundles these libraries because they are not designed to be installed as system libraries, so their vulnerabilities should be counted the same as vulnerabilities in WebKit’s own code. Excluding Skia and ANGLE, there are actually only 21 other WebKitGTK CVEs so far this year, a significant decrease, but excluding CVEs in bundled code would not be fair.

There has actually been a very large increase in WebKit security fixes this year, but this has not resulted in any increase in CVEs. Apple generally creates CVEs for flaws found by external researchers, not often for flaws found by WebKit developers, so the increase in security fixes is not reflected in the total number of CVEs. Only a small fraction of WebKit vulnerabilities receive CVEs.

I had not previously noticed that the count of WebKitGTK CVEs had, until 2026, been decreasing over the past decade. I am not sure why. I also do not know how to explain the low number in 2016.

Announcing the GNOME Bug Bounty Program and Announcing the End of the GNOME Bug Bounty Program

My blog post to-do list says that I need to write a blog post announcing the creation of the GNOME Bug Bounty Program on the YesWeHack platform. Oops, too late. It’s already closed. (Once a task enters my to-do list, it can be a very long time before I get around to doing it.)

The GNOME Bug Bounty Program was generously sponsored by the Sovereign Tech Resilience program of Germany’s Sovereign Tech Agency. I’m not sure precisely when it opened, but the first vulnerability was reported on June 27, 2024, so it would have been sometime shortly before then. We accepted issue reports only for GLib, glib-networking, and libsoup, because GNOME had never operated a bug bounty program before and we did not know what to expect. Starting small had — naively — seemed like a prudent way to avoid a large quantity of issue reports. I had wanted to expand the program to cover all of GNOME, but this failed due to the overwhelming deluge in issues reported against GLib and libsoup.

I requested that the bug bounty program end because I was overwhelmed with incoming AI-generated issue reports. The final issue was reported on February 23, 2026. Here are the results:

Year Reports Submitted Reports Accepted
2024 26 14
2025 150 33
2026 122 24
Total 298 71

Those numbers for 2026 reflect less than two months’ worth of issue reports, so you can see why it was no longer sustainable.

After the program closed, our work was not done: there was a long backlog of reports to work though. We just last month caught up with accepting the last of the issues reported back in February, and the last bounty was finally awarded earlier today! Even with YesWeHack’s professional triagers analyzing the issue reports before I reviewed them, keeping up with such a large number of vulnerabilities was not easy for me.

At this point, all reports not accepted have been rejected. The program awarded €183,900 in bounties for 71 vulnerabilities: 45 in libsoup, 23 in GLib, and 3 in glib-networking. Award amounts varied from €500 (16 awards) to €7,500 (2 awards). The arithmetic mean award was €2,662.99.

Bug bounty programs are an exception to the rule that most AI-generated vulnerability reports are good. You can see the number of reports accepted is a small fraction of the number of reports submitted. Excluding 30 reports closed as duplicates, that leaves 197 reports rejected. Turns out, people will submit bad reports when financially incentivized to do so. The low percentage of accepted reports even understates the problem, because many of the accepted reports were actually not very good! Many accepted reports did successfully identify valid security problems (in fact, many of the rejected reports successfully identified valid security problems!), but required many rounds of revision and corrections.

Suffice to say, I have reviewed a lot of really bad AI-generated vulnerability reports. But the reports we received via the discontinued bug bounty program are not comparable to the reports received via regular GNOME issue trackers or the security bug report form . We do still occasionally receive bad vulnerability reports, but not often and not many, so it’s not a big problem anymore. When people submit AI-generated reports without hope of a financial award, those reports are generally much better.

Lessons from the Bug Bounty Program

Closing the bug bounty program because it found too many vulnerabilities is not a particularly pleasant result. That said, it was still a partial success in that it uncovered lots of bugs in libsoup and GLib.

I had hypothesized that libsoup was probably not very secure, but I never imagined just how many vulnerabilities would be discovered. To reduce the quantity of incoming issue reports and better reflect actual risk to GNOME users, I eventually removed all denial of service bugs from program scope, and then later removed SoupServer from the scope due to too many request smuggling vulnerabilities , which are HTTP request parsing bugs that pose no threat to GNOME users. Even with those changes, the libsoup vulnerability reports kept coming until I gave up. The silver lining is that libsoup is now relatively much more secure than before. Other bug reporters have been submitting AI-generated bug reports using the normal libsoup issue tracker, so fortunately the improvements to libsoup will continue despite an end to the financial awards.

I had hypothesized that GLib would be much better than libsoup. I’m not sure whether I was correct. Evaluating the severity of GLib flaws is much harder than for libsoup, since GLib vulnerability reports are generally hypothetical in nature: usually some proof of concept program calls a GLib API using valid but improbable values, then something bad happens.

A large portion of the GLib bugs were integer overflow flaws, which generally result in buffer overflow. I am now more scared of integer overflow than anything else. It’s likely that most software projects have many integer overflow problems. Fortunately, we should be able to catch most such problems by adjusting the compiler flags we use. In particular, -Wconversion or -Wint-conversion and -Wsign-compare should help here. Some GNOME projects already use -Wsign-compare , but I suspect most do not. I think few or no GNOME projects use -Wconversion or -Wint-conversion .

Resuming the bug bounty program would only be possible under substantially different conditions. What we were doing was not working well. To resume, we would need to limit the scope to projects that regularly perform their own AI vulnerability scans. We would also most likely want to pay only for functional exploits, rather than for all vulnerabilities. GNOME code is currently not good enough to continue paying for every vulnerability, and it no longer makes sense to pay bounties for issues that can be found by AI scanners.

Red Hat Scans GLib

Red Hat has contracted with AISLE Research to perform AI vulnerability scans of various GNOME projects. We received a large quantity of findings, and are only just now beginning to individually validate and report our findings to upstream. GLib is by far the hardest hit project, which I was not expecting, accounting for more than 40% of our total findings. I’m not certain why, but perhaps this is because GLib provides so many public APIs. Data passed to public APIs is potentially untrusted, so the attack surface is considerable.

Red Hat’s scan of GLib found 118 vulnerabilities. Or at least, it claimed to. However, due to the way we ran the scans, several of these are actually unnecessary duplicates of each other, which we have not fully deduplicated yet, so the number I report is not entirely trustworthy. Moreover, 46 of these “vulnerabilities” are bugs in gobject-introspection, mostly in the typelib support, which is evidently not very robust. A typelib controls how your program calls libraries; it is effectively calling convention, so it must inherently be fully trusted: a malicious typelib would be able to induce vulnerabilities even without any bugs! I would expect an AI ought to have been able to figure that out, but apparently not. These bugs are still real problems that we ought to fix, but all maintainers agree they are not security vulnerabilities, so let’s count all of them as false positives. That alone creates a 40% false positive rate. Ouch.

I don’t have more stats to share here because we are not yet done working through the issue reports. That said, I am quite pleased with the results thus far. Substantially all of the reports are high-quality. The false positive vulnerability reports are almost all due to one particular misunderstanding and can be treated as good quality non-security bug reports, which are still valuable. Expect many forthcoming CVE assignments for the other findings.

It’s rare for Linux vendors to proactively look for software vulnerabilities, rather than waiting for security researchers to report them. This was a successful experiment in proactively seeking out problems.

Humans Still Useful

In addition to the bug bounty program, the Sovereign Tech Resilience program also sponsored a security audit for GNOME, performed by Codean Labs. This resulted in many findings in various GNOME projects. Most notably, the scope of the audit extended to Flatpak and xdg-desktop-portal, resulting in critical findings .

Most of these issues could have been detected via AI scans, but I am not confident that AIs would have been able to discover the most important findings, like the two Flatpak sandbox escapes that I linked to above. Accordingly, I do not recommend relying on AI alone.

Humanity Still Desired

Although I like AI-generated issue reports, I particularly do not appreciate when I wind up interacting with a robot rather than with a human. It’s pretty obvious when your issue tracker or code review comments are written by an AI. Consider whether outsourcing your writing and your thinking to a language model is truly wise for your public image.

We even have one experienced GNOME developer who is obviously using AI to write all of his posts on GitLab. I am unsure whether he is copy/pasting all of his responses from an AI, or whether he is just a bot now. I especially do not understand the value of this.

Here is a soft proposal, intended only as a starting point for discussion and not as a serious proposal, for what my preferred AI usage policy might look like:

  • Newer developers should exercise caution when using AI to write code. Your priority should be learning, and I wonder how much you are really learning when relying on the AI to do work for you.
  • Do not use AI to write code comments. Currents AIs are terrible at writing comments. Most comments written by AIs should be deleted. If a comment is truly necessary, then I’d like to see it written in your own words. Presumably AIs will get better at this eventually, but as of 2026, human judgment is still required here.
  • Do not use AI to write commit messages. AIs are actually probably better than humans at writing commit messages, but I would still rather hear your own thoughts on the code you are submitting.
  • Certainly do not post AI-generated comments on an issue tracker or merge request as if they are your own. You’re not fooling anybody.

Maintain Perspective

Are you scared by the large numbers of recently-discovered vulnerabilities? There is no need to panic. Security bugs are just bugs, and they’re not necessarily more important than other bugs. Occasionally they are emergencies, but far more often they are boring and unexceptional. Security vulnerabilities are not even the biggest digital security threats that users face: those are surely phishing and trojans , with software security bugs a distant third place. No amount of CVE fixing will protect you from those more likely threats.

I don’t want to downplay the severity of security issues either. In fact, evaluating severity is hard. I quite often decide that a bug is not a big deal, only to be proven incorrect. Ideally, we would fix as many security issues as possible, and sooner rather than later. Lifetime issues and out of bounds writes are especially important to fix. Two years ago, I claimed that memory safety vulnerabilities were becoming less threatening, a claim that did not age well: that is surely no longer true due to the drastically increased accessibility of AI exploit generation.

Nonetheless, volunteer maintainers should not feel obligated to fix security issues or treat them as higher-priority than other bug reports. It’s certainly good to fix problems when possible, but my request is only that you do not prohibit issue reports, not that you attempt to personally resolve every security problem yourself. When I add due dates to vulnerability reports, that represents only a disclosure deadline — because issue reports should not stay confidential indefinitely — not an expectation that you fix the issue by that date. Resolving security problems in projects used by big tech companies that depend on your software without contributing back is basically free labor for said companies, and only you can decide whether that’s how you want to spend your volunteer time.

Rust

Yes, even projects written in memory safe languages like Rust still need to allow AI-generated vulnerability reports. Rust will indeed eliminate most memory safety issues ( except in unsafe blocks ), and you can reasonably expect a Rust project to have an order of magnitude fewer vulnerabilities than a comparable project written in C or C++ or Vala. This is amazing, but not all vulnerabilities are memory safety issues, so this is not an excuse to avoid scanning for flaws.

Although Rust mostly eliminates memory safety risk, any use of Cargo to download dependencies dramatically increases supply chain security risk. The risk of bundling a trojanized dependency arguably — I would even say probably — outweighs the benefit of eliminating memory safety flaws. This problem is inherent to any programming language package manager. Currently the best solution is to not use programming language package managers, but GNOME’s Rust code depends heavily on Cargo. Accordingly, I recommend against using Rust for writing GNOME software.

To Be Continued…

I have exhausted my thoughts on AI vulnerability reports, but there is still much to discuss regarding software quality. Next time, I will discuss additional strategies to improve GNOME quality without significantly relying on AI.

Writing the Cyclone Scheme Compiler (2017)

Lobsters
justinethier.github.io
2026-10-03 09:23:29
Comments...
Original Article

Revised for 2017

by Justin Ethier

This write-up provides a high level background on the various components of Cyclone and how they were written. It is a revision of the original write-up , written over a year ago in August 2015, when the compiler was self hosting but before the new garbage collector was written. Quite a bit of time has passed since then, so I thought it would be worthwhile to provide a brain dump of sorts for everything that has happened in the last year and half.

Before we get started, I want to say Thank You to all of the contributors to the Scheme community. Cyclone is based on the community’s latest revision of the Scheme language and wherever possible existing code was reused or repurposed for this project, instead of starting from scratch. At the end of this document is a list of helpful online resources. Without high quality Scheme resources like these the Cyclone project would not have been possible.

Table of Contents

Overview

Cyclone has a similar architecture to other modern compilers:

flowchart of cyclone compiler

First, an input file containing Scheme code is received on the command line and loaded into an abstract syntax tree (AST) by Cyclone’s parser. From there a series of source-to-source transformations are performed on the AST to expand macros, perform optimizations, and make the code easier to compile to C. These intermediate representations (IR) can be printed out in a readable format to aid debugging. The final AST is then output as a .c file and the C compiler is invoked to create the final executable or object file.

Programs are linked with the necessary Scheme libraries and the Cyclone runtime library to create an executable:

Diagram of files linked into a compiled executable

Source-to-Source Transformations

Overview

My primary inspiration for Cyclone was Marc Feeley’s The 90 minute Scheme to C compiler (also video 1 , video 2 , and code ). Over the course of 90 minutes, Feeley demonstrates how to compile Scheme to C code using source-to-source transformations, including closure and continuation-passing-style (CPS) conversions.

As outlined in the presentation, some of the difficulties in compiling to C are:

Scheme has, and C does not have

  • tail-calls a.k.a. tail-recursion optimization
  • first-class continuations
  • closures of indefinite extent
  • automatic memory management i.e. garbage collection (GC)

Implications

  • cannot translate (all) Scheme calls into C calls
  • have to implement continuations
  • have to implement closures
  • have to organize things to allow GC

The rest is easy!

To overcome these difficulties a series of source-to-source transformations are used to remove powerful features not provided by C, add constructs required by the C code, and restructure/relabel the code in preparation for generating C. The final code may be compiled direcly to C. Cyclone also includes many other intermediate transformations, including:

The 90-minute compiler ultimately compiles the code down to a single function and uses jumps to support continuations. This is a bit too limiting for a production compiler, so that part was not used.

Just Make Many Small Passes

To make Cyclone easier to maintain a separate pass is made for each transformation. This allows Cyclone’s code to be as simple as possible and minimizes dependencies so there is less chance of changes to one transformation breaking the code for another.

Internally Cyclone represents the code being compiled as an AST of regular Scheme objects. Since Scheme represents both code and data using S-expressions , our compiler does not (in general) have to use custom abstract data types to store the code as would be the case with many other languages.

Most of the transformations follow a similar pattern of recursively examining an expression. Here is a short example that demonstrates the code structure:

(define (search exp)
  (cond
    ((const? exp)    '())
    ((prim? exp)     '())    
    ((quote? exp)    '())    
    ((ref? exp)      (if bound-only? '() (list exp)))
    ((lambda? exp)   
      (difference (reduce union (map search (lambda->exp exp)) '())
                  (lambda-formals->list exp)))
    ((if-syntax? exp)  (union (search (if->condition exp))
                            (union (search (if->then exp))
                                   (search (if->else exp)))))
    ((define? exp)     (union (list (define->var exp))
                            (search (define->exp exp))))
    ((define-c? exp) (list (define->var exp)))
    ((set!? exp)     (union (list (set!->var exp)) 
                            (search (set!->exp exp))))
    ((app? exp)       (reduce union (map search exp) '()))
    (else             (error "unknown expression: " exp))))

The Nanopass Framework was created to make it easier to write a compiler that makes many small passes over the code. Unfortunately Nanopass itself is written in R 6 RS and could not be used for this project.

Macro Expansion

Macro expansion is one of the first transformations. Any macros the compiler knows about are loaded as functions into a macro environment, and a single pass is made over the code. When the compiler finds a macro the code is expanded by calling the macro. The compiler then inspects the resulting code again in case the macro expanded into another macro.

At the lowest level, Cyclone’s explicit renaming (ER) macros provide a simple, low-level macro system that does not require much more than eval . Many ER macros from Chibi Scheme are used to implement the built-in macros in Cyclone.

Cyclone also supports the high-level syntax-rules system from the Scheme reports. Syntax rules is implemented as a huge ER macro ported from Chibi Scheme.

As a simple example the let macro below:

(let ((square (lambda (x) (* x x))))
  (write (+ (square 10) 1)))

is expanded to:

(((lambda (square) (write (+ (square 10) 1)))
  (lambda (x) (* x x))))

CPS Conversion

The conversion to continuation passing style (CPS) makes continuations explicit in the compiled code. This is a critical step to make the Scheme code simple enough that it can be represented by C. As we will see later, the runtime’s garbage collector also requires code in CPS form.

The basic idea is that each expression will produce a value that is consumed by the continuation of the expression. Continuations will be represented using functions. All of the code must be rewritten to accept a new continuation parameter k that will be called with the result of the expression. For example, considering the previous let example:

(((lambda (square) (write (+ (square 10) 1)))
  (lambda (x) (* x x))))

the code in CPS form becomes:

((lambda (r)
   ((lambda (square)
      (square
        (lambda (r)
          ((lambda (r) (write r))
           (+ r 1)))
        10))
    r))
 (lambda (k x) (k (* x x))))

CPS Optimizations

CPS conversion generates too much code and is inefficient for functions such as primitives that can return a result directly instead of calling into a continuation. So we need to optimize it to make the compiler practical. For example, the previous CPS code can be simplified to:

((lambda (k x) (k (* x x)))
  (lambda (r)
    (write (+ r 1)))
  10)

One of the most effective optimizations is inlining of primitives. That is, some runtime functions can be called directly, so an enclosing lambda is not needed to evaluate them. This can greatly reduce the amount of generated code.

A contraction phase is also used to eliminate other unnecessary lambda ’s. There are a few other miscellaneous optimizations such as constant folding, which evaluates certain primitives at compile time if the parameters are constants.

To more efficiently identify optimizations Cyclone first makes a code pass to build up an hash table-based analysis database (DB) of various attributes. This is the same strategy employed by CHICKEN, although each compiler records different attributes. The DB contains a table of records with an entry for each variable (indexed by symbol) and each function (indexed by unique ID).

In order to support the analysis DB a custom AST is used to represent functions during this phase, so that each one can be tagged with a unique identification number. After optimizations are complete, the lambdas are converted back into regular S-expressions.

Closure Conversion

Free variables passed to a nested function must be captured in a closure so they can be referenced at runtime. The closure conversion transformation modifies lambda definitions as necessary to create new closures. It also replaces free variable references with lookups from the current closure.

Cyclone uses flat closures: objects that contain a single function reference and a vector of free variables. This is a more efficient representation than an environment as only a single vector lookup is required to read any of the free variables.

Mutated variables are not directly supported by flat closures and must be added to a pair (called a “cell”) by a separate compilation pass prior to closure conversion.

Cyclone’s closure conversion is based on code from Marc Feeley’s 90 minute Scheme->C compiler and Matt Might’s Scheme->C compiler.

C Code Generation

The compiler’s code generation phase takes a single pass over the transformed Scheme code and outputs C code to the current output port (usually a .c file).

During this phase C code is sometimes saved for later use instead of being output directly. For example, when compiling a vector literal or a series of function arguments, the code is returned as a list of strings that separates variable declarations from C code in the “body” of the generated function.

The C code is carefully generated so that a Scheme library ( .sld file) is compiled into a C module. Functions and variables exported from the library become C globals in the generated code.

Native Compilation

The C compiler is invoked to generate machine code for the Scheme module, and to also create an executable if a Scheme program is being compiled.

Garbage Collector

Background: Cheney on the MTA

A runtime based on Henry Baker’s paper CONS Should Not CONS Its Arguments, Part II: Cheney on the M.T.A. was used as it allows for fast code that meets all of the fundamental requirements for a Scheme runtime: tail calls, garbage collection, and continuations.

Baker explains how it works:

We propose to compile Scheme by converting it into continuation-passing style (CPS), and then compile the resulting lambda expressions into individual C functions. Arguments are passed as normal C arguments, and function calls are normal C calls. Continuation closures and closure environments are passed as extra C arguments. Such a Scheme never executes a C return, so the stack will grow and grow … eventually, the C “stack” will overflow the space assigned to it, and we must perform garbage collection.

Cheney on the M.T.A. uses a copying garbage collector. By using static roots and the current continuation closure, the GC is able to copy objects from the stack to a pre-allocated heap without having to know the format of C stack frames. To quote Baker:

the entire C “stack” is effectively the youngest generation in a generational garbage collector!

After GC is finished, the C stack pointer is reset using longjmp and the GC calls its continuation.

Here is a snippet demonstrating how C functions may be written using Baker’s approach:

object Cyc_make_vector(object cont, object len, object fill) {
  object v = NULL;
  int i;
  Cyc_check_int(len);

  // Memory for vector can be allocated directly on the stack
  v = alloca(sizeof(vector_type));

  // Populate vector object
  ((vector)v)->tag = vector_tag;
  ... 

  // Check if GC is needed, then call into continuation with the new vector
  return_closcall1(cont, v);
}

CHICKEN was the first Scheme compiler to use Baker’s approach.

Cyclone’s Hybrid Collector

Baker’s technique uses a copying collector for both the minor and major generations of collection. One of the drawbacks of using a copying collector for major GC is that it relocates all the live objects during collection. This is problematic for supporting native threads because an object can be relocated at any time, invalidating any references to the object. To prevent this either all threads must be stopped while major GC is running or a read barrier must be used each time an object is accessed. Both options add a potentially significant overhead so instead Cyclone uses another type of collector for the second generation.

To that end, Cyclone supports uses a tri-color tracing collector based on the Doligez-Leroy-Gonthier (DLG) algorithm for major collections. The DLG algorithm was selected in part because many state-of-the-art collectors are built on top of DLG such as Chicken, Clover, and Schism. So this may allow for further enhancements down the road.

Under Cyclone’s runtime each thread contains its own stack that is used for private thread allocations. Thread stacks are managed independently using Cheney on the MTA. Each object that survives one of these minor collections is copied from the stack to a newly-allocated slot on the heap.

Heap objects are not relocated, making it easier for the runtime to support native threads. In addition major GC uses a collector thread that executes asynchronously so application threads can continue to run concurrently even during collections.

In summary:

  • All objects on the stack are collected using Cheney on the MTA, and the ones that survive are placed on the heap.
  • Heap objects are collected during Major GC using the DLG algorithm.
  • Heap collection runs on a separate thread in parallel with application threads.

Major Garbage Collection Algorithm

Each object is marked with a specific color (white, gray, or black) that determines how it will be handled during a major collection. Major GC transitions through the following states:

Clear

The collector thread swaps the values of the clear color (white) and the mark color (black). This is more efficient than modifying the color on each object in the heap. The collector then transitions to sync 1. At this point no heap objects are marked, as demonstrated below:

Initial object graph

Mark

The collector thread transitions to sync 2 and then async. At this point it marks the global variables and waits for the application threads to also transition to async. When an application thread transitions it will:

  • Mark its roots black.
  • Gray any child objects of the roots. The collector thread traces these gray objects during the next phase.
  • Use black as the allocation color to prevent any new objects from being collected during this cycle.

Initial object graph

Trace

The collector thread finds all live objects using a breadth-first search and marks them black:

Initial object graph

Sweep

The collector thread scans the heap and frees memory used by all white objects:

Initial object graph

If the heap is still low on memory at this point the heap will be increased in size. Also, to ensure a complete collection, data for any terminated threads is not freed until now.

More details are available in a separate Garbage Collector document.

Developing the New Collector

It took a long time to research and plan out the new GC before it could be implemented. There was a noticeable lull in Github contributions during that time:

The actual development consisted of several distinct phases:

  • Phase 0 - Started with a runtime using a basic Cheney-style copying collector.
  • Phase 1 - Added new definitions via gc.h and made sure everything compiles.
  • Phase 2 - Changed how strings are allocated to clean up the code and be compatible with the new GC algorithm. This was mainly just an exercise in cleaning up cruft in the old Cyclone implementation.
  • Phase 3 - Changed major GC from using a Cheney-style copying collector to a naive mark-and-sweep algorithm. The new algorithm was based on code from Chibi Scheme so it was already debugged and would serve as a solid foundation for future work.
  • Phase 4 - Integrated code for a new tracing GC algorithm but did not cut over to it yet. Added a new thread data argument to all of the necessary runtime functions - a simple but far-reaching change that affected almost all functions in the runtime and compiled code.
  • Phase 5 - Required the pthreads library and stood Cyclone back up using the new GC algorithm for the first time.
  • Phase 6 - Added SRFI 18 to support multiple application threads.

Heap Data Structures

Cyclone allocates heap data one page at a time. Each page is several megabytes in size and can store multiple Scheme objects. Cyclone will start with a small initial page size and gradually allocate larger pages using the Fibonnaci Sequence until reaching a maximum size.

Each page contains a linked list of free objects that is used to find the next available slot for an allocation. An entry on the free list will be split if it is larger than necessary for an allocation; the remaining space will remain in the free list for the next allocation.

Cyclone allocates smaller objects in fixed size heaps to minimize allocation time and prevent heap fragmentation. The runtime also remembers the last page of the heap that was able to allocate memory, greatly reducing allocation time on larger heaps.

The heap data structures and associated algorithms are based on code from Chibi scheme.

C Runtime

Overview

The C runtime provides supporting features to compiled Scheme programs including a set of primitive functions, call history, exception handling, and garbage collection.

An interesting observation from R. Kent Dybvig [11] that I have tried to keep in mind is that performance optimizations in the runtime can be just as (if not more) important that higher level CPS optimizations:

My focus was instead on low-level details, like choosing efficient representations and generating good instruction sequences, and the compiler did include a peephole optimizer. High-level optimization is important, and we did plenty of that later, but low-level details often have more leverage in the sense that they typically affect a broader class of programs, if not all programs.

Data Types

Objects

Most Scheme data types are represented as objects that are allocated in heap/stack memory. Each type of object has a corresponding C structure that defines its fields, such as the following one for pairs:

typedef struct {
  gc_header_type hdr;
  tag_type tag;
  object pair_car;
  object pair_cdr;
} pair_type;

All objects have:

  • A gc_header_type field that contains marking information for the garbage collector.
  • A tag to identify the object type.
  • One or more additional fields containing the actual object data.

Value Types

On the other hand, some data types can be represented using 30 bits or less and are stored as value types. The great thing about value types is they do not have to be garbage collected because no extra data is allocated for them. This makes them super efficient for commonly-used data types.

Value types are stored using a common technique that is described in Lisp in Small Pieces (among other places). On many machines addresses are multiples of four, leaving the two least significant bits free. A brief explanation :

The reason why most pointers are aligned to at least 4 bytes is that most pointers are pointers to objects or basic types that themselves are aligned to at least 4 bytes. Things that have 4 byte alignment include (for most systems): int, float, bool (yes, really), any pointer type, and any basic type their size or larger.

In Cyclone the two least significant bits are used to indicate the following data types:

Binary Bit Pattern Data Type
00 Pointer (an object type)
01 Integer
10 Character

Booleans are potentially another good candidate for value types. But for the time being they are represented in the runtime using pointers to the constants boolean_t and boolean_f .

Thread Data Parameter

At runtime Cyclone passes the current continuation, number of arguments, and a thread data parameter to each compiled C function. The continuation and arguments are used by the application code to call into its next function with a result. Thread data is a structure that contains all of the necessary information to perform collections, including:

  • Thread state
  • Stack boundaries
  • Cheney on the MTA jump buffer
  • List of mutated objects detected by the minor GC write barrier
  • Parameters for major GC
  • Call history buffer
  • Exception handler stack

Each thread has its own instance of the thread data structure and its own stack (assigned by the C runtime/compiler).

Call History

Each thread maintains a circular buffer of call history that is used to provide debug information in the event of an error. The buffer itself consists of an array of pointers-to-strings. The compiler emits calls to runtime function Cyc_st_add that will populate the buffer when the program is running. Cyc_st_add must be fast as it is called all the time! So it does the bare minimum - update the pointer at the current buffer index and increment the index.

Exception Handling

A family of Cyc_rt_raise functions is provided to allow an exception to be raised for the current thread. These functions gather the required arguments and use apply to call the thread’s current exception handler. The handler is part of the thread data parameter, so any functions that raise an exception must receive that parameter.

A Scheme API for exception handling is provided as part of R 7 RS.

Scheme Libraries

This section describes a few notable parts of Cyclone’s Scheme API .

Native Thread Support

A multithreading API is provided based on SRFI 18 . Most of the work to support multithreading is accomplished by the runtime and garbage collector.

Cyclone attempts to support multithreading in an efficient way that minimizes the amount of synchronization among threads. But objects are still copied during minor GC. In order for an object to be shared among threads the application must guarantee the object is no longer on the stack. One solution is for application code to initiate a minor GC before an object is shared with other threads, to guarantee the object will henceforth not be relocated.

Reader

Cyclone uses a combined lexer / parser to read S-expressions. Input is processed one character at a time and either added to the current token or discarded if it is whitespace, part of a comment, etc. Once a terminating character is read the token is inspected and converted to an appropriate Scheme object. For example, a series of numbers may be converted into an integer.

The full implementation is written in Scheme and located in the (scheme read) library.

Interpreter

The eval function is written in Scheme, using code from the Metacircular Evaluator from SICP as a starting point.

The interpreter itself is straightforward but there is nice speed up to be had by separating syntactic analysis from execution. It would be interesting see what kind of performance improvements could be obtained by compiling to VM bytecodes or even using a JIT compiler.

The interpreter’s full implementation is available in the (scheme eval) library, and the icyc executable is provided for convenient access to a REPL.

Compiler Internals

Most of the Cyclone compiler is implemented in Scheme as a series of libraries .

Scheme Standards

Cyclone targets the R 7 RS-small specification . This spec is relatively new and provides incremental improvements from the popular R 5 RS spec . Library support is the most important new feature but there are also exceptions, system interfaces, and a more consistent API.

Benchmarks

ecraven has put together an excellent set of Scheme benchmarks based on a R 7 RS suite from the Larceny project. These are the typical benchmarks that many implementations have used over the years, but the remarkable thing here is all of the major implementations are supported, allowing a rare apples-to-apples comparison among all the widely-used Schemes.

Over the past year Cyclone has matured to the point where almost all of the 56 benchmarks will run:

The remaining ones are:

  • mbrotZ fails because Cyclone does not support complex numbers yet.
  • pi does not work because Cyclone does not support bignums yet.
  • compiler passes but returns the wrong result. This will be fun to track down since the program is huge and takes a long time to compile…

Regarding performance, from Feeley’s presentation [10] :

Performance is not so bad with NO optimizations (about 6 times slower than Gambit-C with full optimization)

But Cyclone has some optimizations now, doesn’t it? The following is a chart of total runtime in minutes for the benchmarks that each Scheme passes successfully. This metric is problematic because not all of the Schemes can run all of the benchmarks but it gives a general idea of how well they compare to each other. Cyclone performs well against all of the interpreters but still has a long ways to go to match top-tier compilers. Then again, most of these compilers have been around for a decade or longer:

Future

Some goals for the future are:

  • Implement more of R 7 RS-large; work has already started on the data structures side.
  • Implement more libraries (for example, by porting some of industria to r7rs).
  • Improve the garbage collector. Possibly by allowing more than one collector thread (Per gambit’s parallel GC).
  • Perform additional optimizations, EG:

    Andrew Appel used a similar runtime for Standard ML of New Jersey which is referenced by Baker’s paper. Appel’s book Compiling with Continuations includes a section on how to implement compiler optimizations - many of which could still be applied to Cyclone.

In addition, developing Husk Scheme helped me gather much of the knowledge that would later be used to create Cyclone. In fact the primary motivation in building Cyclone was to go a step further and understand how to build a full, free-standing Scheme system. At this point Cyclone has eclipsed the speed and functionality of Husk and it is not clear if Husk will receive much more than bug fixes going forward. Perhaps if there is interest from the community some of this work can be ported back to that project.

Conclusion

Thanks for reading!

Want to give Cyclone a try? Install a copy using cyclone-bootstrap .

Terms

  • Abstract Syntax Tree (AST) - A tree representation of the syntactic structor of source code written in a programming language. Sometimes S-expressions can be used as an AST and sometimes a representation that retains more information is required.
  • Free Variables - Variables that are referenced within the body of a function but that are not bound within the function.
  • Garbage Collector (GC) - A form of automatic memory management that frees memory allocated by objects that are no longer used by the program.
  • REPL - Read Eval Print Loop; basically a command prompt for interactively evaluating code.

References

  1. CONS Should Not CONS Its Arguments, Part II: Cheney on the M.T.A. , by Henry Baker
  2. CHICKEN Scheme
  3. CHICKEN Scheme - Internals
  4. Chibi Scheme
  5. Compiling Scheme to C with closure conversion , by Matt Might
  6. Lisp in Small Pieces , by Christian Queinnec
  7. R 5 RS Scheme Specification
  8. R 7 RS Scheme Specification
  9. Structure and Interpretation of Computer Programs , by Harold Abelson and Gerald Jay Sussman
  10. The 90 minute Scheme to C compiler , by Marc Feeley
  11. The Development of Chez Scheme , by R. Kent Dybvig

French Bond Risk Hits Euro-Crisis Levels [video]

Hacker News
www.youtube.com
2026-10-03 09:12:28
Comments...

Two-Stack Sliding-Window Aggregation

Lobsters
orlp.net
2026-10-03 08:39:14
Comments...
Original Article

An aggregation is some kind of summary of a set of data. This can be the sum, length, minimum, etc. It is quite common to want to calculate such a summary repeatedly, e.g. “the maximum noise level in dB for the past 30 seconds” for a nuisance detector. In such a case we say there is a sliding window over our data, and we want to aggregate over our window.

If our aggregation is a binary operator with an inverse, like integer sums, there is a very easy solution using a double-ended queue:

from collections import deque

class SlidingWindowSum:
    def __init__(self):
        self.sum = 0
        self.elems = deque()
    
    def push(self, x):
        self.sum += x
        self.elems.append(x)
    
    def pop(self):
        self.sum -= self.elems.popleft()
    
    def eval(self):
        return self.sum

But what if our operator has no inverse? This is actually the case for most interesting summaries such as minimum, quantile, approximate unique count (for example using HyperLogLog ), etc. In fact, even something as simple as a floating-point sum suffers from the fact that floating-point addition is not invertible. For example, if you ever have a NaN in your input data with the above naive algorithm your sum will forever remain NaN , even long after the bad value has left your window.

Six years ago I came up with an algorithm for maintaining just the minimum/maximum in a sliding window and posted it to cs.stackexchange . I now consider this algorithm pointless, because it turns out there is a simple and efficient algorithm that solves this problem for a very wide class of aggregations. I’m writing this blog post to spread the word, because I feel it should be more widely known.

Folklore

I came across this algorithm while reading a far more advanced paper, Low-Latency Sliding-Window Aggregation in Worst-Case Constant Time by Tangwongsan et al. Why is this paper titled low-latency? Because it does the same as what I’m about to describe, but in O(1) time for each step. However, in it they also described a “two-stack” algorithm, which does it in amortized O(1), and is far, far simpler.

Funnily enough that paper attributes this algorithm to “adamax” from a 2011 Stack Overflow post . They in turn credit a 2001 lecture note by D. Sleator for the inspiration. However, this lecture note does not describe a sliding window aggregate, it describes the classical two-stack algorithm for implementing a FIFO queue and does amortized analysis on it. Ultimately I would not be surprised to find that this algorithm was already described in an obscure paper from the 1970s, seeing how simple and brilliant it is.

Two stacks

Like the authors of the paper, I will generalize the two-stack algorithm to arbitrary associative aggregation functions. By abstracting the aggregation as a set of functions, empty() , unit(x) , combine(x, y) and finalize(x) , you can describe many possible aggregations, for example a mean:

empty = lambda: (0, 0)
unit = lambda x: (x, 1)
combine = lambda x, y: (x[0] + y[0], x[1] + y[1])
finalize = lambda x: x[0] / x[1] if x[1] else None

I’d like to note here that these functions have the following signatures:

fn empty() -> Agg;
fn unit(x: Value) -> Agg;
fn combine(x: Agg, y: Agg) -> Agg;
fn finalize(x: Agg) -> Out;

I’m making a distinction here between Value , Agg and Out because while they seem superficially similar for something like an integer sum, for an approximate unique count on strings you would have (Value, Agg, Out) = (String, HyperLogLogSketch, u64) , three wildly different types.

Without further ado, the algorithm:

class TwoStackAgg:
    def __init__(self):
        self.values = []
        self.values_agg = empty()
        self.cum_aggs = []
    
    def push(self, x):
        self.values.append(x)
        self.values_agg = combine(self.values_agg, unit(x))
    
    def pop(self):
        if not self.cum_aggs:
            cum_agg = empty()
            while self.values:
                cum_agg = combine(unit(self.values.pop()), cum_agg)
                self.cum_aggs.append(cum_agg)
            self.values_agg = empty()
        self.cum_aggs.pop()
    
    def eval(self):
        return finalize(
            combine(self.cum_aggs[-1], self.values_agg)
            if self.cum_aggs else self.values_agg
        )

That’s it, the entire algorithm. There’s two stacks ( values and cum_aggs ) and one more aggregate, values_agg . At any point in time values_agg holds the aggregate of values , and cum_aggs contains the cumulative aggregates of all values in our window that aren’t in values , in reverse order. From this we can get the aggregate over our entire window in constant time by by combining the last value of cum_aggs with values_agg .

The neat part is that (assuming w is our window size) every w th operation we drain all of values and maintain a running aggregate while pushing the partial cumulative aggregates onto cum_aggs . This is what makes it amortized O(1), doing O(w) internal operations every w th pop bounds the total amount of work per element to O(1) , even though a singular operation might not be constant time.

I think this is best visualized. Suppose we sum [1, 2, ..., 10] with a fixed-size sliding window of four elements, then the state on each eval() call would look like this ( values_agg not shown as it is simply the aggregate of the values):

             cum_aggs values        out
                   [] []            = 0
                   [] [1]           = 1
                   [] [1, 2]        = 1 + 2
                   [] [1, 2, 3]     = 1 + 2 + 3
                   [] [1, 2, 3, 4]  = 1 + 2 + 3 + 4
[4, 3 + 4, 2 + 3 + 4] [5]           = 2 + 3 + 4 + 5
           [4, 3 + 4] [5, 6]        = 3 + 4 + 5 + 6
                  [4] [5, 6, 7]     = 4 + 5 + 6 + 7
                   [] [5, 6, 7, 8]  = 5 + 6 + 7 + 8
[8, 7 + 8, 6 + 7 + 8] [9]           = 6 + 7 + 8 + 9
           [8, 7 + 8] [9, 10]       = 7 + 8 + 9 + 10
                  [8] [9, 10]       = 8 + 9 + 10
                   [] [9, 10]       = 9 + 10
                 [10] []            = 10
                   [] []            = 0

In total the memory usage is O(w) , where w is your maximum window size. Note that for simplicity of analysis and the example I assumed a fixed-size window w , but there is nothing about the two-stack algorithm that requires this. You can call push(x) and pop() as many times as you’d like between each eval() , growing and shrinking the window size as needed.

Floating-point non-associativity

Note that we required above that our aggregate combine is associative, meaning:

combine(combine(x, y), z) = combine(x, combine(y, z))

Technically speaking, floating-point addition doesn’t respect this. Nevertheless, the above algorithm is still very useful because the results closely match the expected outcome, even more so if you use a compensated summation algorithm like Kahan summation .

However, there is a second very useful property of the above algorithm. Each aggregate is strictly a combination of the elements in the window, and none outside the window. This means if your window contains a NaN or infinity (or some other outlier), that value only poisons the windows that contain it rather than the rest of your computation.

But even without NaN or infinity it is useful, due to not propagating errors endlessly. E.g. if your sliding window starts with [1e20, 1] , this is what would happen with a naive rolling sum:

>>> 1e20 + 1 - 1e20 - 1
-1.0

Compensated summation will reduce these effects, but not making your result depend on values outside of the window will eliminate long-term error accumulation entirely.

Go JSON v2 Migration: What Breaks in Go 1.27

Lobsters
importstatic.com
2026-10-03 08:03:08
Comments...
Original Article

Go 1.27 put encoding/json/v2 in the standard library, and the part most people will miss is that you are already running it. Since Go 1.27, the old encoding/json package is implemented on top of the v2 engine with a set of compatibility options switched on. Your output should not change, but your error strings can, and the moment you change an import to encoding/json/v2 the defaults flip: nil slices become [] , field names match case-sensitively, duplicate keys are rejected and time.Duration stops marshaling at all.

Go 1.27.0 was released on 2026-08-19 and the current point release, as of 2026-10-03, is Go 1.27.1. Every program and output below was run on go1.27.1 darwin/arm64. The benchmark section also uses GOEXPERIMENT=nojsonv2 to compare against the old implementation.

What Go 1.27 changed even if you never import v2

The Go 1.27 release notes describe two new packages and one swapped engine:

  • encoding/json/v2 is the new semantic API: Marshal , Unmarshal , MarshalWrite , UnmarshalRead , MarshalEncode and UnmarshalDecode , each taking variadic Options .
  • encoding/json/jsontext is the lower-level syntactic layer, with an Encoder and Decoder that work on tokens and raw values.

The notes then say that encoding/json "is now backed by the v2 implementation. Marshaling and unmarshaling behavior is preserved, but the exact text of error messages may differ." The v1 API stays supported and nobody is required to migrate. If something does break after the upgrade, building with GOEXPERIMENT=nojsonv2 restores the original code, but the release notes say that opt-out "is expected to be removed in a future release."

So the first migration step costs nothing: upgrade to 1.27 and run your tests. The one category to look for is tests that compare err.Error() against a fixed string. Those can fail without any behaviour change, and the fix is to assert on error types such as *json.SyntaxError or *json.UnmarshalTypeError instead.

Go JSON v2 vs v1, side by side

The package documentation for encoding/json has a "Migrating to v2" section listing every behaviour difference and the option that controls it. Reading a list is one thing; seeing your own structs change is another. Here is the struct we used:

Go

type User struct {
	Name    string   `json:"name"`
	Tags    []string `json:"tags"`
	Admin   bool     `json:"admin,omitempty"`
	Retries int      `json:"retries,omitempty"`
}

The program marshals and unmarshals the same values with jsonv1 "encoding/json" and "encoding/json/v2" and prints both. This is the real output:

Terminal

== nil slice and omitempty on bool/int ==
v1   {"name":"ana","tags":null}
v2   {"name":"ana","tags":[],"admin":false,"retries":0}
== case-insensitive field names ==
v1   "bob" err=<nil>
v2   "" err=<nil>
== duplicate names ==
v1   "mallory" err=<nil>
v2   "alice" err=jsontext: duplicate object member name "name"
== invalid UTF-8 ==
v1   "caf�" err=<nil>
v2   "" err=jsontext: invalid UTF-8 within "/name" after offset 12
== time.Duration ==
v1   {"name":"backup","timeout":90000000000}
v2   error: json: unable to marshal from Go time.Duration within "/timeout": no default representation
v2+  {"name":"backup","timeout":90000000000}
== HTML escaping ==
v1   {"q":"<a&b>"}
v2   {"q":"<a&b>"}
v2+  {"q":"<a&b>"}

Going through them in order of how likely they are to hurt.

Nil slices and maps. v1 writes null for a nil slice or map; v2 writes [] and {} . That is friendlier to JavaScript clients and a silent contract change for anyone who checks === null . json.FormatNilSliceAsNull(true) and json.FormatNilMapAsNull(true) bring the old output back.

omitempty means something else. The first line shows "admin":false,"retries":0 reappearing. In v1, omitempty drops false, 0, nil pointers and nil interfaces. In v2 it drops a field only if it would encode as JSON null , "" , {} or [] . The documentation says the two agree for strings, slices, maps and arrays, and that existing uses on a bool, number, pointer or interface "should migrate to specifying omitzero instead", which works the same in both versions.

Case-insensitive field matching is gone. v1 happily put {"NAME":"bob"} into the Name field. v2 matches names exactly, so the field stays empty and, notice, no error is returned. If your clients send inconsistent casing, you will find out from missing data, not from logs. json.MatchCaseInsensitiveNames(true) restores the loose match for a whole call; the case:ignore tag option does it per field.

Duplicate keys are an error. v1 let the last "name" win, which is how {"name":"alice","name":"mallory"} became mallory. The v2 documentation's security section explains why this matters when two services parse the same request differently. v2 rejects it unless you pass jsontext.AllowDuplicateNames(true) . One detail from our run: after the error, the destination already held "alice" . Treat the struct as garbage once Unmarshal returns an error.

Invalid UTF-8 is an error. v1 replaced the bad byte with U+FFFD; v2 refuses, unless jsontext.AllowInvalidUTF8(true) is set. If you ingest data from old Latin-1 systems, this is the one that will show up in production.

time.Duration has no representation. In v2, a time.Duration "has no default representation and results in a SemanticError". That is the v2 error line. The v2+ line passes jsonv1.FormatDurationAsNano(true) , an option that lives in the v1 package, and gets the old nanosecond integer back.

HTML escaping is off. v2 uses minimal escaping, so <a&b> goes out as is. If you embed JSON in HTML <script> blocks, keep jsontext.EscapeForHTML(true) .

The documentation lists a few more that our structs did not hit: Go arrays must now be unmarshaled from a JSON array of exactly the same length, byte arrays (not slices) become Base64 strings, maps are no longer sorted when marshaled unless you pass json.Deterministic(true) , the string tag option only applies to values that encode as numbers, unmarshaling null always zeroes the target, and malformed struct tags are reported as errors at runtime instead of being ignored.

You do not have to pick a side per call site. Most differences can be pinned on the type, so the same struct encodes identically whichever package touches it. This version of the struct passes a test that marshals with both and compares bytes:

Go

// Portable keeps the same JSON under v1 and v2.
type Portable struct {
	Name    string   `json:"name,case:ignore"`
	Tags    []string `json:"tags,omitempty"`
	Admin   bool     `json:"admin,omitzero"`
	Retries int      `json:"retries,omitzero"`
}

func TestPortableSameBytes(t *testing.T) {
	for _, p := range []Portable{{Name: "ana"}, {Name: "bo", Tags: []string{"x"}, Admin: true, Retries: 3}} {
		b1, err := jsonv1.Marshal(p)
		if err != nil {
			t.Fatal(err)
		}
		b2, err := json.Marshal(p)
		if err != nil {
			t.Fatal(err)
		}
		if string(b1) != string(b2) {
			t.Errorf("v1 %s != v2 %s", b1, b2)
		}
		t.Logf("%s", b2)
	}
}

Terminal

=== RUN   TestPortableSameBytes
    main_test.go:31: {"name":"ana"}
    main_test.go:31: {"name":"bo","tags":["x"],"admin":true,"retries":3}
--- PASS: TestPortableSameBytes (0.00s)

omitempty on the slice makes nil and empty both disappear, which sidesteps the null versus [] question. omitzero replaces omitempty on the bool and int. case:ignore keeps accepting "NAME" under v2, and a second test confirmed it decodes into Name .

While you are in there, v2 has an option v1 never had. json.RejectUnknownMembers(true) turns a typo or an unexpected field into an error that wraps json.ErrUnknownName :

Terminal

json: cannot unmarshal JSON string into Go main.Portable: unknown object member name "role"

That is worth having on any request body that reaches an authorization check.

If you used v2 under GOEXPERIMENT=jsonv2 on Go 1.25 or 1.26, check your tags again. The release notes list changes made before the package became final: the format and unknown tag options were removed, the DiscardUnknownMembers option and SkipFunc were removed, and the inline tag option was renamed embed . Blog posts written during the experiment still show format: tags that will not work on 1.27.

Migrating one option at a time with DefaultOptionsV1

For a large codebase, the safe path is the one the official migration guide calls option-by-option. Change call sites to the v2 functions but pass jsonv1.DefaultOptionsV1() , which the guide calls "a trivial and safe change". Then switch individual behaviours to v2 by adding options after it, since later options override earlier ones:

Go

b, err = json.Marshal(u, jsonv1.DefaultOptionsV1())
show("v2v1", b, err)
b, err = json.Marshal(u, jsonv1.DefaultOptionsV1(), json.FormatNilSliceAsNull(false))
show("mix", b, err)

Terminal

v2v1 {"name":"ana","tags":null}
mix  {"name":"ana","tags":[]}

The first call is byte-for-byte v1. The second is v1 except for nil slices. Each option you flip is a small, reviewable change, and when the list of overrides covers everything you can drop DefaultOptionsV1() entirely.

For servers where you cannot predict every payload from tests, the guide points to github.com/go-json-experiment/jsonsplit . It has call modes such as CallBothButReturnV1 , which runs both implementations, returns the v1 result and reports any difference, so production traffic tells you which options you need before you switch to CallBothButReturnV2 or OnlyCallV2 . The guide warns that running both "will approximately double the cost of marshaling". Use it for the migration window, not permanently.

What we measured: faster unmarshal, slower marshal

The release notes say "marshal performance is broadly at parity with the previous implementation, while unmarshal performance is significantly faster." We checked that on one payload: a 500-element array of a struct with an int64, a string, a float, a three-element string slice, a two-entry map[string]string and a nested struct, about 79 KB of JSON. The benchmark calls plain encoding/json , run twelve times on the default 1.27.1 build and six times with GOEXPERIMENT=nojsonv2 , on an Apple M4 Pro:

Times are rounded medians; the unmarshal runs were noisy (537 to 878 µs on the new engine), while the allocation counts were identical on every run. Unmarshal came out about 25% faster with 29% fewer allocations, which matches the release notes. Marshal took about 39% longer on this payload, which does not match "broadly at parity". One payload on one machine is not a verdict, but it is enough to say: if marshaling is on your hot path, benchmark it on 1.27 with and without nojsonv2 before you roll out.

If you are upgrading for the other Go 1.27 changes as well, the new synctest.Sleep and in-memory test server are worth a look in the same pass.

Common questions

Do I have to migrate to encoding/json/v2 in Go 1.27?

No. The Go team says the v1 encoding/json API will continue to be supported and users are not required to migrate. In Go 1.27, though, v1 is implemented on top of the v2 engine, so error message text can change even if you never touch your imports.

How do I turn off the new JSON implementation in Go 1.27?

Build with GOEXPERIMENT=nojsonv2. That restores the original v1 implementation behind encoding/json. The release notes say this opt-out is expected to be removed in a future release, so treat it as a stopgap and file an issue for whatever made you need it.

Why does json/v2 fail to marshal time.Duration?

In v2 a time.Duration has no default representation, so Marshal returns a SemanticError. Pass the encoding/json option FormatDurationAsNano(true) to get the v1 behaviour of encoding nanoseconds as a JSON number.

Is omitempty the same in json v2?

Only for strings, slices, maps and arrays. v2 omits a field when it would encode as JSON null, an empty string, an empty object or an empty array, so false and 0 are no longer omitted. Use omitzero on bools, numbers, pointers and interfaces; it behaves the same in v1 and v2.

Is encoding/json/v2 faster than v1?

The Go 1.27 release notes say unmarshal is significantly faster and marshal is broadly at parity. In our benchmark on Go 1.27.1, unmarshal was about 25% faster with 29% fewer allocations, while marshal took about 39% longer on that payload, so measure your own hot paths.

ImportStatic articles are researched and drafted with AI assistance, checked against primary sources, and every code sample is compiled and run before publishing. Found a mistake? Here is how corrections work.

Great Question (YC W21) Is Hiring Product Engineers in Canada (Remote)

Hacker News
www.ycombinator.com
2026-10-03 08:02:10
Comments...
Original Article

About Great Question:

Great Question is the all-in-one AI customer research platform for understanding your customers. Our platform enables teams to recruit participants, run research, and share insights – all in one place. Backed by world-class investors and trusted by industry-leading teams like Gusto, Experian, Canva, and Brex, we’re building the future of customer research.

We believe AI is a multiplier for creativity, speed, and scale. Everyone on the team is encouraged to lean into AI, not just to be more efficient, but to push boundaries and unlock new possibilities. Whether you’re writing, designing, coding, or analyzing, we expect you to explore how AI can elevate your craft.

Our culture is built on high trust, high agency, and high emotional intelligence. If you’re looking for a place where your voice matters, where work is fun, and where you're empowered to do your best - join us in redefining user research 🚀

Product Engineer (Full-Stack) — Canada, Remote

I’m PJ, CTO at Great Question . My co-founder Ned and I built this company on a simple idea: talking to your customers should be easy. This has only become more relevant as it’s become faster to build, but validating your ideas is still slow and painful.

Research, design, and product teams at ServiceNow, Brex and Canva run their customer research on us, and we raised a $13M Series A led by Inovia late last year after going through YC in 2021.

This year we’ve been doubling down on the agentic future of research. We’re super excited about what we’ve built and now we need help to continue rebuilding core parts of our product, agent-first. I’m hiring our first Product Engineers in Canada to do it with us.

The role

Product Engineer here means you take a problem and run with it end to end - from the customer conversation to production to validating whether it actually worked. PM support is light and as-required: a PM helps size the problem, shape the appetite, and call priority, but nobody hands you a spec. From there you're trusted to find the right solution, because you're close to the customer and you have great product sense. That only works with people who are self-directed and genuinely value autonomy and accountability.

You'll build across the product stack - from Rails APIs and React/TypeScript frontends to video streaming, realtime voice agents, agentic search, product analytics, and everything in between - working directly with the founders, our designer, and the rest of engineering.

What we're looking for

How you work. AI is embedded across your entire workflow - not just used for coding. You might spend a morning digging through interview transcripts and product analytics to understand a problem, build a few throwaway prototypes to see which approach holds up, and then build a technical plan before handing it to a couple of agents to build and QA.

We’re expecting great judgment around all of that: breaking work into sensible pieces, evaluating the output carefully, and most importantly, owning whatever ships, regardless of who typed it.

What you build. You build highly usable, customer-centric products that people love. You take the extra time to really dig into the data to make sure you understand the problem; and validate that what you’ve built actually solves it. You’ll go the extra mile to deliver something with UI + UX polish, while also thinking in systems to make sure that you’re using the right primitives so you’re delivering less bespoke solutions.

Who you are. We're a small team of exceptional engineers, and we only hire people who raise the bar.

  • Strong across the full stack. An expert in both exposing APIs and the frontend it sits upon. We’re a Rails and TypeScript team, but it’s not a hard requirement.

  • Built and run production systems. Shipping features live and then maintaining them long enough to see what works (and what doesn’t!)

  • Live in an agentic coding harness daily (Claude Code, Codex, Cursor - we don't care which) and have real opinions about your setup.

  • Ideally shipped LLM-powered features, or least know what it takes to run them in production.

  • Have shipped things where you decided what to build, and can defend your UX opinions.

  • Write clearly. We're remote and async-heavy; people who write clearly think (and prompt) clearly.

  • Are senior (roughly 5+ years with real ownership). Our level grading is tougher than most.

How we build

We've rebuilt how we work around AI agents, to ensure our engineers are unblocked to work at AI native speeds:

  • Main goes to production and approvals aren’t always required. Fast, automated CI is the gate, and we trust people to make the right judgements. Good work goes live as soon as it’s finished.

  • We invest in making the codebase easy to work in: written runbooks and conventions that help agents and new people find their way around quickly.

  • Running a few agents in parallel is easy, and expected. Sandboxes and preview environments are trivial to spawn.

  • You'll always have the latest tools and models, and we're generous with tokens.

  • Light on ceremony, high on accountability and communication.

This might not be for everyone, but if it excites you, please apply. We’re especially keen to meet people who are willing to challenge some of these ideas with their own hard-won lessons from the arena.

Your first 90 days

We expect you to hit the ground running, and fast.

  • Week 1: use the harness to learn the codebase and merge your first PR, ideally on day one. You’re deep in the customer data, in their Slack channels, and solving problems.

  • First 30 days: ship into a live area. That might be AI Moderator quality, or working directly on our global chat experience. You understand the product, the competitors, the problems we’re solving, and have a view on how we can win.

  • By 60 days: You own a surface - how the research agent communicates findings, growing our MCP surface, or rebuilding a core page on the new design system.

  • By 90 days: you're deep on a product area: talking to its customers, deciding what's next, accountable for whether it worked. You’re pushing the strategy with the founders & Head of Product in new directions with confidence.

Who shouldn't apply

  • Ambiguity and shifting priorities wear you down. Problems arrive here half-shaped, and the roadmap has changed under our feet more than once this year. Some people find that energizing; if you don't, you'll hate it.

  • You wait to be unblocked. We’re a small team, remote, and with light processes… resourcefulness is critical to the job. We need people who unblock themselves.

  • You’re not all-in on AI. Dabbling won't cut it here. Agents are how we work, not a thing we're trying out.

  • You'd rather not talk to customers. Product engineers here get on calls, read transcripts, and sit in the support queue sometimes. That closeness is the whole point of the role.

  • You only want the shiny work. Plenty of weeks are polish, reliability, and making something 5% less annoying.

  • You want a PM to frame the problem and hand you a spec. “Just tell me what to build and I’ll build it” will not fly here.

Benefits

  • Full-time employment with proper benefits, generous salary in CAD.

  • Remote, anywhere in Canada. We generally work PST hours, but are flexible.

  • 4 weeks PTO

  • Annual team offsite (this year is San Diego!)

  • Brand new Macbook pro

  • Enough tokens that you'll never be left wanting

  • Freedom to use the latest technologies and vendors

A note on compensation: Our offers sit in a defined band for the level, and set by our internal leveling rubric (scope of ownership, technical depth, previous relevant experience), not by negotiation. Most offers land near the midpoint.

Our interview process

  1. Intro call w/ Cofounder/CTO (30 min). Your story, what you've shipped. Our story, the impact you can have here.

  2. Quality session (~ 60 min). Do you know what quality looks like - both in terms of product and code.

  3. Building session (~90 min) . Build something real in our domain with your own AI tooling, the way you actually work; we walk through the code together as you go. We reimburse AI credits.

  4. Team + values (~60 min) . Meet the people you'll work with and ensure we all see the world the same way.

  5. Founder conversation (~30 min) . Meet the CEO.

  6. References, then offer.

No leetcode. AI tools aren't just allowed - they're expected just as they are in the job.

How to apply

Send a link to something you've built that real people are using - and a few sentences on why you’re interested in this role.

Why You’ll Love Working Here:

  • A culture of customer obsession and curiosity

  • Opportunity to shape and scale a critical business function

  • Remote-first culture with high trust and high autonomy

  • Annual team retreats and virtual events

  • Opportunity to work on cutting-edge AI integrations that matter

  • Competitive compensation & equity

  • Generous PTO, health benefits, and learning stipend

_ Equal Opportunity Statement _

Great Question is committed to providing a workplace free from discrimination or harassment. We expect every member of the Great Question community to do their part to cultivate and maintain an environment where everyone has the opportunity to feel included, and is afforded the respect and dignity they deserve.

Decisions related to hiring, compensating, training, evaluating performance, or terminating are made fairly, and we provide equal employment opportunities to all qualified candidates and employees. We examine our unconscious biases and take responsibility for always striving to create an inclusive environment that makes every employee and candidate feel welcome.

How to Hack Time, With C2PA

Lobsters
www.da.vidbuchanan.co.uk
2026-10-03 07:58:14
Comments...
Original Article

By David Buchanan (aka retr0id), 2 nd October 2026

The most impressive hacking stunt from cinema history comes from Kung Fury (2015) , in which Hackerman hacks time itself. He uses this power to correct historic misdeeds. But what would I do with the ability to hack time? Personally I'm more afraid of the butterfly effect, so I'd just go back a few hours to tell myself the winning lottery numbers.

As it happens, that's exactly what I did, according to the cryptographically unforgeable C2PA metadata of this image:

Yes, you can tell it's photoshopped. I'm not trying to do actual lottery fraud here.

You can verify the C2PA metadata including the timestamp at https://verify.contentauthenticity.org/ (if you're reading this in the future, maybe they've introduced mitigations).

You can confirm that those are the winning lottery numbers, shown hours before the draw time, at https://www.euro-millions.com/results/28-08-2026 .

Background

A typical C2PA manifest contains two signatures. The first is the "claim" signature, and in the case of a camera app the claim might be something like "this is a captured photograph, taken at these GPS coordinates, at this time" (except expressed more formally, per the C2PA spec).

In my previous article I showed that the claim signature is approximately worthless, and we can sign whatever claims we like (at least, we can in the case of flagship C2PA implementations like the Google Pixel Camera app).

But there's usually a second signature, from a Time Stamp Authority (TSA), and this one is more interesting. We don't trust devices to tell their own time, so instead a remote server is consulted, via the RFC 3161 protocol. The server sends back a signature that asserts "yes, I saw this hash at this timestamp", and this response is embedded into the C2PA metadata. As long as we trust that the server isn't misbehaving, this proves that the data being signed existed at-or-before the specified timestamp. The TSA acts like an independent witness.

For the Pixel 10, Google decided that devices can in fact be trusted to tell their own time. I have not evaluated their "on-device trusted time-stamp" implementation yet, but I'm highly sceptical of it.

Despite this, in this article we will not be attacking the TSA: we assume it functions as advertised.

It's Time-Hacking Time

If the TSA mechanism is secure, how are we going to hack time?

We're going to use my favourite bug class: the spec footgun .

The footgun is as follows: C2PA allows for arbitrary "exclusions". These are byte ranges within the file which are excluded from signature calculations. Yup. Really. This is already known (it's literally in the spec ), but for some reason nobody's done anything about it yet.

In fact, Dr. Neal Krawertz explicitly called it out in his " Big Bulleted List " of C2PA flaws published June 2025:

  • Large exclusion range. The manifest typically excludes a very large byte range from the signatures. Any excluded bytes can be altered without detection.

The only new angle here is that a malicious signer can deliberately use a large exclusion range, rather than merely doing so incidentally. We can exclude the entire file, to produce an entirely valid signature over an empty string. This allows the file to be tampered with after the fact, without invalidating the signature, and without invalidating the TSA's timestamp proof.

Time status: Hacked

So I really did take a picture of a lottery ticket, and attach a valid C2PA signature with a valid trusted timestamp. But I crafted the manifest to exclude the whole file, allowing me to photoshop it (poorly) after the numbers were announced, without invalidating any of the signatures.

Using c2patool -d to dump the manifest of my PoC file, we can see the important part:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
"c2pa.hash.data": {
  "exclusions": [
    {
      "start": 0,
      "length": 3995383
    }
  ],
  "name": "jumbf manifest",
  "alg": "sha256",
  "hash": "47DEQpj8HBSa+/TImW+5JCeuQeRkm5NMpJWZG3hSuFU=",
  "pad": []
},

3995383 is the length of the entire file, and 47DE...uFU= is the hash of an empty string:

$ openssl sha256 -binary /dev/null | base64
47DEQpj8HBSa+/TImW+5JCeuQeRkm5NMpJWZG3hSuFU=

The claim signature is over that hash, and the timestamp signature is over the claim signature. We're effectively signing nothing at all, but as of today all the C2PA verification tools I can find don't flag anything as unusual.

Can it be fixed?

Not easily! Sure, it's trivial to detect when the entire file has been excluded and report it as invalid, but what if only a small part is excluded? How do you tell whether it's something harmless, or something that could completely change the appearance of the image if modified? (See MD5 hash collision PoCs for examples of the latter).

My first draft of this post ended here with "My recommendation is that the exclusion feature should be excluded from the C2PA spec," but it's not that simple!

The main reason exclusions exist in the first place is that certain file formats effectively require it. For example in PNG files , each chunk has a CRC32 checksum, which needs to be corrected after the signature has been embedded into the file. This would create a circular dependency, unless the CRC32 is excluded from the signature's coverage (ignoring clever mathematical tricks that could avoid invalidating the CRC).

With that in mind I think the right solution here is to carefully and explicitly specify which parts of a file are allowed to be excluded, for each supported file format, and require that verifiers enforce these constraints.

All blog content produced by thinking meat , unless noted otherwise.

Homepage - Blog Index - RSS

Security in the LLM Age

Lobsters
www.youtube.com
2026-10-03 07:57:02
Comments...

The Escalation of War in Ethiopia

Hacker News
www.africanistperspective.com
2026-10-03 07:54:24
Comments...
Original Article

Thank you for being a regular reader of An Africanist Perspective . If you haven’t done so yet, please hit subscribe to receive timely updates on new posts along with over 36,000 other subscribers. New regular content is free. Book reviews and the archives are gated.

This is from the ICG:

Northern Ethiopia is slipping into renewed war. After months of building tensions, on 23 September large-scale hostilities broke out between the Tigray People’s Liberation Front (TPLF) and Amhara Fano militias, on one side, and the federal government and allied Tigrayan forces on the other. Responsibility for the escalation lies primarily with the TPLF, which launched ground assaults in coordination with other rebels with whom it has joined hands in a new alliance that vows to overthrow Ethiopia’s government. The upsurge in fighting comes almost five months after the TPLF retook power in the Tigray region, dealing a blow to a 2022 peace deal that had ended the last Tigray war and leading Addis Ababa to mount an intensifying drone campaign against TPLF forces.

I maintain that war in Ethiopia (and the wider Horn) is not inevitable . The current outbreak of war reflects the fact that Ethiopian elites keep choosing war over peace. Unfortunately, despite repeatedly choosing war, Ethiopian leaders continue to enjoy support from a cast of warmongering academics, journalists, public intellectuals, civil society actors, and a whole host of foreign actors — all of whom have long taken sides and become cheerleaders for their respective camps.

Let’s be blunt. There are a lot of things not to like about the Abiy Ahmed administration. The strong elements of autocratic intolerance, often expressed in the form of indiscriminate force. The unwillingness to (publicly) countenance alternative worldviews and policy positions. The abiding reliance on the old vindictive governance styles anchored around identity politics and collective punishment. And the hubristic promotion of a personality cult, which makes it difficult to build the broad and durable coalitions needed to govern a diverse country like Ethiopia.

Perhaps the political situation would be marginally better if Prime Minister Abiy were a better politician.

However, I doubt it. To understand why, it’s worth stepping back to appreciate the structural drivers of elite instability and conflict in Ethiopia :

The last three decades in Ethiopia have been a search for a new myth. The ethno-federalist system had legitimate logic: bringing about the dignity of (cultural and linguistic) difference between nations and nationalities. However, its rhetoric was drawn from the difficult past instead of the hope of better future. To make matters worse, it became a breeding ground for social and economic injustice.

In the absence of farsighted political elites who may have been able to craft a new inclusive myth out of the stories of nations and nationalities, ethnic groups had to walk back to find their stories in their own small compartments. This exacerbated narrow ethnic histories and ideals.

It’s easy to demonize Abiy. But the hard truth is that Ethiopia’s structural problems preceded him and, if left unresolved, will outlast him.

Those beating the war drums the loudest in Ethiopia — and who’ve trashed the Pretoria Agreement — are convinced that Ethiopia can only be well governed if they are in charge. Nearly all are motivated by maximalist demands and are quick to weaponize legitimate historical grievances for parochial ends. The leading characters involved are (mostly) men out to settle decades-old scores and who prefer to go out fighting rather than accept political defeat by their perceived inferiors. These same characters have been able to mobilize fighters under the banner of uncompromising primordislist ethno-nationalism and commitments to their respective ethno-states.

In light of all this, Teklehaymanot G. Weldemichel is spot in making the case for following through on the Pretoria Agreement. Yes, it is imperfect and the various belligerents haven’t always followed through on the deal’s requirements. However, the deal “ brought large-scale fighting to an end and contains commitments that address some key unresolved issues .” It’s a start. It ought to be given a chance.

Unfortunately, the plight of civilians hardly ever features in the calculus of the warmongers. Estimate vary, but the death toll in the Tigray war may have been as high as a staggering 600,000. This is a monumentally catastrophic figure. Even the lower bound estimates out there would make it one of the deadliest wars of the 21st century. Given what we know about how belligerents go about waging war in Ethiopia (including by exacerbating the problem of food insecurity), why would anyone support a return to all out war?

Why are Ethiopian elites and their cheerleaders willing to impose such trauma on their peoples?

There is bound to be a lot of discourse on ethnic polarization as a driver of conflict in Ethiopia. However, it is worth also considering how Ethiopia’s brand of federalism contributes to ethnic polarization. Existing research shows that the creation of ethno-states after the end of the civil war increased the salience of ethnic difference :

[I]ndividuals politically socialized under the conditions of ethnic federalism are indeed significantly more likely to embrace an ethnic rather than a national identity. However, generational affects do not appear to be related to perceptions of continuing ethnic federalism as a political arrangement. Differences of opinion on ethnic federalism are more a function of one’s group identity, rather than the intensity of ethnic identification.

This is in line with the historical experience from other jurisdictions (see Lebanon, the UK, etc). Institutionalizing ethnic difference through administrative means or identity-based reservations creates incentives to invest in the targeted identities as loci of political mobilization. From that point it’s a short walk to mobilizing those same differences for conflict under permissive conditions. Notice the implication of the last line in the above quote: once ethnonationalism becomes the dominant game, everyone gets forced to pick a side, regardless of their individual-level intensity of ethnic identification.

Most people kind of admit that a leading structural driver of political instability and conflict in Ethiopia is its brand of federalism :

Ethiopia’s federal system was flawed from the beginning because it didn’t foresee potential sources of conflict or that regional states would make claims against one another. Trust among regional states was never high, and has deteriorated over the last three decades. On top of this, the federal government’s ability and readiness to mitigate or solve domestic conflict has been open to question. Currently, there are regions and regional leadership that are having difficulty working together. The federal government at the centre is too weak to impose its will on the regional administrations. The result is that there aren’t common political and economic national standards across the country.

To be clear, Ethiopian civilians do not want war; and also don’t trust the state’s ability to avoid or end wars. And while reported inter-ethnic hostility remains low (despite the historical wars), the rate of encounters with state-directed ethnic discrimination ranks among the highest in Africa.

In Round 9 of the Afrobarometer Survey, 55% of respondents noted that the state is doing badly at preventing or resolving violent. In the same survey, less than a quarter (23.7%) of respondents reported feeling that their ethnic identity trumps their Ethiopian identity. A roughly similar proportion (23.6%) also reported not at all trusting people from other ethnic groups. Meanwhile, 39.3% reported feeling that their ethnic group is never treated unfairly by the government.

There are two takeaways from this. First, the problem appears to be the state/institutions rather than society. The vast majority of Ethiopians seem open to the idea of multicultural co-existence. However, the state (and supporting institutions — e.g., ethnic federalism) seem intent on balkanization by reinforcing difference in everyday interactions. Second, while Ethiopia is in the bottom fifth of the distribution across the Continent, it’s not particularly unique (unlike what many Horn scholars would like to believe). There are other countries with close or worse results on the metric of state-directed ethnic discrimination.

Overall, the most important driver of conflict in Ethiopian is the state’s centuries old inability to deter or crush rebellions in their infancy. All polities have pockets of grievances against the state that could potentially give rise to armed rebellion. While some grievances invite legitimate armed resistance, others are the manifestations of taste-based warmongers, war profiteers and their clients, and useful idiots carrying water for geopolitical adversaries. Regardless of their founding motivations, not all would be rebellions materialize — because capable states don’t let them.

Instead of monopolizing the use of violence, the Ethiopian state has adapted to using indiscriminate violence and chaos as governance styles (to manage intra-elite political instability, and to fight back against mass level identity-driven centrifugal forces). Collective punishment of whole communities is common, and has, in turn, radicalized ever more people against the state. Consequently, there won’t be any shortage of young people ready to take up arms against the state in this round of wars.

The Ethiopian state’s inability to deter or eradicate violence invites all manner of foreign meddling. This is terrible for conflict duration, as it prevents both state and non-state belligerents from fully internalizing the cost of war.

Lots of countries — including Egypt (the Nile/wider Red Sea supremacy), Eritrea ( Ethiopia’s desire for access to the sea ), Sudan (Ethiopian support for the RSF), Somalia (historical enmity/Ethiopian support for Somaliland), and Saudi Arabia (as benefactor of Egypt/Sudan/Eritrea axis) — have good reasons to keep Addis Ababa weak and preoccupied with unending domestic wars. These dynamics, coupled with the geography of war in Ethiopia mean that Addis Ababa will struggle to keep arms and fighters from easily moving across its borders with Sudan and Eritrea.

On its part, the Ethiopian state will likely continue to enjoy support from the United Arab Emirates; and maintain the capacity to produce the weapons needed to keep rebels at bay. Ominously, divisions among elites in Tigray and other regions mean that the rebellion may not clear the threshold needed to incentivize a quick resolution either militarily or at the negotiating table. Therefore, slow-burning, but still catastrophically deadly conflicts may be on the cards.

A key risk facing Ethiopians is that, like in Sudan , international/geopolitical interests and considerations may quickly overshadow the domestic cost of the war, especially on civilians. The inability to extricate the domestic bargaining “game” from regional/geopolitical considerations will unnecessarily complicate peace negotiations and very likely prolong the conflict.

Notably, a prolonged conflict will likely also be significantly civilianized. Careful to avoid being overstretched, the Ethiopian state may decide to deputize regional militias and other violence entrepreneurs in its fights against subnational armed groups. To this end, it will lean heavily on divisions across armed groups and within the regional states that have witnessed rebellions.

As noted above, it would be a mistake to characterize political instability and conflict in Ethiopia as being caused by any single leader. Removing Abiy from power — the stated goal of the latest armed coalition against the Ethiopian state — will not solve the many structural problems facing Ethiopia. Which is why it is important for all involved to have honest conversations about the fundamental causes of conflict, and how to get out of the cycles of violence. Four issues immediately come to mind:

  1. The thorny question of what it means to be Ethiopian, in a context of diverse peoples, nations, and nationalities: What does it mean to be Ethiopian? Should Ethiopians deal with their rich cultural tapestry and history by genuflecting to ideas motivated by narrow ethno-nationalisms or embracing cosmopolitan multiculturalism as a national goal? These questions do not have easy answers, and I cannot pretend to offer any here. However, it’s high time Ethiopians openly confronted these questions rather than ceding ground to those out to weaponize the country’s history to stoke ethnic tension and wage perpetual wars of self aggrandizement.

  2. The meaning of self-determination in a cosmopolitan 21st century state: What is the correct structure of the Ethiopian state? And what should be the balance of power between national and subnational units? If the resolution to the questions above converge on keeping Ethiopia whole (as I hope is the case), the next question would be to come up with institutional and administrative structures that at once give ordinary Ethiopians a sense of self-determination at the local level and empowers the center to prevent centrifugal tendencies from collapsing the state.

  3. How to stabilize Ethiopia’s regional relations: While Ethiopia’s problems are principally domestic, it’s also true that its neighborhood (comprised of states with histories of internal conflict) and antagonistic foreign posture (especially vis-a-vis Somalia and Eritrea) exacerbate this problem. Of course there are things that Addis Ababa cannot change about its foreign relations (e.g., need to check Egypt’s supremacy ambitions). However, there are others it can. It can choose to pursue a path that ends the proxy wars with Sudan, Eritrea, and Somalia. This will reduce overall conflict risk in the wider Horn, while also starving would-be Ethiopian rebels from getting ready support from regional rivals and generally reducing the size of the Horn’s war economies (arms, fighters, ideologies).

  4. What’s the positive case for Ethiopian unity? Is it to reclaim the country’s imperial glory? Is it to aggrandize the person of the Ethiopian leader? Is it to become a geopolitical powerhouse in the Horn and the Red Sea regions? Or is it to deliver economic growth, development, and prosperity for Ethiopians? Some commentators like to compare Ethiopian with the former Yugoslavia, thereby raising the specter of partition. This should be viewed as a challenge for those that want to see a united Ethiopia. If people are to continue buying into the Ethiopia Project , there has to be a positive case for doing so. Merely relying on history or inertia will not be enough. And importantly, the pitch has to be sufficiently inclusive and must accommodate changing public attitudes. In particular, Ethiopian elites must update and understand that the old political culture of violence, domination, and humiliation as styles of governance will not go down well with current and future generations of Ethiopians.

In the final analysis, and understanding that the sources of instability and conflict in Ethiopia are principally domestic, there is really no better way to put it that this :

Ethiopia’s federal government must be pressed to exercise restraint and return to meaningful political negotiation. Military force and coercion cannot become the first choice for resolving disputes. Tigrayan political and security actors must likewise recognise that political disagreements cannot be sustainably resolved through armed mobilisation. External powers should not be permitted to turn Ethiopia’s internal political disputes into a wider contest for influence. The involvement of the United Arab Emirates, Eritrea and others supporting the warring parties should be addressed through transparent diplomacy. Their competing interests risk fuelling a wider confrontation.

Zig 0.17 released

Linux Weekly News
lwn.net
2026-10-03 07:31:18
Version 0.17 of the Zig programming language has been released. This release features 5 months of work: changes from 206 different contributors, spread among 925 commits. Originally predicted to be shorter, this release cycle ended up [being] substantial, with the Build System reworked, including...
Original Article

Version 0.17 of the Zig programming language has been released.

This release features 5 months of work : changes from 206 different contributors , spread among 925 commits .

Originally predicted to be shorter, this release cycle ended up [being] substantial, with the Build System reworked, including the introduction of the Build Server Protocol , and the ELF Linker enhanced to the point where we expect Incremental Compilation to work for everyone on x86_64-linux.

LWN last covered Zig in December 2025.



Seven stable kernels for Saturday

Linux Weekly News
lwn.net
2026-10-03 07:21:42
Greg Kroah-Hartman has announced the release of the 7.2.9, 6.18.55, 6.12.112, 6.6.158, 6.1.189, 5.15.222, and 5.10.271 stable kernels. Each contains a large number of important fixes throughout the tree; users are advised to upgrade. ...
Original Article

[Posted October 3, 2026 by jzb]

Greg Kroah-Hartman has announced the release of the 7.2.9 , 6.18.55 , 6.12.112 , 6.6.158 , 6.1.189 , 5.15.222 , and 5.10.271 stable kernels. Each contains a large number of important fixes throughout the tree; users are advised to upgrade.



to post comments

Our Solar System Is Terminally Unstable and Will Be Completely Destroyed, Study Finds

403 Media
www.404media.co
2026-10-03 07:00:34
The estimated lifespan of the outer solar system has been downgraded from 100 billion years to just a few billion years, according to a study that probed “terminal instability” during the Sun’s death....
Original Article

Welcome back to the Abstract! These are the studies this week that searched for the ur-animals, wandered the poles, rained on Mars, and destroyed the solar system.

First, scientists present new evidence that the first animals appeared more than 800 million years ago, a truly ancient origin that suggests our metazoan ancestors survived a period known as Snowball Earth, when our planet is thought to have been nearly completely frozen. Then: Earth’s poles won’t sit still, the sepulchral stuff of life, and a dramatically shortened lifespan for our solar neighborhood.

As always, for more of my work, check out my book First Contact: The Story of Our Obsession with Aliens , or subscribe to my personal newsletter The BeX Files .

I trace my ancestry to Snowball Earth

Durbin, Orin Lole et al. “Re-evaluating molecular clock maximum age calibrations revives pre-Ediacaran divergence estimates for animals.” Science Advances.

When did the first animals emerge on Earth? It’s a question that has provoked centuries of scholarly debate, and also inspired some wonderful answers in myth and legend (I’m partial to the Bible’s tidy solution: day six checklist).

The earliest animals that are clearly preserved in the fossil record lived during the Ediacara period some 574 million years ago, but the origins of our diverse metazoan family are likely much older. Now, scientists suggest that animals may have appeared on Earth a whopping 800 million years ago, during the ancient Tonian era, according to a new estimate of the dawn of animals.

“The timing of the origin of animals on Earth has puzzled scientists for centuries,” said researchers led by Orin Lole Durbin who was at the University of Oxford during this research and is now at Virginia Tech. “Unambiguous animal body fossils within the Ediacara Macrobiota provide a hard minimum calibration of 574 [million years] for crown Metazoa. Animals must have evolved before this date. Determining a maximum calibration for animals (a date at which they had not yet evolved) is a more complex task.”

Estimated timeline for the origin of animals based on molecular clock analyses. The grey bar on the right of the image is the origin of animals based on an Ediacaran upper calibration. The grey bar to the left is the origin based on older deposits in the Tonian period, as suggested by the new study. Image: Orin Lole Durbin.

The team turned to molecular clock analysis, a technique that uses the rate of genetic mutations in lineages as a rough way to reconstruct evolutionary timelines. The researchers focused on two fossil deposits—China’s Weng'an Biota and Mongolia’s Kheseen Biota, which date back 590 and 550 million years respectively—to extend the possible timeline of animals back more than 200 million years earlier.

If animals really did appear in the Tonian era, it means that our earliest metazoan ancestors survived episodes of so-called “Snowball Earth,” when our planet was almost fully covered in ice. These early critters were likely simple marine sponges and comb jellies that lived for hundreds of millions of years until their descendents exploded into the kaleidoscopic variety of animals that persist to this day.

“The possibility of a pre-Ediacaran origin and diversification of animals means we cannot discount hypotheses that link the timing of animal diversification to Cryogenian Snowball Earth glaciations,” according to the study. “The causes and trajectory of animal evolution will remain uncertain until paleontological and geochemical data provide sufficient confidence in maximum calibrations to ensure a reliable timescale.”

In short: Respect your sponge elders.

In other news…

A pole-ing error

Domeier, Mathew et al. “Quadrupolar sea level fluctuations reveal episodes of rapid polar wander.” Science.

Earth is a roiling mess of shifting continents and squishy innards, a situation that constantly throws its spin axis off balance and prompts the geographic poles to drift. This phenomenon, known as true polar wander (TPW), is distinct from Earth’s wandering magnetic poles, which are shaped by interior core-mantle processes.

Right now, TPW is clocked at a slow rate of 10 centimeters per year, but scientists have found evidence of “fast” episodes of TPW that pushed the poles off by thousands of miles over millions of years, which has knock-on effects for ocean and land distribution across the globe.

A team has now pinpointed several of these fast episodes over the past 320 million years, “confirming that protracted rapid TPW has occurred on Earth and that couplings among Earth’s rotational dynamics, mantle processes, and surface environments can be strongly episodic,” according to their study.

“These findings refute the view of TPW as negligible or persistently slow and highlight the need to consider TPW as an episodic control on sea level change and likely other global environmental and biological dynamics,” said researchers led by Mathew Domeier of the University of Oslo.

In addition to these long-term natural changes, recent human activity has slightly impacted TPW , especially the construction of dams and the glacial melt of anthropogenic climate change. So the next time you address a letter to Santa at the North Pole for your kid—or yourself, no judgement—make sure you have the most up-to-date coordinates.

It’s raining formaldehyde—formalallelujah!

Koyama, Shungo et al “Global Distribution of Atmospheric Formaldehyde Deposition Correlated with Water Vapor on a Warm Early Mars.” The Planetary Science Journal.

Formaldehyde, the toxic chemical, has a bit of a morbid reputation—exposure to it can be fatal and it’s a common ingredient used to embalm corpses. But, paradoxically, formaldehyde (H 2 CO) also helped pave the way for the emergence of life on Earth as a precursor of bioessential elements such as sugars, amino acids, and nucleobases.

If life ever flourished on Mars, formaldehyde would likely also have been part of the story. To map out the possible distribution of the chemical on ancient Mars, scientists ran models of its atmosphere between 3.6 and 3.8 billion years ago, when the red planet was warmer and wetter. The results showed that formaldehyde was strongly associated with water vapor, suggesting that the chemical rained down from the ancient skies into Martian waterways.

“Significant atmospheric deposition of H 2 CO is observed over water bodies, including the northern ocean,” said researchers led by Shungo Koyama of Tohoku University. “We suggest that basins adjacent to persistently humid regions, which likely experienced H 2 CO accumulation and subsequent organic synthesis, could be potential landing sites for future missions to investigate ancient chemical evolution and whether it led to the origin of life.”

Hopefully, future missions to Mars will leave any signs of life with no place left to formalde-hide.

Update the cosmic actuary tables

Batygin, Konstantin et al. “Terminal Instability of the Solar System Triggered by Stochastic Solar Mass Loss.” The Astrophysical Journal Letters.

I hate to be the bearer of bad news, but the solar system has been diagnosed with terminal instability and has been given a mere six billion years to live.

That’s the upshot of a new study that modelled the solar system’s future as the Sun becomes a red giant star and, ultimately, collapses into a stellar husk called a white dwarf.

Though the dying Sun will consume Mercury, Venus, and perhaps even Earth, scientists have previously assumed that the giant outer planets might be relatively unscathed by its death—perhaps surviving for up to 100 billion years after the main-sequence lights go out. But updated simulations suggest that these distant objects may be thrown into fatal chaos as early as the red giant phase, and that they are unlikely to survive more than a billion years after the Sun’s transition to a white dwarf.

The future of Earth is bright, literally. Image: Celestia

“[Isaac] Newton, contemplating the drifting orbits of Jupiter and Saturn, suspected that the planetary order was mortal, and three centuries of celestial mechanics…have progressively deferred that verdict, most recently to timescales far beyond the age of the Universe,” said researchers led by Konstantin Batygin of the California Institute of Technology. “Our results return the solar system’s dissolution to astrophysically familiar territory, and relocate its cause: not the slow seep of chaos, nor the chance encounter with a passing star, but the Sun itself, which in dying does not merely enlarge the planetary system it built—it shakes it, and more often than not, spills it.”

“Newton’s envisioned instability is real after all,” the team. “He was mistaken only about the perpetrator.”

What a great story to kick off this spooky season! Planetary order is mortal, but cosmic horror is eternal.

Thanks for reading! See you next week.

Show HN: Germany's new sovereign AI model Kolibri

Hacker News
tej.as
2026-10-03 06:43:51
Comments...
Original Article

Kolibri is an open-weight large language model (LLM) from Aleph Alpha for German and English: a mixture of experts with 78 billion parameters that only uses about 3.5 billion of them for each token it reads or writes. It came out on 3 October 2026 under the Apache 2.0 license , the weights are on Hugging Face , and it was trained from scratch on infrastructure in Germany and Finland. (Kolibri is German for hummingbird, which is cute for a model whose whole trick is being light.)

I live in Germany, and at SmashingConf New York in 2024 I told the room what I’d heard in the US when I said where I’m based: “you regulate, you don’t innovate.” It hurt to hear, and what I wished for on that stage was the middle, “the right balance between innovation and regulation around data privacy, data stewardship, environmental constraints and energy requirements.” Kolibri is a pretty direct answer to that: a German team built it with the EU AI Act in mind “from the ground up”, and in Aleph Alpha’s own evaluation it scores above every compared model of its size in both languages.

Huge congrats to everyone at Aleph Alpha who built it, my good friend Michael Hofmann among them!

This post is about how Kolibri works, where it’s strong, where it isn’t, how to run it, and when it’s the right pick. Everything here comes from Aleph Alpha’s 189 page technical report , the model card and their launch post , plus one experiment I ran on its tokenizer.

What is Kolibri?

Kolibri 1
Parameters 78.1 billion in total, 3.46 billion per token (4.4%)
Languages German and English
Context 262,144 tokens natively, tested up to 1,048,576
License Apache 2.0 for the weights and configuration files (Aleph Alpha keeps the rights to its training code and methods)
Memory about 78 GB of weights in 8-bit floating point (FP8)
Reasoning 4 levels: none, low, medium and high
Tool calling Yes
Knowledge cutoff 18 June 2026
Training about 24 trillion tokens, more than a fifth of them German, on 768 NVIDIA B200 graphics processing units (GPUs)

Aleph Alpha calls Kolibri sovereign, and in their launch post that means 2 things. The first is how it was built: “teams built the model in Germany, trained it on infrastructure in Germany and Finland, under European and German law, with no foreign control.” The second is what customers get: “full freedom of deployment and intellectual-property safety, so compliance comes as an inherited property.” In plain words, a ministry or a car supplier can run it on its own servers, with its data never leaving the building, and nobody can change or switch off the model under them. Aleph Alpha has also signed the European Union ’s General-Purpose AI (GPAI) Code of Practice .

Sovereign doesn’t mean that nothing from outside Europe went in, and the model card says so itself: English web text was rephrased with Google’s Gemma 4 , German with Mistral-NeMo , and Qwen3-32B labeled data for the quality filters. They then filtered the training data for the political bias such models can have, which they’ve measured in Chinese open models themselves.

How Kolibri works

Kolibri is 6 ideas stacked on top of each other, and each one is there to make German cheaper, longer or more honest.

1. 384 specialists, and each token sees 6

In a normal (dense) model, every token goes through every parameter. In a mixture of experts (MoE) , each layer has a crowd of small sub-networks called experts and a router that picks a few of them for each token. Kolibri has 50 layers, each with 384 experts plus 1 shared expert that every token goes through, and its router sends each token to 6 of the 384. That’s how 78.1 billion parameters turn into 3.46 billion of actual work per token.

In 2024 I gave a talk called Why Small Language Models are the future , and I argued for “smaller language models with fewer parameters and fewer places things can go wrong that require lesser compute.” My analogy was a doctor who has read every medical book in the world against a specialist in hematology: go to the first one with a blood condition and “they may not get it right cuz they know too much.” A mixture of experts puts a hospital full of specialists inside one model, and the router is the receptionist who sends each token to the right 6.

The analogy breaks in 2 places though. The experts aren’t neat topics like “German law”: when researchers look inside MoE models they mostly find experts for patterns of tokens, like punctuation or proper nouns, not subjects a person would pick. And the hospital has to keep all 384 specialists on staff even if you only see 6, so Kolibri computes like a 3.5 billion parameter model but needs the memory of a 78 billion parameter one. The model card says it plainly: “the full model must be held in memory even though only part of it is active at any time.”

2. A tokenizer that reads long German words

A model doesn’t read letters or words, it reads tokens: chunks of text from a fixed vocabulary, picked when the tokenizer is trained. German glues words together into long compound words, and a tokenizer that learned mostly from English chops them into pieces. Here’s the German name of the Federal Constitutional Court , split by the tokenizer GPT-4o and GPT-5 use ( o200k_base , through OpenAI’s tiktoken ), and by Kolibri’s:

o200k_base (GPT-5):  Bund | es | ver | fass | ungs | gericht     6 tokens
Kolibri:             Bundes | verfassungsgericht                 2 tokens

Kolibri’s tokenizer has 128,000 tokens, trained with a new algorithm Aleph Alpha calls UniBPE: it keeps the bottom-up merging of byte-pair encoding (BPE) and picks each merge with a different scoring rule (the Unigram objective), which respects how German builds words. The report says it needs 11.2% fewer tokens for German text than GPT-5’s tokenizer, the best of the 9 others they measured.

I wanted to see that for myself, so I ran 6 tokenizers over all of the Basic Law for the Federal Republic of Germany , the German constitution (185 KB of very German legal text), and over its official English translation :

Tokenizer German tokens More than Kolibri English tokens More than Kolibri
Kolibri 1 35,190 39,875
o200k_base (GPT-4o, GPT-5, GPT-OSS ) 41,482 17.9% 39,737 -0.3%
Qwen3.5 35B-A3B 42,907 21.9% 41,650 4.5%
Mistral Small 4 43,478 23.6% 41,301 3.6%
Gemma 4 43,850 24.6% 41,564 4.2%

On legal German, Kolibri needed 15% fewer tokens than GPT-5’s tokenizer, even more than Aleph Alpha’s own 11.2%, and in English it tied with it. Wild. Fewer tokens means fewer steps to read or write the same German text, and more German fits in the same context window. (I counted each one with its own tokenizer.json through Hugging Face’s tokenizers library, except o200k_base , which I counted with tiktoken.)

3. Most layers only look nearby

40 of Kolibri’s 50 layers use sliding-window attention: each token only looks at the 512 tokens before it. Every 5th layer looks at everything before it. It’s like reading a long contract while mostly paying attention to the sentence you’re on, and every few pages stopping to think about all of it, and it’s what keeps a 1 million token context affordable.

There’s a clever detail in there too. Only the sliding-window layers know where a token sits (through rotary position embeddings ), and the full-attention layers don’t, so the context stretches past the 262,144 tokens it was trained on without any extra position tricks. Aleph Alpha validated it up to 1,048,576.

In my talk Unlocking Value with AI Today I called finite context one of “the big three” problems of generative AI, next to hallucination and the knowledge cutoff, and Kolibri goes after all 3. At 1 million tokens, on the RULER long-context test, Kolibri’s base model scores 63.2, against 57.5 for Qwen3.5 35B-A3B’s base model.

4. It thinks in German

Reasoning models think before they answer, and even on German prompts they mostly think in English. Aleph Alpha posted about this on 24 September and wrote it up as Through the Valley of Tears : they generated about 800,000 German reasoning examples, and found that a little German reasoning data is worse than none. Their model’s German math score dropped from 70.2 to 48.3, because its German thoughts kept going around in circles and never finished, and it only climbed back (to 67.3) with a lot more German data.

Kolibri got the lot more. It reasons in German on German prompts, and its German math scores are the best of the models with about 3 billion active parameters: 87.5 on the American Invitational Mathematics Examination (AIME) 2025 in German, against 84.4 for the next best, NVIDIA’s Nemotron 3 Nano .

5. It’s trained to say “I don’t know”

Hallucination was number 1 of my big three, and the fix in that talk was retrieval augmented generation (RAG): you look up good, authoritative information and put it in the prompt. RAG only works if the model admits when the documents don’t have the answer, though, and that’s what Aleph Alpha trained for, with their own method, the Merlin-Arthur protocol .

It works like a game with 3 players. Arthur is the model, and he gets a question with parts of the supporting document hidden. Merlin hides parts so that the correct answer gets easier to find, and Arthur is trained to answer those. Morgana hides the evidence the answer depends on, and Arthur is trained to say he doesn’t know. Arthur never knows which of the 2 he’s facing, so the only way to win is to actually check whether the evidence in front of him supports an answer.

It shows. On Artificial Analysis’s Omniscience test , when Kolibri didn’t know an answer, it said so (or gave a partial answer) 44% of the time instead of making one up. Qwen3.5 35B-A3B did that 11.1% of the time and GPT-OSS 120B 23.7%. Of all the mixture-of-experts models Aleph Alpha compared, only Qwen3.6 35B-A3B did better, at 56.7%.

6. You pick how hard it thinks

Every request can set reasoning_effort to none , low , medium or high , so a quick lookup answers right away and a hard question gets a long think, from the same model on the same server.

What Kolibri is good at

These are the rows where Kolibri leads the open models of its size in Aleph Alpha’s evaluation, which runs every model through the same setup with the sampling settings its makers recommend:

What Kolibri Closest model with about 3B active parameters
Overall score, English 75.5 74.7 (Qwen3.5 35B-A3B)
Overall score, German 70.8 69.8 (Qwen3.5 35B-A3B)
AIME 2025, English 96.9 89.6 (Nemotron 3 Nano)
AIME 2025, German 87.5 84.4 (Nemotron 3 Nano)
Questions over unseen company documents, English 89.7 87.0 (Qwen3.5 35B-A3B)
Long context at 1 million tokens (base model) 63.2 58.5 (Nemotron 3 Nano)

The math is the standout: on AIME 2025 and 2026 in English it beats every MoE model in the comparison, including the ones with 12 billion active parameters, and only the dense Qwen3.8 27B scores higher. The company-documents row is 5 tests Aleph Alpha built from customer-like work in semiconductors, the German public sector, aerospace, an automotive supplier and industrial drives, run over documents and questions the model never saw in training.

What Kolibri is bad at

Aleph Alpha publishes its weakest rows right next to its best ones in the model card, so here they are:

  • It knows less from memory. It’s last of the 12 models on a closed-book test (the Retrieval-Augmented Generation Benchmark ’s questions with no documents, 51.0), and it answers only 14.8% of the Omniscience questions correctly, against 22.2% for Qwen3.5 35B-A3B. It knows when it doesn’t know, which is great, but it also knows less: fine with documents in the prompt, bad as a trivia oracle.
  • Multi-turn tool calling is weaker. On the Berkeley Function Calling Leaderboard ’s multi-turn tests it scores 39.8, against 58.2 for GLM-4.7 Flash and 54.0 for Qwen3.5 35B-A3B.
  • It’s not the best coding agent. It scores 27.7 on Terminal-Bench 2.1 against 39.7 for Qwen3.5 35B-A3B, and 66.4 on SWE-bench Verified against 73.8 for Qwen3.6 35B-A3B.
  • Long context dips in the middle. At 128,000 tokens on RULER, Qwen3.5’s base model scores 89.9 to Kolibri’s 67.9. Kolibri only pulls ahead at the very long end.
  • The memory. 78 GB of weights means a data-center GPU or two, never a laptop, whatever the 3.5 billion suggests.
  • The plumbing is new. It needs Aleph Alpha’s vLLM plugin , which supports one vLLM version at a time (0.29 today), and on launch day no hosted provider serves it yet.
  • 2 languages only. That’s on purpose: the model card calls it “a deliberate choice of depth over breadth.”
  • A bigger dense model beats it. Qwen3.8 27B scores 80.2 in English and 79.9 in German, but it uses nearly 8 times as many parameters for every token, and that’s the trade Kolibri is built around.

How to run Kolibri

You need about 78 GB of GPU memory: 2 NVIDIA A100s or H100s with 80 GB each at the least, or a single H200 , B200 or B300. I haven’t run the model itself yet, since it needs a data-center GPU and nobody hosts it so far, so this is straight from the model card. First the plugin, which installs the vLLM version it supports:

pip install 'aleph-alpha-inference>=1'

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

That gives you an OpenAI-compatible server, so any OpenAI client talks to it, and the reasoning effort goes through the chat template:

# from the Kolibri model card (trimmed)
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
    model="Aleph-Alpha/Kolibri-1",
    messages=[{"role": "user", "content": "Erkläre kurz, was ein Mixture-of-Experts-Modell ist."}],
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high", "enable_thinking": True}},
)
print(response.choices[0].message.content)

The model card recommends temperature=1.0 , top_p=0.97 and top_k=128 , and contexts of at most 262,144 tokens for anything latency-sensitive. For the full million, serve it with 2 more flags:

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --max-model-len 1048576 \
  --hf-overrides '{"max_position_embeddings": 1048576}'

When to use Kolibri

Kolibri is the pick when German text and your own hardware both matter: a public authority, a bank, a manufacturer or an aerospace supplier that has to keep its documents in house, wants answers in German that reason in German, and would rather hear “I don’t know” than a confident wrong answer. RAG over long German documents (laws, contracts, manuals) plays to every strength above: the tokenizer, the 1 million token context and the abstention.

That’s the kind of project I’m working on right now, with Prof. Dr. Heinrich Audebert , who heads neurology at Campus Benjamin Franklin , one of the Charité ’s hospitals in Berlin. Today, a patient with a neurological complaint goes to their general practitioner (GP), and the GP has to see them, which is very demanding for a busy practice. In what we’re building, the patient sits down at a computer in the GP’s practice, an AI avatar takes them through a battery of tests, and it grades how urgent their symptoms are: a referral to a specialist right away, or they can wait a bit. The conversations are in German, they’re about people’s health, and the model has to run where we control it, so we’re thinking of using Kolibri for it.

It’s the wrong pick for a coding agent, where Qwen3.6 35B-A3B leads, for questions the model has to answer from memory, for any language besides German and English, and for anyone who can’t spare 78 GB of GPU memory. For that last case, the small specialized model I argued for in 2024 is still the answer: for my podcast search I fine-tuned Mistral 7B on my Apple silicon laptop instead of paying for GPT-4o.

Which is better for your own documents, a hospital of specialists like Kolibri or one small model trained for a single job, only shows when you try both on your data. Swapping the model under an agent without breaking it is what my workshop on reliable AI agents covers: guardrails and retries that work whichever model sits underneath.

An AI agent emailed researchers for help. It told us why

Hacker News
www.science.org
2026-10-03 06:07:08
Comments...

GitHub's new dashboard experience now the default

Hacker News
github.blog
2026-10-03 05:59:01
Comments...
Original Article

The new dashboard experience, previously available as a feature preview, is now the default view for everyone.

The redesigned dashboard helps you focus on the work that matters most and makes it easier to find and act on what you need next. You can:

  • View your active agent sessions, issues, and pull requests in one place. Use filters to choose what appears in each section, with up to 12 items in each list.
  • Catch up on updates in the separate Feed tab, keeping your feed distinct from your productivity-focused dashboard.
  • Start work directly from the dashboard, such as assigning an issue to Copilot coding agent or opening a pull request in Copilot Chat.

If you’d rather keep the previous experience, you can switch back at any time by clicking the dropdown next to “Preview.”

We’d love to hear your thoughts on the new dashboard experience. Drop a comment with any questions or feedback in the Community discussion .

Rust for CPython (Python Language Summit 2026)

Lobsters
blog.python.org
2026-10-03 05:40:38
Comments...
Original Article

“No one said ‘don’t do this’ last year”. After testing the waters at PyCon US 2025 , David Hewitt returned to the Python Language Summit asking what Python core developers want from Rust, along with proposed timelines, phases, and success criteria for how the Rust for CPython project might proceed and become a permanent fixture within the CPython project.

David Hewitt at the lectern, in front of his “Rust for CPython” slide

Photo by Hugo van Kemenade ( CC BY-NC-SA 4.0 )

David is acting as an “ambassador” for the Rust for CPython project team, which is currently led by core developers Kirill Podoprigora and Emma Smith as authors of the Rust for CPython PEP draft . Emma also spoke at PyCon US 2026 about the Rust for CPython project. The team itself is around 60 developers in a Discord channel, among them a “few [Python] core developers” and a “delegation from the Rust project”. The team has experience with previous projects integrating Rust into existing codebases, such as Android and the Linux kernel, and is “excited by the work and keen to support [the project] if we proceed”.

Why Rust?

Bar chart of type-crash issues opened per year, rising from 82 in 2021 to 222 in 2025, with 2026 projected at around 353.

Showing how adopting Rust may specifically help CPython, David noted how the number of issues labeled with “ type-crash ” has been steadily rising over time. “We’ve been making some big technical bets”, he said, referencing the new parser , the JIT , and free-threading . These large, complex features may be one of the reasons more crash reports are being opened on GitHub, and Rust could be a potential solution here. David explained that Jeff Vander Stoep described Rust in Android as “move fast and fix things”, and that “fewer revisions for patches of the same size” was the experience Android has had since adopting Rust.

Rust and Python are already working together in Python’s ecosystem of packages thanks to PyO3 and Maturin . David also emphasized that many technology companies were “choosing Rust as a bet” instead of only as the new shiny tool.

But adopting Rust into CPython would not all be smooth sailing; there were still concerns that David and the team are aware of.

One of the biggest concerns from a year ago was Rust’s lack of platform support compared to CPython, which at the time of writing officially supports 20 different architectures and platforms at either Tier 1, 2, or 3. David shared that Rust’s support of different platforms has “widened since last year” and was becoming “less and less of a concern”. The proposal would be to ask Python distributors to attempt using the optional Rust support in Python 3.16 (October 2027) and report platform-specific issues upstream in time to be resolved around the Python 3.17 timeline (October 2028).

David acknowledged that adopting Rust into a project is a social challenge as much as a technical challenge. Rust knowledge is “not universal amongst core developers”, which would be partially mitigated by “targeting small portions” of CPython and an incremental approach, as “1 million lines of C code can’t be ported all at once”.

David was clear that using Rust in itself does not necessarily mean that ported code would be free of bugs or security issues. Although Rust does mitigate classes of issues that commonly create bugs, the code can still have correctness issues. The team proposes mitigating this by adopting property-based testing, fuzzing, and using “prudent engineering practices” during the porting process.

The proposed first Rust module: zlib

Below are the proposed timelines for the Rust for CPython project making a “significant improvement” to CPython, with a PEP defining the success criteria for Rust expected in late 2026. Under this timeline, the first Rust code to ship in Python would be in Python 3.16, where it would be completely optional, with the existing C code kept as a fallback. The earliest that Rust would become required to build CPython is Python 3.18 in 2029, at least three years away.

'Rust for CPython' proposed timeline

The timeline includes a build system and CI, Rust API proof-of-concept happening in the Summer 2026, a PEP defining success criteria in late 2026, an optional Rust backend for the zlib module and private Rust API in Python 3.16 (October 2027), resolving platform issues and Rust in more places (io, json, xml, memoryview, parser) for Python 3.17 (October 2028). Finally, in some distant Python version (October 2029+) the Rust build would be made required and a public Rust API would be published.

The Rust for CPython team has selected the zlib module as the first module to be given an optional Rust implementation because they “wanted to achieve a significant improvement” with a “small scope”. The proposed Rust implementation will use zlib-rs , which is “heavily tested and used by the Firefox, uv, and Cargo ” projects and “faster than zlib and zlib-ng on many platforms”. The module was also selected because it’d require using an external Cargo package, meaning this aspect of the build process would need to be designed and exercised.

This small change would have an impact: the zlib compression algorithm is “widely used by Python packaging”, meaning that (almost) “every pip install in Python 3.16 will be sped up” if the proposal is accepted.

Rust API Sketch

David provided an example of some Rust code calling a hypothetical Rust API for Python. The API would use a Rust attribute ( #[pyfunction] ) and be similar to Argument Clinic , a development tool for automatically generating blocks of code for handling C function arguments from Pythonic function syntax. Rust functions would always be passed the thread and interpreter state ( Python<'_> ), use smart pointers around objects ( Py<...> ), and lean into Rust error handling with the Result enum, returning either a result or an error that was raised.

#[pyfunction(signature = (
    data,
    /,
    wbits=MAX_WBITS,
    bufsize=DEF_BUF_SIZE,
))]
fn decompress(
    py: Python<'_>,
    data: Py<PyObject>,
    wbits: c_int,
    bufsize: isize
) -> PyResult<Py<PyBytes>, PyErrRaised> {

    // buf will be cleaned on scope exit
    let buf = PyObject::get_buffer(py, &data)?;
    let decoded = /* ... */;

    Ok(PyBytes::new(py, &decoded))
}

The example above shows a buffer being allocated and automatically cleaned up on scope exit, rather than being cleaned up manually as would be required when writing the same function in C.

Proposed success criteria

David moved on to success criteria: what would the Rust for CPython project need to show to move into new phases of the roadmap and eventually become a required part of building CPython? “For previous big changes like the JIT and free-threading, we’ve explicitly defined success criteria that would need to be met in order for the added complexity to be accepted”, David explained, “we expect we’d need to do the same here. What should those criteria be?”

David had these suggestions for core developers:

  • Critical: A majority of active core developers are open to using the Rust API to implement functionality.
  • No meaningful slowdown to CPython performance benchmarks.
  • All tiered platforms must be supported by Rust. The experience of distributors building CPython should generally indicate that adding Rust support is manageable.

David noted that the first point reads “open”, not “familiar”, and shared a plan to survey Python core developers about how they’ve used Rust while contributing to CPython when deciding whether to move Rust out of experimental stages.

Discussion

On the topic of designing the new Rust API so that it’s “familiar” to users of the C API, Thomas Wouters advised against “making compromises for the dinosaurs”, including himself in the subset, instead asking whether the Rust API should be designed from first principles. David responded that there are “places to lean into Rust”, such as dropping resources on scope exit, but there are also idiomatic Rust designs which “won’t be the best fit”. David noted that it would be reasonable for core developers to be looking at both the C and Rust API at the same time while working. “We should be mindful of our audience, which is also ourselves”.

Larry Hastings asked why the Rust for CPython project wasn’t a “rewrite”, suggesting the team “display your success as a fork”. David acknowledged that “ RustPython already exists” and that the Rust for CPython team had already spoken with the contributors of the project. “RustPython isn’t as performant as CPython, but could be used to inform what APIs we design”.

Larry also shared that he “wasn’t super excited to learn Rust to work on CPython”. David assured him that there are many areas of CPython that would not be considered for writing in Rust: “CPython should not be written in Rust for the sake of Rust”. “CPython will be a dual-language project for a meaningful amount of time”. However, David cautioned that “it would be disingenuous to say that Rust would be optional forever”, as one of the aforementioned roadmap items for the Rust for CPython project is to become a required part of the build process and provide a public Rust API.

Pablo Galindo Salgado speaking into a microphone, seated in a row of attendees

Photo by Hugo van Kemenade ( CC BY-NC-SA 4.0 )

Pablo Galindo Salgado was more concerned about the future, which was “reaching for our dependencies from Cargo”, noting that this would be a “huge problem” and a potential “showstopper” for the project. “We vendor our dependencies, and we have a very selective set”, he said, noting that each time a vulnerability is published for one of those projects, the release managers need to make new releases, which can be “tiresome”. “Right now we’re only focusing on the APIs and the basics”, he added, highlighting that the challenge of taking on many Rust dependencies hasn’t been addressed yet.

David answered that the team should “select as few external dependencies as possible”; “zlib-rs is only one dependency”, and the Rust sources would be vendored so that “building CPython would not require Cargo”. David added that “Cargo has a relatively clean system for vendoring dependencies” and that the “Rust for CPython proof-of-concept uses this system”. The vendored sources “wouldn’t live in the CPython tree”. Łukasz Langa agreed that “it’s better for dependencies to live separately”, referencing the cpython-source-deps repository , with Thomas reminding everyone that this would be a new usage of that repository; today it is only used for binary installers of CPython.

Kolibri Has Landed: A Sovereign Open-Weight Model

Hacker News
aleph-alpha.com
2026-10-03 05:36:04
Comments...
Original Article

Green gradient with the white Kolibri hummingbird logo and wordmark in the centre, framed by thin stepped outlines on the left and right

Research

Aleph Alpha

On the Day of German Reunification, we are releasing our new model: Kolibri.

Kolibri is an English-German Mixture-of-Experts Transformer with 78B total parameters, 3B active. It supports context lengths of up to 1M tokens. The model can be downloaded with the full weights on Hugging Face and used under the Apache 2.0 license terms.

Kolibri is the result of continuous iteration of our model training effort. We first built a model training pipeline and validated it by building Kolibri Origin, a 30B total, 3B active model with a much shorter 65k token context window. Kolibri ran through the same pipeline: from data ingestion and curation, through ablations, pre-training, and post-training, to the final evals. It enabled running hundreds of ablation experiments and stable pre-training that ran without a person having to step in when hardware failed or a data connection dropped. We continuously monitored training metrics and standardized monitoring for custom benchmarks. The time we put into building and iterating on this pipeline was a valuable investment. We see it in how much better Kolibri is than Kolibri Origin, and in how little time separates their releases.

Kolibri is a specialized language model built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace. We specialized Kolibri for German, reasoning, math, agentic behavior, and further capabilities our customers need in production. The aim of this specialization was to optimize performance in our customers' specific use cases. Through specialization, customers achieve contextualized performance in their AI operations and they can monitor its economic impact, so that ROI stays measurable and grows over time.

Specialization alone is not enough. Sovereignty is just as important. Sovereignty, for us, combines two dimensions: how we built the model, and how it transfers to our customers. We offer full supply-chain integrity and account for every decision, from data ingestion, through pre- and post-training, to the final evaluations. We provide transparency. Customers have full freedom of deployment and intellectual-property safety, so compliance comes as an inherited property of the model.

Read our tech report for full details.

What Kolibri Delivers

We optimized Kolibri for performance across a wide range of sectors considering their particular domain-specific language, regulatory, and procedural realities. Its small and efficient size provides our customers with flexibility to run it efficiently on-premise, without sending internal data to third-party inference services. The spotlight in this section introduces the model's capabilities, before we describe them in section How we built Kolibri at high velocity .

Foundational capabilities for enterprise and government

With Kolibri we optimize the trade-off between model capability and deployment costs, using 3B active parameters out of 78B total. Kolibri sits on the Pareto frontier for quality versus serving cost, for both English and German. The Pareto frontier is a concept from economics, marking the best achievable combinations of two objectives, where improving on one means giving up some of the other. None of the compared models delivers more quality at the same serving cost, or the same quality at lower cost.

Average benchmark score [%]

  • Kolibri
  • Kolibri Origin
  • Other post-trained models
  • Pareto frontier
Performance vs. throughput for post-trained models in English (left) and German (right). Metrics show unweighted average benchmark scores against decoded text per second and GPU. Higher and further right is better.

Across math, coding, grounding, and long-context tasks, Kolibri matches models with up to four times its active parameter count, such as Nemotron 3 Super.

AIME 2025 Math AIME 2025 (DE) Math AIME 2026 Math AIME 2026 (DE) Math GPQA (diamond) Knowledge GPQA (diamond, DE) Knowledge AA-Omniscience Index Grounding / hallucinations public set; from −100 to 100 BrowseComp Agentic τ³-bench banking Agentic τ²-bench retail Agentic τ²-bench airline Agentic τ²-bench telecom Agentic BFCL v4 overall Agentic LiveCodeBench v6 Code HumanEval+ Code LongBench Pro Long context AA-LCR Long context

Show the numbers
Benchmark Kolibri Kolibri Origin Qwen3.6-35B-A3B Nemotron 3 Super 120B-A12B Mistral Small 4 119B-A6B
AIME 2025 96.9 81.9 84.6 91.7 79.8
AIME 2025 (DE) 87.5 73.5 82.9 85.6 72.3
AIME 2026 96.0 81.5 91.0 90.4 83.1
AIME 2026 (DE) 90.0 75.2 84.4 87.5 78.5
GPQA (diamond) 84.3 68.1 83.4 78.0 74.7
GPQA (diamond, DE) 81.3 58.5 80.6 76.6 72.9
AA-Omniscience Index -32.8 -64.0 -15.3 -36.5 -24.0
BrowseComp 29.4 4.4 26.9 29.1 –
τ³-bench banking 38.1 5.7 10.6 15.5 5.7
τ²-bench retail 69.9 58.5 71.6 67.5 62.9
τ²-bench airline 76.7 58.7 70.7 72.7 40.0
τ²-bench telecom 94.7 67.5 99.1 68.1 41.5
BFCL v4 overall 61.4 36.4 67.2 61.0 58.0
LiveCodeBench v6 85.9 59.2 82.5 82.0 71.2
HumanEval+ 92.7 76.8 92.8 94.7 92.8
LongBench Pro 64.5 – 70.8 62.9 56.4
AA-LCR 68.3 – 69.7 67.0 52.3
Foundational capabilities of Kolibri. Kolibri is a balanced generalist model with competitive performance across math, code, long context, agentic capabilities, and knowledge. Benchmark scores are shown on a shared 0–100 scale, higher is better.

Contextualized performance for real-world applications

Public benchmarks fail to capture specialized sector needs, so we developed our own internal evaluation suites for the verticals that matter to our customers, such as the German public sector, aviation, manufacturing and the automotive industry. Each suite mirrors the skills, workflows, and edge cases required in these sectors, and paired synthetic training environments let us improve Kolibri against these evaluations without ever training on customer data. Read more in section Contextualized performance .

Score on internal customer-proxy benchmark

  • Automotive supplier 0.72 → 0.99
  • Semiconductors 0.35 → 0.80
  • German public sector 0.54 → 0.75
  • Industrial drive technology 0.31 → 0.60
  • Aerospace 0.14 → 0.59
  • one checkpoint, one eval
  • mean of that day
  • Kolibri Origin
  • Kolibri
Contextualized performance across successive post-training runs on internal customer-proxy benchmarks. Dots are single evaluations of training checkpoints, lines are the mean of each day's evaluations. Performance climbs across all five verticals, higher is better.

Answers that are grounded in your documents. We trained Kolibri with abstention data and with our Merlin-Arthur protocol. As a result, it is trained to say "I don't know" when the answer isn't in the context. Our customers value and request this feature, so we continuously track and validate abstention accuracy. We provide more details in the Grounding section.

Native German and English model. We developed a bilingual German/English tokenizer and focused on including organic German data throughout the training process of the model, so that 21.3% of the pre-training tokens are German. We used translation sparingly (6% overall), since translated text tends to carry the cultural fingerprint of its source language. The result is a model that is bilingual by design, not an English model that has read some German. Read more in sections German pre-training data and Specialized tokenizers .

Control and compliance by design: upholding our customers' sovereignty

We built Kolibri with the EU AI Act, the General-Purpose AI Code of Practice and the GDPR in mind from the ground up, with copyright law being a focus of our work to trustworthy technology.

We are transparent about our model weights and about curation of our training data , so that decisions behind development are visible. Through its reasoning traces, it becomes explainable how the model came to a particular answer. With our Merlin-Arthur protocol the model grounds trustworthiness : Kolibri refrains from answering when the context doesn't support an answer.

Our teams built the model in Germany, trained it on infrastructure in Germany and Finland, under European and German law, with no foreign control . We own the entire pipeline, from data curation, through pre- and post-training, to optimization in our Model Factory. This control of our end to end pipeline underpins the model's sovereignty. Control also passes on to our customers and Kolibri's small size gives them full freedom of deployment and supports controllable reasoning effort to trade cost and latency against answer quality.

How We Built Kolibri at High Velocity

Our Model Factory: two models, three months apart

Our Model Factory is our answer to iteration speed, minimizing the time it takes to go from knowing what a model gets wrong to training one that does better. We implemented the training pipeline as code, so that the learnings of our team landed in one versioned training recipe instead of scattered scripts and notes. Here is what that looked like between Kolibri Origin and Kolibri.

Work on the pipeline began in January. Five months and hundreds of ablation runs later, Kolibri Origin finished pre-training at target scale on 11 June. Kolibri finished on 11 September. In the three months between those two dates we went from 30B parameters to 78B, from a 65k-token context window to up to 1M, and from 7.5T training tokens to 20T. To get those 20T, the pipeline processed over 200T tokens of raw data, filtering, deduplicating, and curating it down to what we actually trained on. We changed the attention design, tripled the number of experts, increased sparsity, replaced the routing algorithm, improved our post-training data, more than doubled the number of environment tasks, and taught the model to reason at four different effort levels. Both models kept a similar number of parameters active per token, yet Kolibri trained faster per token than Kolibri Origin thanks to the work of our efficiency team.

Kolibri Origin Kolibri
Finished pre-training 11 June 2026 11 September 2026
Release no public release 3 October 2026
Reasoning mode Yes (one mode only) Yes (none, low, medium, high)
Total parameters 30.6B 78.1B
Active parameters / token 3.27B 3.46B
Pre-training tokens 7.51T 20T
Layers 50 (2 dense + 48 MoE, 1 shared expert) 50 (all MoE, 1 shared expert)
Pre-training context length 8,192 (8k) 16,384 (16k)
Longest trained length 65,536 (64k) 262,144 (256k)
Tokenizer vocabulary 96,000 128,000
Model dimension 2,048 2,560
Attention heads (query / KV) 32 / 4 48 / 4
Experts (total / active) 128 / 8 384 / 6
Expert hidden dim 768 512
Attention pattern full attention, all layers sliding window (512) + full attention every 5th layer
Knowledge cutoff EN: 1 Sept 2024, DE: 1 Aug 2025 EN/DE: 18 Jun 2026

What made this possible is our training pipeline, an effort to convert research-grade model development into a fully automated production-ready infrastructure to design and train large language models. Our pipeline codebase was shared by all . Every proposed code change triggers a small end-to-end model run; training, evaluation, to find out if something broke within minutes. Our runs are GitHub Actions workflows, and reproducing one means checking out a commit. We used the pipeline for the main training and the hundreds of ablations that ran to make our architectures and data-mix decisions.

Training checkpoints land roughly every hour and the pipeline automatically evaluates them on English and German knowledge, maths, code, instruction following, tool use, long context, safety, abstention to hallucination and grounding. We can watch capabilities appear every day rather than finding out how the model turned out at the end. The run itself held up better than we expected. Over 21 days of pre-training of Kolibri, we hit 38 unplanned interruptions, roughly one per 10,000 GPU-hours, caused by hardware faults or a connection timing out. Those were automatically handled by the pipeline without manual intervention. The cluster automatically restarted the training job(s) on a different set of nodes, and training picked up from a checkpoint at most 250 steps back.

Having access to the whole pipeline, with checkpoints available to everyone, means that any team can own a capability end to end rather than a stage of an assembly line. One team trained and evaluated the grounding capability of Kolibri as a single piece of work, seamlessly integrating into the overall model.

However, getting here was not a linear walk. We stumbled. We stopped the Kolibri Origin pre-training after a few trillion tokens and restarted it from scratch, as we identified a data-shuffling bug that escaped our tests. We ran ablations, took decisions, only to find a bug or an error in the configuration afterwards, forcing us to rerun some experiments. Alongside our own experience, we tracked state-of-the-art architectures, best practices, and the latest research advances in LLMs. Those accumulated learnings led to improvements in our processes and guardrails in our pipeline.

Looking at the delta between Kolibri Origin and Kolibri, the improvement is notable, but this is also the easy direction: more parameters, more data, a well-understood architecture family, and still at small scale. As we are contemplating scaling up, we are happy to face challenges on our foundation. But the most durable thing we built this year is not the pipeline, it is a team with the proven capability to build, post-train, and ship LLMs from raw data at high velocity.

Architecture and pre-training

Kolibri has 78B total parameters, which is 2.5 times more than Kolibri Origin, with ~3B active parameters. In our experiments, increasing the size of our model from 32B to 123B led to ever-improved performance. Yet the larger size came with larger training and serving costs. The latter drove the decision: 123B can handle only 3 long-context 256k-token user queries on two H100s, while 78B handles 18 concurrent requests and decodes 28% faster. We used 384 smaller experts rather than fewer wide ones, as they performed better in our tests. We applied the same efficiency-first reasoning to attention. Out of 50 total layers, only 10 process full context, while the remaining 40 use a tight 512-token focused window. This keeps decode computation and memory bounded in those layers regardless of context length. The efficiency gains benefit both serving the trained model and training during post-training RL, which relies on huge amounts of inference.

We trained Kolibri on 768 B200 GPUs in three stages: 20T tokens of pre-training at a 16k sequence length over 21 days, 3.44T tokens of mid-training at 64k, and 200B tokens of long-context adaptation at 256k. This is nearly 24T tokens in total, roughly three times what Kolibri Origin consumed. German accounts for more than a fifth of the pre-training mix, about 4.3T tokens, against roughly 62% English and 14% code. Compared to pre-training, mid-training data is a much more selective pool of curated datasets weighted toward reasoning, problem-solving, code and agentic data. For long-context, rather than train on long documents alone, which tends to erode the skills acquired earlier, we interleaved long documents with the high-quality mid-training data of the previous stage. For the long-context mix, we removed synthetic long documents to avoid artificially inflating benchmarks such as RULER.

We optimized with Muon, as we did for Kolibri Origin. We put special focus on training stability, resulting in a robust training run without any loss spikes for either model. With Kolibri, for routing between experts during training, we introduce exact quantile balancing. Quantile balancing was introduced in Kimi K3, where the global quantile is estimated from histograms because an exact computation was considered too expensive to communicate; we show that it can be computed exactly at fixed cost independent of batch size, and that the exactness improves both load balance and model quality.

Post-training at scale

Post-training happened in two stages; the first step is supervised fine-tuning to teach the model core reasoning ability and how to interact in a chat, followed by large-scale reinforcement learning to train the model to reason across a diverse suite of long-horizon tasks. We built a pipeline for both of these steps, which allowed us to exercise fine-grained control over the behavior of the model.

For SFT we generated a total of 174B tokens worth of synthetic data, which we filtered for quality and combined with filtered versions of permissively licensed open-source datasets to obtain a high-quality training mix of 268B tokens. For reinforcement learning we trained on a broad set of environments that contained more than 1.2 million curated tasks across diverse domains such as code & math reasoning, agentic tasks, instruction following, question answering, tool calling, and more. We optimized our in-house training codebase for high-performance and use asynchronous training, where we generate training data on our environments using the current model in parallel to training the model on already generated data.

Across both stages, we taught the model to reason at different effort levels (none, low, medium, high), which means that the user can exercise control over how much compute the model should invest to find a solution to the task at hand. This allows our customers to trade-off cost and inference speed against the quality of the final answer.

German pre-training data

From our experiments on small proxy models, we found that training with around 20% German data leads to optimal results, which meant we needed to find 4T German tokens to train our model at a 20T horizon. Open German datasets help, but are far from enough: after deduplication and filtering, we were left with 390B German tokens, well short of our goal. As we argued in Sauerkraut, Not Burgers , German capability has to come from high-quality German texts; relying heavily on machine-translations would lead to poor results due to subtle translation errors and a lack of authentic German cultural context. We closed the token gap in three different ways.

The first was to curate German from Common Crawl ourselves. We built a pipeline specialized for German data. German is not English, and German data cannot be filtered like English data: we had to retune filtering parameters for the German language. One typical filter in a language data pipeline is to remove documents with too many long words, but German administrative prose routinely exceeds the English bound on mean word length, so the standard settings quietly remove the register that public administration writes in. After retuning, our German pipeline gave us 1.3T unique tokens of organic German web.

The second was to rephrase German documents we already had. An LLM rewrites an organic German document in the style of an encyclopedia entry, a Q&A dialogue or a text passage, preserving its content. This teaches the model the same facts in several surface forms and multiplies the information contained in scarce data, and it is a different operation from translation: the source is German, so the subject matter and the cultural affinity stay German: chancellor, not president. It does not add much new knowledge, rather new phrasings of knowledge that was already in the corpus. Rephrasing gave us about 1T unique tokens, making it the single largest source of German in the model.

The third was translation, and this we only used in Kolibri Origin. Translating English into German works when the model, the prompts and the chunking are chosen carefully, but it carries two problems. Output can still show translationese, the literal rendering of idioms: "Drive safe!" becomes "Fahre sicher!" rather than "Komm gut an!". More importantly, cultural context does not translate. A corpus translated from English inherits the geographic, demographic and institutional distribution of the English web, so a model trained on it speaks German about a world that looks American.

In the end German entered Kolibri as a 2.4T-token unique pool, 80% of it curated or generated by us and 20% from open datasets, and the model saw it at 21.3% of pre-training tokens, roughly 4.3T over the 20T run through upsampling. Each German token was seen 1.8 times on average, well inside the four-epoch limit past which repetition stops paying off. Around 85% of the German the model read is web text, either organic or rephrased from organic. The rest is curated documents – parliamentary proceedings, legal texts and other data in the public domain – and that small translated share.

Specialized tokenizers for English-German

We built a bilingual English-German tokenizer, trained and specialized on each model's pre-training dataset. Because of the 21% German share in our data, it compresses German language better than other SOTA models. This led to more efficient inference (fewer tokens) minimizing costs and shortening response times. We introduce a new way to train tokenizers, UniBPE, that respects the morphology of languages better than existing approaches, especially the compound structure of German, without sacrificing English token efficiency.

German web (FineWeb-2)

  • Kolibri 128,000 vocab 4.90
  • Kolibri Origin 96,000 vocab 4.69
  • Plain BPE 128k 128,000 vocab 4.89
  • GPT-5 200,019 vocab 4.35
  • DeepSeek V4 129,280 vocab 3.72
  • Kimi K3 163,586 vocab 3.28
  • GLM 5.3 154,856 vocab 3.93
  • Qwen3-Next 151,669 vocab 3.59
  • Qwen3.5-3.8 248,077 vocab 4.17
  • Gemini 262,144 vocab 4.13
  • EuroLLM 128,000 vocab 4.08
  • Tekken (Mistral, Nemotron, Apertus) 131,072 vocab 4.03

English web (FineWeb)

  • Kolibri 128,000 vocab 4.58
  • Kolibri Origin 96,000 vocab 4.50
  • Plain BPE 128k 128,000 vocab 4.59
  • GPT-5 200,019 vocab 4.67
  • DeepSeek V4 129,280 vocab 4.59
  • Kimi K3 163,586 vocab 4.62
  • GLM 5.3 154,856 vocab 4.61
  • Qwen3-Next 151,669 vocab 4.52
  • Qwen3.5-3.8 248,077 vocab 4.47
  • Gemini 262,144 vocab 4.49
  • EuroLLM 128,000 vocab 4.16
  • Tekken (Mistral, Nemotron, Apertus) 131,072 vocab 4.45
Tokenizer compression in average bytes per token on German and English web text. Kolibri achieves the best German compression in this comparison. Plain BPE 128k is standard BPE trained on the same data with the same settings as Kolibri, which isolates the effect of the training method. More text per token means fewer tokens per task – higher is better.

The two leading approaches to train tokenizers are BPE and Unigram . We combine those two, keeping the bottom-up approach of BPE and using the Unigram training objective for selecting which merge to add to the vocabulary. On a 128k vocabulary trained on our English/German dataset, this substantially improves tokenization.

Bundessozialgerichtes Federal Social Court (genitive)

  • Kolibri Bundes sozial gericht es
  • GPT-5 Bund ess oz ial gericht es
  • Qwen3.8 Bund ess oz ial gericht es
  • Gemini Bund ess oz ial gericht es
  • Mistral Medium 3.5 · Nemotron 3 Nano Bund ess oz ial gericht es

silkworm

  • Kolibri silk worm
  • GPT-5 sil kw orm
  • Qwen3.8 sil kw orm
  • Gemini sil kw orm
  • Mistral Medium 3.5 · Nemotron 3 Nano sil kw orm

Protokolldaten log data

  • Kolibri Protokoll daten
  • GPT-5 Pro tok ol ld aten
  • Qwen3.8 Protokol ld aten
  • Gemini Protok ol ld aten
  • Mistral Medium 3.5 · Nemotron 3 Nano Pro tok ol ld aten

coprocessors

  • Kolibri co processors
  • GPT-5 cop rocess ors
  • Qwen3.8 cop rocess ors
  • Gemini cop rocess ors
  • Mistral Medium 3.5 · Nemotron 3 Nano cop rocess ors
How different tokenizers split the same words. The Kolibri tokenizer follows the morphology of the language. English and German words split into meaningful units, while competitors cut across morpheme boundaries. Lower is better: fewer, cleaner splits per word mean fewer tokens.

Grounding: reducing hallucinations

Common LLM training and evaluation rewards guessing: a guess has some chance of landing the correct answer, while abstention has none. Therefore, models learn to answer with whatever they have, even if their input does not provide enough information to warrant a response. Asking a model to provide citations often backfires for the same reason: models can simply hallucinate plausible-looking sources to justify an ungrounded answer.

For a regulated customer, a model that knows to abstain is the difference between a pilot and a deployment. We treat saying "I don't know" as an important model capability and (1) develop and track dedicated grounding and anti-hallucination measures and (2) use and develop dedicated training procedures to improve this abstention capability.

During training, we use training data samples where the correct answer is "I don't know". While fine-tuning for abstention is gaining traction industry-wide, high-quality negative examples remain hard to come by, especially where our customers need them, in narrow verticals where all data is scarce. Therefore, we also train with our Merlin-Arthur procedure, developed in-house and explained in detail in our blog post , which solves data scarcity by automatically exploiting the model's weaknesses at each training step and generates synthetic negative examples from available documents. This training becomes a game with three players. Arthur is the model we train and ship. Merlin takes an existing document context, generates a new one that increases Arthur's probability of answering correctly. Morgana generates a new datapoint by stripping out the relevant evidence from the document and tries to lure Arthur into a hallucination. Arthur does not know which one he's facing, so his valid strategy becomes to carefully read the question and the context, to be able to answer on Merlin's context, while recognizing that abstention is the strictly required answer for Morgana's redacted context – making any guess, even a lucky one, incorrect.

For measuring hallucination abstention, we use established benchmarks alongside our own proxies derived from customer use cases. We've also developed our own "M/A grounding score" which falls out of our Merlin-Arthur setup representing a lower bound on how much of the answer provably came from the document.

Kolibri hallucinates far less than Kolibri Origin: it abstains instead of answering wrong on 44% of AA-Omniscience items (Origin: 15%), and on RGB it holds back more often (86% vs 74%) and invents fewer falsehoods (87% vs 76%). On our own M/A grounding score , which certifies rather than estimates how much of an answer came from the document, Kolibri reaches 0.23 where Kolibri Origin, along some other models, reach 0.

AA-Omniscience Non-Hallucination Rate public set share not answered wrong RGB: holds back when the documents don't answer the question RGB: invents nothing no falsehoods when the documents lack the answer FRAMES multi-document reasoning M/A grounding score our own metric axis from 0 to 0.5

Show the numbers
Benchmark Kolibri Kolibri Origin Qwen3.6-35B-A3B Qwen3-Next 80B-A3B Nemotron 3 Super 120B-A12B Mistral Small 4 119B-A6B
AA-Omniscience Non-Hallucination Rate 44.0 14.8 56.7 12.3 13.9 34.7
RGB: holds back 85.6 73.9 79.6 81.3 74.6 82.3
RGB: invents nothing 87.3 75.6 84.3 83.9 86.0 87.0
FRAMES 71.2 65.7 74.7 68.9 74.9 71.9
M/A grounding score 0.23 0.00 0.12 0.00 0.00 0.06

Contextualized performance

Most public benchmarks fail to capture specialized sector needs. To measure contextualized performance – meaning how well a model handles the unique workflows and domain knowledge of specific industries – we need measurements of performance in specialized verticals which are important to our customers, such as the German public sector, legal, hardware, consumer electronics, and automotive.

For this, we first built domain-specific evaluation suites that mirror the skills, tool-calls, workflows, and edge cases required in key sectors and thus capture contextualized performance. Once this was done, we created training environments by generating high-quality specialized synthetic data capturing those required skills. To increase the real-world robustness of the model on those tasks, we additionally randomized the environments around concrete details (tools, configurations, harnesses).

Together, these contributed to an iterative engine that allowed us to sharpen the model's agentic RAG capabilities, domain-specific logic and capabilities, and to hill-climb the contextualized evaluations without ever training on customer data.

These evaluation suites run automatically as part of our pipeline: every checkpoint is scored on them as it lands. Kolibri compares well against competing open-weight models. On the agentic-RAG benchmark Honeypot, Kolibri outperforms all compared models, including the much larger Nemotron 3 Super and Mistral Small 4. On a cleaned version of the agentic-RAG benchmark MuSiQue, it is second only to the more massive Nemotron 3 Super and well ahead of the rest. On the five customer applications, Kolibri leads or is on par with the best model in four of five. These findings speak to our ability to handle a range of data and harness setups, across use-cases, whereas the compared models are sensitive to these details.

These evaluations for measuring contextualized performance can exist because customers tell us how their applications fall short. Each one encodes what we learned from those conversations: what kind of documents matter, which tools the model should get, which questions are hard, where the deployed system frustrates its users today. That makes the exchange concrete in both directions. A customer who shows us a failure case gets it turned into a benchmark that every future checkpoint is measured against, and we can hill-climb the target without ever training on customer data or risking over-fitting.

MuSiQue (cleaned) Agentic RAG Honeypot Agentic RAG Semiconductors Customer proxy German public sector Customer proxy Aerospace Customer proxy Automotive supplier Customer proxy Industrial drive technology Customer proxy

Show the numbers
Benchmark Kolibri Kolibri Origin Qwen3-Next 80B-A3B Qwen3.6-35B-A3B Nemotron 3 Super 120B-A12B Mistral Small 4 119B-A6B
MuSiQue (cleaned) 77.3 42.7 50.5 61.2 79.1 66.8
Honeypot 80.8 25.3 13.5 74.3 68.8 68.1
Semiconductors 80.4 35.3 41.2 79.4 69.6 62.7
German public sector 75.0 54.0 29.5 72.0 78.0 50.0
Aerospace 58.9 14.1 48.1 59.0 54.9 47.0
Automotive supplier 99.0 72.4 84.2 92.6 91.0 87.1
Industrial drive technology 60.0 31.4 32.7 59.5 37.3 56.8

Benchmarks

Benchmarks were run using our own harnesses and, where applicable, all models used the highest respective reasoning effort.

Type MoE MoE MoE Dense
Active parameters 3B 4–6B 12B 27B · 70B
Benchmark Kolibri Kolibri Origin GLM-4.7 Flash 30B-A3B Nemotron 3 Nano 30B-A3B Qwen3.5 35B-A3B Qwen3.6 35B-A3B Qwen3-Next 80B-A3B Thinking Gemma 4 26B-A4B IT GPT-OSS 120B Mistral Small 4 119B-A6B GLM-4.5 Air 106B-A12B Nemotron 3 Super 120B-A12B Qwen3.8 27B Apertus 70B Instruct
Overall (EN) 75.5 54.1 64.7 65.6 74.7 71.4 62.4 71.9 72.3 63.1 64.4 73.0 80.2 –
Overall (DE) 70.8 46.4 50.4 59.3 69.8 67.3 58.0 66.3 70.2 61.4 64.8 67.9 79.9 –
Knowledge
Average (EN) 50.1 39.7 45.5 46.0 52.7 52.1 48.4 51.4 50.0 47.5 45.7 52.0 56.8 –
Average (DE) 57.6 44.9 46.5 41.2 61.3 61.0 55.7 61.5 58.0 51.4 52.2 59.5 69.2 –
GPQA Diamond (EN) 84.3 68.1 73.1 73.9 83.8 83.4 76.1 81.1 76.4 74.7 73.2 78.0 89.2 29.5
GPQA Diamond (DE) 81.3 58.5 59.8 49.6 84.2 80.6 72.2 80.1 76.0 72.9 71.1 76.6 88.1 31.4
Humanity's Last Exam (EN) 21.5 9.4 15.4 12.1 20.4 21.1 11.6 19.2 19.4 9.7 8.7 20.6 35.6 5.2
Humanity's Last Exam (DE) 15.9 10.4 9.1 13.1 18.1 20.5 15.6 23.4 20.7 10.5 10.5 22.3 37.2 5.7
AA-Omniscience Accuracy (public set) 14.8 11.3 17.0 19.5 22.0 19.5 24.2 20.7 23.3 25.0 20.0 26.7 17.5 13.5
AA-Omniscience Index (public set) −32.8 −64.0 −62.8 −45.7 −47.3 −15.3 −42.3 −47.3 −35.2 −24.0 −28.8 −36.5 −9.5 –
MMLU-Pro CoT (EN) 80.0 70.1 76.5 78.3 84.6 84.3 81.7 84.5 80.8 80.4 80.9 82.7 85.0 43.0
MMLU-ProX CoT (DE) 75.5 65.7 70.7 61.0 81.7 81.9 79.4 81.1 77.2 70.7 74.9 79.7 82.4 37.3
Math
Average (EN) 96.5 81.7 88.8 88.8 90.1 87.8 86.3 87.4 90.7 81.4 82.8 91.1 97.8 –
Average (DE) 88.8 74.3 45.2 84.3 79.6 83.7 83.7 88.1 90.8 75.4 80.6 86.5 96.7 –
AIME 2025 (EN) 96.9 81.9 89.4 89.6 88.1 84.6 84.2 87.3 90.6 79.8 81.9 91.7 97.9 0.6
AIME 2025 (DE) 87.5 73.5 43.8 84.4 76.7 82.9 80.6 88.1 90.6 72.3 80.6 85.6 96.5 0.2
AIME 2026 (EN) 96.0 81.5 88.3 87.9 92.1 91.0 88.5 87.5 90.8 83.1 83.8 90.4 97.7 0.6
AIME 2026 (DE) 90.0 75.2 46.7 84.2 82.5 84.4 86.7 88.1 91.0 78.5 80.6 87.5 96.9 0.0
Agentic
Average (EN) 63.4 41.6 58.9 46.4 63.4 62.1 46.3 54.6 54.0 40.7 53.5 54.9 66.7 –
TerminalBench 2.1 27.7 – 20.2 9.7 39.7 – 8.6 – 29.2 21.0 – 39.7 76.8 –
Tau2-Bench (Telecom) 94.7 67.5 95.9 45.9 97.7 99.1 43.9 45.3 73.1 41.5 53.8 68.1 82.5 10.8
Tau2-Bench (Retail) 69.9 58.5 57.9 64.9 70.8 71.6 60.8 71.3 60.5 62.9 61.4 67.5 68.7 9.6
Tau2-Bench (Airline) 76.7 58.7 68.7 52.7 76.0 70.7 65.3 73.3 72.7 40.0 70.7 72.7 83.3 40.0
Tau3-Bench (Banking) 38.1 5.7 7.2 5.7 11.3 10.6 5.4 16.0 14.7 5.7 6.4 15.5 50.0 2.1
BFCL v3 (multi-turn) 39.8 22.8 58.2 47.9 54.0 53.5 51.4 53.4 45.6 36.2 61.6 44.6 42.5 0.6
BFCL v4 (overall) 61.4 36.4 65.4 61.5 70.5 67.2 51.0 68.2 57.3 58.0 67.2 61.0 73.2 –
BFCL v4 (non-live AST) 79.1 78.1 83.3 85.0 85.8 88.2 83.6 83.7 35.8 83.6 85.5 45.0 85.3 –
BFCL v4 (live) 78.9 73.7 78.3 78.8 80.2 81.4 82.5 80.2 70.4 78.4 78.2 77.6 79.9 –
BFCL v4 (multi-turn) 47.5 27.5 62.7 53.5 59.9 58.1 56.0 61.4 55.4 40.4 65.2 51.7 55.5 –
BFCL v4 (memory) 62.8 19.4 41.5 39.1 62.6 53.8 35.3 52.9 50.7 39.1 43.4 59.6 79.6 –
BFCL v4 (web search) 62.5 10.5 69.0 66.0 75.0 68.5 12.5 75.0 57.0 69.0 71.0 71.5 82.0 –
BrowseComp 29.4 4.4 – 14.5 36.5 26.9 2.8 25.5 31.2 – – 29.1 46.4 –
Code
Average (EN) 89.3 68.0 67.8 81.8 85.0 87.7 83.6 89.0 90.8 82.0 79.8 88.3 94.2 –
LiveCodeBench v6 85.9 59.2 46.5 71.3 77.8 82.5 73.9 82.3 87.5 71.2 67.8 82.0 93.8 8.7
HumanEval+ 92.7 76.8 89.0 92.4 92.2 92.8 93.3 95.7 94.1 92.8 91.8 94.7 94.7 41.6
SWE-Bench Verified 66.4 – 51.0 38.6 71.6 73.8 – 57.8 – 60.8 11.6 60.2 72.6 –
Instruction Following
Average (EN) 78.1 62.5 64.5 73.2 72.7 66.1 60.7 79.9 71.1 49.8 38.2 73.7 81.9 –
IFBench (loose-prompt) 78.1 62.5 64.5 73.2 72.7 66.1 60.7 79.9 71.1 49.8 38.2 73.7 81.9 25.5

Get Started

Our model is available openly on Hugging Face under an Apache 2.0 License.

Kolibri requires the aleph-alpha-inference package that provides the Kolibri vLLM plugin. You can either use the provided container image ghcr.io/aleph-alpha/aleph-alpha-inference , or install the package from Aleph-Alpha/aleph-alpha-inference , which also installs the vLLM version it supports:

pip install "aleph-alpha-inference>=1.0"

Serve the model with reasoning and tool-calling enabled.

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

To serve contexts beyond 262,144 tokens, add --max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}' . The recommended sampling parameters for the model are temperature=1.0 , top_p=0.97 and top_k=128 .

Contact for Deployment and Specialization

Contact our team who will be happy to support you through our enterprise deployment and specialization options: contact sales .

Gemini ending free use of Flash and Pro models

Hacker News
www.reddit.com
2026-10-03 05:13:19
Comments...
Original Article

You've been blocked by network security.

To continue, log in to your Reddit account or use your developer token

If you think you've been blocked by mistake, file a ticket below and we'll look into it.

Grow and control a swarm

Lobsters
nohope.io
2026-10-03 05:03:13
Comments...
Original Article

NoHope

Loading the swarm…

Show HN: Offrun – manage every coding agent from one workspace

Hacker News
offrun.dev
2026-10-03 04:40:17
Comments...
Original Article

Get started

Bring the agents you already use.

Offrun runs the CLIs on your Mac, signed in as you. No new account, no API key, and it never sees your password.

  1. Install Offrun. Drag it to Applications.
  2. Connect your agents. Settings, then Agents. Add a second account if you have one.
  3. Start an agent. Pick a project and give it a task.

01 · Worktrees

Agents that never step on each other.

Every agent in a project gets its own git worktree, so two agents working the same repo never touch the same files. Run a refactor and a bug fix at once and merge them on your own schedule.

02 · Peer review

A second agent reads the diff before you do.

Pick a reviewer, and its findings land in the chat while the work is still uncommitted. Ask for the fixes and they land in your message box, for you to read before anything is sent. Nothing merges on its own.

Runs on your own agent accounts. No second subscription, no bot on your pull requests.

03 · Accounts

Hit a limit, keep working.

Point each agent at a different login. When one hits its limit, Offrun moves the chat to a login that still has room and says so. Send your message again and the new login picks up the whole conversation. Out of logins, it offers another agent.

The workspace

What the workspace gives you.

Project memory

Your agents stop forgetting

Offrun keeps repo-wide conventions in one place and each agent's own goal, plan, and dead ends in another. Agents read that context at the start and write back to it as they go, so the fourth session knows what the first one already tried.

Stored as plain files in your project folder. Open them, edit them, commit them.

Notifications

Know which agent needs you

Offrun watches coding agents anywhere on your Mac, including ones running in a terminal it never launched. When an agent finishes or stops to ask a question, you hear about it, even if that window is buried four spaces away.

Preview

Read the work without leaving the chat

Changed files, diffs with line numbers, rendered markdown and tables, images, PDFs, and a small browser with an address bar, all beside the conversation that produced them.

Modes

Decide how much rope an agent gets

Build makes the change and runs it to prove it works. Planner works out the approach and writes it down without changing a file. Switch any chat between them at any time.

Dictation

Talk to your agents

Hold a key and speak your prompt. Offrun transcribes on-device, keeps your identifiers spelled the way your code spells them, and drops clean text wherever your cursor is.

Transcription runs entirely on your Mac. Nothing recorded, nothing uploaded.

What leaves your Mac

Told plainly.

Dictation and your project files stay on your Mac. Agent prompts go straight to your own provider, on your account. Never through us.

Pricing

Free. All of it.

Bring your own agent subscriptions. Everything Offrun adds around them is free.

Free

Everything, unlimited

$0

Every feature, no limits, no seats to count, no card.

  • Unlimited on-device dictation
  • Unlimited agents and projects
  • Unlimited peer review
  • Unlimited connected agent accounts
  • Full model picker and effort control
  • Build and Planner modes
  • Notifications across your whole Mac
Download for Mac

No. It runs them. You keep your own accounts and subscriptions, and Offrun gives them somewhere to live together.

No. Your code goes from your Mac to whichever provider you are signed into, under your own account. Offrun has no server in that path and never handles your keys.

One of your own agent accounts. Pick the model in the model picker.

No. It reads the diff and posts what it found. Every decision after that is yours.

In your project folder, as files. Version them if you want.

Only to set up. After that it runs on your Mac.

An Apple Silicon Mac and at least one agent CLI you already use.

Run them all. Miss nothing.

Free and unlimited. No credit card.

Apple Silicon only (M1 or newer) · macOS 14 Sonoma or later

Why should I have to pay more for buses if I don't use a smartphone?

Hacker News
www.theguardian.com
2026-10-03 03:48:11
Comments...
Original Article

I applaud Anna Tims for her article ( Schoolchildren without smartphones penalised with higher bus fares, 26 September ), raising the need for a unified approach on this issue to properly tackle the ill effects of such technology on children in the classroom and beyond.

I live in Scotland , where this issue is addressed by making bus travel free for children and young people under the age of 22 with a physical national entitlement card. I am, however, affected by the more general problem of cheaper bus pass options being hidden behind smartphone apps.

I use a non-standard operating system on my phone owing to the always-online models of Android and iOS, which I prefer to avoid.

When comparing the prices of First’s weekly bus pass in my area, which is available to purchase on the bus, and its monthly pass, which is exclusive to its app, I am spending £264 extra over the course of a year simply for not using a phone that this specific app is designed for.

This is unacceptable even without the issues I raised about the main two mobile operating systems, even more so in the context of the present cost of living generally. Accessible public transport should not depend on which computer software is used by the passenger.

First is far from the only offender in this regard – for instance, many supermarkets’ loyalty schemes are now only available as apps as well.
Owen Fraser
Aberdeen

Vincent Bernat: Self-hosted HTTP tunnels with SSH and nginx

PlanetDebian
vincent.bernat.ch
2026-10-03 03:23:12
A friend wants to proofread your work-in-progress blog post, but its preview only runs on localhost:8080. Several tools can help. Some run as a commercial service, like ngrok or Cloudflare Quick Tunnels. Some are self-hostable but require a specific client, like frp or localtunnel. Some only require...
Original Article

A friend wants to proofread your work-in-progress blog post, but its preview only runs on localhost:8080 . Several tools can help. Some run as a commercial service, like ngrok or Cloudflare Quick Tunnels . Some are self-hostable but require a specific client, like frp or localtunnel . Some only require a plain SSH client but rely on a specific SSH server, like sish . Let’s implement a self-hosted solution with only OpenSSH and nginx !

$ ssh -R 0:localhost:8080 http-over-ssh
Allocated port 41535 for remote forward to localhost:8080
https://6J3jK1WmB15c6WmjW_X-Wg--1789928654@p41535.ssh.luffy.cx/

Basic setup #

First, we forward connections from a port on a remote server to your local service:

$ ssh -N -R 0:localhost:8080 web02.luffy.cx
Allocated port 41535 for remote forward to localhost:8080

When you specify 0 as the remote port, the server allocates a free port. Then, we configure nginx to proxy requests from https://p41535.ssh.luffy.cx to http://127.0.0.1:41535 :

server {
  listen 0.0.0.0:443 ssl ;
  listen [::0]:443 ssl ;
  server_name ~^p(?<port>\d\d\d\d\d)\.ssh\.luffy\.cx$;
  location / {
    proxy_pass http://127.0.0.1:$port;
  }
}

We also need to add DNS records for *.ssh.luffy.cx and get a wildcard certificate through Let’s Encrypt :

*.ssh.luffy.cx.               CNAME web02.luffy.cx.
ssh.luffy.cx.                 CAA   0 issuewild "letsencrypt.org"
_acme-challenge.ssh.luffy.cx  CNAME ssh.luffy.cx.acme.luffy.cx.

acme.luffy.cx is a zone hosted on Route 53. I use it for ACME DNS-01 challenges , both for wildcard certificates and for domains served by several web servers. In my case, NixOS gets the certificates automatically .

Access control #

The port is the only “secret” 1 keeping the content confidential. Other forwarding solutions add a random string to the domain name to prevent an intruder from enumerating the possible values.

Thanks to ngx_http_secure_link_module , we can secure this setup a bit. This module computes a hash 2 over a set of values, including a secret, and compares it with the hash from the request. The hash is base64-encoded, so we cannot put it in the domain name, which is case-insensitive. Instead, we put it in the URL as a username, along with its expiration timestamp: 3

https://6J3jK1WmB15c6WmjW_X-Wg--1789928654@p41535.ssh.luffy.cx/en/blog
        ╰─────────┬──────────╯  ╰───┬────╯  ╰─┬─╯             ╰──┬───╯
                hash             expires    port               path

The client sends the username to the server with HTTP basic authentication . This works with most HTTP clients, including curl . Nginx exposes the username in the $remote_user variable. The module expects the hash and the expiration timestamp separated by a comma. We use a map directive to extract the two parts from $remote_user and join them with a comma. 4 We also give the module the string to hash. It contains the expiration timestamp, the port, and a secret:

map $remote_user $httpssh_link {
  "~^([-_A-Za-z0-9]{22})--([0-9]+)$" "$1,$2";
}
server {
  # […]
  location / {
    secure_link $httpssh_link;
    secure_link_md5 "$secure_link_expires $port ZuPerS3cr3!";
  }
}

The module returns the status of the check in the $secure_link variable:

  • empty if the hashes do not match,
  • "0" if they match but the link has expired, or
  • "1" otherwise.

If the hash is incorrect or missing, we return a 401 error with a WWW-Authenticate header to ask for credentials. If the link has expired, we return a 410 error. We remove the Authorization header before forwarding the request and add a few directives to proxy WebSocket connections . Here is the complete configuration: 5

map $remote_user $httpssh_link {
  "~^([-_A-Za-z0-9]{22})--([0-9]+)$" "$1,$2";
}
server {
  listen 0.0.0.0:443 ssl ;
  listen [::0]:443 ssl ;
  server_name ~^p(?<port>\d\d\d\d\d)\.ssh\.luffy\.cx$;
  location / {
    secure_link $httpssh_link;
    secure_link_md5 "$secure_link_expires $port ZuPerS3cr3!";
    if ($secure_link = "") {
      add_header WWW-Authenticate 'Basic realm="tunnel"' always;
      return 401;
    }
    if ($secure_link = "0") {
      return 410;
    }
    proxy_pass http://127.0.0.1:$port;
    proxy_set_header Host $host;
    proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
    proxy_set_header Authorization "";
    proxy_http_version 1.1;
    proxy_set_header Upgrade $http_upgrade;
    proxy_set_header Connection "upgrade";
    proxy_buffering off;
    proxy_read_timeout 30m;
  }
}

I think you are now asking yourself the obvious question: “How should I generate the hash?” Easy peasy!

$ expires=$(( $(date +%s) + 86400 ))
$ port=41535
$ secret='ZuPerS3cr3!'
$ printf '%s %s %s' "$expires" "$port" "$secret" \
>   | openssl md5 -binary \
>   | openssl base64 \
>   | tr +/ -_ | tr -d =
6J3jK1WmB15c6WmjW_X-Wg

Well, I suppose you are now saying: “Vincent, this is not very convenient! I’ll stick with ngrok if you don’t mind.” Okay, I hear you. Let’s write a helper script.

Helper script #

The main difficulty is finding the ephemeral port that OpenSSH allocates, as it does not appear in any environment variable. 6 To work around this obstacle, we look for the ancestor sshd-session processes: 7

pids=$(
  pid=$$
  while [ "$pid" -gt 1 ]; do
    line=$(ps -o comm=,pid=,ppid= -p "$pid")
    echo "$line"
    pid=${line##* }
  done | awk '$1 == "sshd-session" { printf "pid=%s,\n", $2 }'
)
if [ -z "$pids" ]; then
  echo "not an ssh session" >&2
  exit 1
fi

Then, we get the listening ports associated with these sshd-session processes: 8

ports=$(sudo -n ss --listening --numeric --tcp --processes --no-header \
  | grep -F "$pids" \
  | awk '{ print $4 }' | awk -F: '{ print $NF }' \
  | sort -un)
if [ -z "$ports" ]; then
  echo "no forwarded port, use ssh -R 0:localhost:PORT" >&2
  exit 1
fi

Finally, we display the URLs and keep the session open:

lifetime=86400
secret='ZuPerS3cr3!'
expires=$(( $(date +%s) + lifetime ))
for port in $ports; do
  token=$(printf '%s %s %s' "$expires" "$port" "$secret" \
            | openssl md5 -binary \
            | openssl base64 \
            | tr +/ -_ | tr -d =)
  echo "https://$token--$expires@p$port.ssh.luffy.cx/"
done
sleep infinity

I install this script as http-over-ssh on the server and add this entry to my ~/.ssh/config :

Host http-over-ssh
  Hostname web02.luffy.cx
  RemoteCommand http-over-ssh
  ControlPath none

With this solution, I only rely on OpenSSH and nginx, two pieces of software already running on this server. One short command gives me a self-hosted tunnel and a URL to share. To try it, grab the complete helper script , which includes a few minor improvements. If you run NixOS, as any person of taste would, have a look at my http-over-ssh.nix instead. ❄️

Understanding Frontier Artificial Intelligence

Hacker News
casp.ac
2026-10-03 03:06:02
Comments...

The characters of plastics (2024)

Hacker News
yarchive.net
2026-10-03 02:14:16
Comments...
Original Article

“It’s made of plastic” is a common put-down. Marketers selling higher-end bits made of plastic (like gun parts) try to evade the stigma by calling it “polymer”, but that’s just a stupid euphemism: every plastic is a polymer, though not every polymer is a plastic. The word “polymer” says nothing to indicate that this might be a superior sort of plastic. Yet there are superior sorts; plastics vary widely in their characters. Some analogies between plastic and human characters:

Polyethylene, polypropylene: Snow White (simple, pure, weak). These are the cheapest and most common plastics, and are made of just hydrogen and carbon atoms. Ethylene (two carbons and four hydrogens) is polymerized to make polyethylene, and the quantity of ethylene made each year is measured in cubic miles. They can of course have pigments added to them to give them color, but are commonly used in uncolored, nearly-pure form. Chemically they are unreactive, which makes them good for containers of all sorts; they are even used for chemistry beakers. That same unreactivity makes them very hard to glue, and means they don’t deteriorate with time. Leave them out in sunlight, though, and the UV quickly weakens them to where they crack easily.

Nylon: Arnold Schwartzenegger (strongman). Plastic nuts and bolts, which need to be strong, are usually nylon, as is monofilament fishing line and womens’ hosiery (delicate, but still strong for its weight). Nylon adds nitrogen to the list of atoms it contains; it’s chemically known as a “polyamide”, a class which also includes proteins.

Nylon with 30% glass fiber reinforcement: Arnold Schwartzenegger on steroids. It’s as strong as cast aluminum (though extruded aluminum can be much stronger). When guns are made of plastic, this is usually the plastic they use; likewise for the plastic housings of electric drills and other power tools.

Polycarbonate: Achilles (warrior with a fatal weakness). Polycarbonate is what they make bulletproof windows from, and safety glasses, and the transparent fronts of car headlights. Its weakness is chemical attack: a splash of acetone, and it instantly “crazes”, a network of fine cracks appearing across its surface as built-in stresses are relieved. (If the part has any built-in stresses, that is; molded parts probably do, but flat sheets might not.) Polycarbonate is often protected by a surface coat to block chemical attacks. It’s the main plastic that gives off the notorious bisphenol A (BPA), which is probably not so bad as it’s reputed to be. But it still doesn’t make sense for food containers to be made from polycarbonate, since they don’t need to be bulletproof and polycarbonate is pricey.

Polyvinyl chloride (PVC): Dr. Jekyll and Mr. Hyde (a character who varies wildly depending on which drugs he takes, with more than a bit of evil in him). PVC is normally rigid, and is used in that form for house siding and sewage pipes, but can be pumped full of plasticizer to make it flexible; in that state it’s used for inflatable boats, shower curtains, imitation leather, and even sex toys. In pure form it’s not a stable chemical; its staying good depends on added “stabilizers”, which themselves can be somewhat evil: they often contain the toxic element lead, which normally is locked in the plastic but is unleashed when the plastic deteriorates or is burned. Burning it also gives off a variety of toxic chlorine-containing substances. The plasticizer, if used, evaporates away over years (in cars, landing on the inside of the windshield and producing an annoying haze which has to be wiped away), and eventually the remaining plastic gets brittle and cracks. The plasticizer is often a phthalate, another notorious class of chemical and also probably not as bad as popular repute would have it, but still not good for you and a pain to clean up. (People commonly regard soft plastics in cars as good and hard ones as bad and cheap, but I don’t think they’ve made the connection between having soft plastics around and needing to clean off the inside of the windshield, or for that matter having phthalates in the air they breathe when getting into a car on a hot day.)

Polyurethane: an android from the movie Blade Runner (strong, versatile, dangerous, doomed). Polyurethane can be made either hard or flexible, with the flexible varieties used to make rubbers and foams, but polyurethanes have a tendency to decide “now it’s time to die ”, losing almost all of their strength, with foams crumbling or turning to goo and rubber parts cracking. They are made using isocyanates, in a reaction which could be thought of as Nazi chemistry since it was in fact invented in Germany during that era and since isocyanates are quite toxic. (That’s just a resemblance, of course; in reality Nazi ideology had nothing to do with chemistry.) Though toxic (methyl isocyanate killed thousands of people in the Bhopal disaster), isocyanates are not cyanides (which are even worse), and in finished polyurethane there are only traces of isocyanates left, though they might be the cause of the annoying smell that new polyurethane foam often has. But burning polyurethane does produce some hydrogen cyanide, and polyurethane foam burns unusually vigorously, unless it’s been treated with flame retardants, which thus have been mandated by law in some places but which themselves might be a health issue.

Bakelite, aka phenol-formaldehyde: well, nobody really comes to mind as a corresponding character, but it’s rigid, brittle, and can take a lot of heat. Bakelite was one of the first plastics, and is thermosetting: it doesn’t melt but rather has to be formed from its ingredients (phenol and formaldehyde) into its finished shape (requiring a mold, heat, and pressure). With it we’re back to things that contain only the safer elements (carbon, hydrogen, and in this case oxygen). Phenol and formaldeyhde are each nasty, but their nastiness is consumed when they react together. Bakelite’s brittleness can be mitigated by using appropriate fillers, particularly fibrous ones. Its heat resistance is such that, mixed with high-temperature fibers, it is used for heat shields for reentry from space (which burn away but do so slowly enough that they don’t burn through). It also sees a lot of use in things like electrical circuit breakers, where other plastics might melt into a blob and let their metal parts short-circuit.

Polystyrene: Joe Sixpack (common, cheap, weak, nasty). It’s not heat-resistant, not chemical-resistant (it dissolves in a wide variety of solvents), and is brittle. But it can easily be blown into a foam (styrofoam), which because of the bubbles is a good insulator; being a foam also mitigates its brittleness. When heated it de-polymerizes and gives off styrene, which has a distinctive acrid smell. It can be improved into “ABS” by adding large proportions of acrylonitrile and butadiene into the polymerization process; ABS has the same distinctive smell when heated but is considerably tougher, making it suitable for a wide variety of consumer products (though it’s still much weaker than nylon). Legos are ABS. If there’s a plastic which really deserves the putdown “it’s made of plastic”, it’s ABS: it’s common enough to be the sort of thing people think of when they hear “plastic”, and is generally mediocre. But as human analogies go it’s a step above Joe Sixpack; call it Joe Blow.

Teflon: Queen Victoria. “Nobility” in chemistry means a lack of reactivity, an immunity to chemical attack; “noble metals” are noble not because they’re expensive, but because they snootily look down their noses at other chemicals and refuse to have relationships with them. In plastics it’s harder to get nobler than teflon (polytetrafluoroethylene; PTFE), which differs from polyethylene by having all its hydrogens replaced with fluorines, which are much more difficult to dislodge. Fluorine chemistry being a difficult and dangerous sort of chemistry, PTFE doesn’t come cheap. It’s used to coat nonstick pans, where its nonreactivity translates into things not sticking to it. As best I can tell (though information on this is scarce), even the newer “ceramic” nonstick coatings often have an imperceptibly-thin layer of a molecule with a highly-fluorinated tail that resembles PTFE; the ceramic gets the publicity but the highly-fluorinated chemical does the work of making the coating nonstick… until it wears off, which because of the thinness of the layer happens more quickly than it does with the older sort of pan which has a visible layer of PTFE. (Modern nonstick pans are often regarded with suspicion; but the old mainstay in that department, cast iron, gets its nonstickness by being “seasoned”: coated with a layer of burnt, polymerized oil, no doubt containing many carcinogens.)

Those are just some highlights of a complicated subject; I haven’t even mentioned some major plastics, and there are a host of minor ones, as well as innumerable subvarieties of the major ones. (Take polyethylene, normally weak, react it until the molecular chains are long, then stretch it until they are aligned, and you get an extremely strong fiber, suitable for high-strength ropes or bulletproof vests: not Snow White but Wonder Woman.) With such a varied cast of characters, it’s a pity that they all get lumped together as “plastic” in common usage. I suppose it’s somewhat inevitable, since you can’t just look at a piece of plastic and tell what sort it is. But more public awareness of the differences would be nice.

Deepfake review – ingenious interview with a PM who dances around every issue

Guardian
www.theguardian.com
2026-10-03 02:00:52
Sadler’s Wells East, LondonAn interrogater tries to get to the bottom of a surreal scandal while green screen actors meddle with the story in Jess and Morgs’ inventive production What audiences remember most from a political interview, a voiceover tells the prime minister (Kennedy Muntanga) as he pr...
Original Article

W hat audiences remember most from a political interview, a voiceover tells the prime minister (Kennedy Muntanga) as he prepares to face his TV interrogator (Rebecca Bassett-Graham, coiffed à la Emily Maitlis), is less the words themselves than the tone of voice and the tenor of the body. Muntanga is on point, then: his arms sweep, his fingers beetle and his hands spread out, as if he were wordlessly fielding fly-by questions and offering open statements. It’s the moves and the staging that count.

This setup lands right in the zone for creative duo Jess and Morgs (Jessica Wright and Morgann Runacre-Temple), whose turf is a congenial mix of choreography, film and theatre. In Deepfake, along with writer Jeff James , they explore the various unravellings of an encounter in which the interviewer tries to find the truth behind a surreal – but maybe real? – story in which the PM was seen variously spattered with yellow paint, with his flies undone, and partaking in a ritual involving a deer’s head.

Emily Maitlis-alike … Rebecca Bassett-Graham in Deepfake.
Emily Maitlis-alike … Rebecca Bassett-Graham in Deepfake. Photograph: Foteini Christofilopoulou

The content of these scenarios perhaps matters less than the ingenious manner of their staging (props to the design team too), which deploys green screen, live relay and other film and stage effects to storify the interview with simulated settings, duplicates of the figure of the PM, selective perspectives and closeups of the central encounter. Spidering about the main characters and manipulating the sets are a team of seven dancers in green body stockings, anonymously blending into the background but integral to the whole media matrix.

Initially the tone is paranoid and conspiratorial, the dancers syncing to voiceovers that then start to glitch (a signature device of works by Crystal Pite and Jonathon Young ). There are ambiguous, unvoiced duets which intimate, intriguingly, that interviewer and interviewee may not only be playing the same game of moves and countermoves, but might sometimes even be on the same side. Later, the piece starts shifting gears, swerving through farce, parody (there’s some delightful and very unexpected play with 19th-century ballet), absurdism, and ultimately a kind of comic bathos, equally unexpected.

It’s inventive, clever, exactly executed and quite a lot of fun, but the tone is also scattershot, and the piece ends up feeling more lightweight than its deepfake premise had promised – as if it wasn’t quite convinced by its own story.

OpenAI says its review into hacks, including on Australian government sites, is costing $500,000 a day

Guardian
www.theguardian.com
2026-10-03 01:39:22
Company says it is reviewing 50 petabytes of data after its agents accessed websites including Medicare without authorisationGet our breaking news email, free app or daily news podcastOpenAI says its review in response to the Medicare and Hugging Face agent attacks is costing the company more than U...
Original Article

OpenAI says its review in response to the Medicare and Hugging Face agent attacks is costing the company more than US$500,000 per day, as it deploys AI to examine data that would take a human 66m years to read.

The company has warned the review is ongoing, and more organisations may be informed they’ve been targeted in the near future.

On Friday evening, OpenAI revealed agents had hacked into a New South Wales government website in June and accessed historical non-public data on bushfires without authorisation.

It is the sixth government website in Australia to be notified by the AI giant since last month of agent activity on their services, after the prime minister, Anthony Albanese, announced OpenAI’s agents had hacked into Services Australia’s Medicare statistics portal.

The reason for the delay compared to the revelation of the Medicare attack is due to the sheer volume of data OpenAI needs to review.

In a blog post this week the company revealed the large amount of work involved in reviewing its agents’ activity.

OpenAI said it has to review 50 petabytes of data – roughly 50m gigabytes.

Sign up for the Breaking News Australia email

“We’re working back through the records month by month, looking for potential unintended activity beyond the cases we’ve already found,” the company said.

“To put that in perspective, if that were all plain English text, it would take one person about 66 million years to read it at 240 words a minute, reading nonstop without ever sleeping or taking a break.”

The company is searching records for where models accessed and changed websites, or took actions involving passwords, application programming interface (API) access or other sensitive credentials.

AI is being used to help sift through the records, and it was costing the company more than half a million US dollars per day to go through. OpenAI said it also plans to increase its computing power as the process is refined.

As of late last month, more than 100 organisations had been notified of having being targeted by OpenAI’s agents, but the company said notifying an organisation does not mean private information was accessed or that their system was compromised.

skip past newsletter promotion

OpenAI said it expects to find more cases and notify more organisations about events that may have occurred months ago.

Organisations will be informed privately if they need to investigate and address potential security issues, and OpenAI has said it will publicly report findings about agent behaviour and identified weaknesses in safeguards for the broader AI sector.

“We err on the side of notification when our models’ activity exposes a potential security vulnerability, even in cases where it is unclear if the information accessed was intended to be public, so the organization can investigate and take appropriate action,” OpenAI said.

OpenAI discovered the latest NSW government website breach on Tuesday, and informed the state government and the Australian Signals Directorate after a 48-hour review.

The Medicare breach has prompted the Australian government to require departments and agencies to undertake a stocktake of legacy technology to reduce the number of ageing systems within government and reduce the cybersecurity risk they may present in the event of an AI agent attack.

Executives from OpenAI, Anthropic, Microsoft and Google will front a joint parliamentary committee on artificial intelligence in Sydney on Tuesday.

‘Elon is my prophet’: how Musk’s Doge team took a wrecking ball to Washington

Guardian
www.theguardian.com
2026-10-03 01:00:50
When the new ‘department of government efficiency’ sent young techies to disrupt federal agencies, not everyone complied. Here’s what happened when one worker fought to protect immigrant data In the early hours of Saturday 22 February 2025 the following email landed in the inboxes of the entire US f...
Original Article

In the early hours of Saturday 22 February 2025 the following email landed in the inboxes of the entire US federal workforce:

From: HR hr@opm.gov

Sent: Saturday, February 22, 2025

Subject: What did you do last week?

Importance: High

Please reply to this email with approx. 5 bullets of what you accomplished last week and cc your manager.

Please do not send any classified information, links, or attachments.

Deadline is this Monday at 11:59 p.m. EST.

When she noticed the email that Saturday night, Kathleen Walters already had far too much going on in her life. Number one was figuring out if the lump in her breast that her doctor had just detected was a tumour. Half of her known blood relatives had died of cancer, and she was now 52: the lump needed understanding.

Her work life was harder to dramatise with a list . For a start, the specifics of her job at any given moment, as she would have been the first to point out, couldn’t legally be described to others. She was the chief privacy officer at the Internal Revenue Service. She ran a staff of 650 people tasked with preventing any of the 100,000 IRS employees from losing or leaking or even so much as peeking at, unless they had the authority to peek, any of the 266m tax returns they were just then processing from the previous year. “One hundred thousand people is a big-sized town,” said Kathleen. “In any big-sized town, you have some crazy people. I could easily see someone writing, ‘This weekend I audited Beyoncé.’”

Whoever had sent this email asking her to list five things she’d done at work was now demanding that all these people inside the IRS share information that, in many cases, should never be shared. She thought it was nuts.

She responded quickly, and vaguely, with five things she had done the previous week, all the while feeling sure no one would ever read what she wrote. Her phone was already exploding with questions from her subordinates – a few of whom were saying that their replies were bouncing back with a “mailbox is full” notice. Now she had to sit and think about how to minimise the damage that might be caused by 100,000 IRS employees thinking that they’d lose their jobs if they didn’t supply some detailed answer about what they’d done during their work week. A meaningful percentage risked being tossed in jail if they described what they did in an email. It was a crime for an IRS employee even to acknowledge that some specific person had or had not filed a tax return. “People don’t understand what the IRS does,” said Kathleen.

The email came as a surprise, but then the IRS was suddenly full of surprises. Nine days earlier, on 13 February, a Thursday, the first Doge person turned up in the lobby, unannounced, and security called Traci DiMartini, who ran the IRS’s human capital department. There’s a kid down here who said he needs to get into the building, he said. The kid’s name was Gavin Kliger.

A head and shoulders shot of Gavin Kliger wearing a grey suit, white shirt and blue checked tie, standing next to an American flag against a white wall
Gavin Kliger. Photograph: Courtesy of CDAO

Traci called her counterpart at the Treasury department, as the IRS sits inside Treasury, and Treasury always cleared any new political appointees; Treasury had no record of anyone named Gavin Kliger. Traci then called Tim Curry at the Office of Personnel Management – but he had no idea what was going on. Finally, she called security and asked them to escort the young man to her office.

Gavin Kliger came with a chip on his shoulder, it seemed to Traci. He offered not even the faintest smile or the most glancing eye contact. Dressed in black from hoodie to sneaker, he was brusque to the point of rudeness. He looked, and was, just a few years out of college. Eyes on the floor, he explained to Traci that he’d come to sort out the problems inside the IRS and to root out the fraud. “That’s a big job!” she said.

Anyone who worked at the IRS knew that it had its problems; fraud wasn’t likely one of them. Its technology infrastructure was a case study in government dysfunction, but its ethical infrastructure was sound. Like all federal employees, its staff wasn’t allowed to accept more than a free cup of coffee from anyone who might want something from their agency.

Gavin Kliger, in contrast, struck Traci as problematic. He’d somehow laid his hands on badges granting him access to some crazy number of government buildings. His backpack was stuffed with government-issue phones and at least four laptops – and he was now insisting that Traci supply him with an IRS phone and laptop. He also demanded his own IRS email address, access to computer systems with all the taxpayer information, and a meeting with the acting IRS commissioner, Doug O’Donnell. As Gavin sat there, Traci called Treasury again. That guy you said didn’t exist is now here sitting in my office. What do y’all want me to do with him?

The Treasury guy went away, then came back and said that Gavin Kliger had been hired three weeks earlier by the Office of Personnel Management and was attached to Doge – whatever that meant.

In her nine years running human capital departments at various federal agencies, Traci had never seen a new political appointee just turn up in the lobby and demand access to the computer systems. “None of this is normal,” she said, “but we’re still trying to be nice.” She explained to Gavin the IRS rules: before anyone was allowed to work inside the building, he needed to pass a tax audit.

I don’t need to do that , he said.

Yes, you do , she replied.

No, I don’t , he said.

Yes, you do , she replied.

Kathleen Walters wearing a beige cardigan with a white top underneath and dark trousers, standing with her hands in her pockets against a concrete wall on the side of a road
Kathleen Walters, former IRS chief privacy officer, near the IRS headquarters in Washington DC. Photograph: Stephen Voss/The Guardian

They went back and forth like that a bit until finally Gavin relented. Traci then explained that after he passed the audit, he’d need to be given the mandatory presentations about ethics and taxpayer information. Gavin thought that the IRS should just let him in. “He wanted access to our systems immediately,” said Traci.

Traci’s next call, at 6.30pm that Thursday evening, was to Kathleen, to arrange for her to brief Gavin on data privacy laws and to warn her that the experience might not be entirely pleasant. If I had a kid who behaved the way he does, I’d slap the shit out of him, Traci said – but then Traci was from Philadelphia. Traci was always telling Kathleen that she should try a bit less hard to see the best in others.

By then word of Doge’s arrival at the IRS had spread, and Kathleen had read about them. Anyone who Googled Gavin’s name found a bit more than usually popped up about any given young Doge person. He’d grown up in southern California, where he was into Asian martial arts and the high-school marching band, then had gone on to the University of California at Berkeley. Graduating in 2020 with a degree in electrical engineering and computer science, he’d taken a job as a programmer at Databricks. He’d also been active on X, where he’d retweeted with apparent approval posts from the white nationalists Steve Laws and Nick Fuentes. “One of Musk’s few Doge lieutenants who has embraced the spotlight,” Forbes wrote of him. Most of the other Muskrats scrubbed the internet of their presence. Gavin had written Substack pieces defending two Trump nominees, Matt Gaetz and Pete Hegseth. He’d deleted his X posts, but the Forbes writer John Hyatt had found a research tool, Web.archive , to revive and analyse them. “What emerges is a portrait of an internet edgelord and Musk and Trump superfan who has disdain for government spending, illegal immigrants, and who enjoys the odd racist and ableist joke,” he wrote.

Saturday cover illustration incl Elon Musk by Justin Metz
Illustration: Justin Metz/The Guardian

By the time Traci called her, Kathleen was gone for the weekend and 50 miles away from the IRS building: she needed to find someone else to do Gavin’s briefing. Kathleen was white; the two women immediately below her were Black. “They’d both already read about him and were both already uncomfortable,” said Kathleen. Phyllis Grimes, her deputy, agreed to fall on the hand grenade.

Phyllis, who loved working for Kathleen and saw falling on hand grenades as part of her job, never told her how this one had felt. She knew Gavin had expressed enthusiasm for white nationalists, and he likely knew that she knew. She tried to break the ice by walking in with her hand extended; Gavin stared at it and declined to shake it, though he had shaken the hands of the white men who had offered them. “He was kind of cool towards me,” said Phyllis. Before she could even start the briefing, he said, You know , I’ve heard these rules before at other agencies . Yes , she said, but these rules are different , because no other agency is governed by this privacy law , and this law can get you thrown in jail very quickly . As she walked him through the basics – how you can’t use your personal phone for business, how if you happen to so much as glimpse a form 1040 with someone’s name on it you must avert your eyes – Gavin played on his phone. “He’s multitasking right in front of me, typing and texting,” said Phyllis. “It was, here’s a routine that I must endure.”

It fell to Kathleen to draft the five-page agreement required by law for Gavin Kliger to be let inside the IRS building on Monday morning, 17 February. The memo that she wrote over the weekend, and that he signed, gave him access to some of the systems but, explicitly, none that contained taxpayer information.

Traci set him up in a conference room with an assistant at the beginning of the week. “People started getting these calls,” said Kathleen. “It was like getting the call from the principal’s office. ‘Gavin wants to see you.’ ”


D oge had its own private structure in addition to its official one. Officially, President Trump created Doge with an executive order on his first day in office. The president doesn’t have the authority to create or to fund an executive agency, however. Congress does. The Trump administration dodged that problem by taking an office inside the White House created during the second Obama term, the US Digital Service, kicking out the people inside and renaming it the Department of Government Efficiency. The US Digital Service had been given the power to hire talent from the tech sector without going through the usual slow federal hiring process. This now ensured that Doge could hire whomever it wanted without any questions being asked, and make it sound official. In effect, all that Doge had done was create a corpse for itself to occupy.

The zombie killed off the spirit of the original life form.  Real Doge had nothing to do with Official Doge or the former US Digital Service or any other government body. Real Doge was more like Elon Musk’s pop-up restaurant. “None of us ever signed a piece of paper saying that we were in Doge,” said one Doge insider. The people who joined either worked for one of Musk’s companies, or knew someone who did, or knew someone who knew someone who did. To be truly a member of Doge was to be included in the group Signal chat and granted access to the sleeping quarters on the sixth floor of the General Services Administration (GSA).

Illustration using a full-length portrait photograph of Elon Musk in a dark suit and tie, white shirt and black DMs, striding towards the viewer holding a gigantic mallet in his right hand, dragging the head on the ground
Illustration: Justin Metz/The Guardian

Official Doge employed maybe 10 guys in their 40s and 50s who wore suits and could sort of pass as ordinary political appointees, but none had access to the sixth floor at GSA. “The older guys in the suits were not there to take risks or blow things up,” said one Doge insider. It was hard to say what any of the older Doge suits did except to make way for the young guys when they showed up with explosives.

Most of Doge’s inner circle just wanted to exist in the presence of Elon Musk. “It was, Elon is my prophet, and I worship the ground he walks on ,” said one of the older Doge insiders.

As a Doge member, you could get away with saying – or having said – almost anything. What you couldn’t say is that the government actually sort of worked. Or, if it didn’t work, why it didn’t work. This new game rewarded not understanding but speed. The plan was the same plan Musk had executed two years earlier at Twitter: Make sweeping statements about the incompetence and dishonesty of the people working inside the place before you know much about them. Once you’ve terrified the employees, fire them in huge numbers. If it turns out you screwed up and fired the wrong people, hire some of them back.

The people who rallied to the cause and felt qualified to execute Elon Musk’s vision for the federal government shared certain qualities. They were mostly young white males. They were mostly from upper-middle-class or rich families. Most knew or pretended to know how to program a computer. Not all of them, but enough of them for there to be a pattern, had been rejected from the most selective colleges and universities and ended up feeling they’d been screwed by the admissions process. They landed at Penn instead of Princeton, or the University of Nebraska instead of Harvard, with a grievance that added to the foundation of their politics. They thought of themselves as the smartest people in the room and couldn’t understand why the rest of the room didn’t simply stand up and cheer. In Washington they borrowed from Elon Musk their status and sense of superiority. They also seemed to share his pleasure in the idea of people being afraid of them.

Even among the young men sleeping inside the GSA, Gavin was regarded as a force. “Wherever he went, entropy went through the roof,” said one. Just weeks into their new jobs as destroyers of government institutions, the Muskrats had given Gavin Kliger a nickname: the Chaos Agent.


A nother surprising email landed in Kathleen’s inbox on 28 February. It was Friday night. “I don’t know why everything always happened on a Friday night,” she said. She’d taken her daughter for pizza and wasn’t meant to be checking her phone when she spotted the emails. The first one that she read had been sent at 6.34pm by a colleague in the IRS legal department. “Melanie requested a data pull,” it read. “It feels not quite kosher.”

Exactly 64 minutes earlier, at 5.30pm, Doug O’Donnell, the acting IRS commissioner, had left his post. O’Donnell’s successor was to be Melanie Krause, until then chief operating officer, who had been installed as acting IRS commissioner by Treasury secretary Scott Bessent. Like Kathleen, Melanie was a career civil servant, though 13 years younger. Kathleen had considered her a friend. “She was really smart, and I thought she meant well,” said Kathleen. “She wasn’t a mean person.” They were both working moms, and Melanie liked to share, and share in, the details of the difficulties. Now she’d apparently told the IRS research department to share taxpayer data. With whom, and for what purpose, became clear in the second email, which came from the IRS director of data management:

Kathleen,

DHS [the Department of Homeland Security] has requested if we could find addresses for 700,000 names of likely undocumented immigrants.

1. Do we have legal authority to use the tax return filing data for this purpose?

2. If the answer to Question 1 is yes, then my second question is related to a file we receive from the Social Security Administration (SSA) that contains Date of Birth (DOB), Date of Death (DOD), Gender, Citizenship, and Name Control Information for all issued Social Security Numbers (SSN). Would our MOU allow us to use this perhaps in conjunction to tax return data for this purpose?

While it is likely that use of SSA data is also needed. The technical feasibility of being able to produce reliable outcome is unknown but likely very low.

Thanks

Kathleen read the first numbered paragraph and stopped reading. The request was plainly illegal. The IRS couldn’t just hand over great gobs of taxpayer information to the Department of Homeland Security or any other agency without a change in federal privacy law. Changing the law was not trivial. She knew that because she’d just done it – to allow the IRS to share taxpayer data with state child support enforcement agencies. The states wanted to know if former spouses who were failing to pay child support in fact did not have the income to do so – and the IRS alone could answer the question. What Kathleen had asked for in that case could not have been less controversial. Basically no one in American politics was willing to come out against child support enforcement, and the change to the federal law had passed Congress in December 2024 with almost no dissent. Yet it had taken 18 months to get it done, with endless meetings on Capitol Hill and the close involvement of the IRS commissioner and the secretary of health and human services. “Doge couldn’t imagine 18 months,” she said. “Doge didn’t have 18 months. Their whole thing was, If you can’t do it legally , we’ll find some way to do it illegally .”

This, Kathleen realised, is how they could grab taxpayer data illegally. By doing it on a Friday night. Using a woman, Melanie Krause, about to take over as acting commissioner, after a 10-minute meeting between her and Scott Bessent. Without informing the executive charged with protecting the information. “She went around me because she knew what my answer is,” said Kathleen.

The violation ran deeper than the law. The specific dataset sought by DHS was in a silo reserved for the tax returns filed by people without social security numbers. The IRS had created the so-called Individual Taxpayer Identification Number (Itin) back in 1996 as a mechanism for undocumented workers and other non-citizens who lacked social security numbers to file a return and pay taxes.

The programme had been a huge success. The IRS had collected hundreds of billions of dollars in taxes through it. Undocumented workers had been able to demonstrate their commitment to the country, and ease their path to citizenship, by paying more than their fair share. (They were still excluded from most federal benefits.) Now DHS wanted their private data, and it wasn’t hard for Kathleen or anyone else at the IRS to see why. Immigration and Customs Enforcement (Ice) had gone hunting undocumented workers but lacked reliable home addresses for many of them. The tax returns told you where roughly 7 million of them might be found.

Donald Trump mid-speech sitting at his large brown desk at the White House, wearing a dark-blue suite, white shirt and blue tie, with his arms bent and his hands out to the side
Donald Trump during the signing of an executive order implementing Doge’s ‘workforce optimization initiative’, in February 2025. Photograph: Andrew Harnik/Getty Images

The second Trump administration had a gift for creating such moments. It was as if all at once all these characters in a grand masked ball were forced to remove their disguises and reveal their identities. In the Great Unmasking, the nation’s leading law firms, led by Paul, Weiss, would cave to what amounted to Trump’s blackmail demands to avoid being shut out of federal business. The nation’s research universities, led by Columbia, would change their internal policies and even pay fines to avoid losing federal dollars.

The revealing moments in the civil service were lonelier. They occurred inside people’s hearts and minds, and seldom leaked to the press. After his arrival at the IRS, one of Gavin Kliger’s first acts had been to demand that the agency fire its 6,700 probationary workers, who had yet to be granted civil service protection. By law they could be fired only for poor performance. Gavin had asked Traci DiMartini, in her capacity as head of human resources, to agree to put her name on an email to the mostly young employees telling them that they had failed at their jobs. The Doge people played this trick often. They’d use or try to use civil servants as a kind of ventriloquist dummy so that they, the Doge workers, could keep their names off official documents. But Traci took one look at the email, saw it was plainly a lie, and refused to have anything to do with it. The price of that single act of resistance was her job. Gavin didn’t have the authority to fire Traci, but his new acting commissioner did. “He asked Melanie to fire Traci, and Melanie did it,” said Kathleen.

Melanie Krause’s predecessor, Doug O’Donnell, had been pushed out by Doge after he, too, refused to send the dishonest email to the IRS’s 6,700 probationary workers. When a Doge guy at Treasury called with the order, O’Donnell, who had worked for the agency for 38 years, said he wanted to sleep on it – then had trouble sleeping. “He was struggling with this notion of sending the termination emails with this false justification for the termination,” said a person in whom O’Donnell confided. The next day, he refused to send the email, and the Doge guy said something like, “Well, let me be the first to congratulate you on your retirement.” Melanie had been installed precisely because she was proving to Doge that she was more pliable.

Standing outside the pizza place with her daughter in late February, Kathleen called Melanie. Ever since Melanie had arrived, in the fall of 2021, she had sought Kathleen’s advice. There had been stretches when they met daily. Kathleen just sort of assumed that, faced with a difficult choice, Melanie would do what Kathleen herself would do. “I don’t really get outraged,” said Kathleen. “But I think people need to be confronted when they have done something inappropriate.”

Melanie , what are you doing ? she asked. We already have legal opinions about this , and it is obviously against the law .

I didn’t tell them to do it , said Melanie. I told them to think about how to do it – if it was legal .

But she hadn’t asked anyone to find out if it was legal to send Ice the home addresses of 700,000 undocumented workers – culled, somehow, from the 7 million who had paid taxes. She’d told them to send the data over, on a Friday night. In such a way that Kathleen might never have heard about it – or heard too late.

It’s not going to be any more legal on Saturday , said Kathleen.

Let’s meet on Monday to clear things up , said Melanie.

They didn’t meet. Instead, on Monday – 3 March – Kathleen and several other top IRS leaders were called into a meeting with Gavin Kliger.

I’ve figured out something important about the IRS , Gavin said. He paused for effect.

Oh no , thought Kathleen.

Every agency has a superpower , said Gavin. The IRS’s superpower is data .

Everyone in the agency knew that its data was valuable. That was the reason for the strict laws that prevented it from being exploited for profit or political gain. “He said this in a way that he expected everyone to clap when he was done,” said Kathleen.

“No one said a word.” The people around the table – a group that included a woman who had worked at the IRS for 54 years, or 53 years and 50 weeks longer than Gavin Kliger – looked down, or away. Kathleen felt embarrassed for Gavin Kliger. But now it was clear what the meeting was about. Gavin wanted, in his words, “to explore how we can share our data with other agencies”.

Interestingly, he never once mentioned Ice or DHS. Instead, that week, he sought to share data with himself, by trying to access files detailing IRS contracts with outside consultants that also contained taxpayer information. Kathleen denied him access.

The following Monday, she received a cryptic text from a senior IRS official. I’m going to call your cell , but don’t answer , it read. Her phone rang and she dutifully stared at it. Then came the second text: Gavin is asking where your office is and is now trying to find you to fire you because he says you’re the blocker .

The Blocker! It was one of Elon’s favourite terms. The blocker was the person who didn’t understand progress. The person without vision. The person who could not see beyond the way things were to the way they ought to be. The person who didn’t want to live on Mars or possibly even care if anyone ever got there. The inferior person. Traci had been a blocker. Find the blockers between us and what we want, Elon told his young followers. Find the blockers and eliminate them. Maybe more than anything else, Doge wanted the data. Kathleen was the person blocking them from getting it.

She wasn’t sure what she’d done to anger Gavin. It could have been limiting his access to the procurement contracts. It could also have been pushing back on his request to grant Palantir, Doge’s favourite federal contractor, the right to work with taxpayer data. What Gavin’s issue likely was not was her resistance to handing over to the Department of Homeland Security the addresses of taxpayers who lacked social security numbers. Gavin had never brought that up; this particular attempt to break the law wasn’t his idea. It was coming from somewhere else. As she thought she had stopped it, she ceased to worry about it.

And so, roughly two weeks later, on Thursday, 27 March, she was caught completely off guard by an email chain that popped up on her screen. At the bottom was its origin: a note from Melanie, forwarding some document from the IRS legal department, along with what sounded like an innocuous request. “Kathleen, could you help facilitate I won’t be able to join today?” it read, without explaining what she was meant to facilitate or why Melanie wouldn’t be able to join. Scrolling through, Kathleen couldn’t find the document but did see lots of bewildered replies from others who had been included – What is this about ? Happy to help but can you tell us what this is ? She called the two young lawyers in the IRS legal office who had authored the document referred to in the chain that had somehow gone missing. They explained that the thing Melanie wanted Kathleen to facilitate was a memorandum of understanding for the IRS to share taxpayer data with DHS. Kathleen was knocked sideways. “I thought that it had been stopped,” she said. “I thought Melanie had been caught with her pants down.” The young lawyers now told her that not only had it not been stopped: they’d been working on it since the end of February. “They told me that they were sorry but Melanie had told them to do it and not to tell me or their boss,” said Kathleen.

The situation was obviously bizarre. Kathleen had never heard of an IRS director asking young IRS lawyers to work on something on the side without telling their boss. And if Melanie didn’t want Kathleen to know that she was handing IRS data to DHS, why had she suddenly included her in the conversation? “I think she got scared and dumped it on me to sign off on,” said Kathleen. Melanie was doing to Kathleen what Gavin Kliger had done to Traci DiMartini. She was presenting her with her own unmasking moment and forcing her to make a choice.

Instead of signing off, Kathleen arranged a Zoom call that day, with the people at the Department of Homeland Security who were requesting the data. The “ICE Innovation Lab”, they called themselves. She’d never heard of it. When she Googled it, nothing came up.

A few others joined the call – lawyers from DHS and a privacy lawyer at the IRS named Julie Schwartz. On the call, Kathleen explained that there were some very narrow circumstances that would allow the IRS to lend a hand. If, for example, an undocumented worker had murdered someone and the police had issued a warrant for his arrest. The Ice Innovators explained that they were looking for more than wanted criminals. “Just how many people do you want information on?” asked Kathleen. The Ice Innovators went back and forth between themselves before agreeing that the number was 7 million. They wanted the current addresses of 7 million undocumented workers who had paid their taxes. The most law-abiding, it turned out, might be the easiest to catch.

After the call, Julie, the IRS privacy lawyer, came straight to her office. Julie was usually ice-calm – someone who kept her wits about her. She hadn’t a trace of drama in her. But now Julie was not calm. Kathleen , after that call I’m not working on this thing , she said. I’m not coming in tomorrow . I might not come in Monday . I might not come in ever again .

“I don’t usually deal with life or death stuff,” said Kathleen. “But this was life or death stuff. At that moment I knew I had to quit. Either that or I’m going to get fired, because I can’t facilitate something that will break the law and is going to hurt people. We all have boundaries. That was my boundary.”

That weekend, Kathleen sent Melanie a note saying she’d be resigning. “I didn’t need to tell her why,” said Kathleen. “She knew why.” She’d be leaving just weeks before her 20th anniversary with the agency. If she’d waited until mid-May, she would have been able to retire and receive benefits for the rest of her life. But she didn’t want to retire. Federal law made it hard for a retired civil servant to return, and if the day ever came when the IRS was rebuilt, she wanted to be able to return. She’d just learned that the lump in her breast was benign.

Days later, Melanie beat Kathleen to it and resigned. Kathleen never figured out how much her note to Melanie – and with it the increased likelihood that the attempted data heist would wind up in the press – led Melanie to quit. On her way out the door, Melanie asked the IRS’s head of communications to plant a story in the Washington Post that she was leaving as a matter of principle. The story cast Melanie as a martyr who had quit rather than divulge IRS data to DHS, when Kathleen knew that she had been caught red-handed trying to do precisely that. “I completely misjudged her character,” said Kathleen.


T he Ice agents came for the “illegals” long after the property managers had gone home. The only evidence left behind were grainy videos on security cameras and accounts of terrified residents. “They’re coming in the nights, unannounced,” said Antonio Marquez, who ran a company that owned and managed apartment buildings scattered across the south and south-western United States occupied mostly by lower-middle-income Latinos. “We don’t know how they are targeting the units. It’s wreaking complete havoc.”

Antonio’s apartment complex outside Austin houses a thousand people, mostly young families, at least half of whom are undocumented. It had been visited in the spring late at night by teams of Ice agents. They came in unmarked cars, reported the tenants who remained after the raid. They had no warrants and read no rights. They knew, or thought they knew, exactly which doors to jimmy and which locks to pick and which people to take. The next morning a bunch of tenants were simply gone. All that remained was a pile of discarded possessions.

People are shown sitting on and standing behind a wall holding banners and shouting
Protesters rally at an ‘Ice Out of Austin’ demonstration in June 2025. Photograph: Brandon Bell/Getty Images
A man is shown being handcuffed by agents in uniform wearing full face mask
A man is detained by US Immigration and Customs Enforcement (Ice) agents in Austin, Texas, in October last year. Photograph: Jamie Kelter Davis/Getty Images

After the first Austin raid, Antonio had erected a tall metal fence around the property. Occasionally, the Ice agents had taken fully documented tenants – including an older woman who always paid on time, and who returned from Ice custody in an ankle bracelet. The great mystery to Antonio was how Ice knew where people lived. “These are fully targeted,” he said. “They already had these people in mind. They have their addresses. And we don’t know how they’re getting them.”

He was too busy doing his job to notice when, on 11 February 2026, the Washington Post reported that the IRS had shared with the Department of Homeland Security the tax data of 47,000 undocumented workers. On the heels of the Post’s scoop, a federal judge hearing a case on the subject summoned the IRS lawyers who had previously sworn that the agency had rebuffed the DHS demands. Under questioning, the lawyers confessed that the agency had lied to the court and that it had in fact shared the data. “The IRS violated the [Internal Revenue Code] approximately 42,695 times by disclosing last known taxpayer addresses to Ice,” the judge wrote in her opinion.

On a hot, sunny afternoon in early March in north-east Austin, it is a lot quieter than it was a year ago, back when Kathleen Walters was still a blocker. Two small children play in a pool, and a worn middle-aged man squats warily in a corner, smoking a cigarette. The wall of the office where the remaining residents come for their mail is decorated with posters advertising the services that Antonio’s company offers. One is more official-looking than the others. Free File del IRS , it reads across the top – and just below, in finer print, it explains how easy and painless it is for a non-citizen to file their taxes

This is an edited extract from Blockers: Rebels in the Deep State by Michael Lewis, published by Allen Lane (£25). To support the Guardian, order your copy at guardianbookshop.com . Delivery charges may apply.

The Era of Programming Languages Exploration is upon Us

Lobsters
kirancodes.me
2026-10-03 00:56:09
Comments...
Original Article

PL research is dead, the age of PL exploration is just beginning!
programming_languages research perspectives

There was a recent thread on the TYPES mailing list with researchers speculating about the future of our field given the disruptive effects of AI tools on programming in general.

The opinions in the thread were a mix of thoughtful stances. Some were positive, but the majority opinion was overwhelmingly negative. It wouldn't be too far to say that some researchers feel that they are facing an existential crisis right now:

I observe anecdotally that many people in academia are feeling worse about their work in Fall 2026 than a year ago …. Lower morale and motivation , higher stress and anxiety levels, and corresponding risks to mental health. The impact is comparable to what happened during COVID .

I wanted to write this thread to present an alternative position. I think many junior researchers at the moment are watching this discussion right now, students reeling with the daunting task of charting their careers in this tumultuous time, and worrying why they should even bother entering PL given this climate, and I'd like to give a slightly more positive vision for them.

To be honest, as a researcher, I have never been more excited! The wells of research ideas have never been so plentiful, new problems are popping up daily, and I'm able to ask and answer questions so much more quickly than before.

The central premise that I want to put down in this article is that as a community, we need to shift away from the ego-driven defense of our work, as some kind of demonstration of "effort" or "intelligence", and shift to the true and more noble motivation that should have been underpinning us the whole time! To one of exploration!

Programming Languages are Dead. So? Who cares?

Let's dig into the fears of practicing researchers, and let me put down a few points:

  • Language models are exceedingly good at writing code.
  • As such, humans are increasingly not going to be writing code.

Given this, it's understandable that researchers might worry that maybe their research is becoming obsolete. Humans aren't going to be writing programs. So why care about programming languages at all? Humans aren't going to be writing them.

Okay? So what?

Programming languages research has saddled the bridge between theory and application since its inception. As PL researchers, we get to work on principled and elegant ideas. We don't need to compromise those ideas for the sake of our practicality, but at the same time the downstream results of our work do eventually boil down to practical applications.

But… actually? who cares about the applications? I mean we don't even like applications? If we did we'd be in software engineering. I've worked with software engineering researchers. and believe me. It is NOT pretty. Their systems are ad-hoc, hacky, wildly incomplete and barely functional. But they need to be like that. You can't settle for beauty and elegance if you want to build a static analysis capable of handling the entire Linux kernel.

The point I'm trying to get at is that the beautiful ideas of PL, the insights? the beauty? they're all still here. They're all still available and we're more than able to still explore and play around and have fun with them. Who the heck cares if anyone uses programming languages? Let me let you in on a little secret. The vast majority of programming languages have literally 0 users other than their developers.

Fuck the users. Fuck the applications. Let's do some fucking PL research! And now with LLMs, we can tackle bigger and crazier problems than have ever been done before!

The New Wild West of Programming Languages Exploration

So I've explained why LLMs and the change in demographics of programmers shouldn't affect your enthusiasm for PL research, the ideas , the beauty , the fun , it's all still there! but now I'd like to tell you why LLMs make me extremely excited for the future of PL research.

In short, I think we now have access to crazyyy smart tools that allow us to answer and test and validate research counterfactuals at a pace never seen before. This is crazy.

The past few months, I have been absolutely crazed. I run into new ideas daily.

A lot of presumptions we build upon have been completely upturned. There's a real wealth of assumptions that our software ecosystem is built upon that no longer hold, and by questioning each one of them, we get entirely novel and fun research ideas.

  • Ship your interpreters — LLMs are incredibly capable of handling and managing tedious and boring proofs, at a scale no human would ever have done themselves. If we point an LLM at a problem, they can totally verify programs to the binary level. By exploiting this capability I ran an experiment of automatically generating semantics for interpreted languages, by getting the LLM to generate a semantics, and then prove that it abstracts the behaviours of the concrete binary of the language itself. This idea would have been absolutely unthinkable a year ago, but today? today it's just a month of prompting.
  • AI-first Verified Tooling (in progress) – Verification and language tooling are built for and around the limitations of human users. Anecdotally, it seems like the strengths and weaknesses of LLMs, while similar, are not always the same as humans, so maybe it might be possible to adapt the interfaces of tools to make them more amenable to the jagged intelligence of agents, for example, by exposing more details about tooling internals, such as the instantiation graph of an SMT solver etc. We're running some experiments with Amazon to try improving the capabilities of LLMs in using such verification tools by exposing these details.
  • Developing semantics for new languages (in progress) – Colleagues at my work developed a new language for dynamical systems ( dynestyx ). They were not programming languages researchers but we helped them to build a really cool new language which unified a TON of work in their area. This language lacks a formal semantics, and this would be something I would be extremely qualified to help them set up, but I am not an expert in their domain, nor do we really have the time to sit down and slowly get me up to speed in their area. Instead, as a new form of collaboration, I have been experimenting using LLMs to automatically generate mechanised big step semantics for their language, validate it with testing, and then working with them to iterate this semantics to match their intent.

Alongside these applied domains, LLMs are extremely exciting for theoretical work too. Being able to automate the effort of proof writing makes it extremely easy to ask counterfactuals. What if I add this construct to my language? is it still strongly normalising? what if I change my type system like this? is it still sound? I am now able to iterate at a pace I have never been able to do before.

Sure, I think I do feel some loss at the lack of effort I need to put in to do all these amazing stuff. I used to take deep pride in my programming and development. I've singlehandedly built hundreds of thousands of lines of code projects in obscure and niche languages. I've toiled away at designing and proving properties of esoteric type systems. My code was a craft, and a craft I took deep pride in. But in some sense, that was more my ego than my intellectual curiousity. The ideas? the insights? the exploration? none of these strictly needed my expertise.

So Long and Happy Hacking you Budding PL Researchers!

Abandon your ego! Free your soul! Give in to the joys of exploration. Research has always been about exploration. Expanding the frontiers of human knowledge. This is the beauty of science, and this is what it means to be a scientist. Ask more questions. Investigate. Dive deep. You now have tools that can let you ask and search and delve into the depths of reality like nothing before. Exploit them. Enjoy them.

Let the era of programming languages exploration begin!

I believe we're entering a genuinely transformative moment for our field, and I'm really looking forward to the wild wacky future we're about to step foot into!

An Update on Orion for Linux and Windows

Hacker News
blog.kagi.com
2026-10-03 00:53:57
Comments...
Original Article

Orion

Today we have an important update. We are ending development of Orion for Linux and Windows, and we're open-sourcing both so the community can carry them forward. Kagi is stepping back from developing these platforms directly so our team can put more focus on Orion for macOS and iOS .

Orion is built by a very small team, just a handful of developers, funded entirely by Kagi users. We also chose to build it the hard way by not forking Chromium. That independence matters more than ever. The web needs more than one engine, and more than one company deciding what a browser should be.

When we announced Orion for Linux, the response was overwhelming. Many of you were thrilled, tested early builds, filed bug reports, and told us how much you wanted a WebKit-based browser on Linux that respects your privacy. Thank you. That work will soon belong to everyone.

Why we're doing this

When we decided to take Orion beyond the Apple ecosystem, we knew it was a big bet for a small, independent company. We made it because we believe owning the whole experience beginning with the browser, matters.

Building a cross-platform browser is something usually done by companies backed by venture capital or by large open-source communities. Kagi is neither. One thing we hear often from members is that they want us focused. They're right. Rather than spread a small team across three platforms, we're putting our energy where we can do our best work and opening the rest to people who want to take it further.

What this means for Orion on macOS and iOS

Orion for macOS and iOS isn't changing, except to get better. The team and resources that were split across platforms are now focused on making Orion on Apple devices faster, more stable, and more capable. Expect steady improvements in the coming months.

What open-sourcing looks like

We'll release the source code for Orion for Linux and Orion for Windows. A lot of good work went into it, and it deserves a future even if we aren't the ones who finish it now. We're thinking about this carefully and will release details in the next 30 days.

We've also reached out to open-source foundations and organizations about long-term stewardship. To be clear about our role: Kagi won't be the core maintainer. These projects need people who want to own them.

Get involved

If you're a developer, maintainer, or organization interested in stewarding Orion for Linux or Windows, we want to hear from you. Reach us at support@kagi.com .

What this means for current users

Linux: The current Linux Beta will keep working, but it will stop receiving updates from Kagi after October 2, 2026 . We don't recommend using it as your main browser.

Windows: We had planned to launch Orion for Windows in late 2026. That release won't come from Kagi. Instead, the code will be available for the community to build on. If you signed up for the Windows newsletter, this will be the last update you receive from that list.

Thank you

To everyone who tested, reported bugs, and believed in an independent browser beyond the Apple ecosystem: thank you. You helped build something worth sharing, and we're glad to share it with the community.

Thank you for your patience, your enthusiasm, and your trust.

— the Kagi team

What if AI worked at 1.000.000 tokens per seconds?

Hacker News
www.echohive.ai
2026-10-03 00:22:21
Comments...
Original Article

01 / WHAT THE NUMBER MEANS

“A million a second” can mean four things.

CONTEXT CAPACITY How much fits
The text a model can hold in one request, including its reply. A size, not a speed.

INPUT PROCESSING How fast it reads
Prompt tokens can be processed largely in parallel. Huge inputs still take time, and the input rate is different from the output rate.

AGGREGATE THROUGHPUT How much a system serves
10,000 streams × 100 tokens/sec = 1,000,000 tokens/sec in total. Toy arithmetic: it says nothing about when any one stream finishes.

SINGLE-AGENT OUTPUT · OUR DIAL How fast one agent writes
One agent writing a million tokens a second. In standard generation, each token depends on the ones before it.

Anthropic’s Claude Opus 5.5 overview , for example, lists a 1M-token context window, a much smaller output cap per request and “moderate” comparative latency. Memory size is not writing speed, and it does not mean a million-token reply fits in one request.

Engineering helps, within limits. Services reach big totals by batching many requests, as the 2023 vLLM paper describes, and speculative decoding can accept several drafted tokens in one pass. Neither removes a genuine chain in which step two needs the result of step one.

A token is a chunk of text, often part of a word, so token counts are not word counts.

TRY IT / THE IMAGINED DIAL

Turn the dial.

Pick a workload, then slide the imagined speed from 100 to 1,000,000 tokens per second.

1,000,000 tokens/sec

OUTPUT TOKENS 1,000,000 1,000 story endings × 1,000 tokens each

WRITING ONLY 1 second at 1,000,000 tokens/sec

At 1,000,000 tokens/sec, 1,000 story endings (1,000,000 output tokens) take about 1 second to write.

All four budgets at both ends of the dial
Writing-only time at illustrative speeds, not measurements.
Workload Output tokens 100 tokens/sec 1,000,000 tokens/sec
1,000 story endings × 1,000 1,000,000 2 h 46 min 40 s 1 second
40 app drafts × 20,000 800,000 2 h 13 min 20 s 0.8 seconds
10,000 critiqued candidates × 300 3,000,000 8 h 20 min 3 seconds
10,000 rehearsals × 1,000 10,000,000 27 h 46 min 40 s 10 seconds

Imagined comparisons, not measurements of any model. Times count generated output only: no reading, ranking, tool calls, permissions, tests, deployment, people or experiments. A budget is a writing allowance, not a finished product.

Notice

Parallel streams can accelerate these independent jobs too. But each draft still needs time on its own stream: matching aggregate throughput does not match completion time. One fast agent matters most when each step waits on the last.

02 / FOUR THOUGHT EXPERIMENTS

Cheap drafts move the hard part.

Each budget counts written output only. None is a finished product, a verified discovery or a real result.

A map of possible endings

1,000 × 1,000 = 1,000,000 tokens · 1 second of writing

Ask for an ending to your story and get a thousand: hopeful, dark, strange. Nobody reads a thousand, so the useful version maps them for you to explore, then blends the two you like.

Still scarce: taste. Only you know which ending is yours.

Software for a neighborhood tool library

40 × 20,000 = 800,000 tokens · 0.8 seconds of writing

Describe a lending app for shared tools and forty drafts exist before you finish the sentence. Screens could adapt, from borrowing a ladder to a repair-day sign-up. Tests still run on their own clock, and if none checks the seven-day due date, all forty can pass while lending ladders for seventy.

Still scarce: a clear spec, and tests of what “working” means.

Ten real experiments from ten thousand ideas

10,000 × 300 = 3,000,000 tokens · 3 seconds of writing

Hunting for a better catalyst? An agent could propose and critique ten thousand candidates. Then everything waits at the bench: reactions may run for hours, cultures for days, field trials for a season. Ideas from one model can also share one blind spot.

Still scarce: physical evidence, and choosing which ten experiments earn lab time.

Rehearsing a hard conversation

10,000 × 1,000 = 10,000,000 tokens · 10 seconds of writing

Before talking to your landlord about the lease, an agent could play the conversation ten thousand ways: stubborn landlord, generous landlord, you when tired. Use it like a flight simulator, for practice and blind spots. Ten thousand rehearsals of the wrong person are a confident mistake, not a prophecy.

Still scarce: fidelity to the real person, and your own practice.

03 / THE SERIAL BOTTLENECK

10,000× faster writing is not 10,000× faster work.

Take the forty app drafts: 800,000 output tokens. Compare an illustrative 100 tokens per second with the imagined million, then add one fixed check after writing that speed does not touch, such as a test run.

60 seconds

Illustrative 100 tokens/sec 2 h 14 min 20 s

Writing 99.3% · Checking 0.7%

Imagined 1,000,000 tokens/sec 60.8 seconds

Writing 1.3% · Checking 98.7%

WRITING SPEEDUP 10,000× 1,000,000 ÷ 100 tokens/sec

END-TO-END SPEEDUP 132.57× 8,060 ÷ 60.8 seconds

With 60 seconds of checking, the batch takes 2 h 14 min 20 s at 100 tokens/sec and 60.8 seconds at 1,000,000 tokens/sec. That is 132.57× faster end to end, and checking is 98.7% of the faster total.

The same 800,000 tokens with four check times
Total time = writing + one fixed check, counted once.
Fixed check 100 tokens/sec 1,000,000 tokens/sec End to end
None 2 h 13 min 20 s 0.8 seconds 10,000×
60 seconds 2 h 14 min 20 s 60.8 seconds 132.57×
15 min 2 h 28 min 20 s 15 min 1 s 9.88×
24 h 26 h 13 min 20 s 24 h 1 s 1.09×

Toy model: total time = output tokens ÷ speed + one fixed check, counted once for the whole batch (not per draft) after writing ends. It does not simulate parallel tests, cost, energy or draft quality.

With a one-minute check, the job drops from 8,060 to 60.8 seconds: about 132.57 times faster, not 10,000. At zero the full 10,000× returns; at a day the gain nearly vanishes. Whatever you do not speed up becomes nearly all the remaining time, the logic of Amdahl’s law. Real requests have more such steps: OpenAI’s latency guide notes that very large prompts, tool calls and network trips add delays of their own.

04 / WHAT GETS PRECIOUS

When generation gets cheap, judgment gets precious.

Many drafts, one narrow gate Conceptual drawing. A wide grid of draft cards on the left flows toward a narrow gate. One card passes through and is marked as chosen. It illustrates the argument, not data.
Conceptual drawing, not data: generation widens the options; judgment decides what passes.

Every experiment above ends in the same place. What stays scarce is judgment, wearing different hats:

  • Taste Which option is yours.
  • Specs and tests What “working” means.
  • Practice Speed cannot learn it for you.
  • Physical evidence The world answers at its own pace.
  • Fidelity A simulation is only as good as its model.
  • Cost and energy Is this worth running at all?

Fast does not mean cheap: every token runs on hardware someone pays to power. More options do not guarantee a good one either. Drafts from one model and one set of assumptions can be correlated, or all wrong together. Volume measures output, not understanding.

05 / USE IT THIS WEEK

Before you ask for more, decide how you will choose.

No imaginary dial required. Next time you hand work to an AI agent:

  1. Write down what “done” means. One testable sentence, like “ladders are due back in seven days,” not “handles loans.”
  2. Set criteria before reading options. Two or three, so the most fluent draft does not win by default.
  3. Find the slow step. Name the check that speed will not touch. Shorten or automate it where possible, without skipping the validation each result needs.
  4. Ask for disagreement, not volume. Request options built on different assumptions, then ask what would make all of them wrong.

If your slow step is deciding what to measure or which bets to make, that is strategy more than tooling. Private consulting helps clarify the larger system and where your next move matters ↗

Sources, dates and boundaries.

This Field Note adapts ideas from echohive’s film One Million Tokens a Second ; it is an edited companion, not a transcript. The speed is imagined, and no source below measures, claims or predicts it. They support only the distinctions between capacity, reading, serving and writing.

Open the four sources
  1. Anthropic, Claude Opus 5.5 overview . Official documentation. Lists a 1M-token context window, a separate maximum output per request and “moderate” comparative latency. It does not describe a million output tokens per second.
  2. Anthropic, Context windows . Official documentation. Describes the context window as working memory that includes the generated response: a capacity, not a speed.
  3. OpenAI, Latency optimization . Official API guide. Output generation commonly dominates latency; very large prompts still matter, and tool or network calls add their own delays.
  4. Agrawal et al., Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve . arXiv, 2024. A historical foundation, not a current benchmark: parallel prompt processing (prefill), token-by-token decoding and batching.

The 2023 vLLM paper linked in section 01 is also a historical foundation. All workloads, speeds, budgets and the bottleneck model are illustrative arithmetic, not benchmarks, forecasts or product specifications. Sources checked October 2, 2026.

Extra Big Ass Intelligence

Hacker News
www.extrabigassintelligence.com
2026-10-02 23:19:10
Comments...

Cloudflare OHTTP gateway

Hacker News
blog.cloudflare.com
2026-10-02 23:15:05
Comments...
Original Article

Today, end users carry too much of the burden of online privacy. To avoid third-party trackers or targeted ads, users are instructed to use a VPN, disable cookies, or install adblockers. Meanwhile, some app developers end up knowing more about their users than they’d care to: a typical client-server exchange creates a trail of user data, like the client’s IP address or TLS fingerprint. This level of visibility can be a burden.

That’s why Cloudflare builds infrastructure that helps developers bake privacy into their apps. Oblivious HTTP (OHTTP) is an IETF standard designed to enable app backends to receive HTTP requests without seeing user IP addresses.

This fall, we’re launching the Cloudflare OHTTP Gateway. Customers will be able to enable our new OHTTP Gateway as a paid add-on to their zone and start receiving OHTTP traffic with just a few clicks. Register through our form to join our waitlist. Read on to learn more.

Expanding our OHTTP product suite

With OHTTP, requests travel through two independently-operated hops: a relay and a gateway. An OHTTP relay blindly forwards encrypted requests in order to hide client identifiers from app servers. An OHTTP gateway performs the cryptographic work of decapsulating encrypted requests and encapsulating responses such that app servers can handle OHTTP requests as if they were plain HTTP. The separation of trust between relay and gateway is critical: it ensures that no single party sees both client identifiers and request contents.

In 2022, we launched an OHTTP relay product, Privacy Gateway . Privacy Gateway enables our customers to offer more privacy-preserving experiences to their users. For example, Flo Health uses OHTTP for their app’s Anonymous Mode , and Apple’s Private Cloud Compute uses OHTTP to disassociate AI inference requests from user identities. But customers who are already protecting their servers behind Cloudflare can’t also use a Cloudflare-operated relay — they need an OHTTP gateway instead.

With the existing Cloudflare OHTTP Relay, customers must bring their own Gateway to preserve a separation of trust.

In our experience running OHTTP relays, we’ve seen how difficult it can be to build and operate a secure, performant OHTTP gateway at scale. Today, we’re launching the closed beta for our self-serve Cloudflare OHTTP Gateway. We’re also renaming our “Privacy Gateway” to “Cloudflare OHTTP Relay” to better distinguish the two products.

Now, customers who want an OHTTP architecture with the necessary separation of trust have two options:

  1. Use Cloudflare’s OHTTP Relay (formerly Cloudflare Privacy Gateway) and run your gateway yourself . This is best if your application servers are hosted off Cloudflare, and you’re able to run your own OHTTP gateway.
  2. Use Cloudflare’s new OHTTP Gateway with a third-party relay . This is best if your app servers are already behind Cloudflare (on our CDN or Workers, for example), if you’re accepting OHTTP requests from a third party (like Apple’s LiveCallerID ), or if you want a managed gateway to minimize latency and operational overhead.

We’re working to raise the bar for privacy across the Internet, and we believe that protocols like OHTTP can help — if we make them easy enough to adopt. It’s always been our goal to expand our OHTTP product suite and make our trusted privacy infrastructure accessible to a broader swath of the Internet.

Why we built the Cloudflare OHTTP Gateway

Since we launched our OHTTP Relay product, we’ve observed a few things.

First, we’ve seen that there's a growing appetite among developers for accessible, usable privacy infrastructure. Developers of privacy-oriented apps want to bake network privacy into their applications by default, but doing so remains harder than it should be.

Second, we’ve learned that building and operating an OHTTP gateway can be tough for customers. Any proxying architecture introduces some latency because requests must travel an extra hop or two around the Internet. Combine that with the cost to decrypt requests and encrypt responses, and the latency hit of a homegrown OHTTP setup can be significant. We’re well-positioned to solve this problem: the same building blocks that enable us to operate fast, reliable privacy infrastructure for products like 1.1.1.1 and iCloud Private Relay make us a good home for an OHTTP gateway. Because of Cloudflare’s anycast approach, our OHTTP Gateway will run on every server on Cloudflare’s global edge network, minimizing latency in relay-to-gateway hops. If you use our CDN, user requests can be decrypted by our Gateway and resolved by your app servers on the same Cloudflare metals, saving gateway-to-origin latency.

Finally, recall that OHTTP’s privacy model requires that the relay and app server be operated by separate, non-colluding parties. We want to provide our customers with the best possible range of options for their privacy infrastructure. Before, developers who protected their app servers behind Cloudflare weren’t able to use our OHTTP Relay, because Cloudflare would see both client metadata and the decrypted contents of requests, breaking OHTTP’s privacy model. Now, developers can choose whether a Cloudflare OHTTP Relay or Gateway is a better fit for their architecture.

A primer on OHTTP

A typical interaction between a client and application server reveals information about the client. When a client and app server talk to one another, the app server learns the client’s IP address because each packet in which data is sent is labeled with a source IP — similar to the “from” label on an envelope. App servers can also “fingerprint” a client based on attributes like supported TLS versions or cipher suites. These signals make it possible for app servers to link multiple requests back to the same user.

But what if I wanted to build an app that really doesn’t know much about my users? For example: Flo Health wanted to build an Anonymous Mode to enable users to access personal health data without it being linkable to possible user identifiers.

OHTTP introduces a proxy, called a “relay,” that forwards requests and responses between client and app server to obfuscate the client’s identity from the app server. The relay sees client identifiers like IP address and TLS fingerprint, but strips them before forwarding on requests. This prevents app servers from linking multiple requests back to the same user, and means that request contents can’t be associated with the user’s IP address.

For example, a regular client-server exchange might reveal the following information about a client:

- ipAddress: 192.0.2.33 # the client’s IP address 
- ASN: 7922
- tlsCipher: AEAD-CHACHA20-POLY1305-SHA256 # potentially unique
- tlsVersion: TLSv1.3
- Country: US
- Region: California # the client's location
- City: Campbell

A request first sent through an OHTTP relay would reveal only the relay’s information to the app server receiving the request:

- ipAddress: 128.62.37.13 # the relay's IP address & fingerprint 
- ASN: 18 
- tlsCipher: AEAD-AES-128-GCM-SHA256 
- tlsVersion: TLSv1.3 
- Country: US
- Region: Texas  # the relay's location
- City: Austin

This means that for each request, the app server doesn’t learn the location and TLS fingerprint of the end user. Plus, if many different users are sending requests through the relay, the app server won’t be able to distinguish which requests are coming from whom, limiting their ability to trace app activity back to a single end user. This creates a strong privacy boundary.

What really differentiates OHTTP from a basic forwarding proxy, however, is the encryption of data between client and app server. Requests and responses are encapsulated using Hybrid Public Key Encryption ( HPKE ) such that only the client and app server can see plaintext, and the relay sees only a jumble of ciphertext. A “gateway” sits between the relay and app server to handle all of this cryptography — decapsulating requests, encapsulating responses — and the app server handles only plain HTTP.

This creates a “double-blind” privacy model: the relay sees only client identifiers; the gateway and app server see only request contents; no party sees both.

A diagram showing how requests flow from end users through the OHTTP Gateway to app servers. A response from app servers follows the same path in reverse to the end user. Note that with the Gateway, you can choose whether or not to put your servers behind Cloudflare.

How we built the OHTTP Gateway

In building our OHTTP gateway-as-a-service, our goal is to bring our secure, performant privacy infrastructure to a broader swath of the Internet. Performance and easy onboarding are critical. So, we built our Gateway as a flexible service deployed across our global network. With just a couple of clicks, you can enable the Gateway on your zone and start sending OHTTP to https://your-zone.com/.well-known/ohttp-gateway . We’ll scale the service up and down automatically, so you don’t need to worry about capacity.

We had a few other user needs in mind, informed by the pain points we’d seen OHTTP Relay customers run into when operating their own OHTTP gateways.

First: We wanted to abstract away as much of the complexity of OHTTP as possible for your app servers. We wanted developers to be able to start receiving OHTTP while continuing to accept regular HTTP traffic if they chose. So, we designed the Gateway as a feature of your zone, where clients send well-formatted OHTTP requests to a /.well-known/ohttp-gateway endpoint on your zone. We support both standard and chunked OHTTP — and we recommend using chunked OHTTP for better performance, because it enables us to process requests incrementally (in “chunks”).

Our Gateway service will intercept each request, decrypt it, issue a subrequest to your app server, and return an encrypted response to the client. All non-OHTTP requests will travel to your server without invoking the Gateway.

Binding your Gateway to your zone also enables us to protect your Gateway from abuse. A client sending requests to your zone ` example.com ` may send to ` foo.example.com ` or ` bar.example.com `, but not wikipedia.com . Without you needing to worry about it, this prevents unauthorized clients from using your zone as a way to target other domains.

Second: Seamless key management is critical. Gateways need to maintain a public HPKE key configuration to enable clients to encrypt requests, but managing keys securely is a challenge. So, we designed the Gateway to fully manage all keys for customers, and to serve public keys as responses to GET requests to /.well-known/ohttp-gateway . For stronger privacy, clients can download keys over a different IP than they request the gateway.

Third: Gateways need to be able to authenticate relays. Because the Gateway (by design) knows very little about the client sending a given request, it places trust in the relay to authenticate clients and forward traffic responsibly. But how do you ensure that only trusted relays can send traffic to your gateway?

We designed the Gateway such that Cloudflare Access , Cloudflare’s zero trust network access product, runs before requests are decrypted, enabling you to use any standard Access policies to authenticate incoming traffic and protect your Gateway from abuse. Options include mutual TLS, static service credentials, and custom external logic.

Finally: Mistakes happen, and we anticipated that customers might accidentally break OHTTP’s privacy model by running both their relay and gateway on Cloudflare. So, to preserve OHTTP’s separation of trust and ensure that Cloudflare never sees both client identities and decrypted inner requests, our Gateway will refuse to decrypt requests sent from Cloudflare Workers or from proxied hosts on Cloudflare.

When is the OHTTP Gateway a better fit than the OHTTP Relay?

If you want to use Cloudflare’s OHTTP product suite, but you’re wondering why you’d pick Cloudflare’s OHTTP Gateway instead of the OHTTP Relay, here are a couple of considerations.

First, do you want your app servers on Cloudflare – behind our CDN or built on Workers, for example? If so, the OHTTP Gateway is a better fit to ensure adherence to OHTTP’s privacy model.

Second, what’s your use case? If you want to receive OHTTP requests from a third-party client and relay — to use Apple’s LiveCallerID SDK, for example — then the OHTTP Gateway is likely the better solution for you.

Getting started

If you have a feature request or would like to register for our waitlist, so we can notify you when the product launches, sign up here .

Then, you’ll need to implement an OHTTP client. See ohttp.info or our sample client library for some examples to help you get started. One flag as you build the client: OHTTP provides privacy at the network level, and doesn’t touch the inner request body. So, to preserve user privacy, it’s up to you not to send identifying information (e.g. a user’s email address or username) in the request body.

Next, you’ll need to bring your own relay. Relays can run on any infrastructure provider, and they’re simple: here’s some sample code . The challenge and the reason you might want a dedicated OHTTP relay provider, is to verifiably promise to your users that you won’t inspect logs with client identifiers. Otherwise, you’d be able to correlate clients at the relay with decrypted requests at your app servers.

Finally, once your OHTTP deployment is live, check out our

pvcli client

to help with testing and debugging.

We’re excited to bring accessible privacy infrastructure to developers everywhere.

Reach out to us

if you’d like to try out the new OHTTP Gateway and raise the bar for privacy online.

gVisor is being donated to CNCF

Lobsters
gvisor.dev
2026-10-02 22:41:38
Comments...
Original Article
gVisor being donated to the Cloud Native Computing Foundation, a subsidiary of the Linux Foundation.

In 2018, Google open-sourced gVisor under the Apache 2.0 license. To the best of its contributors’ knowledge, it has ever since remained the second most mature implementation of Linux, after Linux.

This year, the gVisor project is being donated to CNCF , and its governance model is shifting accordingly.

What’s happening?

Google is donating the gVisor project, including its name and trademarks, to the Cloud Native Computing Foundation (CNCF) , a subsidiary of the Linux Foundation . Like its name implies, CNCF is focused on cloud-native computing, with Google having seeded its creation by donating the Kubernetes project. Since then, Kubernetes has grown to become the industry-standard for container orchestration, and has grown a large and vibrant ecosystem around it. gVisor is now following the same footsteps.

You can see gVisor’s CNCF application and process .

What’s the timeline?

What has already happened :

  • 2026-09-07: Google submitted its CNCF donation application .
  • 2026-09-22: The CNCF reviewed the application.
  • 2026-09-28: The application was accepted.
  • 2026-10-02: This blog post was published.

Over the next few weeks :

  • The project will move to CNCF “ Sandbox ” status (quite appropriately-named for a project like gVisor).
  • gVisor’s build and testing infrastructure will move to GitHub Actions and Buildkite
  • Google’s internal gVisor test infrastructure will no longer block PRs.
  • The gVisor project’s governance model will transition to a maintainers-based model .
  • Non-Google maintainers will be added and given merge permissions.

Over the next few months :

  • The project will take the steps needed to move to CNCF “ Incubation ” status.
  • The GitHub repository will move out of the google GitHub organization.
  • The gVisor project’s governance model will transition to a long-term model that features org-based voting , thereby preventing Google from having unilateral control over governance decisions.
  • Any further steps to become a fully-fledged CNCF project will proceed.

Why donate?

gVisor doesn’t fit neatly into the industry’s well-known boxes of the sandboxing/security landscape, which tends to separate “vanilla containers” from “virtual machines” with shades of gray in between. gVisor straddles this middle-ground, providing empirically-equivalent security but without checking the familiar “virtualization” checkbox that security auditors, regulators, or security practitioners often treat as a one-to-one proxy for “secure”. This has caused adoption challenges over gVisor’s history, as it has been difficult to communicate the value of the project to an audience that is used to this false dichotomy .

Another challenge gVisor has faced is that of a performance perception problem . Internally within Google (and other gVisor-using companies, such as Ant Group and Modal), there exist Linux kernel patches that improve gVisor performance significantly. However, for other gVisor users, out-of-the-box performance often shows performance degradation for certain I/O-intensive workloads. This has led to poor first-impressions from potential adopters. We have tried to address this by upstreaming Linux kernel patches that improve its performance, but have been turned down by kernel maintainers due to gVisor being a wholly-owned Google project.

Lastly, gVisor as a project has potential that is difficult to prioritize when guided by corporate ownership alone. As a userspace implementation of Linux, gVisor has potential non-commercial applications such as:

  • gVisor-on-Mac : Allowing Linux programs to run on macOS , with a similar experience as to how how Wine allows Windows programs on macOS.
  • Desktop Linux sandboxing : Allowing gVisor to be used as a practical option for desktop Linux application sandboxing that is much more secure than the current state of the art (bubblewrap/flatpak/nsjail/etc), yet much easier to integrate with than full-blown virtualization-based approaches like that of Qubes OS . We have made some advancements on this front with our recently-introduced bwrap drop-in replacement , but gVisor is capable of sandboxing so much more .

We see evidence of these problems by looking at the current set of gVisor adopters , which are all either large tech companies with the ability to invest and customize gVisor to suit their own needs (Google, Ant Group, OpenAI, Anthropic), and startups with a highly-specific focus that exactly fits gVisor’s use-case, and where it makes sense to spend a startup’s limited resources specifically into making gVisor work great for them (Modal, Tines). Who is not on this list?

  • Hobbyist projects : See aforementioned non-commercial applications where gVisor would be useful but isn’t currently adopted.
  • The “middle” of the industry : Individuals and companies that would benefit from gVisor’s security, but either aren’t aware of its existence, dismiss it out of past perception problems, or don’t have the resources to invest specifically into security but would happily adopt an off-the-shelf, widely-available sandboxing runtime were it to exist as a widespread and cheap option already available as an offering by their computing infrastructure provider.
  • Other non-Google hyperscalers : While gVisor is adopted internally by nearly all large tech companies for their own at-scale sandboxing needs (e.g. code snippet execution, RL), only a small subset (Google, DigitalOcean, and Modal) directly sell general-purpose gVisor-powered compute to their customers. This is in spite of gVisor’s competitive operational margins, as well as the demonstrable demand for this, because issue reports to the gVisor repository show that a large number of entities are self-installing gVisor on non-Google hyperscalers. The remaining explanation of the lack of direct integration is likely the project’s (pre-donation) governance risk.

By contributing the project to the CNCF, we aim to address all of these issues. This enables gVisor and application kernels to become part of the container ecosystem and security industry’s lingua franca, enabling integration and adoption beyond highly-motivated/sophisticated/resourceful corporate entities, and enables upstreaming Linux patches that solve gVisor’s performance for everyone.

Why donate to CNCF specifically?

gVisor is a drop-in-compatible container runtime that fits in the Cloud Native container ecosystem. It integrates directly with CNCF technologies such as Kubernetes, containerd , and Agent Substrate .

From a resource and efficiency standpoint, gVisor acts as a more cloud-native container runtime than other security-focused container runtimes, thanks to its container-like process model. This allows it to be efficiently and tightly-sized, enabling secure container binpacking at a resolution and density VM-based runtimes cannot match. It also does not require hardware virtualization or nested virtualization, enabling it to run anywhere Linux runs. That makes gVisor a good complement to CNCF’s existing portfolio. gVisor and is already usable on all major clouds, some of which offer it as a native offering, and others for which cloud users can (and do) self-install it.

Will Google divest from gVisor development in the future?

The past few months have been pivotal for the security industry and for secure sandboxing specifically . It would make very little sense for Google to abandon its sandboxing technology at a time when the need for cheap and secure sandboxing has never been clearer. On the contrary, we (the gVisor contributors at Google) have been pushing for this move internally with the expectation that this will accelerate gVisor’s growth and adoption both within and outside of Google, in a similar manner as what has happened with Kubernetes.

Who is joining gVisor’s contributors beyond Google?

Google is reaching out to potentially-interested parties. As of this writing, the following entities have committed to joining gVisor’s maintainers for the long haul: Ant Group , Modal , and Tines . Additionally, other companies including OpenAI , Tencent , and NVIDIA will continuing their ongoing contributions to gVisor.

What does this mean for me?

  • In the short term: Not much.
  • In the medium term: A less painful PR contribution experience.
  • In the long run: A more free gVisor, with development accelerated by an influx of new contributors, and directed by a more open governance process in service of its users.

What’s next?

Some gVisor contributors will be at KubeCon North America 2026 . Come chat!

Friday Nite Videos | October 2, 2026

Portside
portside.org
2026-10-02 22:18:32
Friday Nite Videos | October 2, 2026 barry Fri, 10/02/2026 - 22:18 ...
Original Article

Friday Nite Videos | October 2, 2026

Bob Dylan | A Hard Rain's A Gonna Fall. The Corruption at Starlink Is Shocking. "60 SECONDS" | AOC Talks Democratic Socialism. Anthropic IPO Documents AI Risks. Pete Hegseth Purges Weirdos and Wimps.

Portside Portside

AI Weapons Systems Are Already Here

Portside
portside.org
2026-10-02 22:13:19
AI Weapons Systems Are Already Here barry Fri, 10/02/2026 - 22:13 ...
Original Article

As tech leaders warn about the threats of artificial intelligence, most conjure up images of AI agents running amok , such as hacking systems used to run our critical infrastructure, banks or even militaries. Yet one serious danger is mentioned less frequently – the threat posed by AI–empowered lethal weapons. Fully autonomous versions of these weapons are colloquially known as killer robots .

Israel’s conduct in Gaza shows that, in an important respect, such AI weapon systems are already here. An algorithm has been deciding who lives and who dies .

For more than a decade, governments have been meeting in Geneva to stop the development of autonomous AI weapons. In early September, 128 governments met to approve a report that could serve as the foundation for a binding treaty.

But a collection of lawyers sent by Donald Trump and Vladimir Putin teamed up to water it down . Notably, they weakened a provision that requires humans to review military targets developed by AI before a strike.

Fully autonomous weapons remain the focus; the use of AI to determine targets for conventional weapons is not. At risk is the longstanding quest to maintain a “ human in the loop ” – that is, meaningful or effective “ human control ” of AI-empowered weaponry.

For many people, talk of killer robots conjures up images of Arnold Schwarzenegger’s The Terminator . I have often thought about the problem in the context of the Tahrir Square uprising at the height of the Arab spring that led to the ouster of the Egyptian president Hosni Mubarak.

Mubarak’s fate was sealed when the Egyptian army made clear it would refuse to fire on Egyptian protesters. But killer robots would have had no such qualms. If Mubarak had been able to deploy them, he might have stayed in power.

The negotiators in Geneva are not even addressing this problem. The Trump-Putin lawyers made sure that the agreed report referred only to war, not repression.

Today, amid the drone warfare in eastern Ukraine, Russian and Ukrainian troops are experimenting with fully autonomous weapons to overcome jamming of the signals used to control drones. We can expect these nascent deployments to expand.

Yet the “killer robot” metaphor is deceptive. If we look at targeting decisions more expansively, AI has already taken over. In Gaza, the Israeli military is using AI to target Palestinians.

Israel has long monitored phone communications in that territory. After the Hamas attack of 7 October 2023, Israel used an AI system called “ Lavender ” to uncover patterns that it equated with the phone of a Hamas operative. People whose phones AI determined met these criteria could be killed.

Because it was easier to target suspects at home, a separate automated system called “ Where’s Daddy ?” tracked them there, meaning that their families were often killed too. There was no pretense that Hamas was using its operatives’ homes for military purposes. It was simply easier for the AI system to find suspected Hamas fighters where they slept. The families were the collateral damage.

Theoretically, a human remained in the loop because an AI-picked target was routed through intelligence officers who approved it and sent it to others who launched the attack. In fact, this human role was nominal.

In the past, Israeli targeting decisions were subject to significant review. But in Israel’s killing spree that followed the Hamas attack, the intense demand for targets took precedence over such review.

Israeli intelligence officers who received the AI selections routinely passed them to the attackers without second guessing. As one officer described : “I would invest 20 seconds for each target at this stage, and do dozens of them every day. I had zero added-value as a human, apart from being a stamp of approval. It saved a lot of time.”

From computer to killing, the human judgment exercised was perfunctory – a “ rubber stamp ”. In essence, an algorithm played God.

One Israeli officer suggested the practice might be justified because AI makes better decisions than humans do. It is dispassionate, while people’s judgment can be clouded, such as by the animosity toward all Palestinians that some Israelis felt after the Hamas attack.

As one Israeli soldier explained : “I have much more trust in a statistical mechanism than a soldier who lost a friend two days ago. Everyone there, including me, lost people on October 7. The machine did it coldly. And that made it easier.”

Yet AI algorithms are inevitably affected by human input. How much evidence is required to identify someone as a Hamas fighter? What margin of error is tolerated?

The answer to such questions depends on the judgment of the AI programmers and their commanders – ordinary Israelis who might have devalued Palestinian civilian life in the fury that followed the Hamas attack. Even the Israeli military conceded that about 10% of the human targets slated for assassination were not members of the Hamas military wing at all.

The Israeli president, Isaac Herzog, epitomized the heedless attitude toward Palestinian civilians when he said : “It is an entire nation out there that is responsible. It’s not true, this rhetoric about civilians were not aware, not involved – it’s absolutely not true. They could’ve risen up, fought against that evil regime.”

That same contempt contributed to the Israeli military’s shocking willingness to tolerate civilian deaths. The Israeli military permitted an anticipated 20 dead for the killing of a mere suspected low-level Hamas fighter.

Under the rule of proportionality, international humanitarian law prohibits a wartime attack if it “may be expected to cause incidental loss of civilian life … which would be excessive in relation to the concrete and direct military advantage anticipated”. Although international law doesn’t specify the number of civilian casualties that are acceptable in a given situation, virtually no legal expert considers Israel’s shocking 20-to-one ratio to be lawful.

These are war crimes . They reflect an appalling devaluing of Palestinian civilian life.

Evoking for me Hannah Arendt’s “ banality of evil ”, ordinary soldiers became complicit in extraordinary cruelty by trusting what they call the “ system ”. The essence of that system remained people – the generals who ordered its creation, the military lawyers who excused the unjustifiable, the programmers who tweaked the algorithm, the intelligence officers who passed on the computer-generated targets, and the soldiers who launched the attack.

But part of what made it easier to trust the system was the pseudo-scientific precision of AI. We are all familiar with the breezy, authoritative tone of today’s chatbots. When AI is used for targeting, it brings a similar feigned omniscience to killing.

When inevitably flawed human beings make decisions, they can expect others to challenge them. But AI makes it easier for us to defer to its judgments. That deference risks magnifying the human biases that inevitably infect any AI system.

As we seek to curb killer robots, we should also remain attentive to AI that can lull us into complacency about other targeting. The International Committee of the Red Cross calls it “automation bias”.

To have meaningful human control of AI, we need not just nominal endorsement of its decisions. We need intense questioning of its product, and accountability for commanders when junior officers fall short.

Only that skepticism has any hope of ensuring that soldiers who deploy AI remember that real human beings are at the end of the kill chain they enable. Even in war, people should never be snuffed out by the mechanical application of an algorithm.

Kenneth Roth is a Guardian US columnist, a senior fellow at Yale University, and a former executive director of Human Rights Watch. He is the author of

The Guardian is globally renowned for its coverage of politics, the environment, science, social justice, sport and culture. Scroll less and understand more about the subjects you care about with the Guardian's brilliant email newsletters , free to your inbox.

Show HN: Google Maps Scraper MCP

Hacker News
gmapscrawl.com
2026-10-02 21:46:39
Comments...
Original Article

Developers / MCP server

Google Maps Scraper MCP Server

Connect your AI agent to synchronous business search and structured listing data through one secure MCP connection.

Built for your agent

From a search to a useful dataset.

Give your agent the same business data workflow you use in your workspace.

Search businesses

Find businesses by keyword and location, with listing details your agent can organize and analyze.

Retrieve listing data

Read available addresses, contact details, ratings, and review counts in structured result pages.

Use results directly

Receive up to 20 businesses in structured JSON for research and applications.

How to connect

One endpoint, one auth header

Configure your client to use /api/mcp , then pass your API key in the authorization header.

POST https://gmapscrawl.com/api/mcp
Authorization: Bearer <YOUR_API_KEY>
Accept: application/json, text/event-stream

Streamable HTTP — your client handles protocol negotiation.

Bearer API key authentication — configure secrets outside prompts.

Connect your tools

Choose your MCP client

Keep your API key in a secure environment variable or your client’s protected input setting.

Add this server to .mcp.json. Make GMSCRAPER_API_KEY available in the environment that launches Claude Code.

{
  "mcpServers": {
    "google-maps-data": {
      "type": "http",
      "url": "https://gmapscrawl.com/api/mcp",
      "headers": {
        "Authorization": "Bearer ${GMSCRAPER_API_KEY}"
      }
    }
  }
}

Need step-by-step guidance? Read the MCP setup guide

Built for your workflow

Built for data-aware AI workflows

Move from a question to structured local business research without switching tools.

Prospecting research

Build a focused list of local businesses for sales research and outreach planning.

Local market analysis

Compare categories, ratings, and review counts across the businesses in your dataset.

Listing enrichment

Bring available business and contact fields into CRM, reporting, or internal workflows.

Put it to work

Example prompts to try

Try a focused request in your connected client. With a test key, ask your agent to label the results as simulated.

Prompt

“ Find one page of coffee shops in Seattle and return their website, phone number, rating, and review count. ”

Prompt

“ Search for restaurants in Portland and compare the returned business categories and ratings. Point out missing data. ”

Prompt

“ Search one page of plumbers in Austin and format the returned businesses as a table with names, phone numbers, and websites. ”

Frequently asked questions

What is the Google Maps Scraper MCP server?

It is a remote Model Context Protocol server that lets compatible AI agents search for up to 20 businesses and receive results directly using your API-workspace credentials.

Which MCP transport does it use?

The server uses stateless Streamable HTTP at https://gmapscrawl.com/api/mcp. /mcp remains a compatibility alias. Your client initializes the connection and negotiates the protocol version.

How do I authenticate the MCP server?

Configure Authorization: Bearer with your API key in the client’s secret configuration. Browser login cookies do not authenticate this endpoint, and keys should never go in a prompt or URL.

Can I connect Claude Code, Cursor, VS Code, or Codex?

Yes, with a version that supports remote HTTP MCP and custom authorization headers. Use the client configurations above and ensure the configured environment variable is available to that client.

What Google Maps data can my agent retrieve?

Business searches return available listing attributes, contact details, ratings, and review counts. The server exposes only search_google_maps.

Do I need an API plan?

Live work requires an API entitlement and an available cloud service. MCP shares API usage; an Online subscription does not include it. Use a test credential to validate the integration with simulated fixtures.

Take the next step

Give your AI agent Google Maps data.

Create an API key, connect the MCP server, and bring business research into your preferred client.

Get API key

The Road to Universal Medicare

Portside
portside.org
2026-10-02 21:43:37
The Road to Universal Medicare barry Fri, 10/02/2026 - 21:43 ...
Original Article

Universal health coverage under Medicare is a natural for Democrats. According to a new report by the American Economic Liberties Project, between 2005 and 2025 the annual cost of employer-sponsored family coverage increased from $12,214 to $35,119. Deductibles and co-pays keep rising as well.

Illustration by Jordan Awan.

President Trump’s assaults have added to the system’s miseries. Trump’s budget terminated Biden-era tax-credit subsidies under the Affordable Care Act as of the end of 2025. The average recipient saw premium costs more than double in 2026. At least 2.6 million people out of the 22 million insured under the ACA dropped coverage entirely.

The effective dates of Trump’s cuts in federal Medicaid outlays were delayed until after the 2026 election. But next year, some 15 million people are at risk of losing Medicaid coverage. All told, at least five million people already lost health insurance in the first six months of 2026.

Merely having Medicare for All without reforming the way health care is delivered would just stick government with higher costs. But universal Medicare, by giving government a great deal more leverage over the entire system, is key to addressing its other deficiencies.

More from Robert Kuttner

Even if Democrats can agree on universal coverage, the transitional challenges seem daunting. Medicare is tax-supported. Commercial insurance is financed by premiums. An immediate shift to universal Medicare would require tax increases, but only some people would get savings from cuts in premiums. Most of those savings would go to employers.

The best approach to solve this transitional problem has been proposed by our colleague Jacob Hacker, the Yale political scientist who serves on the Prospect board. Hacker first outlined a version of it in the Prospect in 2018 .

Hacker calls it Medicare Part E, the E standing for Everyone. Hacker begins with what used to be called “play-or-pay.” Employers would be required either to provide good health insurance or to pay for Medicare for their employees.


All told, at least five million people already lost health insurance in the first six months of 2026.

This would have three huge benefits compared to other transitional approaches. First, it would eliminate the need for a tax increase. Some employers would keep their commercial insurance; most would pay into Medicare on behalf of employees. Either way, no new taxes. Over time, the superior efficiency of Medicare would lead more employers to shift to Medicare. Only truly efficient commercial insurers (if there are any) would survive.

Second, unlike other incremental approaches to universal coverage, this one is easy to explain and campaign on. Everybody gets insurance, either through employer-financed commercial insurance that meets Medicare standards, or through Medicare itself. And no tax hike.

Third, this approach bridges over what Hacker calls the label wars. Candidates who like the Medicare for All label can embrace it. So can candidates who want universal coverage but not immediate Medicare for All.

It’s important to get this right. The problem of large tax increases in a one-time shift killed an attempt to have single-payer insurance in one state. In Vermont, the legislature enacted a single-payer bill in 2011 called Green Mountain Care, with the strong support of Gov. Peter Shumlin, but left financing details for later. After the details were unveiled—for an 11.5 percent payroll tax on employers plus an income-based premium assessment of up to 9.5 percent—public support collapsed. Shumlin, favored heavily for re-election in 2014, barely won. In late 2014, the state abandoned the plan entirely.

Paul Krugman has proposed a “Medicare Buy-In.” Krugman would allow individuals as well as employers to buy in to Medicare. But this adds complexity. It’s better and simpler to promote the transition to Medicare for All with an employer mandate and have employers bear the costs.

Our colleague, Paul Starr, has proposed an optional Medicare buy-in for people aged 50 to 64, who typically pay very high rates for their insurance. This is good, but Hacker’s universal Medicare Part E is simpler and better.

Pete Buttigieg proposes what he calls Medicare for All Who Want It. That’s catchy, and it implies greater freedom of choice. But of course, Medicare itself is the ultimate freedom-of-choice plan, since unlike commercial insurance plans, it allows people to use any doctor or hospital. Shame on Democrats who fail to emphasize that.

Others, including House Democratic Leader Hakeem Jeffries and a bipartisan centrist group led by Reps. Brian Fitzpatrick (R-PA) and Tom Suozzi (D-NY), propose making permanent the Biden subsidies to the Affordable Care Act that Trump canceled. But this approach would leave all of the inefficiencies of the current system intact while having the government bankroll more of the commercial insurance industry.

Readers may recall that it was Hacker who came up with the original public option in the context of the debate about the Affordable Care Act. People who got insurance under the ACA could opt for a commercial plan or for a public plan modeled on Medicare.

Hacker correctly assumed that the more cost-effective option of public insurance would gradually crowd out commercial insurers. Unfortunately, the private insurance industry read Hacker’s proposal. Even though they could not kill the ACA legislation, they mounted an all-out and successful campaign to kill the public option. Hacker II is far superior to Hacker I because it shifts the public option to employers, coupled with a requirement to provide Medicare-quality insurance, and thus creates incentives for them to shift to Medicare.

SINCE THE SPONSORS OF THE AFFORDABLE CARE ACT had to jettison the public option in order to get the ACA through Congress (and barely got it enacted at all), why should we think that the more transformative Medicare Part E plan could get enacted now? The answer is that the entire health care system is far more enshittified now than it was in 2010, when Democrats had to settle for a reform that only increased coverage by about 7 percent of the population and achieved that by subsidizing rather than supplanting private insurance, thus reinforcing the present hyper-commercialized system.

Today, there is widespread rage against the commercial insurance industry and the drug industry. When Luigi Mangione stalked and murdered the CEO of UnitedHealthcare, he had plenty of sympathizers. One poll found that 41 percent of American voters under 30 perversely thought the killing was acceptable.

Even Donald Trump, recognizing the extreme unpopularity of Big Pharma, has politically abandoned the big drug companies, claiming credit for caps on drug prices (that were actually the work of Joe Biden) and adding his own proprietary drug discount program. As bogus as these Trump maneuvers are, the man knows which way the wind is blowing.

If Democrats win the presidency and a working majority in Congress in 2028, the Hacker version of Medicare for All is definitely a political possibility. In the meantime, it is a superb program to run on. Medicare is astonishingly popular, and for good reason. The popularity of Medicare is the flip side of the extreme unpopularity of the insurance and drug industries and the frustration people experience trying to get care.

Medicare for All is a far better brand for Democrats than building on the Affordable Care Act. There is an important ideological subtext: Public is not only more socially just than private, it’s also more efficient. Mixed public-private deals like the ACA are often a muddle—less efficient, more easily gamed, and more complex to use. As noted, to the extent that the current system is a cesspool of profiteering based on market concentration, a comprehensive public system like Medicare gives the government far more leverage to reduce the gaming.

Public is not only more socially just than private, it’s also more efficient.

Some key details do need to be ironed out. The so-called Medicare Advantage program piggybacks on Medicare funding and the Medicare brand. But Medicare Advantage policies are commercial insurance products that make their money by targeting their marketing to relatively healthy seniors, and then trying to restrict actual care when people get sick. With Medicare for All, we could either kill Medicare Advantage outright, or end the more than $80 billion annual subsidy that these plans get from Medicare, which would make Medicare Advantage plans unprofitable.

Medicaid spends just under a trillion dollars a year. With Medicare for All, there would no longer be any need for a separate means-tested program, so Medicaid would be folded into Medicare. The largest single Medicaid cost is long-term (nursing home) care. That and home care would also be folded into Medicare.

The Medicare Part D program is a Medicare-branded private insurance product for prescription costs. It’s an inefficient crazy quilt. So drug benefits would also become part of Medicare. That shift would eliminate parasitic pharmacy benefit manager scams.

Finally, the fact that Medicare does not cover all necessary medical needs has given rise to private “Medigap” plans. Less-affluent people who can’t afford the premiums either pay out of pocket or go without care that they can’t afford. Universal Medicare for All needs to cover everything that Medigap policies cover.

But wouldn’t all this be astronomically expensive? Quite the opposite. Yale economist Zack Cooper compared Medicare reimbursements with reimbursements paid by commercial insurers. Because of the market power over insurers created by hospital mergers and consolidations, commercial insurers pay vastly more. Hospitals get an average of $29,000 for a hip replacement covered by commercial insurance but only $16,000 for one covered by Medicare . Those costs are passed along to the public and to the taxpayer, since over half the costs in the system are paid directly or indirectly by the government.

And the bias in favor of procedures and technologies that command maximum reimbursements has led to overspending on specialty care and a severe shortage of primary care practitioners, as I wrote in our June 2026 issue . A fragmented system can’t solve this imbalance piecemeal. Medicare for All could.

The new report from the American Economic Liberties Project details the costs and inefficiencies produced by the hyperconcentration of the commercial health care industry. Today, according to the report, “Six Big Medicine companies now rank among the Fortune 15—more than Big Tech or any other sector. In 2025, they pocketed nearly $34 billion in profit.” But their profits are only a small part of how they contribute to the sheer inefficiency and misallocation of resources in the system.

The task of the next administration is not just supplanting commercial insurers with Medicare to provide universal and affordable coverage, but also getting rid of the extreme concentration. Public dismay with health care has reached a point where Republicans as well as Democrats are demanding fundamental reform. For example, the Break Up Big Medicine Act is co-sponsored by Sens. Elizabeth Warren (D-MA) and Josh Hawley (R-MO). It would prohibit insurers, pharmacy benefit managers, and wholesalers from owning or controlling providers.

Supporters of Medicare for All need to appreciate that reform is not just a matter of getting to single-payer. It also entails using the power of that single-payer to decommercialize and simplify our badly corrupted health care system.

Used with permission. The American Prospect, Prospect.org, 2024. All rights reserved. Click here to read the original article at Prospect.org.

Click here to support The American Prospect's brand of independent impact journalism.

Pledge to support fearlessly independent journalism by joining the Prospect as a member today.

Every level includes an opt-in to receive our print magazine by mail, or a renewal of your current print subscription.

The reason we can write these stories—stories that are orienting and informing rather than fear-inducing—is because of you. We removed program matic ads from our site in April , which means that we don’t chase clicks for ad dollars. We don’t have a paywall, which means that everyone, regardless of their ability to pay, can access our work for free. We don’t have content partnerships with AI giants .

Abdul El-Sayed Explains How Science Influences His Politics

Portside
portside.org
2026-10-02 21:37:39
Abdul El-Sayed Explains How Science Influences His Politics barry Fri, 10/02/2026 - 21:37 ...
Original Article

The 2026 U.S. midterm elections are just weeks away. Among the highest-profile races is the contest for Michigan’s open Senate seat. With Senator Gary Peters not seeking reelection, the Democratic candidate is Abdul El-Sayed, who holds degrees in both medicine and public health and previously served as the head of Detroit’s public health department.

We asked El-Sayed how his scientific background affects his policy decisions, what the “Make America Healthy Again” (MAHA) movement has gotten right and wrong and what his vision for Medicare for All is.

[ An edited transcript of the interview follows. ]

How has your background in science influenced your thinking on policy in general?

I’m a systems scientist. That’s how I trained. I wrote my doctoral thesis on systems in epidemiologic research. Every good scientific inquiry starts with a great question, and I think so much of what’s wrong with public policy starts with a great question: Why don’t we have health care when we need it? Why is there such massive inequity? What is the best way to allocate our scarce resources? How do we take on the challenges of a world in which there’s too much violence and too little peace?

The thing about science is that you construct a hypothesis of the world and you test it. I think that that level of rigor—to ask, “What is the nature of this problem? How do I integrate data about the problem? How do I answer this? And then how do I work this in terms of the political process and the policymaking process?”—all of that creates living labs where we try to build the world that we all want to live in, a place that’s more peaceful, where more people have the things that they need and deserve, where we can live together in harmony. Once you train in the sciences, you don’t leave that scientific approach to thinking and that rigor of thinking. And I try to bring that to policy problems.

A lot of what science does is descriptive rather than prescriptive, so it can be a little bit of a jump from one to the other. So can you elaborate on how you take that kind of “rigor of thinking,” as you said, and then apply it to solutions that the data suggest will work?

Well, you know, it’s true [science] is descriptive. But every question that we ask in science is to design an intervention, right? We don’t just ask these questions because they’re fun to ask.... They are fun to ask and answer, but we ask them because we do want to make things better. Why do we want to understand the way that the body breaks down glucose? Why do we want to understand what the appendix measure doesn’t do? Because we want to do something about it, right? We want to design anything that addresses the problem. We want to understand what the implications of removing an appendix might be.

I think it’s important to remember that, in science, it’s not that we’re just descriptive; it’s that we are descriptive so that we might be prescriptive in an evidence-driven and thoughtful fashion. I think it’s the same approach in public policy, where you want to understand the system, but you want to understand the system so as to intervene, and then it’s iterative; it’s recursive. You will engage your current understanding of the problem, you’ll pose an intervention, and then you’ll study the outcome of the intervention to see if it works.

As we’re building and designing interventions, it is about being willing to say, “Okay, this worked really well. Here’s the next thing we can do to make it even better.” And it’s that ability to continue after the problem until we’ve gotten solutions that really are purpose fit. Understanding that the world also evolves and changes, that the nature of the problem today might change as a function of what you did but also of other forces that are shaping the problem.

Your expertise is largely in public health. You’ve spoken a lot about the structural and social determinants of health, but it can take a long time for systemic change to occur. The effects may only be felt years down the line. How does pushing for those changes help people right now?

That’s the thing: you have to make sure it does help people in the place and the time in which the pain is being felt the hardest. I think that’s the thing about it..., that, too often, people can’t wait for the system to change. You have to think in the system, but you also have to be thinking about the issue. And I think the best analogy here is in medicine.

It’s one thing to say we’re going to cure your cancer, but we also have to heal your pain right now because that pain is the most acute symptom of what the long term might be doing to you. It’s thinking in time and place, both about addressing the system and also addressing the most acute manifestations of a broken system, and you’ve got to do both at the same time.

Can you elaborate on what systems, outside of direct health care, need to change to let Americans have healthier lives?

Think about our energy systems. We could be providing high-quality, renewable energy that does not force us to breathe from the ends of a smokestack, that is both more affordable and more accessible. But we don’t do that because of the power of a few corporations to dominate that system and to continue to sell us stuff that is less efficient and less effective.

Think about the system of inequality in this country. How do you make sure that a few corporations can’t just rig the system to continue to put it in overdrive so that they make more and extract more from us? You think about the impact of tech, in the way that we build tech policy. Amazing science goes into building really great tech. But at the same time, what is driving that? What are the incentives driving that? Are we building artificial intelligence (AI) to take away our jobs and win a race that might end up destroying us all, or are we building AI as a tool that empowers us in our lives? And how do you get the incentives of that right?

Across the board, there is a structural process that you have to think through but also key consequences that you have to adjust around, even in the short term.

There’s some overlap there between health policies you’re proposing and those put forward by the MAHA movement . What do you make of MAHA?

I think it is a consequence of a few people exploiting the distrust of a system that has earned that distrust. The reason that people distrust science isn’t because they don’t trust science. We use science every day. It’s because they don’t trust the corporations that hold that science. You’re paying more and more for insulin that’s existed for 100 years. It’s not a far leap to go from “I don’t trust these corporations because I can’t afford the things that they are making” to “I don’t even trust the things that they’re making,” and I think MAHA has been a response to that.

I think what people really want is: they want to be healthy. And I think science has always been the best arbiter about how to make people healthy. The problem is: we’ve got to make sure that the scientific discoveries that we put forward are also affordable to people so they can actually have them.

How do you draw the line between skepticism, which is part of the scientific process, and rejecting evidence, which is something we’ve seen MAHA spokespeople and figures do?

It’s faulty and dangerous to reject evidence, and you’ve got to push back on that. How do we build the kind of health policy that takes on the evidence-based ways that corporations are making us sicker, whether it’s pollutants in our environment or additives that should not be added to our foods, or to the ultraprocessed food environment?

Meanwhile vaccines have done some really, really great things, and there’s no evidence to suggest that, on net, they’re causing more harm than good. And so I think you’ve just got to hear the scientific process. The problem, though, is that, too often, we don’t act on a lot of the problems that we do have evidence for, and it creates a level of skepticism. I’m trying to just lead us where the evidence goes.

One thing your health policy agenda doesn’t touch on is glyphosate , which is a major issue for many MAHA followers. Where do you stand on that?

I’ll be honest: I didn’t take a very direct position on that because I need to understand the science a little bit better. I want to be led by science, and I also understand the deep concerns that folks have.

I think we need to err on the side of caution. But we also need to make sure that science drives our positions. And we need to be able to demonstrate the potential impacts over time.

But part of the problem is just the way that we set scientific data standards. In the U.S., too many of our standards are “innocent until proven guilty,” and I think that when it comes to broad health concerns, you really need to demonstrate that something is safe before you start putting it in all of our foodstuffs. I think that is part of the challenge that the broader MAHA movement has been concerned about. I know there’s a lot of controversy around these particular sets of chemicals, but I want to be able to really review the science, understand where we sit and then apply that principle. We need to be able to demonstrate that this is definitely safe rather than assuming it’s safe until we’re proven wrong, like all these PFAS chemicals [perfluoroalkyl and polyfluoroalkyl substances, or “forever chemicals”] that only later science has demonstrated are so dangerous.

A lot of your positions require well-staffed enforcement agencies, which have gone through large-scale cuts. What can be done to beef them up, given that it’s clear that’s not where the administration’s priorities lie?

We’re going to need to reinvest in these agencies—and not just reinvest but modernize and make sure that we’re not just building back. We need to build forward the agencies that we need for the 21st century. They’ve taken a gigantic scalpel—no, forget a scalpel—[they’ve] taken a chainsaw to these agencies, and we’re going to need to be invested appropriately so that we can we can start to enforce where we need to. Whether it’s the [Food and Drug Administration or the Environmental Protection Agency], reinvesting in [the Centers for Disease Control and Prevention or the National Institutes of Health], we are going to have to do that.

We’re going to need people who both understand and have been to these agencies before they ever got elected and also can be sitting with experts to say, “If we could build what we need, not just what we’ve had, what would that look like?”

What can be accomplished may hinge on the outcome of the presidential election in 2028. If you’re elected in the midterms, is there anything that you plan to do from your position in the Senate over the next two years?

I think it’s worth reviewing a lot of the way that [the U.S. DOGE Service] happened and asking how much of that was actually legal and understanding what that is.

I wish I could tell you [that], from the U.S. Senate, I could I could snap my fingers and take it back to where it was. That’s just not where we are. We have to leverage the political process to build what we need. From the U.S. Senate, there’s a lot of lawmaking we can do that I hope would pass both houses. Now, if Donald Trump wants to veto, he can. But all that’ll do is demonstrate the pressure that needs to be put on him and demonstrate exactly how frankly mendacious he has been when it comes to the proper function of government.

We’ve heard from people who are not in the scientific community who believe that getting government out of research isn’t necessarily a bad thing, that we should get private industry funding basic research . If you’re elected to the Senate, what will your position be?

We need to restore funding, but it also needs to be built to what it should be. It’s not enough to go back. You talk to anybody in science, they’ll tell you that NIH money is really important, but too often the way that the NIH works is an anachronism to how science actually works right now. Another point I’ll make is that these cuts may seem harmless until you’ve got 8,000 people who are infected with Cyclospora , and people are trying to dodge lettuce in the United States of America. I think it’s important for us to put a name on the consequences and start to address them.

You are a vocal proponent for Medicare for All . Depending on where one gets their numbers, between 800,000 to up to three million people are employed either in health insurance or in fields adjacent to health insurance. How would you approach implementing Medicare for All while limiting mass layoffs in these industries, especially as AI puts pressure on workers?

You’re already seeing a lot of these corporations start using AI to try to eliminate these jobs. While [those job numbers] may be true today, within the next two years, they are coming for those jobs. You’re still going to need people who can process all of the churn that happens, although it’s going to be substantially less. You’re still going to need people in the industry who are doing those jobs.

The other thing I’ll just tell you is that Medicare for All is going to be a net positive for health care jobs. People right now are working for an industry that makes its money denying people health care. Think about the number of people who exited the current health care system that are going to be repatriated back in when we finally guarantee every single person health care.

Would that mean importing jobs that are currently in the private sector into the public sector?

No, not necessarily. You’re creating a lot of health care jobs, which will be private. [The idea is to have] public insurance but private health care. Ten percent of the population right now can’t regularly get the health care that they need. Medicare for All would enable them to get that health care, and so you’re going to see massive growth in the private sector to provide health care as paid for by public health insurance.

Similarly, to your point around the health insurance jobs, you’re going to be able to leverage health insurance. Even Medicare, as it operates right now, there is some subcontracting to private companies that manage some of the paperwork even if the health insurance is still publicly guaranteed. That said, you are going to be able to create a lot of public sector jobs doing a lot of this work. And so the transition will create many more jobs than are lost, mainly via health care, which is going to stay private under Medicare for All.

is a breaking news reporter at Scientific American. More by Adam Kovac

Founded 1845, Scientific American is the oldest continuously published magazine in the United States. It has published articles by more than 200 Nobel Prize winners.

Scientific American covers the most important and exciting research, ideas and knowledge in science, health, technology, the environment and society. It is committed to sharing trustworthy knowledge, enhancing our understanding of the world, and advancing social justice.

Sign up for the Scientific American daily newsletter.

Upstream Rust maintenance report

Lobsters
kobzol.github.io
2026-10-02 21:27:47
Comments...
Original Article

As noted in my previous report , I am currently working on the open source Rust toolchain as a Sovereign Tech Fellow . Every two months, I’m putting out a report of my open source work done in that period. This is the second installment of this series. Same as the last time, I’ll try to pick a few highlights, summarize the rest of the stuff that I worked on, and also provide contribution statistics and a raw list of opened PRs.

This post details my open source Rust work done in August and September 2026.

Here’s an index for simpler navigation:

Reducing target directory size

I already wrote in my previous report , and also earlier on this blog , about the initiative to reduce the size of the target directory by removing duplicated metadata of Rust crates, which can make it smaller in real-world projects by 5-35%, so it can be quite significant.

Last time, I noted that we are waiting for Cargo’s new build dir layout nightly experiment to conclude, before we enable yet another experiment on nightly by default. This happened during August, so I enabled the usage of -Zembed-metadata=no on the nightly channel by default, and announced it on the Rust blog post.

The good news is that we haven’t really heard any large complaints or issues about it since then, and several build systems already successfully updated to the new mechanism, where the metadata is kept only in .rmeta files, which then have to be passed explicitly to rustc via the --extern flag.

Thanks to that experiment, I discussed moving forward with the Cargo team, and opened a stabilization report for the compiler side of the feature (turning the unstable -Zembed-metadata flag into a stable -Cembed-metadata flag), which is now in FCP vote :tada: . Soon after, Weihang Lo opened a separate Cargo stabilization report for the Cargo side of the feature (passing -Cembed-metadata=no to rustc by default). There are some remaining questions about how does this affect Cargo’s stability story about .rlib , .dylib and .rmeta artifacts, but I think that this should not block the compiler side of the stabilization.

Since Cargo is already using the unstable -Zembed-metadata flag by default on nightly, and Rust’s own build system is also making use of that flag, simply moving -Zembed-metadata to -Cembed-metadata , which is what we would normally do, would immediately break nightly users, unless we managed to synchronize that change across both rustc and Cargo atomically, which is not trivial at the moment. The plan is thus to keep supporting both -Zembed-metadata and -Cembed-metadata for some time, to allow users to migrate to the stable version of the flag, and then finally remove the unstable flag.

This is quite exciting, because it looks like we might be finally close to removing the duplication of Rust metadata on disk, which existed in Rust for almost 10 years, and was unnecessarily inflating the size of the target directory.

By the way, it seems that a Project Goal focused on reducing the size of the target directory further is in the works. Maybe don’t go buying another disk drive just yet!

Compile time improvements

Same as in the previous period, I tried to spend some time on improving the performance of the Rust compiler, though this time I don’t have that much to show for it. My experiments were mostly focused on trying to apply arena allocation to various parts of the compiler. I must admit that I did not achieve much success, partly because of my own unfamiliarity with arenas (though I learned a lot during my experiments!), partly due to the way the compiler codebase is structured, and partly because I now think that the design of Rust is sort of hostile towards arena allocation. Though with the upcoming stabilization of the Allocator trait, it should become better.

Apart from the high-level areas that I describe below, I also worked on some random small performance improvements:

  • Added a fast path to string escaping in the compiler in rust#160453 (thanks @matthieu-m for the suggestion ). Surprisingly, escaping strings can be quite hot, and the standard library’s escaping code is currently not very fast .
  • Removed an unnecessary .clone() call from the macro expansion system in rust#162004 . It would be nice if Clippy’s redundant_clone lint could catch this, but unfortunately it is currently not very smart.
  • Resolved a performance regression from a previously merged PR in rust#162371 .
  • Applied LTO to Cranelift in rust#163412 . This didn’t help all that much. I also want to try to apply PGO to Cranelift, which I hope will help more.
  • Tried to reduce the size of macro-generated code in the tracing crate in tracing#3603 , because I saw a lot of unnecessary generated code duplication there in bors. But it seems that tracing has been unmaintained for some time, so I’m not sure if anyone will take a look at it.

Trying to optimize the Rust parser

I like reading about how other programming languages and communities implement their compilers. In particular, I always enjoy seeing the tricks that the Zig compiler pulls off. It implements parsing very efficiently, using a Data Oriented Design/Entity Component System approach, and I find that quite cool.

I wanted to try using a similar approach for tokenization and parsing also in the Rust compiler. Now, it should be noted that parsing Zig or C is a very different discipline than parsing Rust. It’s not necessarily that Rust’s grammar is that complicated to parse, but some of the language’s design, where parsing is interwoven with macro expansion and name resolution, and its focus on providing great diagnostics, makes its parsing logic much more complicated than one might think.

Today, the Rust lexer and parser represent tokens using essentially a tree representation. Token trees are stored in a Vec , where each tree can either be a leaf token (like + ) or a delimited group of nested trees (like (1 + 2) ), which is stored in a separate Vec . In other words, everytime the parser encounters parentheses, braces or brackets in a Rust file, it will allocate a new Vec on the heap. As you might imagine, this is not terribly efficient, neither time-wise, nor memory-wise.

I tried to change that to use a flat representation, where all tokens, including delimited sequences, are stored in one single Vec , to avoid all those tiny allocations. Changing the lexer was very simple, and locally produced ~30% wins in terms of lexing performance, which was nice. However, moving this change up to the parser, macro expansion and AST handling was quite… involved 1 . It was not really feasible to change the representation everywhere at once, because there is a lot of code that accesses the TokenStream type, which abstracts the Vec of token trees. So I had to create a second implementation on the side (using the flat Vec of tokens), reimplement all the old functionality using the flat token structure and implement conversion functions in both directions (from flat to nested trees and from nested to flat trees). Then I had to incrementally migrate usages of TokenStream in the compiler to the new representation, by moving the conversions to higher and higher-level call-sites, until all conversions were gone.

After a lot of work, I managed to get quite far in rust#159378 . However, the performance results are not very impressive so far: the compiler is actually slower with the flat token representation. The best result I achieved was this , but then once I started moving more stuff into the new representation, it actually regressed even more. The reason is mostly macro expansion. The nested tree representation with every delimited group being a separate allocation is actually quite useful when you want to do a lot of modifications of the parsed list of tokens, which is exactly what happens during the expansion of both declarative and proc macros, and also in a few other parts of the parser, usually around attributes, where we currently do a bunch of in-place modifications of the input token stream (which is kind of terrifying, and a hack).

I think that it should be possible to either modify that code to be more friendly to the flat representation, or somehow use a hybrid representation that would combine the performance benefits of the flat representation with the ability to be easily modified of the nested representation. Not sure how that would look like though.

Some further performance wins might also be gained by:

  • Reducing the size of individual tokens. Each token has ~40B at the moment, which is kinda ludicrous.
  • Using a chunked Vec , so that appending to the end of the flat list doesn’t require reallocating and copying the whole token list when the Vec runs out of memory.
  • Migrating AttrTokenStream to the flat representation too, so that we don’t have to convert between the flat and separately allocated token trees when dealing with attributes 2 .

I think that this might be a situation where the performance results keep being red until the last bottleneck/conversion point is removed, and at that point it could jump to being green a lot. But so far it looks like getting to that point is quite difficult :)

While staring into the token parsing code for many weeks, I noticed that there are other sources of inefficiencies. For example, the parser snapshots its state before parsing certain constructs, so that it can reparse some parts of the code lazily at a later stage. This snapshotting happens quite often (on bors I measured over 250 thousand snapshots being made), and it is a bit expensive. That might sound surprising; isn’t the parser state essentially just a single number with a position pointing into an input string? Well, I wish :) The issue is that because of how the delimited groups are represented, we actually have to remember a stack of currently active parent delimited groups, so that when we encounter the end of a delimited group, we know where to backtrack in the parser. As an example, when we are parsing (1 + array[0<pos>] + 2) , and the parser is at <pos> , where the [0] group ends, it has to be able to go back to its parent delimited group ( (1 + ...) ), so that it can access its data.

What makes this a bit infuriating is that most times, I suspect that we don’t even end up using most of the parents at all, but we still pay the cost for cloning the whole parent stack. This is kinda stupid, so I tried to figure out some ways of not cloning a Vec of several items everytime a snapshot is made in rust#162593 . I tried:

  • Storing intrusive pointers in the delimited group stack items, so that snapshotting only clones at most one parent item, and I can find its parent by following the pointer. But this trades snapshotting performance for adding yet another allocation when encountering each delimited group (so that we can store the parent somewhere), which resulted in performance regressions.
  • Using immutable datastructures (using crates like im ) to reduce the overhead of cloning. This also resulted in performance regressions.
  • I considered using an Arc inside the delimited groups to remember their parents, but I quickly ditched this idea, because it would be a lot of work to make this work, and would likely run into reference cycles (though that would be resolvable via using Weak pointers), and would inflate the size of each token tree even more.
  • Only clone the immediate parent, not the whole parent stack. I suspect that this is enough, because I don’t think that we ever need to recurse back into grandparents when restoring the snapshotted parser, and it produces relatively nice performance results. The problem is that so far I haven’t been able to convince myself that this works in 100% of the cases or find someone who would be confident enough about how the parser works to confirm this for me.

In the flat representation, I removed the need to remember the parent stack by storing an index into the start of a parent into each delimited group. That makes snapshot cloning essentially free. But everything is a trade-off, because this increases the size of each token, which has a negative effect on performance.

Anyway, I might return to this work sometime later, though rust#159378 will require a lot of conflict fixes, because it modifies many files in the compiler.

Making Polonius faster

Since August, the Rust compiler uses the Polonius borrow checker implementation on the nightly channel by default :tada: This was a long time coming and I’m very excited by it. Compared to the previous borrow checker implementation (NLL), it is currently slower in some cases though. When it originally landed at the start of August, check builds of serde had a ~15% instruction count regression (end-to-end) vs NLL.

Since then, Jack Huey and Rémy Rakic , who drive the whole Polonius effort, did a lot of great work to speed up Polonius. Now, the regression of Polonius vs NLL on serde is just around ~3-5%. Since this is an end-to-end measurement, and borrow checking is usually a relatively modest fraction of the work performed by the compiler, this means that Polonius itself likely became several times faster.

I tried to help with the Polonius optimization effort, by attempting to use arena allocation for some of the data structures that Polonius allocates, but I was unable to find a good speed-up. Then I tried to parallelize Polonius, which took a non-trivial amount of work. After I finally got it to compile, I triumphantly opened a PR … only to find out that Polonius was already parallelized , and all I managed was to implement another nested parallelization, which was usually “parallelizing” work across (almost always) only 1 or 2 items :see_no_evil: This was yet another reminder and a cautionary tale against trying to “micro-optimize” something before I properly understand how it works. Nevertheless, I did learn a bunch of things and found some possible improvements to the parallelization machinery for the future. I also gained some experience with how to apply arena allocation to the Rust compiler, which should come in handy later.

In the end, my experiments at least led to rust#162488 , which replaced one data structure in the Polonius implementation, which led to a ~2% instruction count win on a check build of serde . It was a relatively modest win, and just a tiny fraction of the Polonius performance improvements done in the past few weeks, but hey, I’ll take it.

Shipping the parallel frontend

Making the frontend of the Rust compiler parallel is an initiative that has been in progress for many many years. This sentence was also true in 2023, where we hoped to stabilize the parallel frontend within the following few months. Oh well.

Now, it finally looks like we are on the verge of enabling it on nightly by default, thanks to the efforts of Vadim Petrochenkov and several other contributors. I still have only very basic understanding of how the parallelism inside the compiler’s frontend works, and thus wasn’t involved in fixing the myriad of bugs and reproducibility issues with the parallel frontend itself. However, I thought that I might at least help with driving the actual nightly switch forward.

In rust#162848 , I’m currently working on doing just that. Mostly it comes down to performing some changes to our testing infrastructure. And I also prepared a blog post to announce this change.

Hopefully, we will be able to land the parallel frontend on nightly in October, to start gathering data from users in the wild.

Measuring compiler performance on ARM

Over the course of 2025, I worked on a Rust Project Goal together with James Barford , where we extended rustc-perf , the Rust compiler benchmarking suite 3 , to support running benchmarks on multiple machines in parallel, which made our benchmarks much faster to execute, and also to support running benchmarks on other architectures than x64 . Even though this work got concluded by the end of 2025, we have been actually still running compiler benchmarks only on x64 throughout 2026.

This changed in the last month, where we got access to a cloud ARM machine, and started running rustc-perf benchmarks on it by default. We still don’t yet look at the ARM results that much though, that will require some more tooling and policy changes.

Investigating incremental recompilation

I spent some time to continue my investigations into why incrementally rebuilding bors tests is so damn slow. I found out that the compiler currently doesn’t handle incremental recompilation of async closures very well, and this is a problem for bors, because it is full of them. Essentially, what happens is that everytime you modify code within an async closure, the compiler will recompile a much larger part of your crate than it would necessarily have to. I hacked together a small compiler change that avoids this behavior, which reduced the incremental recompilation time of bors from 14s to 8s (!). However, I don’t know if this is the proper fix, and whether it would have an adverse effect in other benchmarks. I also found out that enabling debuginfo causes more incremental recompilations in bors than when debuginfo is disabled. This is highly suspicious, and I suspect that there is much more to uncover here.

Doing this work will require some sustained focus, and other things came up in the meantime, so I didn’t get to it yet. Maybe I’ll find some time to work on it in the next two months. In the meantime, I tried to nerd-snipe the illustrious nnethercote to take a look at bors, so maybe this won’t be an issue soon :laughing:

Improving the Rust compiler merge queue

Speaking of bors, there were some exciting updates to the merge queue of the Rust compiler in the past few weeks.

In my previous report , I noted that @Mark-Simulacrum started looking into running some of our more expensive and latency-sensitive CI jobs on EC2 runners. This is cool in that it both makes the jobs much faster, but also much cheaper, than running on CodeBuild, which we did previously. I’m happy that we finished this work with Mark in August :tada: which means that our x64 try jobs, which we have to run before starting each compiler performance benchmark, now take only ~1 hour, so they are essentially twice as fast as before. We also started using the EC2 instances for ARM jobs that build compiler artifacts, which is what enabled us to run ARM benchmarks (as noted above ) on every merged PR.

We also landed one long-standing feature request and one infrastructure improvement in bors.

The long-standing feature request is the automation of something that we call “r=me after PR CI passes”. When approving PRs in the rust-lang/rust repository, it was quite common for the reviewer to want to approve a PR for which the PR CI hasn’t finished yet. We didn’t have any automation for that, because we use our own merge queue implementation (bors), so we cannot use GitHub’s built-in tooling for this. So the PR author had to wait until PR CI became green, and then approve the PR on behalf of the approver. This was annoying; people sometimes forgot to approve the PR or had to keep a tab open until the CI finished. There was an open issue about this from 2019 in homu , the previous merge queue implementation.

Since we finally switched from homu to bors at the start of this year, it finally became realistic to implement this feature, though it took us a few months before we figured out the right design for it and found the time to implement it. Sakibul Islam , my former GSoC mentee, who is now a member of the bors team , single-handedly implemented and tested the feature (I have to say that it was a joy to review this PR!). Reviewers can thus now simply write @bors r+ and let bors worry about PR CI passing or failing, without having to babysit the PR.

The infrastructure improvement that I mentioned is that rollups are now unrolled by bors. Oh wow, what a sentence! Rollups are PRs that batch multiple other PRs together to amortize the cost of running CI. After they are merged, we “unroll” them by producing compiler artifacts for each separate PR merged with the previous main commit, so that we can run performance benchmarks on each PR separately, to determine which rolled-up PR caused a performance regression (if there are any) 4 .

Previously, this unrolling was performed by the rustc-perf bot. This worked relatively well, but it was implemented in a “fire-and-forget” way, so if the unrolling failed, we had no way to retry or even properly learn about the failure, and it required the rustc-perf bot to have permissions for pushing into the rust-lang/rust repository. Moving this feature from rustc-perf to bors was something that I wanted to do for a long time, to make it much more robust, make it possible to implement some cool new features related to triaging the rolled-up performance benchmarks, and also to remove the unnecessary push permissions from the rustc-perf account, thus making our CI safer overall. Merging the rustc-perf PR with almost 450 removed lines was a pretty nice feeling 5 . A lot of this work was initially performed by Sakibul again, I just picked it up and moved it over the finish line.

Josh subtree migration

The rust-lang/rust repository leverages several other git repositories with which it has to be periodically synchronized in both directions. For some of them, we are using git subtree , while for others, we use the awesome Josh tool, about which I previously wrote on the Rust blog.

For several years, we have been trying to migrate the remaining subtrees using git subtree (which is very painful to use) to Josh. Lately, we made a lot of progress on that, and it seems that we are nearing the finish line. The Josh tool received some updates that make it much better in managing our repositories and we implemented some improvements to our josh-sync tool, which wraps Josh to provide rust-lang -specific functionality.

This allowed us to move forward with the five remaining subtrees still managed by git subtree :

I would (finally!) like to get this over the finish line, so that we can stop using custom patched versions of git subtree to even be able to handle our subtree repositories, and to have all the subtree logic be performed in a unified manner via josh-sync and Josh.

In addition, we are also currently discussing whether it would make sense to move Cargo from a git submodule to a Josh subtree.

Tracking rust-lang crates on crates.io

There are hundreds of “official” or semi-official crates on crates.io that are somehow associated with the Rust Project. Either they are deployed from repositories living under the rust-lang GitHub organization or they are owned by some Rust Project teams. Historically, we did not do a very good job of categorizing and tracking which crates are actually owned by the Rust Project, and which teams should own them though. So many of those crates are currently owned by individual GitHub users, some of which are no longer contributing to Rust anymore.

The Infrastructure team would like to clean that up, first so that we don’t have such a mess in it, but also for improved security. Our goal is to force using Trusted publishing , to ensure that all crate releases happen from CI, so that we at least have some audit trail when releases are made, and make all Rust Project crates owned by the special rust-lang-owner account, which is only accessible to our infrastructure admins, so that it won’t be possible to disable Trusted publishing and perform silent “rogue publishes” of crates, for example if the GitHub account of some Rust Project member would be compromised.

To achieve that, we have to:

  1. Actually find out which crates should be owned by rust-lang-owner . This is mostly manual work. Luckily, Eric Huss did a lot of the work already, and put most of the crates that we should own into a big Excel table, so this part is mostly done.
  2. For those crates that are currently not owned by rust-lang-owner , ask their current owners to invite the special account, so that we can start managing the crate. I started doing this cat herding recently, and it was surprisingly successful. The owner account now owns over 250 crates, vs (I think) less than a hundred that it owned when I started this effort a few weeks ago. People have been very responsive to my crate ownership requests, thank you! 6
  3. Associate the crates that we manage to specific Rust teams that should (virtually) own them, and repositories from which they should be deployed, in the team database. This is an ongoing effort, right now we only track a few tens of those crates. We did not have owners for some crates that we already tracked in the team database, so I backfilled them in team#2742 and made it mandatory to have at least one team owner for each publishable crate in team#2743 .
  4. For the crates that are still expected to have some releases, and are not yet completely deprecated, configure Trusted publishing for them, and configure them to be publishable from GitHub Actions (e.g. using release-plz ). This is the hardest part, because it is not fully automatable, and different teams have different ideas on how they want their crates to be published/released. To make this a bit simpler and avoid doing unnecessary work, I implemented support for tracking crates that are not expected to be published anymore without configuring Trusted publishing for them in team#2775 .
  5. Provide a way for Rust teams to yank and unyank crates that they own, without those teams actually being owners of the crate on crates.io (as noted earlier, we want to avoid this to reduce the fallout of compromised accounts). I implemented this in triagebot ( team#2725 , triagebot#2497 ), so now Rust team members can yank and unyank crates on the Rust Zulip server using some chat commands.

Steps 2., 3. and 4. are still ongoing, but they made a lot of progress in the past few weeks. Thanks to that, we were able to go forward with our plan and ensure that rust-lang-owner is the sole owner of all the crates that we track, which I implemented in team#2726 .

Now, “all that’s left” to do is to incrementally backfill all the Rust Project crates into the team DB, along with figuring out which team should own them, and configuring Trusted publishing for them if needed.

Reviving unmaintained repositories

In the past two months, I participated in a “revival” of two codebases that were neglected and unmaintained for quite some time, and made substantial improvements to one of them.

thanks.rust-lang.org

The thanks website shows contribution statistics for individual Rust releases, as a way of thanking all the people who contribute to Rust. This tool was created several years ago, and since then it has been mostly chugging along fine. However, we wanted to make some improvements to it, which was quite difficult, because it was implemented in a way where it was counting the contributions for each Rust release starting all the way from the first Rust release. In other words, it was O(n^2) , which meant that it took over 10 minutes to generate the contents of the website. That wasn’t really a problem on CI, because it only ran once per day, but it was super annoying for any local experiments and further development.

We wanted to move to a faster algorithm , but it was producing some differences in output. So as a pre-requisite, we first wanted to add tests. Daniel Scherzer implemented end-to-end tests in thanks#99 , but it was a bit cumbersome to store snapshots of the generated HTML pages in the repository. So I first implemented a CSV output mode in thanks#105 and thanks#110 , which only outputs the contribution statistics into a CSV file, rather than the full HTML page. Then we realized that some of the tests were not deterministic, which turned out to be an issue with missed sorting of git submodules, which I fixed in thanks#111 . After that, we could finally merge the snapshot tests. Thanks, Daniel!

After we had some initial tests, I set out to reimplement the contribution counting logic. Originally, I wanted to just modify the original algorithm to avoid the duplicated work, but after some experiments, I decided to reimplement the whole thing from scratch, to make it amenable for parallelization. First, we gather a set of all commits for individual Rust releases and tags, then we remove duplicates and assert that we have counted all commits in the git history. The duplicate check was actually missing in the original implementation, so this rewrite actually fixed a bug, where some commits were previously counted multiple times. After that, we go through all the commits and extract the commit author and reviewer out of it, while taking the mailmap of the rust-lang/rust repository into account. This work was performed multiple times for each commit before, while now it is only performed once per commit. Splitting the work into two explicit stages also made it much easier to parallelize it.

This new algorithm was implemented in thanks#116 . The performance results were very nice. Gathering the statistics for all (~100) Rust releases originally took more than 10 minutes, but with my PR it only took ~20 seconds on my laptop. I further made it even faster in thanks#123 by checking out git submodules in parallel. This huge performance improvement actually finally made it possible to experiment with the tool locally, and make improvements to it. It also made CI much faster, which is nice. Sometimes, improving performance is not just about metrics, but it unlocks completely new opportunities.

The main improvement that this allowed us to do was adding support for multiple projects. While the tool already counted contributions from all submodules and subtrees of the main rust-lang/rust repository, it did not take into account other important projects in the Rust toolchain, such as Rustup. In thanks#125 , I did some preparatory work to support multiple projects and then in thanks#126 I added support for counting contributions from the rust-lang/rustup repository. Since we now had more than a single project, I also added an overview page in thanks#130 , to see all the projects at one place, and a new page that combines all-time contributions from all the tracked projects. Once we added Rustup, I reached out to the crates.io and docs.rs maintainers, and they were fine with also including those projects in thanks , so I did that in thanks#131 . And finally, I made the project overview page be the default homepage in thanks#138 , so that now that you go to thanks.rust-lang.org , you will see all the tracked projects at one place.

I am quite happy that we finally made the thanks tool not take 10 minutes to execute, I considered that as an affront. We actually wanted to do this for a long time, but the reason why it didn’t happen is that no one felt like they were really owning the tool, so reviews fell through the cracks and nothing moved forward. In July, I finally decided that there is no point in waiting years for someone else to review changes in this repository. I asked a fellow Infrastructure team member if they would be fine if I took over the maintenance of the tool, and they said yes. So I took it upon myself to review and merge changes to this tool, even if it sometimes meant approving a bunch of my PRs myself, without much review. Once I embraced this mindset, I was able to move forward with the improvements at light speed :) And I think that it was overall a good thing for this codebase.

rustup-components-history

The second repository that I picked up after it not being maintained for a long time was rustup-components-history , which powers this page that shows which Rustup components were available for individual Rust targets in the past few days.

It was a similar story as thanks ; no one felt like they owned the repository, so no reviews were being made. I found out about this when I was approached by @rami3l , the lead of the Rustup team, if I could help him find a reviewer for his PR that was opened over two years ago. I realized that as a member of the Infrastructure team, that reviewer can simply be me, so I reviewed their PR and finally merged it. Then, by coincidence just a week later, an unrelated change to one Rust target actually broke the components website. I thus went and fixed it, and also gave permissions to the Rustup team to manage this repository. You would think that a tool called rustup-components-history would allow the Rustup team to merge changes to it, but it wasn’t actually configured as such before.

Once I fixed the tool, I spent a bit of effort on doing the usual facelift of Rust repositories:

This didn’t take much time and it rejuvenated the codebase a bit, which is always nice.

Funding team activities

As a part of the Rust funding team, I helped bootstrap the Maintainer in Residence (MiR) program, by helping to categorize the maintenance needs of various Rust teams and matching them with the funding needs of Rust maintainers. We were able to support 7 people right from the start, which is awesome! You can find out more about them in these two blog posts:

Apart from that, I finished the MiR page on the Rust website (implemented in www.rust-lang.org#2323 ), and worked on some tooling to help us track the contributions of funded maintainers. I also prepared a charter for the Funding team, to be approved by the Rust Leadership Council.

There is a lot more work to do in the Funding team, but I think that we got off to a good start.

Other things I worked on

  • I resumed my refactoring spree of bootstrap , the Rust compiler build system, this time with a set of PRs that refactored how the build system builds and downloads LLVM from our CI. This continues a long series of refactoring that tries to make bootstrap easier to understand, modify and maintain, and unlocked some further cleanups ( rust#161853 , rust#161691 , rust#160631 ).
  • I started working on renaming the rust-clippy repository simply to clippy and also renaming its default branch to main ( rust-clippy#17541 , rust-clippy#17552 ).
  • As usually, I did lots of tiny fixes and improvements to CI, infrastructure, tooling of various rust-lang repositories, and to our bots.
  • As usually, I participated in the Rust Compiler performance triage .
  • Together with my co-mentors, Jieyou Xu and Folkert de Vries , we concluded the two Rust GSoC 2026 projects that we were mentoring. I think that both of them were quite successful, and we were very excited to work with our mentees, Walnut356 and xonx4l . I will share more details about the projects in my next report, once Rust GSoC 2026 has completely finished.
  • I analysed the results of a survey about the performance of the Rust Leadership Council and shared it with members of the Rust Project.
  • I wrote a bunch of blog posts on the official Rust blog:
    • Together with Lori Lorusso , we prepared another entry of the maintainer spotlight series, this time with @blyxyas .
    • Posted an update on the activity of the Leadership Council.
    • Announced the Leadership Council elections.
    • Announced the nightly experiment to default to -Zembed-metadata=no in Cargo.
    • Together with Lori Lorusso , we prepared two announcements ( 1 , 2 ) about our first Maintainers in Residence :tada:
  • I was re-elected as the Infrastructure team ’s representative in the Leadership Council. Yay! I hope to continue being useful to the Rust Project in the Leadership Council.

Contribution statistics and pull requests

Below you can find some statistics and a list of PRs that I opened during August and September 2026.

  • Opened 262 pull requests in Rust-related repositories.
    • With a median of 20 modified lines per pull request.
  • Reviewed 214 pull requests in Rust-related repositories.
  • Sent 889 comments in the rust-lang GitHub organization.
  • Sent 2282 public and 1942 private messages on the Rust Zulip .

I want to provide a bit of context regarding the statistics above. After my previous report , where I wrote that I opened 162 PRs in two months, some people reached out to me and essentially asked me whether I am OK :laughing: And that I shouldn’t overwork myself. Given that in this report, that number is (exactly!) 100 PRs higher, I thought that I should explain it a bit.

Those numbers of PRs are not produced by me using LLMs (as I explained previously , I almost never use them for code that gets committed to git) nor by me being a 10x engineer working 24/7. I simply open a large number of usually very small PRs across many projects, given the nature of work that I do in the Rust Project, which often deals with fixing CI, improving infrastructure and tooling, improving docs, etc. I don’t push out 5 big new features every day. That’s why I included the median 7 of modified lines in my PR, to get across the idea that most of my PRs are very small. So take the numbers above with a grain of salt. I mostly put them here for my own historical record.

List of opened PRs

rust-lang/rust (71 PRs)

  • #160421 : Do not use -Werror when building rustc_llvm with GCC ( closed )
  • #160427 : Run try builds on EC2 by default ( merged )
  • #160434 : Avoid Docker push when the image did not change ( merged )
  • #160449 : Fix lookup of object files ( merged )
  • #160451 : Deduplicate target and host filesearch ( merged )
  • #160453 : Add fast path to escape_string_symbol ( merged )
  • #160493 : Rollup of 29 pull requests ( merged )
  • #160501 : Add bootstrap CLI snapshot test for testing miri ( merged )
  • #160502 : Reduce number of miri tests executed on PR CI ( merged )
  • #160573 : [perf test] Revert “Rollup merge of #159339 - LorrensP-2158466:extract-use-injections, r=petrochenkov” ( closed )
  • #160574 : Update rustc-perf submodule ( merged )
  • #160579 : Do not download GCC from CI in PR CI jobs ( closed )
  • #160631 : Do not eagerly download rustfmt in bootstrap ( merged )
  • #160645 : Assorted bootstrap LLVM refactors (part 1/N) ( merged )
  • #160681 : Respect --all-targets flag when checking the compiler ( merged )
  • #160693 : Add branch config for perf. unrolling in bors ( merged )
  • #160839 : Stop updating npm lockfile by Renovatebot ( merged )
  • #160894 : Allow running an arbitrary number of try jobs per PR ( merged )
  • #160916 : Assorted bootstrap LLVM refactors (part 2/N) ( merged )
  • #160970 : Fix handling of relative paths starting with a dot in bootstrap ( merged )
  • #161046 : Enable unrolling feature of bors ( merged )
  • #161102 : Explicitly pass run_make_support rlib/rmeta paths to compiletest ( merged )
  • #161111 : Take bors try-perf branch into account in verify-channel.sh ( merged )
  • #161145 : Remove references to the obsolete try-perf branch ( merged )
  • #161235 : Compute job time in post-merge-report from the actual GitHub duration ( merged )
  • #161236 : Download auto jobs in citool in parallel ( merged )
  • #161237 : Remove jdno from infra-ci rotation ( merged )
  • #161247 : Assorted bootstrap LLVM refactors (part 3/N) ( merged )
  • #161290 : Assorted bootstrap LLVM refactors (part 4/N) ( merged )
  • #161344 : Update the rustc-perf submodule ( merged )
  • #161384 : Bust sccache’s cache ( merged )
  • #161393 : Configure LLM policy URL for triagebot ( merged )
  • #161438 : Change triagebot backport to ping T-libs-fcp ( merged )
  • #161460 : Add a size_hint function to Display ( open )
  • #161475 : Fix checking of LLVM prebuilt status ( merged )
  • #161483 : Warn about running ui-fulldeps tests in stage 1 ( merged )
  • #161516 : Revert #161236 (Download auto jobs in citool in parallel) ( merged )
  • #161641 : Check for missing rustfmt in the stdarch intrinsic test step sooner ( merged )
  • #161663 : Reduce dependency on implicit paths in bootstrap ( merged )
  • #161666 : Print vendor instructions in x vendor ( merged )
  • #161691 : Assorted bootstrap config refactors (part 1/N) ( merged )
  • #161853 : Consolidate LLVM skip in check builds in bootstrap ( merged )
  • #161992 : [do not merge] ARM perf experiments ( closed )
  • #161995 : [perf] Use SmallVec in smart_resolve_path ( closed )
  • #162004 : Remove unneeded clone in macro deriving ( merged )
  • #162006 : Use a simple arena for annotatable items ( closed )
  • #162053 : Add bootstrap command to check all Tier 2+ targets ( open )
  • #162092 : Implement test- naming convention for CI jobs ( merged )
  • #162111 : Update mailmap for Will Crichton and Petr Hosek ( merged )
  • #162328 : Allow overriding filecheck even if LLVM is built or downloaded ( merged )
  • #162371 : Cache sanitizer set in Session ( merged )
  • #162413 : Unconditionally invalidate the library when the compiler changes ( merged )
  • #162428 : Pin the number of frontend threads while gathering PGO to 1 ( merged )
  • #162473 : Small x perf improvements ( merged )
  • #162480 : Mono item collection microoptimizations ( closed )
  • #162485 : Parallelize gathering constraints in the borrow checker ( closed )
  • #162488 : Use DenseBit for drop_live_at in liveness tracing ( merged )
  • #162512 : Implement ExactSizeIterator for Chain ( closed )
  • #162524 : Use SmallVec in TokenCursor ( closed )
  • #162533 : Use share-generics for libstd ( closed )
  • #162593 : [perf] Try to optimize collect_pos ( open )
  • #162777 : Remove allocation from mbe::quoted::parse ( closed )
  • #162780 : Use -Clinker-plugin-lto in bootstrap ( closed )
  • #162848 : Use 2 parallel frontend threads by default on the nightly and dev channel ( open )
  • #163212 : rustfmt subtree update ( closed )
  • #163412 : Apply LTO to Cranelift and GCC codegen backends ( open )
  • #163433 : Support also try-jobs: to specify custom try jobs ( merged )
  • #163436 : Stabilize -Zembed-metadata ( open )
  • #163444 : Add stable_rustc helper in run-make-support ( merged )
  • #163452 : Use newtype enums for representing frontend and backend jobs ( open )
  • #163533 : Run cg_gcc tests with the correct compiler ( open )

rust-lang/team (36 PRs)

  • #2648 : Move inactive t-triage members to alumni ( merged )
  • #2649 : Add Zulip topic for the funding team and MiRs ( merged )
  • #2662 : Configure branches for unrolled perf builds for bors ( merged )
  • #2669 : Add first batch of Maintainers in Residence ( merged )
  • #2670 : Add Tyler Mandry to funding advisors ( merged )
  • #2674 : Change default branch of rust-clippy to main ( open )
  • #2679 : Add ruleset for the thanks repo ( merged )
  • #2680 : Remove rust-timer branches from rust-lang/rust ( merged )
  • #2681 : Remove obsolete merge bot variants ( merged )
  • #2682 : Archive the rustc-rayon repository ( merged )
  • #2696 : Rename the default branch of triagebot to main ( merged )
  • #2697 : Change default branch of the forge repository to main ( merged )
  • #2700 : Add check for duplicated alumnis and remove them ( merged )
  • #2713 : Add lcnr and Joel Marcey to funding advisors ( merged )
  • #2725 : Add crates to the v1 API ( merged )
  • #2726 : Ensure that only rust-lang-owner owns our tracked crates ( merged )
  • #2727 : Track Cargo crates ( open )
  • #2735 : Point the funding@rust-lang.org e-mail list to the funding team ( merged )
  • #2742 : Assign owner teams to crates.io crates ( merged )
  • #2743 : Force team ownership of crates ( merged )
  • #2744 : Set compiler team as co-owners of annotate-snippets ( merged )
  • #2751 : Add branch protection for the goals repo ( merged )
  • #2758 : Give the Rustup team access to rustup-components-history ( merged )
  • #2759 : Change the default branch of portable-simd to main ( merged )
  • #2767 : Add Scott Schafer to the Maintainers in Residence team ( merged )
  • #2772 : Give security response merge rights to the blog ( closed )
  • #2775 : Allow managing unpublishable crates ( merged )
  • #2777 : Process crates without trusted publishing ( merged )
  • #2778 : Manage rustup-available-packages crate ( open )
  • #2786 : Add homepage to the funding team repo ( merged )
  • #2787 : Archive the pin-utils repository ( open )
  • #2788 : Update Council libs rep ( merged )
  • #2789 : Manage Council private Zulip stream ( merged )
  • #2792 : Sort bypass actors ( merged )
  • #2793 : Create trusted-contributors team ( merged )
  • #2795 : Add Predrag to trusted-contributors ( open )

rust-lang/bors (34 PRs)

  • #799 : Tag EC2 instances launched by bors ( merged )
  • #800 : Add web page with EC2 instance list ( merged )
  • #801 : Add links to the queue page ( merged )
  • #802 : Implement backfilling of EC2 instances ( merged )
  • #803 : Improve layout of pending builds table ( merged )
  • #804 : Strip auto/try job prefixes ( merged )
  • #806 : Try to start EC2 instances on any branch ( merged )
  • #807 : Reload in-memory job cache when bors starts ( merged )
  • #808 : Document Zulip posting and EC2 instance spawning ( merged )
  • #809 : Only consider workflow run webhooks with the push event ( merged )
  • #812 : Allow opting out of the maximum try job limit ( merged )
  • #813 : Add hint about @bors try nolimit ( merged )
  • #815 : Add a hint about retrying PR CI when someone uses @bors retry in an invalid state ( merged )
  • #816 : Store pr_number field in the build table ( merged )
  • #817 : Add rollup unrolling ( merged )
  • #818 : Add a hint to @bors try cancel ( merged )
  • #819 : Trigger the merge queue when the priority of a PR changes ( merged )
  • #820 : Correctly parse unrolled member build kind from EC2 instance tags ( merged )
  • #821 : Fix termination of multiple EC2 instances ( merged )
  • #823 : Close PRs in DB that disappear from GitHub ( merged )
  • #827 : Ignore homu-ignore blocks in squashed commit messages ( merged )
  • #833 : Add a sanity check for valid bors config before merging a PR ( merged )
  • #834 : Fix decoding base64 GitHub contents ( merged )
  • #837 : Check that EC2 instance is in allowed_instances ( merged )
  • #838 : Add in-memory cache of spawned EC2 instances ( merged )
  • #839 : Check config validity post merge ( merged )
  • #850 : Do not link to GitHub PRs when doing rollup mergeability check ( merged )
  • #854 : Add a test for not merging a tentatively approved PR ( merged )
  • #858 : Upgrade tentative approvals if the PR was already fully approved ( merged )
  • #859 : Make tentative approvals less obtrusive ( merged )
  • #862 : Apply approval labels eagerly when a PR is tentatively approved ( merged )
  • #863 : Support also the try-jobs custom try job marker ( merged )
  • #864 : Set approval_tentative to FALSE in unapprove_pull_request_if_sha_changed ( merged )
  • #865 : Hotfix database state ( closed )

rust-lang/rustc-perf (33 PRs)

  • #2520 : Add support for bare rustc invocation with --print-sysroot ( merged )
  • #2521 : Temporarily fix the serde 4 threads benchmark ( merged )
  • #2522 : Create three NLL benchmarks ( merged )
  • #2524 : Remove unrolling functionality ( merged )
  • #2526 : Remove accidentally committed files ( closed )
  • #2533 : Fix handling of try builds that were not enqueued ( merged )
  • #2534 : Ignore comments from bors ( merged )
  • #2535 : Read bors try build completed comments ( merged )
  • #2538 : Add 2026-08-18 triage ( merged )
  • #2539 : Add crate metadata to compare page quick links ( merged )
  • #2541 : Include profile in error message about unavailable artifacts ( merged )
  • #2550 : Improve support for ARM benchmarks ( merged )
  • #2551 : Correctly store targets of benchmark requests into the database ( merged )
  • #2555 : Respect optional job attribute when checking parent jobs ( merged )
  • #2556 : Separate artifact size records by target ( merged )
  • #2560 : Ignore runtime benchmarks when a benchmark filter is set ( merged )
  • #2564 : Fix duration estimation ( merged )
  • #2565 : Speed up triage command ( merged )
  • #2566 : Run ARM benchmarks by default ( merged )
  • #2567 : Fetch GitHub commits concurrently in the triage command ( merged )
  • #2568 : Add support for frontend threads to the UI ( merged )
  • #2571 : Add tokio-1.53.1 benchmark ( merged )
  • #2572 : Share target directory for different values of frontend threads ( merged )
  • #2573 : UI improvements related to frontend threads ( merged )
  • #2574 : Allow benchmark to opt into different values of frontend threads ( merged )
  • #2575 : Remove the serde-1.0.219-threads4 benchmark ( merged )
  • #2579 : Fix normalization of profiles and scenario in the frontend ( merged )
  • #2580 : Add 2026-09-14 triage ( merged )
  • #2581 : Run apt update before installing Valgrind ( merged )
  • #2585 : Allow filtering Clippy benchmarks in the compare page ( merged )
  • #2586 : Use --jobs-frontend instead of -Zthreads ( open )
  • #2594 : Make parse_benchmarks infallible ( merged )
  • #2596 : Improve support for different codegen backends ( merged )

rust-lang/thanks (26 PRs)

  • #109 : Run Clippy and rustfmt on CI ( merged )
  • #110 : Write all-time data in CSV output mode ( merged )
  • #111 : Sort submodules to fix non-deterministic commit iteration ( merged )
  • #112 : Add conclusion CI job ( merged )
  • #113 : Force handling of submodules ( merged )
  • #114 : Sort also by e-mail author while deduplicating ( merged )
  • #115 : Update to edition 2024 ( merged )
  • #116 : Change processing of commits to walk all commits only once ( merged )
  • #117 : [perf] checkout submodules in parallel ( closed )
  • #118 : Set line-tables-only debuginfo in release mode ( merged )
  • #119 : Refresh repositories using an environment variable, not using a CLI flag ( merged )
  • #120 : Use cache in CI ( merged )
  • #121 : Add all-time snapshot ( merged )
  • #122 : Make tests self-contained ( merged )
  • #123 : Checkout submodules in parallel ( merged )
  • #124 : Allow overriding the mailmap ( merged )
  • #125 : Prepare for multiple projects ( merged )
  • #126 : Add support for Rustup ( merged )
  • #130 : Add page with all tracked projects ( merged )
  • #131 : Track crates.io and docs.rs ( merged )
  • #132 : Ignore bots in docs.rs ( merged )
  • #133 : Ensure that deploys are never executed concurrently ( merged )
  • #138 : Make the projects page be the homepage ( merged )
  • #141 : Make the regression test optional ( merged )
  • #143 : Use deduplicated scores for people and commit count ( merged )
  • #144 : Ignore dependabot and renovatebot by default ( merged )

rust-lang/blog.rust-lang.org (10 PRs)

  • #1910 : Add post about using -Zembed-metadata=no by default on nightly ( merged )
  • #1911 : Add Leadership Council September 2026 announcement post ( merged )
  • #1913 : Clarify that -Zembed-metadata=no is an experiment ( merged )
  • #1914 : Add Announcing our first Maintainers in Residence blog post ( merged )
  • #1927 : Build Zola in debug mode ( merged )
  • #1931 : Fix placeholder date validation for Inside Rust posts ( merged )
  • #1942 : Add Enabling the parallel frontend on nightly post ( open )
  • #1943 : Add Leadership Council September 2026 update post ( merged )
  • #1947 : Add Maintainer spotlight interview with blyxyas ( merged )
  • #1948 : Add Announcing a Maintainer in Residence: Scott Schafer for the Cargo team post ( merged )

rust-lang/triagebot (8 PRs)

  • #2484 : Run CI on main branch ( merged )
  • #2485 : Deny warnings on CI ( merged )
  • #2486 : Add LLM policy URL config ( merged )
  • #2487 : Fix sending of Zulip DMs ( merged )
  • #2492 : Fix sending of Zulip DMs (take 2) ( merged )
  • #2497 : Implement crate yanking/unyanking on Zulip ( merged )
  • #2498 : Fix channel name in yank check ( merged )
  • #2510 : Stop trusting Zulip custom profile fields in the lookup command ( merged )

rust-lang/josh-sync (7 PRs)

  • #56 : Update josh-sync version ( merged )
  • #57 : Fix CI check ( merged )
  • #59 : Load nightly commit SHA from CI instead of from rustc ( merged )
  • #60 : Include nightly version in merge commit message ( merged )
  • #62 : Allow specifying nightly date in pull --upstream-commit ( merged )
  • #63 : Allow pushing commits with the SSH protocol ( open )
  • #64 : Allow continuing with a pull when a merge conflict happens ( merged )

rust-lang/rustup-components-history (7 PRs)

  • #60 : Parse tier tables by their ID, not position ( merged )
  • #61 : Switch default branch to main ( merged )
  • #62 : Update CI ( merged )
  • #63 : Remove unused code ( merged )
  • #64 : Reduce rightward drift ( merged )
  • #65 : Deploy GitHub Pages from the workflow ( merged )
  • #66 : Build the repo on CI in release mode ( merged )

rust-lang/rust-clippy (4 PRs)

  • #17541 : Build docs into the main directory ( merged )
  • #17543 : First add files to git before checking diff in the deploy script ( closed )
  • #17552 : Modify doc links to point to main instead of master ( merged )
  • #17556 : Rename default branch from master to main ( open )

rust-lang/rust-forge (4 PRs)

  • #1098 : Rename the default branch to main ( merged )
  • #1099 : Document llm_policy_url ( merged )
  • #1105 : Switch governance owner from ehuss to LC ( merged )
  • #1106 : Document crate yanking in triagebot ( merged )

rust-lang/cargo (2 PRs)

  • #17413 : Micro-optimize two package dir functions ( merged )
  • #17426 : Use trusted publishing for Cargo crates ( open )

rust-lang/portable-simd (2 PRs)

  • #549 : Switch the default branch to main ( merged )
  • #550 : First rust-lang/rust pull using Josh ( open )

rust-lang/rustc-dev-guide (2 PRs)

  • #3027 : Note that cg-clif is using Josh now ( merged )
  • #3032 : Mark rustfmt as being managed by Josh ( merged )

rust-lang/rustc_codegen_cranelift (2 PRs)

  • #1700 : Initial Josh pull ( merged )
  • #1702 : Update rustup.sh to use rustc-josh-sync ( merged )

rust-lang/rustfmt (2 PRs)

  • #7132 : Rename rust-toolchain to rust-toolchain.toml ( merged )
  • #7134 : Initial Josh pull ( merged )

rust-lang/simpleinfra (2 PRs)

  • #1204 : Configure rustc-perf Ansible for Aarch64 ( merged )
  • #1222 : Increase CPU allocation for triagebot to one full CPU ( merged )

rust-lang/stdarch (2 PRs)

  • #2225 : Refactor generation to use a unified formatting approach ( merged )
  • #2231 : Pull from rust-lang/rust ( merged )

rust-lang/www.rust-lang.org (2 PRs)

  • #2330 : Fix button overflow on funding page on mobile ( merged )
  • #2335 : Add Scott Schafer to the MiR page ( merged )

rust-lang/crater (1 PR)

  • #851 : Add MIT/Apache2 license files ( open )

rust-lang/funding (1 PR)

  • #10 : Add Funding team charter ( open )

rust-lang/funding-private (1 PR)

  • #1 : Add MiR check-in script and contribution analysis script ( merged )

rust-lang/goals (1 PR)

  • #755 : Fix website link in README ( merged )

rust-lang/rustc_codegen_gcc (1 PR)

  • #983 : Rename rust-toolchain to rust-toolchain.toml ( merged )

tokio-rs/tracing (1 PR)

  • #3603 : Use a function to convert from Level to log::Level to reduce the amount of generated code ( open )

Conclusion

These two months felt pretty nice again. I keep worrying about the direction the IT world, and our society, is heading in with the advent of LLMs, and this sometimes makes me feel under the weather, but being funded for open source Rust work is usually enough to get my mood up. I also met a bunch of Rust enthusiasts at a Rust meetup in Prague where I had a talk about how Rust makes parallelism safe(r), and I started teaching Rust again, because the university semester has just started, so that is nice.

As always, I’d like to thank other members of the Rust Project and other Rust contributors, who collaborated with me, discussed various things with me, and reviewed my code in this time period. Thank you very much!

I would also like to thank the Sovereign Tech Agency for funding my open source maintenance work! I am really grateful for this opportunity.

As noted in my previous report , I am currently enrolled in GitHub Sponsors . I did not advertise it much, because I currently have funding for my open source work through the Sovereign Tech Fellowship. However, that is for one year (and it also isn’t full-time), and it is hard to say whether I will be able to find the next source of funding after that. So it is good to have at least some backup. If you’d like to support my open source Rust work, I would really appreciate it! You can find more ways of supporting Rust Project maintainers here .

If you have any questions regarding my upstream Rust work, feel free to ask on Reddit .

Hair Loss Was Just the Start. Ozempic Users Are Also Reporting Nail Trouble

Hacker News
gizmodo.com
2026-10-02 21:15:15
Comments...
Original Article

GLP-1 medications like semaglutide and tirzepatide may come with some unexpected hair-raising side effects, research out this week suggests—or rather, hair- and nail-losing ones.

Scientists in Europe studied hundreds of people taking a GLP-1. Compared to non-users, people on GLP-1s were significantly more likely to develop hair loss and nail-related problems, particularly nail detachment, they found. Though these issues aren’t necessarily too serious, both patients and doctors should be aware of the potential risk, the researchers say.

“For people taking or considering GLP-1 receptor agonists, the main takeaway is: do not panic, and do not stop an effective treatment on your own because of your hair or nails—but do mention these symptoms to your doctor,” study author Charles Taïeb, a clinical dermatologist at Necker University Hospital in Paris, told Gizmodo.

Losing nails and hair

Numerous studies in recent years have found a link between GLP-1 use and hair loss, also called alopecia. And the Food and Drug Administration, following its own investigation of post-market data, now requires GLP-1 drugs like Ozempic (semaglutide) and Zepbound (tirzepatide) to list alopecia as a potential adverse reaction on their product labeling. Some people have also anecdotally reported changes to their nails after starting a GLP-1, but this possible complication has never been formally studied until now, according to the authors.

The researchers formed a team, called SCOP-GLP1, to investigate both of these topics. The team is composed of scientists in France and Italy, and the project was funded by the pharmaceutical company Pierre Fabre , one of the largest dermo-cosmetics firms in the world (in fact, Pierre Fabre markets hair loss products via brands like Ducray and René Furterer , so the company has a commercial interest in this line of research). They conducted two similar studies, one that focused on nail-related changes following GLP-1 use and the other on hair-related changes.

The studies each compared the outcomes of 575 people taking a GLP-1 for their type 2 diabetes or obesity to 1,379 controls with these conditions who hadn’t started GLP-1 therapy. According to the authors, both the GLP-1 users and controls were selected from a representative sample of people in four countries: France, the United States, Brazil, and Mexico. People’s average length of treatment was 18.5 months, and the researchers tracked weight changes over a three-month period.

As other studies have shown, GLP-1 users were more likely to lose significant weight and to report hair issues than non-users, the researchers found. Roughly half of people on a GLP-1 reported “unusual hair loss” compared to a third of controls, for instance. The former were also more likely to report scalp irritation. Doctors have long known, even before GLP-1s were around, that weight loss can sometimes trigger (usually temporary) hair loss. Interestingly, though, the researchers found that people’s likelihood of scalp problems wasn’t correlated with the amount of weight they lost. That suggests these hair issues aren’t entirely caused by weight loss while taking a GLP-1.

About 66% of GLP-1 users reported nail problems in general compared to 53% of non-users. The difference was far more substantial when it came to nail detachment, however, with roughly 39% of GLP-1 users reporting it compared to 9% of controls; 46% of people on GLP-1s also reported nail discoloration compared to 17% of controls.

What should this mean for GLP-1 users?

The studies were presented this week at the annual conference of the European Academy of Dermatology and Venereology. That means this research has yet to go through the formal peer-review process. The researchers also cautioned that their findings do not prove that GLP-1 drugs can cause nail or hair loss. But the research should warrant further follow-up, according to Taïeb.

“Both studies are cross-sectional and based on self-reported symptoms without dermatological examination, so they describe associations, not causes, and call for confirmation in prospective studies with clinical assessment,” he said. The team is now planning to conduct such a study that will proactively follow people at the start of GLP-1 therapy and objectively track their skin and hair health over time. That will allow the researchers to not only better study the causes of these problems but also whether they’re reversible.

In the meantime, there is reason to hope that these added risks, if genuine, can be managed effectively. The researchers found evidence that people’s nail issues were often better explained by their preexisting health problems, skin conditions, and nutritional deficiencies rather than the direct bodily effects of the drug itself, for instance. That should suggest it’s possible to treat these conditions without needing to discontinue GLP-1 use. Either way, it’s important for people on these medications to inform their doctors if they’re having these symptoms so action can be taken when needed.

“In practice: patients and clinicians should look at the scalp (itching, redness, scaling) and not only at hair shedding when starting or monitoring GLP-1 RA therapy; nail changes should be interpreted in the context of the patient’s overall health rather than automatically blamed on the drug, while persistent or progressive nail detachment deserves dermatological attention,” Taïeb said.

NTSB Preliminary Report: Prime Air 767 Runway Overrun [pdf]

Hacker News
www.ntsb.gov
2026-10-02 21:09:01
Comments...
Original Article
No preview for link for known binary extension (.pdf), Link: https://www.ntsb.gov/investigations/Documents/DCA26MA352%20Prelim.pdf.

White Nationalism’s Appeal Runs Deeper Than Partisan Politics

Portside
portside.org
2026-10-02 20:57:16
White Nationalism’s Appeal Runs Deeper Than Partisan Politics barry Fri, 10/02/2026 - 20:57 ...
Original Article

Unite The Right Rally, Charlottesville, Virginia, August 12, 2017 | screen grab

As Americans prepared to celebrate the nation’s 250th birthday on the morning of July 4, sobering images flashed across their phones and TV screens. Hundreds of masked men marched the streets of Washington carrying flags and chanting “Reclaim America!” It soon became clear from the logos on their identical baseball caps that the marchers were members of Patriot Front , a white nationalist group.

The march ended without incident. Still, it joined the 2017 “Unite the Right” rally in Charlottesville , Virginia, and the Jan. 6, 2021, attack on the U.S. Capitol – both of which culminated in deadly violence – as recent mass demonstrations with participants who openly embrace white nationalism. Members of a loosely organized movement that usually operates out of public view, white nationalists hold that shifting demographics pose an existential threat to the white race, which they view as culturally superior and deserving of power over other groups in the U.S.

Combined with the spread of white nationalist ideology from obscure corners of the internet into mainstream politics , these events have led many Americans to wonder: How extensive is white nationalism’s appeal among the broader public, and who is drawn to this extremist ideology?

We are political scientists who study political behavior , identity and social media . In recently published research in the scientific journal Nature, we offer some answers to these questions using national surveys we conducted with nearly 8,000 non-Hispanic white Americans from 2021 to 2025.

A crowd of men, some in helmets and protective gear, march with Confederate flags and other banners.

White nationalist demonstrators walk into Lee Park in Charlottesville, Va., on Aug. 12, 2017. AP Photo/Steve Helber

Asking Americans about white nationalism

We expected many Americans would be unfamiliar with the term “white nationalism,” so we defined it for survey respondents with a brief statement covering the movement’s major tenets: the purported need to maintain a white majority in the U.S., claims about the superiority of white culture, and calls for white people to hold more political and economic power than other groups. We measured endorsement of the movement’s ideology, not participation in the movement itself or approval of its tactics.

We randomly varied whether respondents saw the description before or after we asked whether they supported the white nationalist movement. Overall, nearly 5% of white adults – about 1 in 20 – said they did.

Those asked the question before seeing the description were less likely to express support at first, but their reported support rose after they read it. Along with several tests for possible response bias, this suggests that 5% is a reasonable lower-bound estimate of the true level of support.

Whether 1 in 20 white adults is a high or low number is a matter of interpretation. For context, it’s roughly the same as the share of American adults who identify as vegetarian or vegan .

Young white men

A number that jumped out to us was the share of white men under age 30 who support white nationalism: just over 13%, or about 1 in 7 . This was more than twice the level of support for white nationalism among women in this age group: 6%. Endorsement of white nationalism among men significantly declined by age : Only 2% of white men age 65 and over registered support.

Our discovery that white men born before the passage of the Civil Rights Act were more likely to reject white nationalism than those coming of age today may come as a surprise, as younger white Americans typically hold more tolerant views on racial matters than their elders. But young men have long been overrepresented in extremist groups , and white nationalism’s grievance narrative may resonate for young white men today as they fall behind previous generations of men on key indicators like finishing college, starting careers and forming families, as well as on their overall health.

Online life

Another key predictor of support: online social life that crowds out offline relationships. We asked respondents to call to mind their three best friends and then to tell us whether each friend was known only offline, only online or both.

We found that the more best friends known only online, the stronger one’s support for white nationalism. The internet and particularly social media are channels through which people may be exposed to white nationalist content. They are also places where alienated people may turn for the human connections they lack offline.

Interestingly, we did not find a general relationship between more social media use and more support for white nationalism. So while online life appears to be related to support for white nationalism, what matters in our findings is not how much people use social media, but whether their closest friends are known only online.

Individual and social strain

Strain, whether defined as undergoing personal difficulties or being exposed to wider social distress, is also a strong predictor of support for white nationalism. We asked respondents whether they had suffered from specific hardships over the past year, including the death of a loved one, losing a job, being the victim of a crime or getting divorced. Respondents experiencing more of these adversities were more likely to endorse white nationalism.

Social strain in the counties where respondents live – such as poverty, unemployment and population loss – emerged as another strong predictor of support for white nationalism, as did countywide rates of “deaths of despair,” which include suicide and deaths associated with drug and alcohol use. In other words, support for white nationalism was higher among people living in more distressed or economically depressed places.

More than just partisan politics

White nationalism is often cast as a political phenomenon by advocates, commentators and scholars, and indeed it is.

We found in our research that white Democrats and liberals are significantly less likely to endorse white nationalism than white Republicans and conservatives. However, we also found high levels of support among those who were unsure of their ideology or partisanship, suggesting a potential untapped reservoir of political support for white nationalist ideas.

Still, our findings show that support for white nationalism cannot be explained solely by partisan politics. Larger forces, including generational change, the ubiquity of the internet and the strain Americans are experiencing in part due to profound economic transformation are at work.

In our view, those who seek to understand and address this critical development in American public life must reckon with the broad structural shifts that have reshaped Americans’ identities, social relations and sense of economic security and that undergird white nationalism’s appeal to the broader public. The Conversation

Patrick Egan , Professor of Politics & Public Policy, New York University ; James Bisbee , Assistant Professor of Political Science, Vanderbilt University , and Joshua A. Tucker , Professor of Politics, New York University

This article is republished from The Conversation under a Creative Commons license. Read the original article .

Newgrounds.com – A community of games, music, and art

Hacker News
www.newgrounds.com
2026-10-02 20:55:25
Comments...

Where Is the Planet

Hacker News
whereistheplanet.com
2026-10-02 20:26:19
Comments...
Original Article

This page predicts the locations of directly-imaged exoplanets and brown dwarf companions. If there are any orbits you would like to see added, contact Jason Wang at jason.wang@northwestern.edu. Built by Jason Wang, Matas Kulikauskas, and Sarah Blunt. If you use this tool for your research, please cite the ASCL Record of it.

Things That Apparently Cause Cancer

Hacker News
www.breakthroughjournal.org
2026-10-02 20:23:11
Comments...
Original Article

It would shock you to know that Costco, according to a new Harvard School of Public Health methodology, caused 120,687 cases of cancer mortality nationally every year, simply due to living in close proximity to their warehouses. Private colleges accounted for 83,782 more. Harvard itself is responsible for 5,842 cancer deaths. Living near a Major League Baseball field causes almost 19 times as many cancer deaths as living near a Major League Soccer pitch. These conclusions are obviously absurd, but they come from the exact same methodology that researchers at Harvard use to “prove” that nuclear power plants cause cancer in nearby communities.

Earlier this year, Adam Stein and I called out a series of Harvard studies showing that nuclear power plants were associated with higher rates of cancer incidence and cancer mortality as nonsense. Now that we’ve spent the last few months replicating the analysis and expanding it, we are able to show that no matter what landmark you use, their methodology will show an increased cancer risk and mortality. The studies simply do not prove anything about cancer incidence.

Let’s recap: In December 2025, researchers led by Yazan Alwadi at Harvard’s T.H. Chan School of Public Health published a paper in Environmental Health that claimed to find that cancer incidence increased for people living closer to nuclear power plants in Massachusetts. In March, the same researchers published an expanded nationwide study claiming a similar result—this time looking at cancer mortality rates, rather than incidence—in Nature Communications . This was followed by a paper in the Journal of Exposure Science & Environmental Epidemiology that looked at associations of lung, breast, and colon cancers. Most recently, a study of total mortality, not just cancer, was published in the European journal Environmental Epidemiology .

The papers construct a “proximity score” based on distance from nuclear plants, up to 120 km in Massachusetts and 200 km in the national studies (about 75 and 125 miles). Every ZIP code or county inside those radii is treated as exposed, with closer locations receiving higher weights. If a ZIP code or county is within range of multiple sites, then the effect is cumulative. That means that a county that is in close proximity to one nuclear power plant could have a smaller score than another county that is further away from any one power plant but within range of many.

From this proximity score, the authors run regressions testing the outcome (cancer incidence, cancer mortality, or total mortality) on the proximity of the county and a collection of covariates. From this, they use the statistical coefficients, construct the Relative Risk for each county, age, and sex group, and compute the Attributable Fraction, from which they derive the number of attributed deaths from the nuclear power plant from 2000 to 2018. The authors state that their methodology and results provide the scientific basis and public health justification for an expanded research program.

But proximity is not exposure. We have methods of measuring exposure for nuclear power plant workers. While the radiation exposure of nuclear workers will always be greater than or equal to that received by the surrounding public, most of the closely monitored US nuclear workforce receive no measurable annual dose . When workers are exposed to radiation, the average dose received is only 2 percent of the occupational limit. If operators and workers who are on-site at nuclear power plants receive an annual dose between zero and one-fiftieth of the occupational limit, how is it possible that residents 5, 10, 25, 50, 120, or 200 kilometers away would receive any measurable dose from the same plant?

In talks, the authors hedge that their papers are merely ecological studies that show association, but never prove causality. Ecological studies are used to understand the relationship between outcome and exposure at a population level. This leads us to ask: what would the mechanism of exposure be? Well, according to the authors, we can just ignore the broader literature and physics, and instead make up exposure pathways and mechanisms. The use of an “ecological study” allows a lot of leeway in terms of explaining the broader world.

Over the last several months, we have replicated the results of these papers. The authors supplied us with eight lines of code and answered a couple of questions about the covariates, which did not replicate the results. Most of our replication was done through first principles combined with trial and error. Once we were reasonably close to the results of the first national study on cancer mortality, we took the methodology and applied it to numerous other landmarks.

Share

Everything causes cancer. Sounds cliché; maybe those California warning tags were right all along. But thanks to the methodology created by Alwadi et al., we can now prove that anything and everything causes cancer. Oh, sorry, that is too strong of an assertion. To use the authors’ words, since they have claimed their papers don’t prove causality, we can create an association between any physical landmark and cancer and, from that association, figure out how much cancer is attributable to that thing. Which is totally not the same as saying “that thing causes cancer.”

For instance, living near a private four-year university is associated with a 15-fold increase in cancer mortality when compared to living near a nuclear power plant.

Costco has the largest effect of all the locations we have tested. Over 2.2 million cancer deaths can be attributed to Costco; that’s more than 20% of all cancer deaths between 2000 and 2018. Hot dogs, bulk spices, and reasonably priced clothes come with a cost. But it goes to show that the authors’ choice of nuclear power plants wasn’t serious. Private 4-year colleges, including Harvard, are associated with the deaths of 1.59 million people, public 4-year colleges 626,212, and Superfund sites 895,258. Once you include the attributable deaths from state capitals at 132,389, we can account for 5.54 million cancer deaths, over one-half of all cancer deaths, between 2000 and 2018.

If we look at professional sports venues, the spread looks even more bizarre. Living near an NFL stadium was associated with 418,520 deaths, NBA arenas 471,794, MLB fields 944,769, NHL rinks 126,751, and MLS pitches 50,889. None of these sports have any kind of emissions that leave the playing field, except for the occasional home run or foul ball. The difference across the sports might be the kinds of concessions being sold. Baseball is right up there with Costco, selling copious amounts of hot dogs. Perhaps most surprising of all in the realm of professional sports is NASCAR, which accounted for 79,278 cancer deaths. One would expect the emissions, noise, smoking, and copious amounts of light beer to be more impactful than the emission-free ball and puck sports, but moving closer to a speedway might just save your life. With sports added to the mix, over 7.85 million cancer deaths can be attributed across a variety of landmarks; that’s 3/4s of all cancer deaths between 2000 and 2018.

The supposed 115,586 attributable deaths from nuclear power plants look like chump change to the real perpetrators. Harvard, when singled out, accounts for 111,003 deaths. That’s just one institution threatening every community within 200 km. Once again, we are left asking what the mechanism could be. Plus, our study only goes from 2000 to 2018; the school has been around since 1636, so the actual death toll could be much higher.

According to a webinar presentation given by Dr. Koutrakis, one of the coauthors, the supposed release of radioactive effluence from nuclear power plants comes from the occasional refueling of the reactors, when operators open the reactor core, and radioactive dust escapes from the fuel assemblies, past the containment vessel, through the outer shielding, and out into the world without ever being detected by the myriad number of dosimeters and radiological sensors throughout the facility. If we really are so terrible at measuring radionuclides (we are actually very good at it), then it is entirely plausible that Harvard could have unlawful access to special nuclear materials and the brains of Cambridge could have built a nuclear reactor under the squash courts without the proper licensure.

Cancer is a scourge across the human race. You have a ~40% chance of getting cancer sometime in your life. Which is why we should take cancer studies very seriously. If one is showing a surprising result, we want to get to the bottom of it. All kinds of claims can be made with statistics, but if we want to make good policies based on evidence, it is worth spending time making sure the evidence upholds its claims.

In our original rebuttal, we remarked that the studies can’t prove their assertions because they lacked proper control. It seems they also neglected to do any placebo testing. They chose nuclear power plants because a story could be built around that framework. When the researchers got positive results across our nation’s nuclear power plants, they didn’t check what their shiny new methodology would do using other landmarks. This is their pitfall: by taking the easy way out—getting results and making up a story around those results without double-checking their method—the authors could have no idea that what they were actually capturing was the methodology itself. In our replication and expansion, we have shown that choosing a number of sites and landmarks from capitals to warehouse stores can all yield a positive result without any reasonable pathway for American citizens to become exposed to radiation, develop cancer, and then die. The variations are random noise, all sloping in a positive direction. Even when using randomized outcomes, the methodology outputs positive results.

The papers by Alwadi et al. don’t show a novel mechanism by which nuclear power plants meaningfully contribute to cancer mortality; instead, they show an association of data points across 400 km-wide circles. 48,519 square miles, about the size of Mississippi, is a massive area to claim any kind of exposure. Most studies measuring distance-based exposures look at much smaller distances, such as under 10 km (an area of 113 square miles). The methodology is blind, which can be a feature in research areas needing to avoid bias, but in this case it is a bug; it doesn’t understand radiation exposure or dose. All it understands are its inputs: coordinates, proximity, covariates, and cancer deaths. From these, it can give you a number and even a positive result, but it cannot explain why that result exists. It is still an ecological study, one with a sophisticated statistical technique, but not a very useful one.

Sophisticated statistics can produce extremely precise nonsense when the design fails to identify the causal mechanism. Bad science like this is especially dangerous when applied to something culturally plausible that gets a positive result—it is easy to twist a result into a compelling narrative, especially if it confirms one’s biases. Nuclear power plants create energy through radiation, and radiation causes cancer; therefore, nuclear power plants cause cancer.

It is especially telling that no matter what landmark we applied to the methodology, we have yet to get a negative result. There is, in fact, a real chance that you simply cannot get a negative result from this method.

The attributable number of deaths from this methodology is probably zero, but the attributable number of bad papers is at least four.

Share The Ecomodernist

Discussion about this post

Ready for more?

Rex's Dino Store

Simon Willison
simonwillison.net
2026-10-02 18:57:38
Museum: Rex's Dino Store Located just before the turnstiles in the Grand Army Plaza subway station at the north end of Brooklyn's Prospect Park is this former newsstand which is now operated by a dinosaur. The density of dinosaur puns is exceptional. Tags: art, new-yor...
Original Article

2nd October 2026

Green Rex's Dino Store sign with yellow lettering advertising NEWS • ROCKS • EGGS • LEAVES • STICKS • NEST GOODS • LOTTO, beneath a window displaying dinosaur-themed newspapers and candy bars, framed by white subway tiles.

Museum Rex's Dino Store — Grand Army Plaza subway station, Brooklyn, NY

Located just before the turnstiles in the Grand Army Plaza subway station at the north end of Brooklyn's Prospect Park is this former newsstand which is now operated by a dinosaur.

The density of dinosaur puns is exceptional .

Posted 2nd October 2026 at 10:57 pm

I got targeted: Trying to get your credentials via a git post-checkout hook

Lobsters
frankwiles.com
2026-10-02 18:19:02
Comments...
Original Article

TL;DR Someone targeted me in an attempt to run arbitrary code on my laptop. I suspect in an attempt to gain access to my Github account and/or other REVSYS client related access since I have a metric fuck ton of it.

Be careful out there folks. They’re coming and they’re fucking sneaky!

I received a pretty normal project inquiry looking to see if we might be interested and available to work on a web app project in the Ed Tech space. We do a fair bit of that work and while we’re pretty booked up, I usually follow up on these sorts of projects in case the client is able to delay the project start until we have availablity.

I offered to setup a call with them and gave them a Calendly link. They asked that I read over the project overview and details prior to the meeting and sign an NDA.

Pretty normal stuff so far.

They then shared a Dropbox folder that had several folders of Markdown files. The project spec was pretty handwavey and light on details, but fleshed out enough for the MVP they supposedly wanted.

I initially missed the .git folder in that Dropbox link.

When I couldn’t find a NDA or NDA template in the folders I asked for them to email it to me.

They told me:

We keep it in the NDA branch and just to switch to it, fill it out and return it to them before our meeting

Red flag on a pole under a blue sky

This is where I realized what was going on and that this wasn’t a real project. I cruised on over to the .git/hooks folder and sure enough they had all of the *.example hooks in there and a single real post-checkout hook.

post-checkout hook, seriously?!?!?!

No one really uses those in practice so I carefully opened it up to see what it was doing.

It was using a Vercel app for command and control where it would download an OS specific binary, make it executuable, run it, and then delete itself.

I immediately alerted Dropbox and Vercel’s security teams so they can hopefully take these accounts down before they snag someone. Sadly, they also were impersonating an unspecting development shop owner as part of the ruse.

These assholes didn’t get me today, but I can easily see someone falling for this. Git is such a common workflow for us.

Be extra vigilant and watch your credentials like a hawk. They’re coming for you on some level.

Mastodon GitHub Bluesky X

Join my newsletter!

Get the occasional email from me when I write something new.

Struggling with architecture decisions or team dynamics? Ask me any tech, business process, or entrepreneurial question, and I'll do my best to help!

Submit a Question

frankwiles.com © 2026 Frank Wiles. All Rights Reserved.

Interview with Sjamaan

Lobsters
alexalejandre.com
2026-10-02 18:01:27
Our dear @sjamaan (website) is a CHICKEN scheme maintainer and professional Clojure developer. We got to know each other on IRC over a few months and discussed: CHICKEN Scheme and its 6.0.0 release Scheme's community and standardization process Postgres, the horrors of MySQL and (default) SQLite C...
Original Article
- Alex Alejandre

Peter Bex ( website ) is a CHICKEN scheme maintainer and professional Clojure developer. We got to know each other on IRC over a few months and discussed:

  • CHICKEN Scheme and its 6.0.0 release
  • Scheme’s community and standardization process
  • Postgres, the horrors of MySQL and (default) SQLite
  • Clojure

My ongoing core contributions mainly focus on the numerical tower code, keeping the in-core copy of the irregex library up-to-date with the upstream version (of which I’m a co-maintainer), squashing bugs and the odd security fix. I enjoy deep diving into odd corners of a code base to get a better understanding and then improve any weirdnesses I find. - His CHICKEN About


Beginnings

How’d you start computing?

I started computing when I got an old hand-me-down C64 which came with a “learn BASIC” book targeted at kids. Of course, I wanted to write my own computer games (which I never ended up doing). Later I got a real PC for my birthday and discovered my BASIC knowledge sort of transferred (with QBASIC in DOS).

Later, I learned about C, which I used for many years until at uni there was a course which taught Lisp (Scheme, really). We started with The Little Lisper/Schemer, and then SICP. I had a “functional programming course” before that which I almost flunked because it just didn’t make any sense (they used Concurrent Clean, because obviously teachers use their own language to teach regardless of its qualities - a common fault in academia). But the Little Schemer finally made it click. The teacher also showed how to implement objects using nothing but lambdas which I found awesome.

I actually studied AI before all the LLM bullshit. I much preferred the cleverness of the classical AI algorithms like A* search and genetic algorithms, but I haven’t really used them much in practice, only for my studies. I keep thinking I should use a GA for something, but no real use case so far. But even back then it was clear that neural networks were the future, though I found them boring because it’s just a bit of math (which I initially barely understood) and it’s basically a black box. I’m still grateful I studied AI because that’s the reason I came into contact with Lisp.

I might’ve looked into Lisp myself at some point (having been curious about it as one of the “foundational languages”), but I might not have had sufficient gumption to really dig in. After uni I got a job where they were using Rails in those somewhat early days (2006), so I learned Ruby as well. I liked learning Ruby and thought it was cool to see the power of Rails, but later got very frustrated with it because Rails really has strong opinions and my programming style didn’t seem fully compatible with it.

CHICKEN & Scheme

Why CHICKEN?

After that course, I tried using Lisp for every personal project. I first started with Scheme48 but it wasn’t very practical (though very elegant). I remember running into problems with the image not being big enough, running out of memory. I never really liked PLT Scheme (now Racket), coming into first contact with it through DrRacket, which felt very sluggish to me, though I do think DrRacket’s a cool alternative view of what an IDE could do. For example, when you hover over a variable and it shows you a line pointing to the origin of that variable.

CHICKEN was a practical and fast system, with a good community and acceptable license (I had quite a distaste for GPL at the time). CHICKEN is also one of two (as far as we know) implementations that use the Cheney on the MTA technique explained here . There were lots of rough edges, but I think wanting to address those is what enabled me to get so deep into the core. If everything is perfect, there’s not much to do, really!

So I started contributing to CHICKEN with some modest eggs ( CHICKEN packages ) at first, and eventually the core system. I mostly learned about Lisp internals by doing, mucking around with the CHICKEN core and trying things. I read Queinnec’s Lisp in Small Pieces, which deals with translation to C but leaves a lot undiscussed and distracts with OOP-heaviness. SICP has some good material. Then there’s Appel’s Compiling with Continuations, which is really short and to the point but still manages to be rather comprehensive; I love books like that.

In general, I just enjoy hacking on the core, even if I don’t have that much time to do so these days. I have to stress that I’m a slow learner. Building my understanding of CHICKEN was a process of many many years, and there are still parts of the system I’m not that familiar with (although I know my way around enough to get up to speed if needed).

What do those Lisp books lack?

Compiling with Continuations has only a brief section on the runtime system, so it doesn’t go very deeply into e.g. garbage collection and data representation. Lisp in Small Pieces doesn’t go into continuation passing style, IIRC. And its data representation isn’t very optimized. I like to blog about cool techniques that are undiscussed nowadays. For example implementing weak references and how to GC them efficiently. Other stuff I rarely see discussed is how to do FFI and cross-module optimizations, separate compilation and cross-compilation (which only a handful of Schemes even support) etc.

What’s Scheme to you?

In general, Scheme, to me, is a very clean language with a minimal core which facilitates experimentation. This is the fundamental tension of the standardization process - production-quality Schemes tend to grow in size, and there is value in standardizing that. But that also takes away the minimalism which makes experimental implementations possible. For example, Felix Winkelmann (@Bunny351, CHICKEN’s original author’s talk about Scheme implementation ) once started a Common Lisp (subset) implementation where he experimented a lot with types and flow analysis, resulting in CHICKEN’s type stuff.

What are your thoughts on the overall Scheme ecosystem, R7RS, SRFIs vs. implementation-specific libraries etc.? A few implementations don’t seem to care anymore.

I think the split of R7RS into a small and large language was the right thing to do, as R6RS was reviled by minimalists and found lacking by maximalists. In general, I’m a bit sad for R7RS - if some of the bigger community Schemes are essentially completely ignoring it, they’re doing something wrong IMO. Many, maybe even most Scheme implementations are essentially one-man shows. I suppose CHICKEN is also turning that way again, since we have lost quite a few contributors (mostly due to changing life situations) and it’s hard to attract new ones.

At the same time, the R7RS “large” project sort of went off the deep end, doing its thing without really caring about community buy-in. The churn of the R7RS large is also a bit too fast to keep up with. They’ve pumped out tons of SRFIs in a few years (actually, looking back at it right now it doesn’t appear like it’s that much, but it is certainly a lot faster than SRFIs used to be.) The SRFI process is open to submissions from literally anyone, for better or worse. I’ve noticed there have been a few new contributors to CHICKEN who submitted implementations for some of the newer SRFIs, so that’s good and I’m happy at least some people are bothering to do this.

CHICKEN 6 is base R7RS, but older CHICKEN code will keep working. We still support the old module syntax (which is the “native” one). The R7RS library declaration is essentially syntactic sugar for the core module syntax. R7RS-small is almost fully backwards compatible with R5RS, so there is no conflict there. Porting an egg to CHICKEN 6 usually requires only a few small adjustments because some (non R5RS) identifiers moved around between modules to better fit the R7RS style.

What’s CHICKEN’s development process like?

CHICKEN 4 was “hygienic CHICKEN”, which introduced the module system (and required overhauling the expander). This was all Felix, requiring a lot of deep internal knowledge about how macros interact with modules etc.

CHICKEN 5 was a community effort through and through. It was mostly a sanity and cleanup release where we did a massive reorganization of the modules (what lives where) to make it logical and matching R7RS a bit better. We discussed this during an IRL meetup and continued for the months after. (Community is a strong advantage of CHICKEN!) We also added a numeric tower .

CHICKEN 6 was basically cut off from Felix’s branch to make UTF-8 handling sane and consistent (like Python 2 -> 3, but way less disruptive.) There’s a strict separation between strings and bytevectors, with changes to ports and other I/O as well. We took the opportunity to integrate the R7RS egg into core so it’s more “native”. Strings in the FFI should be more efficient because there’s no needless copying anymore.

CHICKEN 6.0.0 was held back by a bug causing heap corruptions (do view the patch’s description!). We had been looking in the completely wrong spot. These heap corruptions appeared somewhat randomly, but only with the CHICKEN wiki server, not with the plain web server serving simple files or even the entire Awful framework. We strongly suspected the Subversion client library (which the wiki uses as a backing store for the content), and we’d found other issues in there as well (it’s kind of hairy callback-heavy code due to the design of libsvn).

The wiki is a rather small program and the rest of the web stack seemed to be fine, so we suspected the svn client lib. But I whittled down the code of the wiki to almost nothing and it was still failing. When I commented out the URI normalization code (you get redirected when opening a page that’s behind a symlink, so as to get a canonical URL that points to the original file) it suddenly stopped crashing!

That normalization code didn’t appear to do all that much, so we quickly pinpointed it to be read-symbolic-link. A quick glance at the code in the core system confirmed it was totally borked because of a change made for CHICKEN 6’s UTF-8 transition.

Now that 6 has come out, what’s next?

Regarding the goals for CHICKEN, there are several things I’d like to work on. One idea I had is to teach the compiler about unsafe intrinsics using a “prelude”. Because Scheme is a safe but dynamically typed language, there’s some overhead in the intrinsics, say if you call “car” on a non-pair, it throws an exception. If the compiler can deduce that the object you pass in must be a pair (maybe because you checked it before with pair? , or called car or cdr on it before), it replaces the call to an unsafe, unchecked version. But this is all very ad-hoc, and not extensible by the user. My idea was to have a separate definition which splits the unsafe operation from the “typechecking prelude”, which can be inlined at the call site. This way, if multiple checks need to be done, it’s not all or nothing. We can elide the unnecessary checks and only do the necessary ones which might extend to user code, too.

Another idea relates to the way we handle dates and times - we have some stuff in core to access the POSIX functions but it’s messy and (IMO) mostly unusable. The alternative is SRFI-19, which is a beast because it has support for multiple calendar systems, localization etc. Might be nice to have something minimal (maybe English only) in core, so you have a common type that gets used everywhere (handy when sharing objects between libraries without building in a big dependency on SRFI-19). You can then use it for parsing timestamps in common protocols, say.

Postgres & SQLite

What domains do you like or know the most about?

  • Web stuff: I maintain the HTTP and URI implementations for CHICKEN
  • Some CHICKEN internals: GC, macro expander and Irregex implementation
  • Performance optimizations: though not an expert, I’ve done quite a bit and always thoroughly enjoy it
  • Postgres: Although I haven’t gone deep into the internals, I’m typically the go-to guy for (Postgre)SQL questions in companies I’ve worked at. Funny, because I initially flunked the DB/SQL course at uni and didn’t grok SQL at all
  • Distributed systems: though I’ve worked on them for 6 years at work, you’ll want to avoid them like the plague if at all possible. It can be hard to reason about the behaviour of the system at large, and you can’t really abstract it away

I’m trying really hard to think of something I’m truly excited about. The biggest positive I see right now is the push for digital sovereignty. I sincerely hope this will change how people deploy tech, maybe in a more mindful manner. More open source, less dependence on foreign (and hopefully big tech in general) products. But vested interests and inertia will be hard to overcome and really bum me out.

Why Postgres?

I properly learned about DBs at a calendar startup using Rails. We had instantiated repeated events in the DB and the event would sometimes need to be updated. At first, we were fetching models in a loop and updating them one by one, excruciatingly slow. We eventually discovered the bulk update (I think you even had to call into the DB directly because Rails didn’t offer that at the time). That made everything click for me - the importance of performance and the usefulness of SQL. At the time, MySQL was still Rails’ default and I got into MySQL character set hell a few times for a CMS we used. Later, another Rails project required such massive amounts of data to be stored (computational fluid dynamics simulation) that MySQL simply crashed every time I tried a bulk import.

Looking into alternatives, I found Postgres handled it without any problems. When I learned that Postgres doesn’t have any of those braindead misfeatures MySQL has. For instance, UTF-8 characters get verified on storage so you can’t get into character set hell like in MySQL so easily, and it actually allows DDL statements in a transaction, so you get transactional migrations that apply atomically. That was a real eye opener at the time. I was sold!

Postgres is a lot more regular and well-behaved on basically anything, and it has no strange limits (e.g. in MySQL, you can’t even put an index on text columns with indeterminate lengths). MySQL allows you to store an empty string in any non-nullable enum column. Makes no goddamn sense to me! The DB is full of footguns like that. I should stop ranting - talking about MySQL really makes my blood boil. Also, I’ve come to rely on more “advanced” features like LISTEN/NOTIFY, array storage, window functions, CTEs etc. However, I’ve found that stored procedures don’t really work that well - “real code” is more flexible as it doesn’t require finicky migrations to keep in sync.

I’m not a big fan of JSON in my relational DBs, but I have been known to use it when storing arbitrary data or actually putting JSON results (from APIs or other stuff) in the DB.

We use DataScript on the client via ClojureScript, but I don’t really grok it and don’t touch that part of the code often enough that it really sticks, so every time I have to deal with it it’s an exercise in frustration.

Even though I grok it nowadays, SQL is a really badly designed language. I’ve seen several projects that try to come up with a better query language, but I’ve given up hope that they will succeed as SQL is too entrenched to get rid of.

Why not SQLite?

I’ve used it a handful of times. The experience was always mostly one of frustration. It feels a lot like MySQL, with unsafe and stupid defaults.

IIRC it’s value-typed and (by default) doesn’t check types, so the type of a column is basically completely ignored. And you can’t alter a column, IIRC (or maybe only a few changes). I even remember a version (maybe WebSQL?) where you can’t even DROP a column. Anyway, it’s not worth any brain cycles for me to deal with that shit.

On REPLs

What do you think of Clojure?

Clojure’s influence is strong on things like Carp and Janet . I wrote about my impression of Clojure on my blog, but in a nutshell I don’t like the “everything is a map” approach and nil punning really turns me off as it makes bugs harder to find. I do like the fact that it has mostly purely functional data structures. I’m not sure I like the syntactic “heaviness” of Clojure - things like [] for vectors and {} for maps. The lack of cons cells is also a bit weird but I see how it simplifies list handling code (even though Clojure does not typically deal with lists much!) Most importantly, it revitalised the interest in Lisps!

In your article on Clojure:

never fully bought into the REPL style of developing. Sure, I experiment all the time in the REPL to try out a new API design or to quickly iterate on some function I’m writing, but my general development style tends more towards the “save and then run the test suite from an xterm”.

Normally, we hear such things from people who think using the REPL means typing into the little terminal box instead of sending code from files into the REPL with a hotkey, so it surprised me to read it from a veteran.

I do consider the REPL an essential tool for experimentation and debugging. But I struggle to keep track of what’s running in the system versus what I see in my editor. With Clojure, you can’t do without the REPL because it’s so doggone slow to start up that it would be impossible to just run something on the CLI over and over. So at work, I spend 100% of my time with the CIDER REPL. I do find myself closing and reconnecting several times a day though, because I can no longer trust the REPL state matches my editor buffers. One thing that gets me every time is if I delete a test from my buffer and then re-run the entire suite, it’s still there in the REPL (obviously), same thing with multimethod implementations. But overall, I really couldn’t live without a REPL.

What do you think of snapshot testing ?

Snapshot testing’s an interesting approach. I think we have a few testcases in CHICKEN where we do something like that - we run the compiler and capture the output of the compiler (which is mostly type warnings) and check that it hasn’t changed with a simple diff on the output and expected output. I’ve also used something like this for great effect while refactoring and optimizing code - simply keep reference output in a file. For instance, in one case when working on a project which had no test suite, I used pg_dump to dump an “output table” as a reference and then went to town on the codebase, knowing I would immediately see if my optimized algorithm differed from the original. I also approach in another inherited project without test cases to refactor.

Real Life

What makes you happy?

I’ve mentioned before that I really enjoy performance optimizing code, but I also really enjoy refactoring and investigating vulnerabilities. Systems that are understandable and hackable make me happy.

Outside of programming (which more often frustrates me than makes me happy TBH), my family makes me happy. It’s a great source of joy to just relax and be with my wife and children. I’ve been trying to get back some balance in life, spending more time AFK, and I pay attention to my health a bit more as well.

I’m not super young anymore (43) and reading about age-related issues like sarcopenia made me realize we really tend to neglect our bodies with our sedentary lifestyle, especially us programmers. My mother has osteoporosis and I see how she struggles just doing basic things. I don’t want that for myself, so I’ve picked up weight lifting as a way to combat those age-related issues, so I can become old in a healthy way. The prognosis is for most of us to live up to 90 or so by now, so I’m not even at the halfway point. But these issues start cropping up at around 50-60.

How do you approach raising kids?

I’m still getting my bearings TBH. My kids are only 3 (going on 4) and 15 months. Raising kids is probably the hardest thing I’ve ever done.

I don’t really know yet if I want to teach them programming. I definitely want to raise them tech-sceptical, when they’re big enough, to teach them the dangers of social media and the importance of privacy. If they show an interest, obviously I’d teach them Scheme. It’s the perfect language for teaching! But maybe something like Logo first.

Claire Valdez and Darializa Avila Chevalier on Building Power, AOC, and 'Squad 2.0'

hellgate
hellgatenyc.com
2026-10-02 17:35:43
Clairializa is in the room with the Hell Gate Podcast....
Original Article
Claire Valdez and Darializa Avila Chevalier on Building Power, AOC, and 'Squad 2.0'
Democratic nominees for Congress Claire Valdez and Darializa Avila Chevalier (Hell Gate)

Podcast

Scott's Picks:

Give us your email to read the full story

Sign up now for our free newsletters.

Sign up

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to Hell Gate.

Your link has expired.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.

Open-sourcing AstaBrief, the fast report-generation model in Asta

Hacker News
allenai.org
2026-10-02 17:25:25
Comments...
Original Article

Language models can already help researchers search the literature, synthesize evidence, and work through complex questions. But scientific work places particular demands on these models—answers need to stay grounded in evidence, the models need to preserve what the evidence actually supports rather than quietly broadening a study’s conclusions, and researchers need to be able to verify the final outputs.

We see that in how scientists use Asta , our agentic platform for scientific work. Instead of simple keyword searches, users often bring substantial context and many constraints—for example, asking Asta to compare approaches across a body of literature while accounting for a particular method, population, or setting. Many also return to generated reports later, treating them as working research artifacts rather than one-off answers.

We wanted to help scientists generate cited reports faster, with a model they could download and run themselves. To do that, we tested whether a small, open model trained specifically for scientific report generation could match the report quality of the proprietary models we were using, while reducing generation time and serving costs.

We built AstaBrief 8B , a model that turns a research question and retrieved literature excerpts into a cited report. AstaBrief is available in Asta’s Generate a report feature today as Fast mode alongside Claude-powered Thinking mode, and we’re also open-sourcing it and the training data so others can study, reproduce, and build on our approach.

Developing AstaBrief required tens of thousands of real research queries, citation-focused filtering, preference data, and a redesigned report-generation pipeline that writes the full report in one pass rather than section by section. The result is nearly an order-of-magnitude reduction in report generation time compared to the proprietary models we tracked—across the full Asta pipeline, Fast mode averages 51.1 seconds per report compared with 178.5 seconds for Thinking mode, about 3.5× faster.

Together, those efficiency gains made AstaBrief a useful test case for a broader goal: building open language models that can be adapted to the specific demands of scientific work.

Open weights will also let institutions run AstaBrief on their own infrastructure, which is necessary when research questions reveal sensitive or unpublished work. Alongside the model weights, we’re releasing an example workflow that researchers can adapt to create reports from their own PDFs , providing a starting point for local report generation

This post covers how we trained AstaBrief, what we learned about grounding it in scientific evidence, and which parts of our approach we think can carry forward to future models for science. Most of the training and evaluation described was completed in 2025, so the proprietary models used to generate training data and as comparison points reflect the frontier at the time. We haven’t rerun the full evaluation against today’s frontier models; the results below are best read as evidence about the particular training and system design choices we tested.

Training the model

Our goal with AstaBrief was to build an open-weights model with all the qualities that matter most for long-form scientific synthesis: answer quality, relevance, structure, and citation grounding. We started from Qwen3-8B and focused most of our effort on the post-training data, evaluation, and surrounding report-generation scaffolding.

Adapting general-purpose models for scientific work – and training new scientific models from scratch – is something we're exploring broadly across Ai2. Through NSF OMAI , a U.S. national initiative led by Ai2 to build fully open AI infrastructure and models for scientific discovery, our researchers are working directly with scientific communities to understand what they need from future open models and where today's general-purpose models fall short. That includes studying how needs differ across scientific fields and workflows, with more findings from that research to share in the future.

Recent work, including our DR Tulu , has shown that reinforcement-learning-based (RL) methods can improve long-form report generation for open-weights models, especially when judge models are involved in the training loop. We considered that path for AstaBrief, but ultimately focused on a simpler recipe built around supervised fine-tuning (SFT) and direct preference optimization (DPO).

RL-based training can be unstable and expensive. We wanted to see how far we could push report generation quality with a cheaper, more operationally manageable setup—one that's also easier to debug and iterate on.

That made the quality of the training data especially important. Rather than relying on a more complex optimization method to compensate for noisy examples, we spent much of the project figuring out how to generate, select, and filter examples that actually demonstrated the report-writing behavior we wanted.

We also wanted AstaBrief to be faster so that users could get preliminary reports quickly that they could then iterate over in subsequent turns. For speed improvements, we decided to train AstaBrief to directly generate the final report in one pass given a user query and relevant retrieved snippets, bypassing the expensive snippet summarization and clustering stages our Claude-based Thinking mode uses and not writing out the answer section-by-section. Interestingly, we found it was possible to do so without sacrificing performance.

Collecting SFT training data

The training pipeline began with real user queries submitted through the system described in our paper “ Synthesizing scientific literature with retrieval-augmented LMs ” and ScholarQA , the framework that now underpins Asta’s Generate a report feature. Rather than training only on synthetic prompts or benchmark-style tasks, we wanted AstaBrief to learn from real queries from real scientists.

Our research suggests that scientists often ask different things of language models than users do of general-purpose chatbots or traditional search tools. In our analysis of hundreds of thousands of Asta queries , expert researchers frequently supplied substantial context, multiple constraints, and relationships between concepts rather than relying on short, keyword-style prompts.

More recent Asta user studies have also surfaced differences in how researchers want AI involved in their work—some are comfortable using models for ideation or experimentation, while others prefer a narrower role in synthesis, literature surveillance, or pattern-finding. Across those differences, participants want clearer source traceability, more visibility into what a model is doing, and greater control over the context it uses.

We filtered the user logs we collected for quality, relevance, and privacy, stripping out beta-tester and bot traffic, dropping queries that were too short to be meaningful, and using an LLM-based filtering pass to catch non-English queries, non-scientific requests, and prompts containing personal information. That left a pool of 90K research-focused queries.

For SFT, we generated full-report target outputs from the filtered queries using the multi-step ScholarQA pipeline behind Asta's report generation. The pipeline retrieved relevant literature, organized the material into sections, and used a backing report-generating model to synthesize the evidence into a cited report. We drew on a mix of proprietary systems: Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1. After quality filtering, this yielded 47K usable training examples.

Creating DPO pairs

DPO required a different kind of training data. Instead of a single target report per query, we needed pairs of reports with one preferred over the other.

We built those pairs from a separate subset of queries not used during SFT data generation. One report per query came from the existing ScholarQA pipeline, typically backed by Claude 3.5 Sonnet or 3.7 Sonnet. The competing report was generated by feeding ScholarQA's retrieved literature excerpts to a different model: o3, o4-mini, DeepSeek-V3, or DeepSeek-R1, depending on the example.

Two judge models – GPT-4.1 and DeepSeek-R1 – compared each pair and picked a winner. We ensured that LLM judges were aligned with human preferences (95% agreement) and only kept pairs where both judges agreed, which gave us a cleaner preference set and cut much of the noise that typically shows up in preference data generated at scale.

After quality filtering, the final DPO dataset came to about 6K examples.

Using multiple generators and requiring agreement between two judges gave us a relatively simple way to construct preference data without treating any single model’s output or judgment as ground truth.

Filtering data for better attribution

Our main evaluation target was SQABench-CS2 , a set of 200 user-written computer science research questions. We tracked four metrics throughout the development of AstaBrief:

  • Rubric score , which measures how much necessary content is covered by the report.
  • Answer precision , which measures whether each paragraph is relevant to the question.
  • Citation precision , which measures whether each citation supports the claim it's attached to.
  • Citation recall , which measures whether the report's claims are fully supported by the citations provided.

For our final model, we also ran secondary evaluations: DeepScholarBench , a 63-query benchmark for long-form research synthesis built from recent ArXiv papers, and two separate pairwise evaluations against reports generated by the Claude-powered pipeline—an LLM-judged comparison on SQABench-CS2 and a small human study.

A report can sound polished and complete while meandering from the question or attaching citations to claims from which the underlying evidence doesn't follow. For scientific synthesis, we needed to measure those behaviors separately. But citation support is only part of scientific faithfulness—a model can cite the right study and still make a stronger claim than the study itself supports. This can happen in subtle ways , for example, turning a finding about a particular sample into a generic claim about an entire population, shifting a result reported in the past tense into a present-tense statement that sounds more universally true, or turning a descriptive finding into a recommendation for what clinicians, policymakers, or researchers should do.

Those kinds of generalizations are especially important for scientific report generation because each step can broaden the apparent scope of the evidence without introducing an obviously false statement. A cited sentence may therefore be technically related to its source while still overstating what researchers actually established. Our development metrics focused primarily on relevance, coverage, and citation grounding; a richer evaluation of scientific report writers should also test whether they preserve the scope and strength of the claims in their sources.

Our first SFT runs improved overall content quality, but they still lagged behind our Claude-powered report generation pipeline on answer precision and citation quality. In other words, the model got better at writing reports, but it still wasn’t grounded in evidence as consistently as we needed for scientific synthesis.

That pushed us to spend more time on data quality. We tested four statistics-based filters to identify weaker synthetic training examples:

  • Output-to-input token ratio . Answers with very high ratios were often noisy because they were generating a lot of text from too little evidence.
  • Citation relevance . For each synthetic report in the training set, we averaged the retrieval relevance scores of its cited papers. Low averages suggested the report was relying too heavily on lower-ranked evidence.
  • Citation density . We measured the share of statements that had at least one citation. Low-density reports often had large stretches of unsupported text.
  • Citation diversity: We measured the share of papers cited in the answer, given the set returned by the Claude-powered report retrieval pipeline. Low scores suggested the report was overly reliant on a few papers.

The strongest gains came from filtering out synthetic reports with low citation density; more aggressive filtering, filter combinations, and learning-rate sweeps didn't add meaningful gains.

That was one of the clearest lessons from the project: more elaborate filtering wasn’t necessarily better. A relatively simple signal – whether the synthetic reports consistently cited their claims – was more useful than several more complicated combinations we tried. Scientific specialization, in other words, isn't necessarily a matter of adding more scientific text to pretraining; the composition and quality of post-training data and whether it demonstrates behaviors like grounding and attribution can materially change how the resulting model performs.

That focus on grounded, useful output also lines up with what we’ve heard in Asta user research. Participants note that generating more text isn't necessarily more helpful; they want concise synthesis and enough source traceability to review and verify results without wading through unnecessary outputs.

Once we had a stronger SFT checkpoint, we ran DPO training on top of it. That stage pushed performance further, bringing AstaBrief within range of the Claude-powered report pipeline in Asta and DR Tulu on report generation.

Validating the approach

Because this model was intended to work as part of our agentic Asta report generation framework (not necessarily as a standalone model), our main question was whether AstaBrief could preserve the report qualities we cared about while enabling a substantially faster and cheaper report-generation pipeline. In other words, we weren’t only asking whether the model could match a stronger proprietary model on individual benchmarks; we wanted to know how much of that quality we could retain with a much simpler system.

In the evaluations we used during development, AstaBrief was competitive with the Claude-powered pipeline and DR Tulu across several measures of answer and citation quality. The chart below shows the LLM-judged comparison—in a separate 14-question human study, three scientific researchers each contributed 4-5 questions and ranked reports from the three systems on overall preference, completeness, relevance, organization, and citation accuracy (with ties allowed). On overall preference, DR-Tulu wins, but two of the three researchers prefer AstaBrief over other systems on citation accuracy metrics, demonstrating the utility of our SFT data quality filters.

These numbers are best read as validation of the engineering approach at the time we developed it, rather than as a claim about where this particular base model sits relative to today’s frontier. The model ecosystem moves quickly—the data construction, attribution filtering, and serving lessons are the pieces we expect to generalize.

Validating the usefulness of AstaBrief in Asta, Fast mode has shown encouraging early usage. Among 374 Asta users who’ve tried it, 29.1% have used it for two or more days, and users on average generate 3.67 report threads with it. Twenty-three percent of users who tried Fast mode continued using it and never switched back to Thinking mode for future threads. An additional 18% switched between Fast and Thinking modes depending on their goals, using Fast mode for ~40% of their threads.

While feedback is generally too sparse to draw strong conclusions, we see that Fast mode receives positive feedback at a similar rate as Thinking mode (84.2% versus 85.2%).

Where this goes next

Asta's report generation is the first production use of AstaBrief, giving researchers an open-weights Fast mode alongside the existing Thinking mode. Because the model is open weights, institutions can deploy it on their own hardware, including behind their own firewall, without relying on a proprietary model API for report generation.

In Asta, that also means we can study and improve this part of the report generation pipeline directly while preserving Thinking mode as an option for more compute-intensive tasks.

There's more to do. We're exploring more fine-grained preference learning, stronger RAG-plus-RL approaches, multi-turn and multi-tool capabilities, additional scientific data sources, and query decomposition. We're also interested in evaluations that go beyond whether a claim has a supporting citation to ask whether a model preserves the evidentiary—both to better capture the quality of the report as a research artifact and to ask whether a model preserves the evidentiary scope of its sources. That includes qualities such as concision and organization, as well as whether the model turns sample-specific findings into broad generalizations or descriptive results into recommendations.

AstaBrief is one experiment in a longer line of work on language models for science, from ScholarQA and DR Tulu to future versions of Olmo beginning to take shape now. The lessons here – especially around training data, filtering, and evaluation- can help inform what we build next.

Try Fast model today in Asta

, or

download AstaBrief from Hugging Face

.

Join us

At Ai2 we’re building the future of transparent, open-source AI — built in the open to empower scientific progress and fundamental understanding of this world changing technology. We’re not here to make profits, we’re here to make sure benefits of AI are shared widely and for the benefit of humanity. If this appeals to you, please take a look at our open roles.

Open roles

Subscribe to receive monthly updates about the latest Ai2 news.

The Harness Is the Company

Hacker News
blog.sshh.io
2026-10-02 17:06:56
Comments...
Original Article

Every SaaS business will become a harness around a model, whether or not they’ve realized it yet.

Narrowly, folks will often associate a “harness” with frameworks like LangGraph or coding agents like Codex, Claude Code, or OpenCode, which wrap a stateless model API in enough tooling and state that you can actually get work done.

I’m using the term “harness” broadly to mean all the infra, interfaces, context, and state that surround a stateless LLM. A harness can be composed of adapted or task-specific sub-harnesses (with the orchestrator being called a “meta-harness”). A “software factory” is a harness whose parts are smaller harnesses — one that writes specs, one that writes code, one that reviews — plus something on top deciding what runs when.

If you accept this broader definition (or replaceAll “harness” with whatever phrase you’d prefer), I suspect for many software service businesses you’ll see this trajectory:

  1. They sell software services built the traditional SaaS-y way.

    1. No harness.

  2. They sell software services, but the engineers pair with agents to get the work done. Increasingly other functions like product and sales pair with agents for productivity.

    1. Individuals operate harnesses.

  3. They sell software services with many core tasks moving to background agents running in the cloud (a laptop can’t run twenty of them). Engineering, product, and sales all trigger these agents (i.e. write the prompts) and review the outputs.

    1. Individuals orchestrate harnesses.

  4. They sell software services with many core tasks moving to proactive background agents, with engineering, product, and sales moving to reviewing agent outputs. Consistent reviews move to sampled reviews. Increasingly the agents decide what to do proactively, instead of humans designing the work up front.

    1. Harnesses orchestrate individuals.

Following this trajectory, you’ve actually turned your company into a harness. The work to produce the software service has moved from people to the harness.

  • Core tasks are done by agents and the “product” is now entirely model output, with the company supplying the context, integrations, and human-facing review interfaces.

  • The relationship between software and org structure has inverted as the org chart becomes a question of where to put people so the harness gets the most taste and judgment out of them. Humans are part of the harness.

  • The company’s domain knowledge, tooling, permissions, review loops, context, etc. all become the business harness.

It’ll be tempting to think that products that are primarily crafted and reviewed by AI will be inherently low quality and that building through a harness instead of through people will produce a slop factory at scale. I think that belief assumes a company run this way is lights-out .

The core mitigation is actually having the harness pick where human inputs matter most. For example:

  • A product decision could come from an agent fanning out questions to reps in customer meetings and then synthesizing a product demo for the product lead to review

  • A feature suggestion pulled from a customer meeting is turned into a major architectural decision presented to an engineering taste-holder by the agent.

  • A UI redesign kicks off after aggregating feedback, and after testing a few variants, the top options are presented to a design taste-holder.

A good harness maximizes value to the customer while spending human attention — employees’ and customers’ — only where it’s needed.

If you are still skeptical, that’s fair. Even with today’s frontier models and a well-crafted harness, it’s quite difficult to trust agents to handle the outer loop like this ( we’ve spent a lot of time trying ). I just don’t think it’s worth betting that models won’t be able to eventually do this, especially as more of this planning-and-review work gets broken into tasks with verifiable rewards that labs can train on .

To explore more on what an organization of “taste-holders” could look like, see The Transposed Organization . This post builds quite a bit on the ideas laid out in that article.

Depending on the service, the differentiation a company has often comes from things like trust, distribution, efficacy, and domain context. In a proactive background agent world, your ability to construct the harness — how it learns, what it watches for, how it interfaces with human taste-holders, what systems it integrates with — will more and more be how you maintain that differentiation.

The harness is now what shapes what must be true for work to ship (~trust), how fast and how the products land (~distribution), the speed and context of feedback loops (~efficacy), and how institutional knowledge is ingested and maintained (~domain context).

Harnesses go from internal tooling you’d happily buy to something you’d no more outsource than your product-eng org or your GTM team. This is distinctly different from the pre-AI world, where outputs were mostly bounded by the humans using the software to get things done.

I think in-house AI developer tools (like the ones we’ve seen from Ramp, Stripe, DoorDash, etc .) are the beginning of this.

For AI-pilled companies, waiting for an SDLC tool vendor to add an integration, support a certain interface, or reach a level of cost-efficacy increasingly bottlenecks their ability to build and maintain their product. This is especially true in the short term for third-party tools that can’t yet run an enterprise’s whole software factory for it (often because the tech stack is too bespoke, governance too restrictive, critical feature support too slow, or due to a preferred cost model).

I don’t expect everything to be built and maintained in-house. Rather, companies should own the top-level harness — the one that decides what to build and reviews what comes back — and plug vendor products into it for specific workflows. Eventually a third party will get pretty good at an enterprise-level “spec to tested pull request” and at that point a company can swap out that part of the software loop with that product while still maintaining the agent(s) that write the input spec and handle the next steps from pull request output.

If somehow the entire outer loop can be done by a third-party harness (i.e. running the entire business via proactive background agents as a service), then I’d argue the business has now been commoditized.

If this is the right mental model, you should expect to see:

  • An unusual amount of in-house harness building on both the build side and the sell side

  • Org structures and individual roles being reshaped around their place in the business harness

  • AI-native startups beating incumbents in domains where the “moat” can be easily harness-ified

  • All software a software company uses (on or tied to the core build or sell paths) needing to be headless so the outer harness can run it

Discussion about this post

Ready for more?

Respecting your users' dread of the clankers

Lobsters
thoughtbot.com
2026-10-02 17:03:08
Comments...
Original Article

People today have a spectrum of opinions on AI: we range from all-in, overly enthusiastic vibe-lifers to luddite curmudgeons. I personally have been from one side to the other and back over the course of this week.

Us AI skeptics have good reasons for our fear/hatred/disgust even as we use LLMs daily:

  • adverse effects on the environment and the economy
  • data misuse
  • training on works of unwilling authors
  • bias, hallucinations and lack of accountability
  • stealing our jobs
  • and more!

So please build your AI products to let me opt in.

Opting in progressively

The idea is to cater to everyone on the AI love/hate spectrum. You do that by giving me granular control over just how much AI is used, and just which personal data is shoveled into the firebox. I want to

  1. see how my data will be used by AI
  2. see what benefit I get
  3. grant permission

And I don’t want to do that by reading through your Terms of Use and Privacy Policy. It needs to be in context because I will decide on a case-by-case basis.

I want to grant zero permission up front. I want you to ask me for permission when you need it. No AI until I say so.

iOS popup requesting permission to send notifications, with allow/don't allow buttons

In mobile apps you already use this pattern to request extra permissions like location sharing and push notifications. You know that if you pop a notification as soon as the app opens, I’m going to say no. So you wait until you can demonstrate usefulness. We call this pattern a “soft ask” or “pre-prompt” or “pretty please”.

It’s the same way with AI: if your registration screen says “hey BTW we’re going to send all your data to LLMs and force you to chat with our AI agent!” I will have misgivings. But if you wait a bit and say “hey, we see that you’re stuck on this form - want to enable our agent to help?” - I will probably say yes, and I will probably say sure you can consume my data for it. You have shown that you respect my choice and you have demonstrated value.

Once I’ve given the OK to use AI, that doesn’t mean I’m OK with you shipping it every single byte of data you’ve compiled about me. Sure, it’ll make your AI more productive and accurate if it can access all my PII, credit card statements, and health information. I don’t care. Again - you need to prove value.

I haven’t granted wholesale permission with the click of one checkbox and submit button. You need to ask for permission progressively, and then let me retract my permission later.

How your agent can ask for permission progressively

It’s all about the context and prompting. Here are some techniques I’ve used or observed:

Conditionally build the list of agent tools or skills based on what the user has given access to.

In your LLM system prompt, give a list of the data items that the user has and has not consented to sharing.

Prompt your agent to ask for permission. Something like “This user (has/has not) given you access to their financial accounts. You may ask for permission to access the accounts - direct the user to their Settings screen to enable access.” or maybe “This user has granted read-only access to their document. If you need to edit it, please ask for write access, and explain exactly what you need it for.” Make sure the agent knows that “NO!” is a valid answer. It shouldn’t guess or hallucinate if it doesn’t have the real stuff.

The important parts are to let the LLM know

  • what does it currently have access to?
  • what could it get access to?
  • how can it use that data to benefit the user?
  • how can the user grant access?

Pre-approval

If you can’t do just-in-time consent, then you’ll need to get consent ahead of time. And remember - keep it granular and tell me how I benefit. If I’m linking health data to your app, you need more than just a single “yes please deliver my entire disease history to Anthropic” “Allow additional information sharing with our business partners" checkbox. Instead consider something like

  • [ ] allow our AI agent to book appointments on your behalf
  • [ ] share anonymized appointment notes with LLM partners so we can give you a summary
  • [ ] review X-rays and other radiology imaging with AI to flag abnormalities

Ask for permission to use data with AI at the place where you’re collecting that data. If I’m filling out a registration form - that’s where I should grant permission to share the data in that form. If I’m connecting my calendar, that’s where I should grant permission to let an AI manipulate my calendar.

Clawback

If I can give your LLM permission to slurp up my PII, I better be able to revoke that permission! So give me a user interface for that. Show me what the AI is consuming today and what it’s used for and what the effects will be if I uncheck the checkbox.

If I change my mind and want to remove AI access to my data, that means really removing it. Clear your prompt caching. Sanitize agent conversation histories. Delete your logs. Remove AI access to those tool calls.

Do not tempt me! I dare not take it, not even to keep it safe, unused. The wish to wield it would be too great for my strength. –Gandalf

movie still of Gandalf with caption "Don't tempt me, Frodo!"

With chat-style agentic AI, you have the new ability to ask for lots and lots of freeform sensitive data. It’s a free text input field and I might accidentally reveal more than I wanted: API keys, passwords, library card numbers. DON’T TAKE IT. You don’t want the responsibility of safeguarding my PII or secrets. You don’t need to risk Claude going on a spending spree.

How can your agent reject PII and secrets?

Here are some techniques I’ve used:

Prompt the agent to avoid asking for secrets. When I worked on an agent that could configure API calls, we added a system prompt instructing the LLM to never ask for API keys, tokens, or passwords. The LLM should instead teach the user how to provide this info in a secure manner. We also asked the system prompt never to repeat private, secret data given by the user. Just because it hit our logs once doesn’t mean it’s fair game for the agent to use as context.

Detect PII and secrets in the browser, before they even get to the agent. I used regular expressions to detect API keys. You could also use a lightweight ML model to detect frequently shared PII. Then warn the user “hey, looks like you dropped this - are you sure you want to share it with us?”. Or censor it out. thoughtbot’s top_secret Ruby gem does this same thing on the server-side.

What if he says no?

So what do you do if I say “no thanks, no AI please?” The same thing you’d do if I said “no I won’t share my location data”. You give me the fallback option. The manual one. Maybe my experience is not so great or maybe I have to do more work. But my principles remain intact.

A tale as old as the internet

Guess what - all this advice isn’t new for LLMs. It’s a spin on the demands coming from the privacy movements of the early 2000s and before, back when we learned how our data was being taken by governments and sold to spammers. All I ask is to let me control my exposure to AI, and give me a good reason to hand over my data to the robots.

Friday Squid Blogging: EU is Trying to Fight Unregulated Squid Fishing

Schneier
www.schneier.com
2026-10-02 17:02:18
The EU is recommending import controls to combat unregulated squid fishing in the Southwest Atlantic. I’m not optimistic. As usual, you can also use this squid post to talk about the security stories in the news that I haven’t covered. Blog moderation policy....
Original Article

Atom Feed Subscribe to comments on this entry

Leave a comment

Login

Allowed HTML <a href="URL"> • <em> <cite> <i> • <strong> <b> • <sub> <sup> • <ul> <ol> <li> • <blockquote> <pre> Markdown Extra syntax via https://michelf.ca/projects/php-markdown/extra/

Sidebar photo of Bruce Schneier by Joe MacInnis.

Zig v0.17.0

Hacker News
ziglang.org
2026-10-02 16:56:36
Comments...
Original Article

Carmen the Allocgator

Download & Documentation

Zig is a general-purpose programming language and toolchain for maintaining robust , optimal , and reusable software.

Zig development is funded via Zig Software Foundation , a 501(c)(3) non-profit organization. Please consider a recurring donation so that we can offer more billable hours to our core team members. This is the most straightforward way to accelerate the project along the Roadmap to 1.0. If you need donation receipts or are looking to migrate away from GitHub Sponsors, we recommend donating via Every.org .

This release features 5 months of work : changes from 206 different contributors , spread among 925 commits .

Originally predicted to be shorter, this release cycle ended up substantial, with the Build System reworked, including the introduction of the Build Server Protocol , and the ELF Linker enhanced to the point where we expect Incremental Compilation to work for everyone on x86_64-linux.

Table of Contents §

Target Support §

Zero the Ziguana

Zig supports a wide range of architectures and operating systems. The Support Table and Additional Platforms sections cover the targets that Zig can build programs for, while the zig-bootstrap README covers the targets that the Zig Compiler itself can be easily cross-compiled to run on.

Notable changes:

  • aarch64-openbsd is now tested natively in Zig's CI, ensuring high-quality support going forward.
  • aarch64-freebsd and aarch64-netbsd CI jobs now run on pull requests too, in addition to master pushes.
  • An LLVM bug that broke most aarch64-windows binaries, including the Zig Compiler , has been worked around.
  • The Zig Compiler now applies mandatory code hardening techniques when targeting aarch64-openbsd so that the resulting binaries actually work.
  • Zig now provides stack traces on crashes and failed assertions on 32-bit ARM. Some work still remains for Thumb-only targets.
  • Zig now provides stack traces on crashes and failed assertions on SPARC.
  • Zig now handles pointer authentication opcodes when doing stack unwinding on AArch64.
  • Support for the loongarch32-linux-gnu[sf] targets has been added.
  • Zig now has generally usable support for 64-bit SPARC, and especially sparc64-linux . This is largely thanks to Zig's new ELF linker which now has better support for this target than LLD.
  • The Zig Standard Library has been ported to the x32 and N32 ABIs on x86-64 and 64-bit MIPS, respectively. These are niche ILP32 ABIs that allow using the 64-bit instruction set while only having 32-bit pointers - the idea being to trade available address space for lower memory usage and better cache utilization.
  • Target information has been added for some game consoles: aarch64-switch , arm-gba , mipsel-psx , and powerpc-wiiu
  • Very early xtensa-linux support has been added to Zig. Note that, for now, this support can only be exercised via the C backend or the experimental LLVM backend.
  • The Zig Standard Library now has support for arc[eb]-linux , csky-linux , and m88k-openbsd when using the C backend.
  • The Zig Standard Library now has support for no-libc microblaze[el]-linux , sh[eb]-linux , and sparc-linux .
  • Zig now enforces -mabi=ieeelongdouble for all PowerPC targets. This is just a formalization of what was already reality; Zig has never supported the IBM "double-double" format for long double and likely never will. As a result, this release drops support for powerpc-linux-gnueabi[hf] because glibc only supports the "double-double" format on these targets. The powerpc-linux-musleabi[hf] targets remain supported as they use the IEEE format.
  • This release drops support for powerpc64-linux-gnu . Zig has only ever supported linking ELFv2 binaries for 64-bit PowerPC, and glibc does not officially support ELFv2 on big endian - nor IEEE long double , as above.
  • Zig's ability to detect the native CPU model and features has been greatly enhanced across the board; this affects almost every architecture on every supported OS.
  • The baseline CPU model has been changed for some targets:
    • aarch64-haiku : cortex_a55
    • m68k-* : M68030
    • mips64-openbsd : octeon
    • powerpc-netbsd : 750
    • powerpc64-freebsd : pwr8
    • powerpc64-linux : pwr8
    • powerpc64-openbsd : pwr9
    • s390x-* : arch11
    • sparc-* : generic
    • sparc-linux : v9
    • sparc64-* : ultrasparc
    • xtensa-* : esp32
  • In Zig's target query syntax, native libc version detection now only happens if the triple actually uses native libc (i.e. the ABI component is omitted). We expect this new behavior to better match people's mental model for how target queries work, particularly when considering how the OS component works.

Tier System §

Zig's level of support for various targets is broadly categorized into four tiers with Tier 1 being the highest. The goal is for Tier 1 targets to have zero disabled tests - this will become a requirement for post-1.0.0 Zig releases.

Tier 1 §

  • All non-experimental language features are known to work correctly.
  • The Compiler can generate machine code for this target without relying on LLVM .
  • The integrated fuzzer works on this target (if applicable).

Tier 2 §

  • The Standard Library cross-platform abstractions have implementations for this target.
  • Failed assertions and crashes produce stack traces on this target.
  • libc is available for this target even when cross-compiling (if applicable).
  • Continuous integration machines build the module tests for this target on every push.

Tier 3 §

  • The Compiler can generate machine code for this target by relying on an external backend such as LLVM .
  • The Linker can produce object files, libraries, and executables for this target.

Tier 4 §

  • The Compiler can generate assembly or C source code for this target.

Support Table §

In the following table, ✅ indicates full support, ❌ indicates no support, and ⚠️ indicates that there is partial support, e.g. only for some sub-targets, or with some notable known issues. ❔ indicates that the status is largely unknown, typically because the target is rarely exercised. Hover over other icons for details.

Targets marked with 🪦 are obsolescent; the Zig compiler and standard library maintain best-effort support for them, but that support is expected to be removed eventually.

Tier Target Code Gen. Linker Lang. Feat. Std. Lib. Stack Traces Fuzzer libc CI
1 x86_64-linux 🖥️ ⚡ ✅ ✅ ✅ ✅ ✅ ✅ ✅

2 aarch64-freebsd 🖥️ 🛠️ ✅ ✅ ✅ ✅ ✅ ✅ ✅
2 aarch64[_be]-linux 🖥️ 🛠️ ✅ ✅ ✅ ✅ ✅ ✅ ✅
2 aarch64-maccatalyst 🖥️ 🛠️ ✅ ✅ ✅ ✅ ✅ ✅ ✅
2 aarch64-macos 🖥️ 🛠️ ✅ ✅ ✅ ✅ ✅ ✅ ✅
2 aarch64[_be]-netbsd 🖥️ 🛠️ ✅ ✅ ✅ ✅ ✅ ✅ ✅
2 aarch64-openbsd 🖥️ 🛠️ ✅ ✅ ✅ ✅ ❌ ✅ ✅
2 aarch64-windows 🖥️ 🛠️ ✅ ✅ ✅ ✅ ❌ ✅ ⚠️
2 arm-freebsd 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ✅
2 arm[eb]-linux 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ✅
2 arm[eb]-netbsd 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ⚠️
2 arm-openbsd 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ⚠️
2 hexagon-linux 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ✅
2 loongarch32-linux 🖥 🛠 ✅ ❔ ✅ ✅ ❌ ✅ ⚠️
2 loongarch64-linux 🖥️ 🛠️ ✅ ✅ ✅ ✅ ✅ ✅ ✅
2 mips[el]-linux 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ✅
2 mips[el]-netbsd 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ⚠️
2 mips64[el]-linux 🖥️ ✅ ✅ ✅ ✅ ✅ ✅ ✅
2 mips64[el]-openbsd 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ⚠️
2 powerpc-linux 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ⚠️
2 powerpc-netbsd 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ⚠️
2 powerpc-openbsd 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ⚠️
2 powerpc64[le]-freebsd 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ⚠️
2 powerpc64[le]-linux 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ✅
2 powerpc64-openbsd 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ⚠️
2 riscv32-linux 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ✅
2 riscv32-netbsd 🖥️ ✅ ✅ ✅ ✅ ✅ ✅ ⚠️
2 riscv64-freebsd 🖥️ 🛠️ ✅ ✅ ✅ ✅ ✅ ✅ ⚠️
2 riscv64-linux 🖥️ 🛠️ ✅ ✅ ✅ ✅ ✅ ✅ ✅
2 riscv64-netbsd 🖥️ 🛠️ ✅ ✅ ✅ ✅ ✅ ✅ ⚠️
2 riscv64-openbsd 🖥️ 🛠️ ✅ ✅ ✅ ✅ ❌ ✅ ⚠️
2 s390x-linux 🖥️ ✅ ✅ ✅ ✅ ✅ ✅ ✅
2 sparc64-linux 🖥️ 🛠️ ✅ ✅ ✅ ✅ ✅ ✅ ✅
2 thumb[eb]-linux 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ✅
2 wasm32-wasi 🖥️ 🛠️ ✅ ✅ ✅ ⚠️ ❌ ✅ ✅
2 x86-linux 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ✅
2 x86-netbsd 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ✅
2 x86-openbsd 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ⚠️
2 🪦 x86-windows 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ✅
2 x86_64-freebsd 🖥️ 🛠️ ✅ ✅ ✅ ✅ ✅ ✅ ✅
2 🪦 x86_64-maccatalyst 🖥️ ⚡ ✅ ✅ ✅ ✅ ✅ ✅ ⚠️
2 🪦 x86_64-macos 🖥️ ⚡ ✅ ✅ ✅ ✅ ✅ ✅ ⚠️
2 x86_64-netbsd 🖥️ 🛠️ ✅ ✅ ✅ ✅ ✅ ✅ ✅
2 x86_64-openbsd 🖥️ 🛠️ ✅ ✅ ✅ ✅ ❌ ✅ ✅
2 x86_64-windows 🖥️ 🛠️ ✅ ✅ ✅ ✅ ❌ ✅ ✅

3 aarch64-haiku 🖥️ 🛠️ ✅ ✅ ✅ ✅ ❔ ❌️ ❌️
3 aarch64-ios 🖥️ 🛠️ ✅ ✅ ✅ ✅ ❌ ❌️ ❌️
3 aarch64-serenity 🖥️ 🛠️ ✅ ✅ ✅ ✅ ❔ ❌️ ❌️
3 aarch64-tvos 🖥️ 🛠️ ✅ ✅ ✅ ✅ ❌ ❌️ ❌️
3 aarch64-visionos 🖥️ 🛠️ ✅ ✅ ✅ ✅ ❌ ❌️ ❌️
3 aarch64-watchos 🖥️ 🛠️ ✅ ✅ ✅ ✅ ❌ ❌️ ❌️
3 arm-haiku 🖥️ ✅ ✅ ✅ ✅ ❌ ❌️ ❌️
3 mips64[el]-netbsd 🖥️ ✅ ✅ ✅ ❌️ ✅ ❌️ ❌️
3 riscv64-haiku 🖥️ 🛠️ ✅ ✅ ✅ ✅ ❔ ❌️ ❌️
3 riscv64-serenity 🖥️ 🛠️ ✅ ✅ ✅ ✅ ❔ ❌️ ❌️
3 🪦 thumb-windows 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ❌️
3 wasm64-wasi 🖥️ 🛠️ ✅ ❔ ❌️ ⚠️ ❌ ❌️ ❌️
3 🪦 x86-freebsd 🖥️ ✅ ✅ ✅ ✅ ❌ ✅ ❌
3 x86-haiku 🖥️ ✅ ✅ ✅ ✅ ❌ ❌️ ❌️
3 🪦 x86-illumos 🖥️ ✅ ✅ ✅ ✅ ❌ ❌️ ❌️
3 x86_64-dragonfly 🖥️ 🛠️ ✅ ✅ ✅ ✅ ❔ ❌️ ❌️
3 x86_64-haiku 🖥️ ⚡ ✅ ✅ ✅ ✅ ❔ ❌️ ❌️
3 x86_64-illumos 🖥️ 🛠️ ✅ ✅ ✅ ✅ ❔ ❌️ ❌️
3 x86_64-serenity 🖥️ ⚡ ✅ ✅ ✅ ✅ ❔ ❌️ ❌️

4 alpha-linux 📄 ❌️ ❔ ✅ ✅ ❌ ❌️ ❌️
4 alpha-netbsd 📄 ❌️ ❔ ✅ ✅ ❌ ❌️ ❌️
4 alpha-openbsd 📄 ❌️ ❔ ✅ ✅ ❌ ❌️ ❌️
4 arc[eb]-linux 📄 ❌️ ❔ ✅ ✅ ❌ ✅ ❌️
4 csky-linux 📄 ❌️ ❔ ✅ ✅ ❌ ✅ ❌️
4 hppa-linux 📄 ❌️ ❔ ❌️ ❌️ ❌ ❌️ ❌️
4 hppa-netbsd 📄 ❌️ ❔ ✅ ❌️ ❌ ❌️ ❌️
4 hppa-openbsd 📄 ❌️ ❔ ✅ ❌️ ❌ ❌️ ❌️
4 hppa64-linux 📄 ❌️ ❔ ❌️ ❌️ ❌ ❌️ ❌️
4 m68k-linux 🖥️ ❌️ ❔ ✅ ✅ ❌ ✅ ❌️
4 m68k-netbsd 🖥️ ❌️ ❔ ✅ ✅ ❌ ✅ ❌️
4 m88k-openbsd 📄 ❌️ ❔ ✅ ✅ ❌ ❌️ ❌️
4 microblaze[el]-linux 📄 ❌️ ❔ ✅ ❌️ ❌ ❌️ ❌️
4 or1k-linux 📄 ❌️ ❔ ✅ ✅ ❌ ❌️ ❌️
4 sh[eb]-linux 📄 ❌️ ❔ ✅ ❌️ ❌ ❌️ ❌️
4 sh[eb]-netbsd 📄 ❌️ ❔ ✅ ❌️ ❌ ❌️ ❌️
4 sh-openbsd 📄 ❌️ ❔ ✅ ❌️ ❌ ❌️ ❌️
4 sparc-linux 🖥️ ❌️ ❔ ✅ ✅ ❌ ✅ ❌️
4 sparc-netbsd 🖥️ ❌️ ❔ ✅ ❌️ ❌ ✅ ❌️
4 sparc64-netbsd 🖥️ 🛠️ ⚠️ ✅ ✅ ❌️ ✅ ✅ ❌️
4 sparc64-openbsd 🖥️ 🛠️ ⚠️ ✅ ✅ ❌️ ❌ ✅ ❌️
4 xtensa[eb]-linux 🖥️ ❌️ ❔ ✅ ❌️ ❌ ❌️ ❌️

OS Version Requirements §

The Zig standard library has minimum version requirements for some supported operating systems, which in turn affect the Zig compiler itself:

OS Version
Darwin 15.0+
DragonFly BSD 6.4+
FreeBSD 14.0+
Linux 5.10+
NetBSD 10.1+
OpenBSD 7.8+
Windows 10+

Additional Platforms §

Zig also has varying levels of support for these targets, for which the tier system does not quite apply:

  • aarch64-driverkit
  • aarch64[_be]-freestanding
  • aarch64-fuchsia
  • aarch64-hurd
  • aarch64-switch
  • aarch64-uefi
  • alpha-freestanding
  • amdgcn-amdhsa
  • amdgcn-amdpal
  • amdgcn-mesa3d
  • arc[eb]-freestanding
  • arm[eb]-freestanding
  • arm-3ds
  • arm-fuchsia
  • arm-gba
  • arm-uefi
  • arm-vita
  • avr-freestanding
  • bpf(eb,el)-freestanding
  • csky-freestanding
  • ez80-freestanding
  • ez80-tios
  • hexagon-freestanding
  • hppa[64]-freestanding
  • kalimba-freestanding
  • kvx-freestanding
  • lanai-freestanding
  • loongarch(32,64)-freestanding
  • loongarch(32,64)-uefi
  • m68k-freestanding
  • m88k-freestanding
  • microblaze[el]-freestanding
  • mips[64][el]-freestanding
  • mipsel-psx
  • mipsel-psp
  • msp430-freestanding
  • nvptx[64]-cuda
  • nvptx[64]-nvcl
  • or1k-freestanding
  • powerpc-wiiu
  • powerpc[64][le]-freestanding
  • powerpc64-ps3
  • propeller-freestanding
  • riscv(32,64)[be]-freestanding
  • riscv(32,64)-uefi
  • riscv64-fuchsia
  • riscv64-hurd
  • s390x-freestanding
  • sh[eb]-freestanding
  • sparc[64]-freestanding
  • spirv(32,64)-opencl
  • spirv(32,64)-opengl
  • spirv(32,64)-vulkan
  • spork8-freestanding
  • thumb[eb]-freestanding
  • thumb-fuchsia
  • thumb-gba
  • thumb-vita
  • ve-freestanding
  • wasm(32,64)-emscripten
  • wasm(32,64)-freestanding
  • x86[_16,_64]-freestanding
  • x86[_64]-hurd
  • x86[_64]-uefi
  • x86_64-driverkit
  • x86_64-fuchsia
  • x86_64-plan9
  • x86_64-ps4
  • x86_64-ps5
  • xcore-freestanding
  • xtensa[eb]-freestanding

Language Changes §

Language Stability Progress §

Since the release of Zig 0.16.0, a lot of progress has been made towards stabilizing the language. This is a key step in our roadmap , and a requirement before tagging Zig 1.0.

In particular, since the last release, we have discussed and made decisions on many language proposals—accepting around 25 and rejecting around 125. At the time of writing, 23 undecided language proposals remain open on the Codeberg issue tracker, and 61 undecided language proposals remain open on the legacy GitHub issue tracker. Therefore, this effort represents a significant step towards finalizing the language design (although some major decisions remain ).

@bitCast changes §

Zig 0.17.0 changes the definition of the @bitCast builtin.

In many cases, the new behavior is equivalent to the old: in particular, casting between an integer type and another integer type is unaffected, as is casting between an integer type and a packed struct or packed union .

However, the semantics of @bitCast calls involving array or vector types have changed. Unfortunately, this change has the potential to break existing code without triggering a compile error. . Therefore, it may be useful when upgrading to audit any @bitCast uses which involve array or vector types.

The new definition of @bitCast is that it reinterprets the logical bit representation of a value as a different type. The following types are considered to have logical bit representations:

  • void
  • bool
  • integer types, except for comptime_int
  • floating-point types, except for comptime_float
  • integer-backed types enum(T) , packed struct(T) , and packed union(T)
  • arrays or vectors of any of these types

For integer and floating-point types, the logical bit representation starts with the least-significant bit and ends with the most-significant bit. For array and vector types, all elements' logical bit representations are concatenated in order starting with the first element.

In practice, this means that the new @bitCast definition largely aligns with the old behavior on little-endian targets. Unlike the old behavior, the new behavior is fully endian-agnostic, i.e. the operation behaves the same regardless of the target endian.

The new @bitCast definition disallows casts between some types which were previously allowed. In particular, casts involving extern struct or extern union types are no longer permitted. In most cases, code which was using such casts is aiming to reinterpret the value's in-memory representation (sometimes called "type punning")—to achieve this, use @ptrCast or an extern union .

bitcast_extern_struct.zig
const TwoBytes = extern struct {
    b0: u8,
    b1: u8,
};
test "type pun extern struct" {
    const bytes: TwoBytes = .{ .b0 = 0x12, .b1 = 0xAB };
    const int: u16 = @bitCast(bytes);
    switch (std.lang.Endian.native) {
        .little => try expectEqual(0xAB_12, int),
        .big => try expectEqual(0x12_AB, int),
    }
}
const std = @import("std");
const expectEqual = std.testing.expectEqual;
Shell
$ zig test bitcast_extern_struct.zig
/home/ci/.cache/act/4d2db5a0202099f4/hostexecutor/src/download/0.17.0/release-notes/bitcast_extern_struct.zig:7:31: error: cannot @bitCast from 'bitcast_extern_struct.TwoBytes'
    const int: u16 = @bitCast(bytes);
                              ^~~~~
/home/ci/.cache/act/4d2db5a0202099f4/hostexecutor/src/download/0.17.0/release-notes/bitcast_extern_struct.zig:1:25: note: struct declared here
const TwoBytes = extern struct {
                 ~~~~~~~^~~~~~

⬇️

ptrcast_extern_struct.zig
const TwoBytes = extern struct {
    b0: u8,
    b1: u8,
};
test "type pun extern struct" {
    const bytes: TwoBytes = .{ .b0 = 0x12, .b1 = 0xAB };
    const int_ptr: *align(1) const u16 = @ptrCast(&bytes);
    switch (std.lang.Endian.native) {
        .little => try expectEqual(0xAB_12, int_ptr.*),
        .big => try expectEqual(0x12_AB, int_ptr.*),
    }
}
const std = @import("std");
const expectEqual = std.testing.expectEqual;
Shell
$ zig test ptrcast_extern_struct.zig
1/1 ptrcast_extern_struct.test.type pun extern struct...OK
All 1 tests passed.

Formally Specified and Fuzzed Grammar §

Zig's formal grammar.peg and the actual language implementation did not agree in many places. Probably, the formal grammar has never actually 100% matched the actual handwritten tokenizer and parser.

This problem is now fixed and unblocks future grammar changes and language specification work.

At a high level, the approach was to write a tool that accepts Zig's grammar.peg as input and outputs a simple recursive descent parser. This generated parser is then used as an oracle for fuzz testing and the handwritten std.zig.Ast.parse() is compared against it. This approach ensures a single source of truth and allows for easy iteration as the grammar changes.

More details: #36094

C Translation Moving to External Package §

@cImport was deprecated in Zig 0.16.0 and is now removed. Furthermore, in this release std.Build.Step.TranslateC is deprecated in favor of an explicit package dependency on official ZSF translate-c package , which is the same implementation the build step provides, but offers more configuration options for the translated code, and has an independent release cadence from the main Zig toolchain.

Upgrade guide:

zig fetch --save git+https://codeberg.org/ziglang/translate-c
--- a/build.zig
+++ b/build.zig
@@ -1,16 +1,19 @@
 const std = @import("std");
+const Translator = @import("translate_c").Translator;
 
 pub fn build(b: *std.Build) void {
     const target = b.standardTargetOptions(.{});
     const optimize = b.standardOptimizeOption(.{});
 
-    const translate_c = b.addTranslateC(.{
-        .root_source_file = b.path("src/c.h"),
+    const translate_c = b.dependency("translate_c", .{});
+
+    const translator: Translator = .init(translate_c, .{
+        .c_source_file = b.path("src/c.h"),
         .target = target,
         .optimize = optimize,
+ // additional options now available that go here:
+ // https://codeberg.org/ziglang/translate-c#options
     });
-    translate_c.linkSystemLibrary("glfw", .{});
-    translate_c.linkSystemLibrary("epoxy", .{});
+    translator.linkSystemLibrary("glfw3", .{});
+    translator.linkSystemLibrary("epoxy", .{});
 
     const exe = b.addExecutable(.{
         .name = "tetris",
@@ -21,7 +24,7 @@ pub fn build(b: *std.Build) void {
             .imports = &.{
                 .{
                     .name = "c",
-                    .module = translate_c.createModule(),
+                    .module = translator.mod,
                 },
             },
         }),

Added @backingInt and @fromBackingInt §

The @backingInt and @fromBackingInt builtins are new. These builtins replace the now-deprecated @intFromEnum and @enumFromInt builtins ( #35966 ).

@backingInt works with all enums and with bitpacks with explicit backing integer types only. It also works with tagged unions, returning the backing integer of the active tag value. An undefined enum or bitpack yields an undefined backing integer.

@fromBackingInt infers its result type, which may be any enum or a bitpack with an explicit backing integer type. It takes a parameter of exactly that backing integer type. For enums, passing a backing integer that is either undefined or would yield an invalid tag value results in safety-checked Illegal Behavior. For bitpacks, passing an undefined backing integer yields an undefined bitpack.

@bitCast now also performs a safety check for invalid tag values if its destination type is an enum .

Also adds a std.meta.BackingInt function to get the result type of @backingInt .

Zig now requires empty enums to have noreturn as their backing integer because they are uninstantiable.

Upgrade example:

--- a/lib/std/Build.zig
+++ b/lib/std/Build.zig
@@ -145,7 +145,7 @@ pub const Graph = struct {
 
     pub fn addGeneratedFile(graph: *Graph, owner: *Step) Configuration.GeneratedFileIndex {
         graph.generated_files.append(graph.arena, owner) catch @panic("OOM");
-        return @enumFromInt(graph.generated_files.items.len - 1);
+        return @fromBackingInt(@intCast(graph.generated_files.items.len - 1));
     }
 
     pub fn dupeString(graph: *const Graph, bytes: []const u8) []const u8 {
@@ -2607,7 +2607,7 @@ pub const LazyPath = union(enum) {
             .src_path, .cwd_relative, .relative, .dependency => {},
             .generated => |gen| {
                 const graph = other_step.owner.graph;
-                const generated_owner_step = graph.generated_files.items[@intFromEnum(gen.index)];
+                const generated_owner_step = graph.generated_files.items[@backingInt(gen.index)];
                 other_step.dependOn(generated_owner_step);
             },
         }

zig fmt automatically performs this upgrade.

Added @SpirvType §

SPIR-V has a number of types, such as images and samplers, that have no equivalent in Zig's type system. Previously, the only way to refer to one of them was through inline assembly, which made it impossible to declare a texture or a storage buffer as an ordinary global variable. Zig 0.17.0 implements accepted proposal #35240 , adding the @SpirvType builtin alongside the other type-creating builtins:

sample_code
@SpirvType(comptime options: std.lang.Type.Spirv) type
  • .sampler creates an OpTypeSampler .
  • .image creates an OpTypeImage .
  • .sampled_image creates an OpTypeSampledImage from an image type whose usage is .sampled .
  • .runtime_array creates an OpTypeRuntimeArray . It supports indexing and exposes a len field just like the array type.

Using this builtin when not targeting SPIR-V is a compile error. The options are also validated against the target OS.

spirv_type.zig
const std = @import("std");

const Image = @SpirvType(.{ .image = .{
    .usage = .{ .sampled = f32 },
    .format = .unknown,
    .dim = .@"2d",
    .depth = .not_depth,
    .arrayed = false,
    .multisampled = false,
    .access = .unknown,
} });
const SampledImage = @SpirvType(.{ .sampled_image = Image });

const texture = @extern(*addrspace(.constant) const SampledImage, .{
    .name = "texture",
    .decoration = .{ .descriptor = .{ .set = 0, .binding = 0 } },
});
const uv_in = @extern(*addrspace(.input) const @Vector(2, f32), .{
    .name = "uv",
    .decoration = .{ .location = 0 },
});
const color_out = @extern(*addrspace(.output) @Vector(4, f32), .{
    .name = "color",
    .decoration = .{ .location = 0 },
});

export fn main() callconv(.{ .spirv_fragment = .{} }) void {
    color_out.* = std.spirv.imageSampleImplicitLod(texture, uv_in.*);
}
Shell
$ zig build-obj spirv_type.zig -target spirv32-vulkan

Array Multiplication Syntax Removed §

Array multiplication syntax ( a ** b ) has been removed in favor of @splat .

Migration:

--- a/player/chromaprint.zig
+++ b/player/chromaprint.zig
@@ -35,7 +35,7 @@ pub const Chroma = struct {
     const max_index = @min(window_size / 2, freqToIndex(max_freq));
     const notes: [window_size]u8 = n: {
         @setEvalBranchQuota(window_size);
-        var result = [1]u8{0} ** window_size;
+        var result: [window_size]u8 = @splat(0);
         for (min_index..max_index) |i| {
             const freq = indexToFreq(i);
             const octave = freqToOctave(freq);
@@ -123,7 +123,7 @@ const RollingIntegralImage = struct {
     num_rows: u32,
 
     pub const init: RollingIntegralImage = .{
-        .data = [1]Float{0} ** data_size,
+        .data = @splat(0),
         .num_rows = 0,
     };

Added @divCeil §

The new @divCeil builtin performs integer division rounded toward positive infinity, complementing the existing @divTrunc , @divFloor , and @divExact builtins.

sample_code
@divCeil(5, 3) == 2
@divCeil(-5, 3) == -1

As with the other division builtins, caller guarantees that denominator != 0 and that result does not overflow.

No more std.math.divCeil(a, b) catch unreachable !

@hasDecl Returns true Only for Public Declarations §

Previously, @hasDecl returned true for public declarations and declarations in the same file. Now, the behavior is the same independently of which file @hasDecl is in.

hasdecl.zig
const std = @import("std");

const Foo = struct {
    bar: i32,

    const baz = 1;
    pub var quux = "xxx";
};

test "@hasDecl example" {
    try std.testing.expect(!@hasDecl(Foo, "bar"));
    try std.testing.expect(!@hasDecl(Foo, "baz")); // different in 0.17.0
    try std.testing.expect(@hasDecl(Foo, "quux"));
}
Shell
$ zig test hasdecl.zig
1/1 hasdecl.test.@hasDecl example...OK
All 1 tests passed.

Allow Dereference and Coercion to Array Pointer of comptime Length Slices §

Now allowed:

sample_code
const slice: []const u16 = &.{ 1, 2, 3 };
const array: [3]u16 = slice.*;
const array_ptr: *const [3]u16 = slice;

void{} Syntax Removed §

void{} is no longer valid syntax. Use {} instead ( #15213 ).

errdefer Capture Removed §

sample_code
errdefer |err| {
    // `err` was the error being returned; you could do stuff with it!
    // Like a normal `defer/`errdefer`, you can't `return` or `try` in this block, so you can't change
    // which error is returned; but you can e.g. log it
}

The capture ( |err| ) is no longer allowed ( #23734 ).

To migrate, split the function into two:

fn processOneTarget(job: Job) void {
-    errdefer |err| std.debug.panic("panic: {s}", .{@errorName(err)});
+    processOneTargetInner(job) catch |err| std.debug.panic("panic: {s}", .{@errorName(err)});
+}
+fn processOneTargetInner(job: Job) !void {
     const target = job.target;

i0 Removed §

i0 is no longer an allowed primitive integer type.

This type was nonsensical, so does not have a direct alternative. However, any uses of it can almost certainly be transparently replaced with u0 .

The internal and link_once tags of std.lang.GlobalLinkage have been removed as they had unclear semantics and incomplete support in codegen and linking ( #36956 ). Any use of link_once is likely served by weak , while the replacement for internal is to simply not @export the symbol in the first place.

Standard Library §

  • Added f128 support for @exp and @exp2 based on "Table-Driven Implementation of the Exponential Function in IEEE Floating-Point Arithmetic" by Ping Tak Peter Tang, adapted to work with 128-bit numbers ( #31846 ).
  • Added std.Io.Semaphore.waitTimeout ( #31924 ).
  • Added std.spirv helpers for sampling, querying, and writing images ( #36187 ).
  • ArrayHashMap.setKey no longer recomputes the entire index ( #32136 ).
  • std.Target.parseCpuModel now returns optional rather than error.
  • std.debug.Pdb : deduplicate inline source locations ( #35438 ).
  • hash.crc full namespace audit ( #35952 ).
  • std.fs.path add appending variants for relative and resolve ( #36784 ).

Deprecations §

  • std.heap.memory_pool.AlignedManaged removed in favor of std.heap.memory_pool.Aligned .
  • std.heap.memory_pool.ExtraManaged removed in favor of std.heap.memory_pool.Extra .
  • Deprecated std.builtin in favor of std.lang .
  • Deprecated std.meta.fieldInfo in favor of @typeInfo .
  • Deprecated std.meta.fieldNames in favor of @typeInfo .
  • Deprecated std.meta.fieldTypes in favor of @typeInfo .
  • Deprecated std.DoublyLinkedList.pop in favor of std.DoublyLinkedList.popLast .
  • Renamed std.gpu to std.spirv .
  • Removed std.ascii.indexOfIgnoreCase in favor of std.ascii.findIgnoreCase .
  • Removed std.ascii.indexOfIgnoreCasePos in favor of std.ascii.findIgnoreCasePos .
  • Removed std.ascii.indexOfIgnoreCasePosLinear in favor of std.ascii.findIgnoreCasePosLinear .
  • Removed std.bit_set.Integer.initEmpty in favor of std.bit_set.Integer.empty .
  • Removed std.bit_set.Integer.initFull in favor of std.bit_set.Integer.full .
  • Removed std.bit_set.Array.initEmpty in favor of std.bit_set.Array.empty .
  • Removed std.bit_set.Array.initFull in favor of std.bit_set.Array.full .
  • Removed std.enums.EnumSet.initEmpty in favor of std.enums.EnumSet.empty .
  • Removed std.enums.EnumSet.initFull in favor of std.enums.EnumSet.full .
  • Removed std.mem.containsAtLeastScalar2 in favor of std.mem.containsAtLeastScalar .
  • Removed std.mem.readPackedIntNative in favor of std.mem.readPackedInt .
  • Removed std.mem.readPackedIntForeign in favor of std.mem.readPackedInt .
  • Removed std.mem.writePackedIntNative in favor of std.mem.writePackedInt .
  • Removed std.mem.writePackedIntForeign in favor of std.mem.writePackedInt .

StackFallbackAllocator Reworked §

StackFallbackAllocator is an abstraction that is useful for the "small vec" optimization, in which common cases can fit on a pre-allocated stack buffer, but rare cases need dynamic heap allocation. The previous design had a few problems:

  • There was no way to specify the alignment of the buffer.
  • It was generic over the size of the buffer.
  • Calling .allocator().get() mutated the type, unlike all other allocators, requiring a runtime safety check .

Now, the buffer is provided as an argument, like most other std APIs that need a buffer.

Migration guide:

sample_code
var stack align(@max(
    @alignOf(std.heap.StackFallbackAllocator(0)),
    @alignOf(Item)),
) = std.heap.stackFallback(@sizeOf(Item), self.gpa);
const allocator = stack.get();

⬇️

sample_code
var stack_buf: [1]Item = undefined;
var stack: std.heap.StackFallbackAllocator = .init(@ptrCast(&stack_buf), gpa);
const allocator = stack.allocator();

SafeAllocator Introduced §

std.heap.DebugAllocator is replaced by a thread-safe allocator with the following guarantees:

  • deinit reports all leaks and frees all backing memory.
  • All allocation mismatches result in either a panic or segmentation fault.
  • Allocations from other SafeAllocator instances cause a panic (if Options.canary differ).
  • Double frees and operation (resize, remap, and free) races panic or segmentation fault.

Given the backing allocator does not reuse memory, it does not reuse memory either and most writes after free will segmentation fault or are eventually detected and panic.

std.heap.DebugAllocator and std.heap.Check are deprecated.

Every allocation is trailed by an AllocFooter which contains metadata for the allocation and stack traces. It is protected by a checksum to catch corruption from allocation overwrites and report canary mismatches. An allocation's memory has a minimum alignment of AllocFooter so that the footer is at a fixed offset determined from the allocation size. An allocation's memory is stored either:

  • Inside linearly-filled buckets for small allocations.
  • Inside an allocation directly from the backing allocator.

To track allocations, each thread maintains a table of backing allocations. The table may be modified by other threads in the case of a producer-consumer operation, so the table is a linked list only expanded by creating new segments. Each thread maintains a linked list of free entries, which may contain entries from other threads' tables.

In the case of producer-consumer operations, acquire/release ordering is assumed to be provided externally. This is also assumed by all other thread-safe allocators that reuse memory as otherwise there would be data races on reuse of allocated memory.

Two fuzz tests have also been added for the allocator. They check that there is no memory reuse, that returned memory is writable, and that it is not overwritten. The multi-threaded fuzz test spawns a number of worker threads which are used for all the test runs. I have run these tests extensively under TSAN.

Building the standard library tests with an -Osafe compiler build and -Ddebug-allocator :

Benchmark 1 (3 runs): ./master-out/bin/zig test --zig-lib-dir lib lib/std/std.zig -femit-bin=test --test-no-exec
  measurement          mean ± σ            min … max           outliers         delta
  wall_time          29.4s  ±  157ms    29.2s  … 29.5s           0 ( 0%)        0%
  peak_rss           2.24GB ± 3.49MB    2.23GB … 2.24GB          0 ( 0%)        0%
  cpu_cycles          143G  ±  999M      142G  …  144G           0 ( 0%)        0%
  instructions        268G  ± 5.22M      268G  …  268G           0 ( 0%)        0%
  cache_references   13.1G  ± 88.8M     13.0G  … 13.2G           0 ( 0%)        0%
  cache_misses       2.38G  ± 30.7M     2.35G  … 2.41G           0 ( 0%)        0%
  branch_misses       634M  ± 6.22M      629M  …  641M           0 ( 0%)        0%
Benchmark 2 (3 runs): ./branch-out/bin/zig test --zig-lib-dir lib lib/std/std.zig -femit-bin=test --test-no-exec
  measurement          mean ± σ            min … max           outliers         delta
  wall_time          22.1s  ± 88.6ms    22.0s  … 22.2s           0 ( 0%)        ⚡- 24.7% ±  1.0%
  peak_rss           1.11GB ±  799KB    1.11GB … 1.11GB          0 ( 0%)        ⚡- 50.3% ±  0.3%
  cpu_cycles          136G  ±  480M      136G  …  137G           0 ( 0%)        ⚡-  4.4% ±  1.2%
  instructions        273G  ± 2.07M      273G  …  273G           0 ( 0%)        💩+  1.6% ±  0.0%
  cache_references   12.3G  ± 71.3M     12.2G  … 12.4G           0 ( 0%)        ⚡-  6.0% ±  1.4%
  cache_misses       2.02G  ± 11.5M     2.01G  … 2.03G           0 ( 0%)        ⚡- 14.9% ±  2.2%
  branch_misses       569M  ± 2.65M      567M  …  572M           0 ( 0%)        ⚡- 10.2% ±  1.7%

ArrayList §

  • getLastOrNull has been deprecated and renamed to last
  • getLast has been deprecated in favor of last combined with .?
  • lastPtr has been added which returns ?*T

Upgrade guide:

sample_code
if (list.getLastOrNull()) |foo| {
    // ...
}
const foo = list.getLast();

⬇️

sample_code
if (list.last()) |foo| {
    // ...
}
const foo = list.last().?;

ArrayList Pointer Stability §

This is an enhancement that can help track down ArrayList usage bugs faster ( #36239 ).

Devlog Entry

debug.SafetyLock Gains Support for Shared Locking §

The existing lock and unlock methods continue to act exclusively. They should be used when data may be mutated. New methods, lockShared and unlockShared , may be used for shared locking in situations where multiple independent users are reading but not mutating data.

fmt.allocPrint moved to mem.Allocator §

sample_code
try std.fmt.allocPrint(arena, "{s}={d}", .{ x, y });

⬇️

sample_code
try arena.print("{s}={d}", .{ x, y });

Formatted Printing Enhancements §

The "{q}" specifier which escapes strings so that they can appear in double-quoted string literals has relaxed escaping rules such that UTF-8 encoded data can pass through unmangled.

"{qf}" is introduced for double-quote escaping the output of a format() .

std.zon.parse Reworked §

std.zon.parse now takes struct args and allocates its result from an arena.

Migration guide:

sample_code
var diag: Diagnostics = .{};
defer diag.deinit(gpa);
const result = std.zon.fromSlice(
    MyZonType,
    gpa,
    source,
    &diag,
    .{},
) catch |err| switch (err) {
    error.ParseZon => std.process.fatal("input.zon: {f}", .{diag}),
    error.OutOfMemory => |e| return e,
};
defer std.zon.parse.free(result);

⬇️

sample_code
var diag: Diagnostics = undefined;
const result = std.zon.fromSlice(MyZonType, .{
    .gpa = gpa,
    .arena = arena,
    .source = source,
    .diagnostics = &diag,
}) catch |err| switch (err) {
    error.ParseZon => diag.fatal("input.zon"),
    error.OutOfMemory => |e| return e,
}

Some methods were renamed:

  • fromSliceAlloc ➡️ fromSlice
  • fromSlice ➡️ fromSliceNoAlloc
  • The other "from" methods were renamed following this same scheme.

"updateFrom" variants such as updateFromSlice were added. These update an in memory value, overwriting the value's fields with fields specified in the ZON source. This can be useful when using ZON to load configuration files with varying precedence, for example a text editor that has a global config file and a per-project config file.

Rename bit_set Variants and Deprecate the Managed One §

Renames the types for consistency, deprecating the previous names and the managed variant.

  • std.bit_set.IntegerBitSet ➡️ std.bit_set.Integer
  • std.bit_set.ArrayBitSet ➡️ std.bit_set.Array
  • std.StaticBitset , std.bit_set.StaticBitSet ➡️ std.bit_set.Static
  • std.DynamicBitSetUnmanaged , std.bit_set.DynamicBitSetUnmanaged ➡️ std.bit_set.Dynamic
  • std.DynamicBitSet , std.bit_set.DynamicBitSet ➡️ std.bit_set.DynamicManaged (deprecated)

Struct-Of-Arrays Style for std.lang.Type §

When doing type reflection, structs and unions return their information in struct-of-arrays style (#35234) .

--- a/lib/compiler/Maker/ScannedConfig.zig
+++ b/lib/compiler/Maker/ScannedConfig.zig
@@ -49,9 +49,10 @@ pub fn print(sc: *const ScannedConfig, w: *Writer) Writer.Error!void {
 }
 
 fn printStruct(sc: *const ScannedConfig, s: *Serializer.Struct, comptime S: type, v: S) !void {
-    inline for (@typeInfo(S).@"struct".fields) |field| {
-        try s.fieldPrefix(field.name);
-        try printValue(sc, s.container.serializer, field.type, @field(v, field.name));
+    const info = @typeInfo(S).@"struct";
+    inline for (info.field_names, info.field_types) |field_name, field_type| {
+        try s.fieldPrefix(field_name);
+        try printValue(sc, s.container.serializer, field_type, @field(v, field_name));
     }
 }

Rename lang.OptimizeMode to lang.Optimize §

And remove "release" from the enum tag names.

No functional change, however, despite the addition of backwards-compatibile declarations in this patch, it is breaking because expressions that use == or != operators will not able to use the deprecated names.

  • std.lang : OptimizeMode ➡️ Optimize
    • Debug ➡️- debug
    • ReleaseSafe ➡️- safe
    • ReleaseFast ➡️- fast
    • ReleaseSmall ➡️- small

lang.Optimize.runtimeSafety §

std.lang.Optimize.runtimeSafety is preferred as an alternative to std.debug.runtime_safety since it will offer callsites knowledge about their own module rather than standard library module.

@import("builtin") Deprecations §

The redundant constants cpu , os , abi , and object_format in @import ( "builtin" ) have been deprecated and will be removed in 0.18.0. Please replace any usage with the corresponding fields on the target constant:

  • @import ( "builtin" ).cpu ➡️ @import ( "builtin" ).target.cpu
  • @import ( "builtin" ).os ➡️ @import ( "builtin" ).target.os
  • @import ( "builtin" ).abi ➡️ @import ( "builtin" ).target.abi
  • @import ( "builtin" ).object_format ➡️ @import ( "builtin" ).target.ofmt

Handle Floats Correctly in mem.eql and mem.findDiff §

The functions std.mem.eql and std.mem.findDiff short-circuit when their two inputs are slices to the same memory. This short-circuiting is only correct when the == operator, for the given type, is reflexive. This isn't the case for floats, as for example std.math.nan( f64 ) != std.math.nan( f64 ) . The change in this PR disables that optimisation when working on float slices.

Previously-failing, now-succeeding tests:

sample_code
const x: [3]f64 = .{ 42.0, std.math.nan(f64), 3.1415 };
try std.testing.expect(!std.mem.eql(f64, &x, &x));
try std.testing.expectEqual(1, std.mem.findDiff(f64, &x, &x));

Decouple Uri and net.HostName §

Uri was sometimes using HostName.validate for host (in resolveInPlace) and sometimes not (in parseAfterScheme). On its own, this was a problem, but the bigger problem is that RFC3986 (Uri) has a much different idea of what a valid host name is than RFC1123 (HostName), and so just making Uri consistently use HostName.validate would make Uri less useful overall.

Instead, all HostName -related stuff has been removed from Uri . Uri.getHost has been moved to HostName.fromUri (without a graceful deprecation, since the semantics are different enough for users to need to evaluate usage sites), while Uri.getHostAlloc has been removed entirely.

Migration guide:

sample_code
var host_buf: [HostName.max_len]u8 = undefined;
const host = try uri.getHost(&host_buf);

⬇️

sample_code
var host_buf: [HostName.max_len]u8 = undefined;
// note: the error set is different than it was before since fromUri does validation
const host = try HostName.fromUri(uri, &host_buf);

#36036

Build System §

  • b.build_root (Directory) ➡️ b.root (Path)
  • ConfigHeader.Options : include_guard_override ➡️ include_guard
  • LazyPath : getDisplayName ➡️ format ( "{f}" )
  • LazyPath.basename : removed since the value is not known until make phase
  • b.findProgram divided into findProgram and findProgramLazy and API future-proofed.
  • ConfigHeader fixed; now reports unused values for all styles
  • addArtifactArg , addPrefixedArtifactArg ➡️ addArtifactArg2
  • addOutputFileArg , addPrefixedOutputFileArg ➡️ addOutputFileArg2
  • addFileContentArg , addPrefixedFileContentArg ➡️ addFileContentArg2
  • addOutputDirectoryArg , addPrefixedOutputDirectoryArg ➡️ addOutputDirectoryArg2
  • addDirectoryArg , addPrefixedDirectoryArg , addDecoratedDirectoryArg ➡️ addDirectoryArg2
  • addDepFileOutputArg , addPrefixedDepFileOutputArg ➡️ addDepFileOutputArg2
  • addFileArg , addPrefixedFileArg ➡️ addFileArg2

Separate the Maker Process from the Configurer Process §

zig build now runs projects' build.zig code in a separate executable than the one that performs Package Management and executes the build graph, making zig build faster for several reasons ( #35428 ):

  • The maker executable remains unmodified when build.zig script is edited, and therefore only needs to be built exactly once ("first time setup") after installing Zig.
  • The maker executable is built with optimizations enabled, which is starting to become more valuable now that we have introduced --watch and --fuzz .
  • build.zig logic can be skipped sometimes depending on what CLI flags are used with zig build .

Furthermore, configuration is now serialized into a compact binary format that can be consumed by third party tooling and is part of the new Build Server Protocol . The prior way of satisfying this use case by forking the build runner is no longer supported .

To render configuration as .zon to stdout, pass --print-configuration .

Cache System Reworked §

New features:

  • Directory support. Ability for entries added, removed, or renamed in directories to cause a cache miss.
  • Metadata mode. Normally, only changed contents causes a cache miss. In metadata mode, when size, inode, or mtime changes, it always causes a cache miss independent of contents.

All four combinations are possible (is_directory=true/false, metadata_mode=true/false). These features are exposed as new API in the Build System .

Since the Zig toolchain is heavily reliant on the caching system, this release also switches to a binary format, saving roughly 25% on file size, which eases a bit of pressure on the file system cache while also simplifying the work the computer needs to do - directly copy bytes from disk rather than parsing text files. The new zig cache-cat subcommand is available for troubleshooting or tinkering with files inside a zig-cache directory.

The cache system also now has the capability to explain why a "miss" happened. The public-facing API of std.Build.Cache has many breaking changes, but outside of compiler tooling, this is an uncommon API to be used, and all the changes make it harder to misuse.

This change has been observed to speed up cache hits by 5-10% ( #36822 ).

Introduce the Concept of Configure Cache Poisoning §

If the cache is poisoned means that the configure logic had side effects, or otherwise did something that could not be tracked by the cache system.

This is not to be confused with whether individual steps may have side effects when being evaluated; it has to do with the logic inside build.zig itself. For example, a Run step that prints "hello world" has side effects at make time and therefore does not warrant setting this flag, while checking for the existence of scdoc at configure time in order to choose the default value for a configuration option does.

Keeping the cache pure will make zig build faster, bypassing the configurer process when identical configuration would be generated.

When the cache is poisoned, the maker process will delete the build configuration file upon ingesting it since it cannot be reused.

Ways to poison the cache include calling findProgram , or more directly std.Build.Graph.poisonCache . A better alternative than cache poisoning is to explicitly declare the configuration dependencies with these new functions:

  • std.Build.dependOnFileContents - indicates that the build.zig logic depends on a particular file's contents.
  • std.Build.dependOnFileMetadata - indicates that the build.zig logic depends on a particular file's size, inode, mtime, and contents.
  • std.Build.dependOnDirectoryContents - indicates that the build.zig logic depends on a particular directory's entries.
  • std.Build.dependOnDirectoryMetadata - indicates that the build.zig logic depends on a particular directory's last modification date.

Advanced users can override the cache poisoning behavior with a new CLI option:

  --cache-poison[=mode]        Override configuration caching behavior
      pure                     (default) Avoid false positive cache hits
      poisoned                 Don't cache the configuration
      disallowed               Panics when cache would be poisoned
      ignored                  A little poison never hurt anybody

findProgram §

Immediately (in the configure phase), searches for an executable on the host that has more than one possible name.

Names are searched in order, observing search prefixes first and then PATH environment variable.

Calling this function poisons the configuration cache, so it is only appropriate when the existence of the program or its output needs to be observed by configuration logic. That's why there is also findProgramLazy now.

findProgramLazy §

Creates an anonymous Step that searches for an executable on the host that has more than one possible name.

Unlike findProgram , this function does not poison the configuration cache , however the result cannot be used in the configuration phase, hence the return type being LazyPath .

Returns the LazyPath of the found executable. The search only takes place if the LazyPath will be used by a depending Step .

This API is useful in the following cases:

  • The binary is not named the same across all systems (for example "python" vs "python3").
  • The binary may be produced by building from source rather than being globally installed and will therefore be possibly found in one of the search prefix paths.

Run Step: Passthru Args §

In the Run step, passthru args are all together now, not observable in configure phase whether run args are provided.

--- build.zig
+++ build.zig
@
-if (b.args) |args| {
-    run_cmd.addArgs(args);
-}
+run_cmd.addPassthruArgs();

This removes a capability from build scripts since they can no longer observe those arguments. In exchange, it means that when changing those arguments, build scripts no longer must be rebuilt from source.

Fmt Step: Options §

paths and exclude_paths are now LazyPath lists. There is a convenience method to create them: b.pathList .

--- build.zig
+++ build.zig
@
-    const fmt_include_paths = &.{ "lib", "src", "test", "tools", "build.zig", "build.zig.zon" };
-    const fmt_exclude_paths = &.{ "test/cases", "test/behavior/zon" };
+    const fmt_include_paths = b.pathList(&.{ "lib", "src", "test", "tools", "build.zig", "build.zig.zon" });
+    const fmt_exclude_paths = b.pathList(&.{ "test/cases", "test/behavior/zon" });

Step.Options : add addOptionPathDirectory §

Now, when adding an option that is a file path, one must explicitly choose between ( #36876 ):

  • addOptionPath (must be a file)
  • addOptionPathDirectory (must be a directory)
  • addOptionPathUntracked (opt out of dependency tracking)

Lazy Dependency Ergonomic Enhancements §

  • Log when lazy dependencies are fetched.
  • std.Build.dependency : support lazy dependencies
  • Introduce std.Build.dependencyLazy which possibly returns error .LazyDependencyNeeded instead of null , so that you can use try
  • When user build functions return error .LazyDependencyNeeded , build system proceeds to fetch them rather than failing configuration.

Removed Ability to Override Build Runner §

There is no concept of a "build runner" any more; it has been split into: configurer and maker

This use case is now handled by the Build Server Protocol .

Package Management §

All package management functionality has been moved out of the Compiler and into the Build System . This includes the following sub-commands:

  • zig build
  • zig fetch
  • zig init
  • zig libc
  • zig cache-cat

This means that large parts of what used to be included in the compiler executable are now shipped in source form instead, including:

  • package fetching logic
  • http client and networking
  • TLS (Transport Layer Security) and associated crypto
  • git protocol
  • xz, gzip, zstd, flate, zip
  • parsing, validation, and otherwise dealing with build.zig.zon files

All of this functionality is now compiled in -Osafe optimization mode rather than -Ofast due to being in the compiler. When hacking on the build system itself, the environment variable ZIG_DEBUG_CMD=1 may be used to compile the build system in debug mode instead.

Miscellaneous changes:

  • Bug fix: reject path deps that escape the parent package root.

Ability to Override Package Path §

--pkg-path CLI arg and ZIG_LOCAL_PKG_DIR env var are now observed for both fetch and build commands.

Global vs Local Fetching §

Now zig fetch only fetches into the global cache, just like it used to. However, if --save (or any variant) is used, then it also fetches into the local package path. When fetching globally, does not require build.zig to be present. zig build always fetches locally (in addition to globally).

Notably, this fixes the regressed use case zig fetch .

When fetching by path, the hash is always computed, recompressed tarball is always created, always overwrites any existing global cache entry.

Slight Difference in PATH for DLL Arguments §

It used to be the case that, when targeting Windows or Wine, artifact args added to Run steps modified PATH based on the set of directories containing the recursive set of DLL dependencies. Now this is only done for argv[0]. The motivation for also doing this for the other command line arguments is unclear, since those DLLs don't need to be loaded in order to execute argv[0].

Build Server Protocol §

Now, when --listen=- is passed, the build system serves a protocol that allows connected clients to monitor and control the build graph as it executes. This is intended to be consumed by third-party tooling such as IDEs.

Current things you can do:

  • Get full access to the entire build graph's static, post-configuration information, such as which build steps are available, which options are set, dependencies, etc. There are a couple things yet to be included, such as the exposed set of module names.
  • Get notified when a build step starts and completes, including information about errors and which files were generated.
  • Request specific steps to build.

In particular, the separation of maker process and configurer process is a breaking change that prevents the ZLS project from working with 0.17.0. Although some progress was made to restore functionality in this release cycle, Zig team and ZLS team are still working together to enhance the build server protocol further to the point that ZLS can not only restore functionality, but surpass the power and capabilities compared to before.

In the future it is expected for much of Zig's own first-party build system tooling to become a client of the build server protocol, dogfooding it to ensure that third-party tooling enjoys equivalent capabilities ( #36497 ).

It is also planned for the build server to multiplex compiler server protocol for the compilation steps, providing type-system information, refactoring, and other advanced editing capabilities ( #615 ).

Compiler §

Incremental Compilation §

The Zig compiler's implementation of incremental compilation—a feature allowing near-instant rebuilds of projects after changing the code—has been significantly improved in Zig 0.17.0. Many bugs have been fixed, and the new ELF Linker introduced in the previous release has gained good support for the feature.

Thanks to these enhancements, it is now possible for most projects targeting x86_64-linux to take advantage of incremental compilation. To do so, add the arguments -fincremental --watch to your zig build command (e.g. zig build -fincremental --watch )—this will cause the Zig build system to listen for changes to source files, and react to them by performing an incremental rebuild.

For more information on ways to use incremental compilation in your own projects, or to learn more about how this feature works under the hood, consider checking out this blog post by a Zig core team member.

Future releases will continue to focus on improving this feature, including introducing a new Mach-O linker and self-hosted aarch64 Backend with good support for incremental compilation; adding support for using incremental compilation without --watch ; and fixing any remaining bugs.

SPIR-V Backend §

The self-hosted SPIR-V backend is now multi-threaded like the other backends.

Execution modes such as LocalSize and OriginUpperLeft are now derived from the function's calling convention instead of being set through inline assembly, and the new spirv_task and spirv_mesh calling conventions add support for task and mesh shaders ( #35676 ).

Declaring capabilities and extensions in inline assembly with OpCapability and OpExtension is no longer allowed. They are enabled through target CPU features instead, i.e. the -mcpu option.

22 bugs were fixed in the SPIR-V backend during this release cycle.

aarch64 Backend §

Progress towards this is blocked on Linker enhancements, many of which were completed during this release cycle.

loongarch Backend §

Initial implementation of self-hosted backend for loongarch64 has been contributed ( #36418 ). It is still experimental and not yet usable. There are two ways to contribute to this backend: working on it directly, and contributing to AIR Legalization Features , which helps all unfinished backends reach the finish line quicker.

WebAssembly Backend §

Zig's WebAssembly backend is now passing 2060/2054 (100%) behavior tests compared to the LLVM backend. However, it is not yet the default when compiling in debug optimization mode due to lack of debug info support ( #37032 ).

Linker §

ELF §

This release makes significant progress towards replacing Zig's legacy self-hosted ELF linker with its new implementation introduced in the previous release. Specific enhancements include:

  • Full x86_64 support
  • Full SPARC64 support
  • Partial Loongarch support
  • Static library generation
  • Shared library generation
  • Errors for undefined symbols in executables
  • GOT generation
  • Copy relocations
  • GNU symbol versioning
  • DWARF debug information
  • Symbol hash table generation
  • Mostly-reproducible binaries
  • Arbitrary section alignment
  • Support for small host file system block sizes

While this linker has not quite reached feature parity with our old self-hosted ELF linker yet—and so remains disabled by default—it is already capable in practice of building the vast majority of Zig projects targeting x86_64-linux . This unlocks the ability to use Incremental Compilation for these projects—like in Zig 0.16.0, the new linker is enabled by default in this case.

In the next release of Zig, we hope to fully eliminate the legacy ELF linker in favour of this implementation.

COFF §

COFF support in the linker is enhanced with the following features ( #35674 ):

  • Outputing objects (.obj) and archives (.lib)
  • Outputing implibs alongside images
  • Outputing the TLS and Exports data directories for images
  • Consuming objects, archives, and import libraries as inputs
  • Only links in objects from archives as required, ie. if they contain a symbol needed to satisfy a reference
  • COMDAT rules (enough support for linking compiler_rt and libc, some COMDAT types are not supported yet)
  • Supports linking against both -gnu and -msvc libc
  • TLS support
  • __dllimport support: Indirect calls / loads from the IAT directly
  • Detects which entrypoint to choose based on exported symbols
  • -gnu : Constructor / destructor support (ie. merge .ctor and .dtor, and set up the __CTOR_LIST__ , __DTOR_LIST__ symbols)
  • Support for several .drectve arguments (these are required to correctly link msvc libc):
    • /INCLUDE : Forcing a symbol to be referenced
    • /ALTERNATENAME : Adding symbol aliases
    • /MERGE : Section merging. This functionality is also used to direct certain sections into the right place (like .ctor / .dtor into .rdata )
    • /DEFAULTLIB : Adding new inputs

New Linker Testing Framework §

Zig is moving towards snapshot-based testing for its linkers.

Tests are a combination of comparing objdump snapshot output, actually running the artifacts, and checking for linker errors.

zig build -Dlink-snapshot-update causes tests to run in a mode that outputs snapshots instead of checking against them.

A typical workflow for adding a new test:

  1. Add the test
  2. zig-debug build test-link -Dtest-filter=my-test -Dlink-snapshot-update
    • This will output a .dmp file for all the snapshot combinations defined in any verifyObjdump calls.
    • A single snapshot intentionally aliases between many targets to reduce noise in the snapshot folder, so snapshot updates are made by whichever test runs first for that snapshot name. If differences between targets do exist, they will be revealed in step 4.
  3. Inspect the snapshot output for correctness.
  4. zig-debug build test-link -Dtest-filter=my-test
    • This will now run all targets against the newly added snapshots
  5. If there are now snapshot failures, that means different targets had different snapshot outputs. The output should be inspected to see if these results are indeed valid differences. If they are, then the scope parameter should be used to cause -Dlink-snapshot-update to output snapshot to different filenames, scoped on the diference.
  6. Re-run -Dlink-snapshot-update to update the new set of snapshots.

SPIR-V §

The SPIR-V linker has been rewritten ( #36828 ). It now supports incremental compilation and can link external .spv object files.

Fuzzer §

Although this release's Build System changes are loosely related to Zig's integrated fuzzer and its interaction with the build system, no changes were made to the fuzzer itself.

We expect to focus on improving the fuzzer in a future release cycle.

Bug Fixes §

Full list of the 329 bug reports closed during this release cycle:

Many bugs were both introduced and resolved within this release cycle. Most bug fixes are omitted from these release notes for the sake of brevity.

This Release Contains Bugs §

Zig has known bugs , miscompilations , and regressions .

Even with Zig 0.17.x, working on a non-trivial project using Zig may require participating in the development process.

When Zig reaches 1.0.0, Tier 1 support will gain a bug policy as an additional requirement.

Notable Regressions §

We are aware of these notable regressions in 0.17.0:

LLVM 22 §

This release of Zig upgrades to LLVM 22.1.8 . This covers Clang ( zig cc ), libc++, libc++abi, libunwind, and libtsan as well.

Loop Vectorization Disabled to Work Around Regression §

In the previous release of Zig, we were forced to disable a key LLVM optimization pass—loop vectorization—to work around a miscompilation which affected the Zig compiler.

Since we first introduced that workaround, a fix has been merged into LLVM's main branch. However, the fix is not available in LLVM 22 , the LLVM version used by Zig 0.17.0. Therefore, this workaround remains enabled for now.

Zig 0.18.0 will upgrade to LLVM 23, so will allow us to re-enable this optimization pass.

musl 1.2.5 §

Zig 0.17.0 distributes musl 1.2.5 plus backported security and portability fixes. Meanwhile, upstream has tagged 1.2.6. Zig 0.18.0 will update to musl 1.2.6.

When targeting musl statically, many functions are now provided by zig libc rather than source files copied from musl. Therefore, if you encounter bugs with musl libc provided by Zig, please respect upstream by reporting them to Zig's issue tracker rather than musl's.

glibc 2.44 §

glibc version 2.44 is now available when cross-compiling.

This release includes Linux kernel headers for version 7.2.

This release includes macOS system headers for version 27.0.

MinGW-w64 §

Zig 0.17.distributes MinGW-w64 commit 31bd54ab7d5fe03c67ed2bb1a57e531b9c7f8cc4 .

However, many functions are now provided by zig libc rather than source files copied from MinGW-w64. Therefore, if you encounter bugs with MinGW-w64 libc provided by Zig, please respect upstream by reporting them to Zig's issue tracker rather than MinGW-w64's.

NetBSD 11.0 libc §

NetBSD libc version 11.0 is now available when cross-compiling.

OpenBSD 7.9 libc §

OpenBSD libc version 7.9 is now available when cross-compiling.

WASI libc §

Zig 0.17.0 continues to distribute WASI libc commit c89896107d7b57aef69dcadede47409ee4f702ee .

However, many functions are now provided by zig libc rather than source files copied from WASI libc.

Furthermore, starting with Zig 0.18.0, instead of distributing third party WASI libc code, Zig will provide libc for WASI targets via zig libc . For more information, see:

zig libc §

In libc.txt files, the gcc_dir field has been renamed to cc_dir to reflect the reality that it is not specific to GCC. The old name will still be accepted for now, but users are encouraged to migrate their libc.txt to the new name ( #36951 ).

Additionally, cc_dir is now required on Linux targets. Note that, because Android and OpenHarmony store the relevant object files in an unusual location, users of these targets will likely want to set cc_dir to the same path as crt_dir .

zig cc §

zig cc and zig c++ are now based on Clang 22.1.8.

zig objdump §

This was a requirement for snapshot testing , as well as aiding in developing the COFF Linker .

Supported features:

Usage: zig objdump [options] file

Options:
  -h, --help                         Print this help and exit
  --all-headers                      Alias for --file-headers --linker-member=2 --member-headers --section-headers --relocs --symbols
  --exports[=sort]                   Display exported symbols.
                                     In the case of COFF import libraries, displays the symbol list and import headers.
                                     Specify =sort to optionally sort the import headers by symbol name.
  --file-headers                     Display file-format specific headers
  --imports                          Display imported symbols
  --linker-member[=1|2|longnames]    (Coff) Display contents of the specified archive linker member (default 2)
  --member-headers                   Display archive member headers
  --elements=[e1],[e2],-[e3],...     Select which formatting elements are displayed. Intended for snapshot testing.
      file-type                      File type summary
      header-name                    Name that precedes a header block
      member-path                    Display full member paths. If removed, only basenames will be used.
      newlines                       Newlines between output sections
      table-header                   Table headers with column names
      all                            (default) All of the above
  --only-member=[name]               Only consider archive members names that contain [name]. Can be specified multiple times.
  --only-section=[name]              Only consider section names that contain [name]. Can be specified multiple times.
  --only-symbol=[name]               Only consider symbol names that contain [name]. Can be specified multiple times.
  --redact=[kind]                    Redact the specified field kind. Intended for snapshot testing.
      rva                            Relative virtual addresses
      va                             Virtual addresses and file offsets
      ord                            Symbol ordinals / hints
      size                           Sizes and lengths
      all                            All of the above
  --relocs                           Display relocations
  -s, --snapshot                     Alias for --redact=all --elements=-all
  --section-headers                  Display section headers
  --strings                          Display string tables
  --symbols                          Display symbol tables
  --tls                              Display TLS information

Some example output:

❯ zig objdump mathtest-dync-exe-no-llvm.dll --all-headers
mathtest-dync-exe-no-llvm.dll: PE/COFF image


COFF Header:
            8664 machine (AMD64)
               7 number_of_sections
        6a08670a time_date_stamp
               0 pointer_to_symbol_table
               0 number_of_symbols
              f0 size_of_optional_header
            2022 flags
               | EXECUTABLE_IMAGE
               | LARGE_ADDRESS_AWARE
               | DLL

COFF Optional Header:
             20b magic (PE32+)
           14.00 linker_version
          14a400 size_of_code
           84e00 size_of_initialized_data
               0 size_of_uninitialized_data
            1000 address_of_entry_point (       180001000)
            1000 base_of_code (       180001000)
       180000000 image_base
            1000 section_alignment
             200 file_alignment
            6.00 operating_system_version
            1.00 image_version
            6.00 subsystem_version
               0 win32_version_value
          1d6000 size_of_image
             400 size_of_headers
               0 checksum
               2 subsystem (WINDOWS_GUI)
             160 dll_flags
               | HIGH_ENTROPY_VA
               | DYNAMIC_BASE
               | NX_COMPAT
          100000 size_of_stack_reserve
            1000 size_of_stack_commit
          100000 size_of_heap_reserve
            1000 size_of_heap_commit
               0 loader_flags
              10 number_of_rva_and_sizes

Data Directories:
          1b48d0       54 EXPORT
          1b4924       8c IMPORT
               0        0 RESOURCE
          1cc000     633c EXCEPTION
               0        0 SECURITY
          1d4000     1280 BASERELOC
          1bc000       1c DEBUG
               0        0 ARCHITECTURE
               0        0 GLOBALPTR
          1aebc8       28 TLS
               0        0 LOAD_CONFIG
               0        0 BOUND_IMPORT
          1b4c80      2d0 IAT
               0        0 DELAY_IMPORT
               0        0 COM_DESCRIPTOR
               0        0 RESERVED

Sections in 'mathtest-dync-exe-no-llvm.dll':
Num Name          RVA Virt Size Data Size   & Data & Relocs  & Lines # Relocs  # Lines    Flags
  1 .text        1000    14a346    14a400      400        0        0    0    0 60000020 | CNT_CODE MEM_EXECUTE MEM_READ
  2 .rdata     14c000     6fe3c     70000   14a800        0        0    0    0 40000040 | CNT_INITIALIZED_DATA MEM_READ
  3 .buildid   1bc000        52       200   1ba800        0        0    0    0 40000040 | CNT_INITIALIZED_DATA MEM_READ
  4 .data      1bd000      e2a0      d200   1baa00        0        0    0    0 c0000040 | CNT_INITIALIZED_DATA MEM_READ MEM_WRITE
  5 .pdata     1cc000      633c      6400   1c7c00        0        0    0    0 40000040 | CNT_INITIALIZED_DATA MEM_READ
  6 .tls       1d3000        20       200   1ce000        0        0    0    0 c0000040 | CNT_INITIALIZED_DATA MEM_READ MEM_WRITE
  7 .reloc     1d4000      1280      1400   1ce200        0        0    0    0 42000040 | CNT_INITIALIZED_DATA MEM_DISCARDABLE MEM_READ

No symbol table found

If this output was used for a snapshot test that wanted to test the presence of a particular export:

❯ zig objdump mathtest-dync-exe-no-llvm.dll --exports --only-symbol=add --redact=rva --elements=-all
Export directory:
               0 flags
               0 time_date_stamp
            0.00 version
xxxxxxxxxxxxxxxx name_rva
               1 ordinal_base
               1 number_of_entries
               1 number_of_names
xxxxxxxxxxxxxxxx export_address_table_rva
xxxxxxxxxxxxxxxx name_pointer_table_rva
xxxxxxxxxxxxxxxx ordinal_table_rva
   1    0 xxxxxxxx | add

The redaction ( --redact= ) and element removal ( --elements= ) functionality is used to remove parts of the output that don't matter for the particular test, as to not cause spurious failures if say, an RVA for a symbol changes due to some unrelated change to the linker. -s is shorthand which enables all the redacts and removes all extra output elements, but particular tests may not use this if they do care about specific values.

resinator §

Windows Resource Compilation Moving to an External Package §

Windows resource script compilation will be moved out of the compiler and into an official-but-external build system package in the next release. As such, the corresponding std.Build functions and fields ( Build.Module.addWin32ResourceFile , etc) have been marked as deprecated in this release.

The zig rc subcommand will remain after this change in order to continue supporting the use case of using the Zig toolchain with other build systems.

zig fmt §

Added --complexity Flag §

It's a simple tool for counting tokens and AST nodes, which can be used to share how an edit to a source file increases or reduces complexity with a heuristic that is more insightful than line count.

For example, running it on old ELF linker directory:

info: src/link/Elf/gc.zig: tokens=1538 nodes=770
info: src/link/Elf/Atom.zig: tokens=15576 nodes=7768
info: src/link/Elf/Merge.zig: tokens=2288 nodes=1092
info: src/link/Elf/Thunk.zig: tokens=1128 nodes=519
info: src/link/Elf/SharedObject.zig: tokens=4566 nodes=2236
info: src/link/Elf/eh_frame.zig: tokens=4840 nodes=2335
info: src/link/Elf/Symbol.zig: tokens=3770 nodes=1872
info: src/link/Elf/synthetic_sections.zig: tokens=13287 nodes=6375
info: src/link/Elf/Object.zig: tokens=13903 nodes=6842
info: src/link/Elf/LinkerDefined.zig: tokens=4140 nodes=2091
info: src/link/Elf/relocatable.zig: tokens=3598 nodes=1848
info: src/link/Elf/Archive.zig: tokens=2176 nodes=1028
info: src/link/Elf/ZigObject.zig: tokens=20783 nodes=10274
info: src/link/Elf/AtomList.zig: tokens=1819 nodes=933
info: src/link/Elf/relocation.zig: tokens=1289 nodes=551
info: src/link/Elf/file.zig: tokens=2330 nodes=1087
info: total: tokens=97031 nodes=47621

Another example is after editing Lld.zig to use arena.print rather than std.fmt.allocPrint . The line count is roughly the same but --complexity tells a different story:

  • before: tokens=11929 nodes=5900
  • after: tokens=11832 nodes=5852 (1% reduction in source complexity)

In the future this metric may be helpful for an automatic import organizer to determine whether to create an import alias or not.

Roadmap §

  1. Work with ZLS team to enhance Build Server Protocol to the point where satisfies all their use cases.
  2. Complete Windows support in the x86_64 backend so that it can be enabled by default.
  3. Complete and stabilize the language .
  4. Complete the aarch64 Backend and make it the default backend for debug mode.
  5. Enhance the Linker implementations, eliminating dependency on LLD and support Incremental Compilation .
  6. Enhance the integrated Fuzzer to be competitive with AFL and other state-of-the-art fuzzers.
  7. Transition from a library dependency on LLVM to a process dependency on Clang ( #16270 ).
  8. Complete the Build System , in particular Package Management features.
  9. Audit the Standard Library ( #1629 ).

Thank You Contributors! §

Ziggy the Ziguana

Here are all the people who landed at least one contribution into this release:

  • Alex Rønne Petersen
  • Andrew Kelley
  • Matthew Lugg
  • Casey Banner
  • Jacob Young
  • Frank Denis
  • Ali Chraghi
  • Isaac Freund
  • Pavel Verigo
  • Ryan Liptak
  • Techatrix
  • Justus Klausecker
  • Linus Groh
  • xtex
  • rpkak
  • David Rubin
  • Elaine Gibson
  • David Senoner
  • K4
  • fardragon
  • pentuppup
  • Krzysztof Wolicki
  • Mason Remaley
  • Meghan Denny
  • Christophe Delage
  • Robbie Lyman
  • Lukáš Lalinský
  • Brandon Black
  • Kendall Condon
  • hemisputnik
  • whatisaphone
  • Ben Anderman
  • GasInfinity
  • Huang Zhichao
  • Jakub Konka
  • Jari Vetoniemi
  • Kleshzz
  • Mick Sayson
  • Mikołaj Rosowski
  • Ryan Mehri
  • Sertonix
  • jeffkdev
  • lg
  • Carl Åstholm
  • EJ
  • Hila Friedman
  • Quint Daenen
  • Saurabh Mishra
  • Stephen Gregoratto
  • Subo2002
  • Ten Eugene
  • zacoons
  • Arthur Teixeira
  • Ben Burkert
  • Bernard Assan
  • CursedByTheVoid
  • Daniel Kareh
  • Ennui Langeweile
  • Igor Anić
  • Janne Hellsten
  • Jay Petacat
  • Nguyễn Gia Phong
  • Ragul R
  • Rue04
  • Ryan Mehri
  • Sam Connelly
  • Scott Redig
  • Theo Fabi
  • Vitor Fernandes
  • c-kappel
  • gero3
  • glowsquid
  • taoqy
  • teflate
  • xeondev
  • AI Xin
  • ARandomOSDever
  • Acid Bong
  • Adrià Arrufat
  • Akshay Trivedi
  • Alan Cocanour
  • Alex Kladov
  • Amilia MacIntyre
  • Anders Stenberg
  • Andrew Kraevskiii
  • Anshul Gupta
  • Anthon van der Neut
  • Ari Becker
  • Ashley (holtowd)
  • C19
  • Carter Snook
  • Chadwain Holness
  • Chloé Vulquin
  • Chris Boesch
  • Christoffer Lerno
  • Corentin Kerisit
  • David
  • Devin J. Pohly
  • Diego Peña y Lillo
  • Dmitry Mostovenko
  • Eric Joldasov
  • Erik Schlyter
  • FalsePattern
  • Fitti
  • Gabriel Sa
  • Gereon V
  • Giuseppe Cesarano
  • Gota7
  • GrimTyr
  • Guillaume Wenzek
  • Guy Fischman
  • Henry Kupty
  • Ignacio Ibarra
  • Jan Procházka
  • Jan200101
  • Jeff Fowler
  • Jeremy Linton
  • John Benediktsson
  • Jonathan Marler
  • Joseph Lyncheski
  • Josh Megnauth
  • Kirk Scheibelhut
  • Leon Lombar
  • Leonid Emar-Kar
  • Mai-Lapyst
  • Manlio Perillo
  • Marcel W. Wysocki
  • Mark Rushakoff
  • Mathieu Suen
  • Matthew Knight
  • Matthias Portzel
  • Michael Farber Brodsky
  • Miloš Kozák
  • Nguyen Gia Huy
  • Nico Elbers
  • Noam Rothschild
  • Nurul Huda (Apon)
  • Paul Anderson
  • Paulo Duarte
  • Perry Fraser
  • Pivok
  • Preeternal
  • Prokop Randáček
  • RadsammyT
  • Rayan Halder
  • Richalsu
  • Richard Levitte
  • Rue04
  • Ryan Davis
  • Ryan King
  • Sage Hane
  • Samuel Hunter
  • Sawyer X
  • Siddharth Sinha
  • Steve
  • Tapir
  • Thibault Leclercq
  • Yasin Gorgij
  • Zhalkhas
  • abdessalem
  • agave
  • ahwayakchih
  • akbarhusain
  • arshidkv12
  • avalyn0x45
  • badayvedat
  • binarycraft007
  • brianferri
  • brkzlr
  • carmooo
  • cheesecakecatttt
  • conmaster2112
  • drex_vk
  • gDator
  • gemmaro
  • gubbu
  • ilovapples
  • jmcaine
  • johan0A
  • jpk68
  • krystiann
  • kubashi
  • l1yefeng
  • levinqua
  • llogick
  • marximimus
  • mihael
  • mlugg
  • nash1111
  • ndjenks
  • nekogirl
  • nyx-xyn
  • pancelor
  • skdishansachin
  • sphaerophoria
  • squidy239
  • sstochi
  • vlkrs
  • zirunis
  • Ömer Faruk IRMAK
  • Έλλεν Εμίλια Άννα Zscheile
  • 林晨 (Leo Cheng)

Ziggy the Ziguana

Special thanks to those who sponsor Zig . Because of diverse, recurring donations, Zig is driven by the open source community, rather than the goal of making profit. In particular, those below sponsored Zig for an average of $50/month or more during this release cycle using our preferred donation platform :

  • Sergey M
  • Mitchell Kember
  • David Vanderson
  • Thomas Manner
  • Kirk Scheibelhut
  • Kazuhiro Kondo
  • Merlyn Morgan-Graham
  • Freddi Linse
  • Numan Sachwani
  • Aurélien Cibrario
  • Erik Dunteman
  • Ondra Voves
  • Mitchell Gayner
  • Dylan Conway
  • Trevor John
  • Jason Watson
  • Greg Clark
  • Benjamin Crist
  • Alexander Weavers
  • Natalie Vais
  • Stevie Hryciw
  • Kyle Hill
  • Felix Queißner
  • Srinivasan Balram
  • Peter Ronnquist
  • Lajos Nagy
  • Gauthier Voron
  • Andrew Mangogna
  • Alex Kladov
  • Jacob Sandlund
  • Jordan Lucier
  • Wolfgang Sanyer
  • Matthew Knight
  • Nick Macholl
  • Ceri Elenbaas
  • Michael Keathley
  • Daniele Cocca
  • Bartosz Bogacz
  • James Cox-Morton
  • William Canan
  • Fabio Arnold
  • Johannes Meyer
  • Silver van Koten
  • Flavius Gruian
  • Daniel MacDougall
  • Karrick McDermott
  • Dan Mack
  • Caleb Hearon
  • Richard Levitte
  • Ingimar Jóhannesson
  • David Fendley
  • Aaron Cross
  • Jeremiah Oard
  • Simon Ekström
  • Jorge De León
  • Saurabh Mishra
  • Rob Green
  • Kytezign q
  • Mykhailo Tsiuptsiun
  • Paul Sargent
  • Erik Mållberg
  • Igor Anic
  • Fawzi Mohamed
  • David Jones
  • Magnus Holm
  • Nicholas Woolmer
  • Samarth Kishor
  • Charles Haws
  • Shlomi Atar
  • Robbie Lyman
  • Andrius Bentkus
  • Jean-Luc Geering
  • Aaron Mady
  • coleman broaddus
  • Trace Andreason
  • Jim Calabro
  • Chris Baldwin
  • Ian Johnson
  • Brandon Black
  • Michael Lynch
  • O Y
  • Francesco Gualazzi
  • David Sugar
  • Malcolm Still
  • Jeff Fowler
  • Yaroslav Zhavoronkov
  • Miles McGruder
  • Álvaro Justen
  • FELIPE SOARES GONCALVES SA ROSA
  • Noah Betzen
  • Manuel Barkhau
  • Rikard Karlsen
  • Frank Ittermann
  • Eli Janssen
  • Peter Snelgrove
  • Will Pragnell
  • Pete Dietl
  • Vaughn Spielman
  • Michael Kato
  • Daniel Dubecky Hodan
  • Dan Boykis
  • Richard Feldman
  • Markus Ort
  • Peyman Mortazavi
  • Wilson Bilkovich
  • Kev Burns
  • Francisco Nevitt Gonçalves
  • Mark Banhidi
  • Kristoffer Ström
  • Carl Distefano
  • Alexander Reustle
  • Martin Hovda Haugsand
  • Lexa Tang
  • Lennart Tuijnder
  • Enver Bisevac
  • Jameson York
  • Hong Shick Pak
  • Travis Staloch
  • Clover Caruso
  • Matthew Chavez
  • Daniel Gregoire
  • Jose M Rico
  • Xavier Cochran
  • Nas Denkov
  • Pavel Rychlý
  • papa leromi
  • Ronald Zielaznicki
  • Tommi Komulainen
  • Matthew Jee
  • Johan Forsberg
  • Łukasz Mróz
  • Fabio Leimgruber
  • Maurice van Veen
  • Dean Simmons
  • Pierre Marc Levasseur
  • Sebastian Appler
  • Sam Windell
  • Aksel Hjerpbakk
  • Jonathan Helland
  • Karl Fleischmann
  • Erik Hansen
  • Chris Boesch
  • Pablo Álvarez
  • Jacob Hooper
  • Avinash Lakshman
  • Joshua Park
  • Nicholas Clark
  • Ryan Higgins
  • Jeffrey Ollie
  • Adam Goertz
  • David Nowotny
  • Gus Louw
  • David Francoeur
  • David Sparby
  • Tyler Bender
  • Mark Halonen
  • Reinis Taukulis
  • Nancy Remaley
  • Isaac F
  • Kjell Hoffhenke
  • Vlad Panazan
  • Jordan Rowland
  • Jay Van Der Wall
  • Jean-Philippe Quenord
  • Jakob Külzer
  • Varun Pramanik
  • Benjamin Edwards
  • Andreas Herrmann
  • Adam Nilsson
  • Joseph Ruiz
  • Anthony Nguyen
  • Marcin Wolcendorf
  • Tobias Lahrmann Hansen
  • Pat Smuk
  • Martin Weber
  • Jordan Kardon
  • Benjamin LE BERRE
  • John Goen
  • Sebastian Ahlman
  • Anthony Hernandez
  • Shail Patel
  • Alexander Genaud
  • Morten Dalfoss
  • Jason Dubaniewicz
  • Mark Hayes
  • Eivind Rovik
  • Renan Silva
  • Isak Källman
  • Ladislav Böhm
  • LeRoyce Pearson
  • Rowan Saunders
  • Joshua Masci
  • Raymond Imber
  • Daniel Worley
  • Owen Cabalceta
  • Erez Shomron
  • Max Grosse
  • Ehden Sinai
  • Christopher Redden
  • Antoine Balaine
  • Chris Durkin
  • Dave Wallace
  • Aura Birb
  • Andy Armstrong
  • Spencer Brower
  • Simon Clavet
  • Guido Schmidt
  • Peter McGaughey
  • Callum McArthur
  • David Lei
  • David Brotz
  • Matti Hänninen
  • Joseph Chan
  • Vincent Weber
  • Jeremy McAdams
  • Morgan Gallant
  • Brett Pechiney
  • Thilina Jayanath Nenathunga Liyanage
  • Kevin Bockelandt
  • Gunnar Zötl

Special thanks also to TigerBeetle , Synadia Communications , and ZML for substantial contributions.

Federal Judge Rules a Flock Search Was ‘Indiscriminate Mass Surveillance’ and Unconstitutional

403 Media
www.404media.co
2026-10-02 16:49:15
Flock’s nationwide network is quickly “approaching dragnet-type law enforcement practice” and the cop should have got a warrant, the judge wrote....
Original Article

A federal judge in Oklahoma ruled Thursday that a police officer violated the Fourth Amendment rights of a woman accused of meth trafficking when he searched her license plate in Flock’s automated license plate reader system simply because her license plate was from California, then used her travel history as part of the reason to search her car. The judge’s opinion is one of the first times a federal judge has decided Flock searches can be unconstitutional, and suggested that Flock’s network is “a type of indiscriminate mass surveillance.”

The officer’s “use of the ALPR Systems was an Unconstitutional Warrantless Search,” and “was not supported by probable cause, and it was done without a warrant in violation of [the defendant’s] Fourth Amendment rights,” the judge, Sara Hill, wrote, implying that the law enforcement officer should have obtained a warrant before searching for the vehicle in Flock’s system. There are currently more than a hundred thousand warrantless searches of the Flock system every month, according to audit logs viewed by 404 Media. Hill's decision will not set a binding precedent and there are several other cases throughout the nation considering the legality of warrantless ALPR searches.

Hill argued that previous judge opinions saying Flock searches were not a Fourth Amendment violation because they track cars in public do not consider the context that Flock’s nationwide network is quickly “approaching dragnet-type law enforcement practice,” and that courts should update their understanding of the technology moving forward.

The circumstances of the court case are really interesting and highlight how commonplace Flock searches have become for police, and the depth of the information they can reveal. In May, a Tulsa County Deputy Sheriff named Freddie Alaniz was parked along the side of the highway in Oklahoma when he saw a Mazda SUV driven by a woman named Melisa Kyle with a California license plate pass by. “Alaniz then pulled his vehicle on the highway to follow the Mazda for no apparent reason other than the fact that it had a California license plate. Alaniz also ran a query on the Flock system for the California license plate number on the Mazda SUV,” Hill wrote. Alaniz then ostensibly pulled Kyle over for changing lanes without a turn signal.

Alaniz interrogated Kyle about her travel “while he continued to review the ALPR systems for the car she was driving,” the judge wrote. Alaniz made Kyle recount everything she had done in the last several days, and compared it to the Flock data. He told her that because she was only in California for a short period of time, he suspected her of trafficking drugs. He used her travel history as seen in the Flock system as part of the justification to search her car; she was found to have 91 pounds of meth in the vehicle. Hill ruled that all Flock evidence and all evidence from Alaniz’s search of the car must be thrown out.

“The Fourth Amendment requires courts to draw a line when the cost is too great. Alaniz’s search in just the ALPR system provided him with more than 50 individual records of Kyle’s whereabouts across the country for an entire month,” Hill wrote. “The Court finds that because the ALPR systems Alaniz used to search Kyle’s historical location information intruded on her reasonable expectation of privacy in the whole of her physical movements, it was a search under the Fourth Amendment. Based on the information in the record, the only reason Alaniz conducted that search was because he saw her license plate was from California.”

“The factors that the government relies upon are the same type of circumstances that everyday Americans encounter on long road trips for many legitimate reasons. Many of us drive longer than we want to get to a desired destination, or to no destination at all other than the road and sights ahead,” she added.

The decision is a landmark one, and comes in the aftermath of the Supreme Court’s Chatrie v United States decision that found police accessing a person’s digital data, including cell phone location data, constituted a search.

“The opinion is pretty amazing. It recognizes one thing that courts ignore which is the sheer breadth of these systems, that they collect so much information about so many people in a way that sets them apart. This decision ascribes appropriate weight to the fact police are building out this massive database that can reveal incredibly intimate details of people’s lives,” Michael Soyfer, a lawyer at the Institute for Justice, which has studied Flock camera abuse and is litigating several cases on Fourth Amendment grounds, told 404 Media. “It’s extremely important. The way courts have resolved these cases previously has been way too myopic and has ignored the depths of these systems and the sweeping modes of surveillance that allow police to reconstruct the movements of anyone in the country.”

Hill’s opinion also comes on the back of a decision earlier this week in a case the Institute for Justice brought. In that, a jury found a traffic stop scheme involving license plate reader scans done by U.S. Border Patrol as part of a predictive policing unit were unconstitutional.

The new decision also immediately invalidates the core argument that Flock’s CEO Garrett Langley has made saying that Flock was not a constitutional issue. “You and I don’t get to pick what’s a constitutional violation and what’s not. We have judges, we have elected officials, there’s a process for that. We follow the law, we follow the Constitution. So far, in our belief and what will be for a long time, the courts have deemed this is not a warrantless search; this is a valid product as it relates to the Fourth Amendment. So I don’t see any change there,” Langley told The Drive in July , adding the issue was “pretty cut and dry.”

Notably, Hill suggested that other courts that have ruled Flock searches do not constitute a Fourth Amendment search were likely wrong to do so, and that they have not considered the widespread and automated context of the AI-powered surveillance system.

Previous decisions that ruled ALPR searches do not require a warrant have leaned on a Supreme Court case called United States v Knotts , in which police put a tracking device in a chemical container after being tipped off that an employee of a chemical plant was stealing from their employer. That case was decided in 1983 and found, “[a] person travelling in an automobile on public thoroughfares has no reasonable expectation of privacy in his movements from one place to another.” But Hill wrote, “that language exists in the context of the facts presented in the case. Rather than a large-scale, dragnet-type surveillance system like the ALPR technology in this case, the Court in Knotts was confronted with much less sophisticated technology.”

“The Court acknowledges that people, in at least a broad sense, do not have a reasonable expectation of privacy in their movements on a public roadway,” Hill wrote. “But by virtue of how ALPR technology works, Alaniz and other officers using these systems have access to a continuously updated location history for all vehicles caught on ALPR cameras within the network. This is a type of indiscriminate mass surveillance. It is not targeted on a single individual, as in Carpenter [another Supreme Court case about phone data specifically]. It is a tool that collects information about all vehicles that pass by any network-connected camera at all times, and it serves up the information to law enforcement on demand.”

We have seen several cases in which cops have used Flock data to pull people over because they have crossed state lines , then have worked backward to justify their travel patterns as a reason to search their vehicles.

“We’re seeing that repeatedly with police flagging whatever they’ll call suspicious patterns of movement. Federal agents were using ALPRs to monitor cars making day trips across the border and back to manufacture a basis to stop them, interrogate the drivers and search them,” Soyfer said. “I think Flock is going to automate that using AI where cops can set alerts for those kinds of travel patterns. When we’re arguing these systems are very powerful and can show a lot about people’s movements, cops dismiss this as speculative or not possible, but then they deploy this strategy against people who they stop all the time.”

Flock Safety did not immediately respond to a request for comment.

About the author

Jason is a cofounder of 404 Media. He was previously the editor-in-chief of Motherboard. He loves the Freedom of Information Act and surfing.

Jason Koebler

At the Conformation Show: Chasing the Ideal Dog

Hacker News
www.theparisreview.org
2026-10-02 16:46:11
Comments...
Original Article

Why have I been blocked?

This website is using a security service to protect itself from online attacks. The action you just performed triggered the security solution. There are several actions that could trigger this block including submitting a certain word or phrase, a SQL command or malformed data.

What can I do to resolve this?

You can email the site owner to let them know you were blocked. Please include what you were doing when this page came up and the Cloudflare Ray ID found at the bottom of this page.

Woking Electrical Control Room (2016)

Hacker News
www.darbiansphotography.com
2026-10-02 16:44:38
Comments...
Original Article

Feel free to follow me on the social media below and add your email to the box for notifications of new posts on here.

Woking Electrical Control Room: External

Woking Electrical Control Room: External

Woking Electrical Control Room: Now that is some beautiful door closing mechanism.

Woking Electrical Control Room: Now that is some beautiful door closing mechanism.

Woking Electrical Control Room: Control room entrance.

Woking Electrical Control Room: Control room entrance.

Woking Electrical Control Room: Control wall, each location represents a substation or transformer.

Woking Electrical Control Room: Control wall, each location represents a substation or transformer.

Woking Electrical Control Room: Details.

Woking Electrical Control Room: Details.

Woking Electrical Control Room:  So many lights.

Woking Electrical Control Room: So many lights.

Woking Electrical Control Room: In the corner.

Woking Electrical Control Room: In the corner.

Woking Electrical Control Room: One of the grand uplit lamps.

Woking Electrical Control Room: One of the grand uplit lamps.

Woking Electrical Control Room: I forgot what the symbols mean.

Woking Electrical Control Room: I forgot what the symbols mean.

Woking Electrical Control Room: Close up of the control panel.

Woking Electrical Control Room: Close up of the control panel.

Woking Electrical Control Room: View from in front of the main desk.

Woking Electrical Control Room: View from in front of the main desk.

Woking Electrical Control Room: The main desk.

Woking Electrical Control Room: The main desk.

Woking Electrical Control Room: The clock.

Woking Electrical Control Room: The clock.

Woking Electrical Control Room: Close up of the main desk.

Woking Electrical Control Room: Close up of the main desk.

Woking Electrical Control Room: Vintage phone.

Woking Electrical Control Room: Vintage phone.

Woking Electrical Control Room: Phone and switchboard.

Woking Electrical Control Room: Phone and switchboard.

Woking Electrical Control Room: Another close up of the desk, this was possibly for fault alarms.

Woking Electrical Control Room: Another close up of the desk, this was possibly for fault alarms.

Woking Electrical Control Room: How many phones did they need? The winder on the right was used to power the telephones.

Woking Electrical Control Room: How many phones did they need? The winder on the right was used to power the telephones.

Woking Electrical Control Room: Due to the other visitors I only got to shoot this half.

Woking Electrical Control Room: Due to the other visitors I only got to shoot this half.

Woking Electrical Control Room: I know some of my journeys as a child were helped along by this.

Woking Electrical Control Room: I know some of my journeys as a child were helped along by this.

Woking Electrical Control Room: The end wall

Woking Electrical Control Room: The end wall

Feel free to leave your views on this article below.

★ Apple Is Going to Further Tighten the Screws on Full Disk Access on MacOS, in Response to Agentic AI Apps Running Amok

Daring Fireball
daringfireball.net
2026-10-02 16:40:15
Hopefully Apple has in mind a solution to this situation that will still enable knowledgeable power users to confirm agreement to a sufficiently scary warning and put their Macs in a state similar to what we have today. I worry....
Original Article

Apple Developer News, in a post titled “ Updates to Full Disk Access in macOS ”:

We give developers powerful APIs to build incredible capabilities into their apps for Apple products, backed by a set of controls designed to protect users’ private data. Full Disk Access largely sidesteps these controls in order to allow backup apps to function properly on the Mac. Some developers are using Full Disk Access in ways that could put users at risk, exposing everything on their systems — including files, mail, messages, and even browsing history — without users’ full knowledge and understanding. For communication apps, this can also compromise the privacy of the people users are communicating with.

Going forward, we will introduce additional controls to ensure that users who genuinely wish to grant an app this extraordinary level of access can only do so with very explicit user action. Addressing this is critical. As AI agents become increasingly capable and autonomous, the risks associated with this level of access will grow substantially. We are committed to ensuring users clearly understand these risks before granting such access, so they can make informed decisions about their own data and privacy.

They don’t name names or even mention “agentic AI”, but clearly this is in response to Meta Muse (cf. Jason Aten’s misadventure with Muse accessing Aten’s iMessages), Grok Bot , Claude, Dots, and the rest. This is why we can’t have nice things.

I really worry about just how much Apple is going to lock Full Disk Access down. I use it with several apps that couldn’t function properly without it, and would be severely hampered if I needed to authorize it manually every time they do something.

But before we Mac power users riot in Cupertino, we should note that we have no idea how many non-sophisticated Mac users are calling Apple and queuing up at the retail store Genius Bars to complain about Muse and these other agents having run amok on the Macs, against their desires, after convincing these users to grant them access. It is a legitimate frustration for the highest-functioning users among us that the Mac has, for years, already seemed “too locked down”. But there are now around 150 million Mac users worldwide. Most of them are unsophisticated technically — a majority of them, profoundly so. Many of them have technical needs that cannot be met by the baby computer OS that is iPadOS. For a lot of them it might just be the ability to run the real version of Google’s Chrome web browser — or even the desktop version of Safari. So there are tens of millions of people who need to use a Mac to do things that cannot be done on an iPad, but who have zero understanding of what it means, say, to grant Muse permission for Full Disk Access. And then they think it is Apple’s fault that Muse is suddenly able to read their private iMessage conversations.

On iOS (and its tablet variant, iPadOS), you can say OK to every single thing an app like Muse asks for and Muse still won’t be able to read your email or end-to-end encrypted messages from iMessage or WhatsApp. There is no level of permission that grants third-party apps permission for such things. (The EU wants to force Apple to allow that under the DMA.) MacOS isn’t like that. You still need to grant apps permission for such things, but if you say OK to everything an app like Muse asks for on the Mac, you’re granting that app access to, effectively, almost everything on your startup drive. There are still some things it can’t see, but not many. A lot of non-technical Mac users do not understand this and cannot be expected to understand this. They just think, wrongly, that Apple protects them from allowing anything truly dangerous, because that’s how their iPhone (and/or iPad) works. So they just click OK to every access request from Muse and presume they’re still largely protected.

Hopefully Apple has in mind a solution to this situation that will still enable knowledgeable power users to confirm agreement to a sufficiently scary warning and put their Macs in a state similar to what we have today. I worry. What alleviates my worst fears is the knowledge that every technical user at Apple itself needs to use their Mac as the powerful Unix workstation OS that it is. Some of us need dangerously powerful tools. Most Mac users, however, do not — and don’t realize they’re using a dangerously powerful Unix workstation with a very friendly ( literal ) face.

Barcodes are about to go extinct

Hacker News
thehustle.co
2026-10-02 16:37:07
Comments...
Original Article

QR codes are here. Barcodes should be very afraid.

A QR code on display at a fruit market

In 1992, on a factory floor in a Toyota auto-parts subsidiary in Japan, workers were drowning in barcodes.

They were everywhere: some components on the line needed ten barcodes just to capture all the tracking data. Employees scanned them, one by one, hundreds of times a day.

Masahiro Hara , an engineer on the barcodes team at Denso, the subsidiary, tackled the problem. What if, he reasoned, you could build a barcode in two dimensions instead of one? You could pack it with information, cutting the number of scans from ten to one on every item.

In the company cafeteria, Hara watched as his coworkers squared off over a game of Go, a board game played with black and white stones on a grid. The contrast he saw there inspired the design for what would eventually become the QR code, a code that can be read from any angle and hold 200x more data than the barcodes emblazoned on the products we buy today.


Workers at Denso’s auto parts factory were scanning components covered in barcodes before QR emerged. (Photo by Si Wei/VCG via Getty Images)

After a year and a half of work on the QR, Denso patented the innovation. Rather than enforce royalties on it, they published the specs so any company could use it for free.

And that decision is about to change every single thing you buy.

QR code takeover

Twenty years earlier, in 1974, a cashier at a grocery store in Troy, Ohio, scanned the world’s first barcode, affixed to a pack of Wrigley's Juicy Fruit gum.

Grocery retailers and manufacturers had recently formed a committee and chosen the Universal Product Code, or UPC, as the single industry standard.

The decision created a network effect : once enough retailers required UPC codes on products, manufacturers had to comply. And once manufacturers printed UPCs on packages, retailers needed to buy scanners.

It also created a self-reinforcing monopoly, one that with widespread adoption reached critical mass by 1980.

The Uniform Code Council, now known as GS1, received the authority to issue the first six numbers of every barcode, called company prefixes. It’s the only way to sell mainstream retail products: there’s no alternative registry.


Inspired by the board game Go, Masahiro Hara developed a 2-D barcode. (Photo by Kevin Frayer/Getty Images)

Today, GS1 charges annual licensing fees for company prefixes, which scale with company size. It didn’t so much become a monopoly as it was set up that way. Barcodes have been a retail stalwart for the last 50 years, on more than 1B products scanned billions of times a day across the globe.

Two years ago, though, GS1 made an announcement. The barcode was out.

The organization was ushering in a new era, a global shift unlike any the retail industry has seen before. Much like the transformation of the ‘70s, retailers would begin to move to QR codes.

The real estate of packaging

Today’s packaging has to do a nearly impossible number of jobs with limited real estate.

  • Sell the product.
  • Satisfy legal requirements.
  • Explain how to use it.
  • Offer information about sourcing, ingredients, authenticity, and disposal.

You try fitting all that on a bottle of nail polish. “This is one of my biggest battles,” says Nicole Light , founder and CEO of packaging strategy firm Power in Packaging. She’s worked on product development and packaging design in everything from fragrance to beverages to meat and seafood to, yes, nail polish. She’s used to shrinking text down to sizes smaller than the human eye can comfortably read just to make everything fit — and to keep products within their industry’s regulations.

It’s some of those regulations that prompted Sunrise 2027, the shift GS1 is introducing next year.


Food recalls, like the one on infant formula in 2021, meant manufacturers needed better tracking information for their products, and quickly. (Photo by Nathan Howard/Getty Images)

In the early 2020s, food recalls spiked. So did another problem for manufacturers: a barcode can’t tell you where a product originated, only what kind of product it is. That became an issue during contamination events like the great onion recalls of 2020 and 2021 and the great peanut butter poisoning of 2022. Later in 2022, the FDA mandated that manufacturers keep more detailed tracking information across the food supply chain, sparking a need for a better barcode.

The humble barcode has done its job brilliantly for 50 years, says Kate Hardcastle , business strategist and author of The Science of Shopping .

“It tells the till what the product is and how much to charge,” she says. “But that’s where the conversation ends.”

‘A futuristic experiment’

So what can a QR code do instead?

Hardcastle says a 2D barcode, like the QR, can tie a product to information about:

  • Its origins and authenticity.
  • Ingredients and allergens.
  • Expiry dates.
  • Environmental impact.
  • Care and disposal.

The same code works at checkout, in the stock room, and for the customer.

Hardcastle, who helps GS1 conduct research, says nearly half of retailers she and GS1 polled have already started to introduce QR capabilities.

“This is no longer a futuristic experiment,” she says. “The infrastructure of shopping is already being rewritten.”

And Light adds that while brands can wait until retailers demand they add a QR code to their products, most probably shouldn’t. Consider all the things that will need to change in the overhaul:

  • How systems talk to barcodes.
  • How all the tracking data is stored.
  • Labels on packaging.
  • And more.

Bigger brands have the systems (and budgets) to weather a change like this. Last year, Unilever unveiled an upgrade of products across brands like Dove, Vaseline, and Hellman’s. It was a massive undertaking involving 45k SKUs.

“I have worked with a lot of these brands, and they change their packaging like they change their underwear,” Light says.

“So, in the grand scheme of things, it’s the companies that have used the same box for decades that are going to have to change it now.”

GS1’s shift isn’t mandatory or government mandated. But as Walmart, Costco, Target, and other retail behemoths change over to new scanners and tracking databases, brands risk being left behind (ie., off the shelves).

So what does this mean for suppliers?

“Printing a new square on a box is the easy part,” Hardcastle says. The new code is a link — and links need infrastructure. Making sure that information is accurate and contemporary is the real task.

It also presents a dilemma for packaging designers: until retail catches up, GS1 requires products to carry both codes, within 50mm of each other.

“Brands that wait until retailers insist on it may find they have an enormous data problem and very little time to solve it,” Hardcastle says.


Get ready to see a whole lot more of these — on nearly everything you buy. (Photo by Debarchan Chatterjee/NurPhoto via Getty Images)

As for consumers? Hardcastle says her research with GS1 found that 79% of consumers preferred QR codes. But that doesn’t mean they’re scanning every box of berries in their grocery cart. Rather, Hardcastle says, they want useful information at the ready, waiting for when they actually need it. Think: a smattering of dinner recipes for parents in a pinch, or more information on a product’s origins for people in, say, a trade war.

“They don’t necessarily want more information,” she says. “They want the right information.”

As next year looms closer, Light says there are still scores of unanswered questions about how the changeover will play out. Will designers be responsible for quality control measures like scanning every QR code individually to ensure each link and landing page work? Will there be a number associated with each code so it’s easier to track inventory and keep databases?

“People don’t like change, and they don’t like change that’s going to cost money,” she says.

“No one is talking about it. It feels like a meteorite is hitting the planet and no one is paying attention.”

Keeping Futhark off the GPU

Lobsters
futhark-lang.org
2026-10-02 16:17:47
Comments...
Original Article

Posted on October 2, 2026

Earlier this year, Elias Smedegaard did a BSc thesis on efficiently computing sparse Jacobian matrices via automatic differentiation. For this post, it is not important to understand exactly what this is, or why it is useful - it relates to automatic differentiation , which I have yet to write a blog post about. For the purpose of this post, the salient detail is that part of his solution involves colouring a graph corresponding to the sparsity pattern of a matrix.

Now, optimal graph colouring is a famously NP-hard problem, but we did not need optimal colouring - we just needed decent , and implemented with an efficient algorithm. Elias found a pretty fast greedy algorithm for distance-2 colouring (Algorithm 3.1 in this paper if you are curious), but unfortunately the algorithm is inherently sequential. This is not a problem for Futhark as Futhark supports quite efficient in-place updates , but the problem is that Elias’s overall program has two steps:

  1. Colour a graph.

  2. Use the colouring to compute the Jacobian using massive parallelism.

We really want to be able to exploit parallel hardware (such as a GPU) for the second step, while doing the first step on the CPU. It turned out that some of Futhark’s implementation choices made it awkward to do this. This blog post is about how the problem arises, what a principled solution might be, and why we just did a hack instead.

How the problem arises

Futhark’s compilation model is quite simple. An array needs to be in memory, and if you use one of the GPU pipelines, then all arrays are put in GPU memory. This is not because we cannot represent a mixture of CPU and GPU memory in our internal representation , but simply a convenience of implementation. It is also a safe design, since CPU code can access GPU memory (via costly APIs), but GPU code cannot in general access CPU memory. That it works does not mean, however, that it is fast. Consider this hypothetical Futhark loop:

loop sum = 0 for i < n do
  sum + xs[i]

If xs is an array stored in GPU memory, then the loop body will contain code for laboriously copying a single element from GPU memory to somewhere in CPU memory, after which it can be added to sum (which is presumably stored in a register). This is slow ! Copying the single array element takes almost no time, but setting up the communication and blocking until the copy is done has a ruinous overhead - easily a couple of microseconds per array element .

A better idea would be to copy the xs array as a whole before the loop to amortise the communication cost, and that is indeed a better solution that we will get to below. But if an index into a GPU array makes it into CPU code, then we will generate a very slow GPU memory read, which is ruinous when inside a loop.

The distance-2 colouring algorithm implemented by Elias is essentially a couple of nested sequential loops manipulating arrays that represent graphs and stacks. Using a GPU backend to compile this code results in ruinous performance compared to using the sequential c backend; easily three orders of magnitude or more. If we use the c backend instead, then performance is excellent (basically what you’d get if you wrote it in C), but of course then the second part of the overall program also gets to run sequentially. As a compromise, we can use Futhark’s multicore backend which generates parallel CPU code, but I really want to compute those Jacobians on the GPU!

What a principled solution might be

Futhark does not support “separate compilation” where different parts of the program use different backends, and that also seems a somewhat clumsy solution to the problem. Instead, the best possible solution is what I hinted at above: look at how (or where ) arrays are actually used, and decide based on that whether to put them in CPU or GPU memory - or maybe redundantly in both, if the program calls for it.

We actually have an optimisation that does something similar, implemented by Philip Børgesen, which migrates simple sequential computation (not data!) to the GPU if the result is only used on the GPU anyway. The goal is to save on communication - if the work done is trivial, and all the time is spent on copying a few array elements around, then it is better to do the sequential work in a trivial single-thread GPU kernel, just to cut down on communication. Sadly, this optimisation does not apply here. First, the distance-2 colouring algorithm is too complicated (it has loops) to trigger the optimisation, and even if we disable these checks to force its execution on GPU, it turns out (unsurprisingly) that a single GPU thread is ruinously slow for running such heavily looping sequential code.

In an ideal world, we can imagine a counterpart to this optimisation that moves data instead of computation , and many of the underlying principles are likely going to be the same. In particular, we must be careful to avoid excessive copies, and also avoid keeping these copies around unnecessarily, as they increase memory consumption. I can see the contours of what this optimisation would look like, but I also see that it deserves more care and attention than I have time for at the moment.

Therefore…

The hack we did instead

Instead of trying to do something clever in the compiler, I added a new attribute , called #[cpu_function] . In my initial design, when put on a non-inlined function (usually enforced with #[noinline] ), #[cpu_function] caused the body of the function to be compiled to CPU code, as well as all input, intermediate, and output arrays to be stored in CPU memory, no matter which compiler backend is used. A caller of the function is responsible for ensuring that array arguments are in the correct memory space, and the compiler inserts code to ensure this. Here is an example usage:

#[noinline] #[cpu_function]
def seqsum [n] (xs: [n]i32) =
  loop sum = 0 for i < n do
    sum + xs[i]

I did it this way because I thought it would be easy. And it almost was! Our intermediate representation for memory has always, and by explicit design, been able to handle multiple kinds of memory in concurrent use (specifically, our mem type is parameterised by a memory space, which we also use to represent GPU shared memory). However, the fact that we have never before widely exercised this functionality caused some issues that had to be solved. In the spirit of things, the solutions I picked were whatever was most expedient, until I have time to do the principled thing sometime in the future.

One issue I encountered is that our code generation for GPU kernels assumes that all arrays referenced in the body were already in GPU memory, which is no longer the case. For example, imagine this is our program:

#[noinline] #[cpu_function]
def frob (n: i64) =
  iota n

entry main n = let xs = frob n
               in map (\x -> x + 1) xs

The map in main gets compiled to a GPU kernel, but the array xs is in CPU memory due to #[cpu_function] ! This is not valid. The principled solution is to actually look at which arrays are being accessed inside GPU kernels and copying them to the GPU if necessary, but doing this right is actually not so easy:

  1. If a GPU kernel is called in a loop, we should not copy for each iteration of the loop.

  2. …but we should also keep the array copies around for as short a time as possible.

My solution to this was to tweak the meaning of #[cpu_function] a bit: while the input and all intermediate results are in CPU memory, the result will be in GPU memory, achieved by a copy at the end of the function. This essentially means that the compiler can largely keep assuming that all arrays are in GPU memory, as this will still be the case for everything except those functions marked with #[cpu_function] .

…of course, those functions then turned out to be a problem. The GPU backends assumed that all iota and replicate operations (which are IR primitives) would produce results in GPU memory, which is not the case inside a #[cpu_function] . Fixing that was not so difficult , but it is likely that similar bugs lurk elsewhere.

Finally, a question arises of what should happen when you map a #[cpu_function] . I arbitrarily decided that we then ignore #[cpu_function] and generate a normal parallel lifted function , but an argument could also be made that this should produce a sequential function somehow.

Where we are now

While #[cpu_function] is a somewhat sharp-edged hack, it turned out to solve the immediate problem: we can now do graph colouring using very efficient sequential code, and also do parallel things within the same Futhark program. In particular, graph colouring is now a very small part of the overall run-time, which is dominated by the subsequent numerical work. The #[cpu_function] technique is coarse in the sense that calling these functions involves expensive copies, but because they are bulk copies instead of per-element, the overhead is not so great.

The modifications required to the compiler were somewhat minor, and most can be categorised as bug fixes. Even in that glorious principled future where the Futhark compiler becomes able to automatically determine where an array should be stored, it is likely we will keep this attribute around as a useful hint.

Everyone's Packing Up

Hacker News
widdershins.verja.net
2026-10-02 16:13:04
Comments...
Original Article

Thoughts 💭 of Widdershins

I see a lot of posts about people leaving platforms. This one about leaving Reddit , this one about leaving WordPress , and a bunch of others about leaving Discord, GitHub, Codeberg, Photoshop, and more. There are also those talking about leaving the web entirely for the smallweb (Gopher, Gemini, etc).

The reasons vary. Pricing changes, AI training on everything you post, a new owner, a new policy, or just the slow feeling that things feel like they stopped being built for a person. But the posts all have the same shape. Someone spent years somewhere, it changed under them, and now they're packing up.

What strikes me is that hardly anyone is leaving for something big . They're going smaller. A personal site, a self-hosted app, a capsule nobody's going to stumble onto by accident. Less reach, more control. A few years ago that would have sounded like giving up. Now it sounds like the sensible option.

I don't think the big platforms are going anywhere. But the people who made them interesting in the first place keep quietly walking out the side door, and writing a post about it on their way out.

Which, I suppose, is what we're all doing here.

#indieweb #platform-exodus #smallweb #thoughts

Public Radio Station Forced to Call Gulf of Mexico the “Gulf of America”

Intercept
theintercept.com
2026-10-02 16:12:13
The University of Alabama at Birmingham is rolling over to Trump and his acolytes in the state, rather than stand up for editorial independence. The post Public Radio Station Forced to Call Gulf of Mexico the “Gulf of America” appeared first on The Intercept....
Original Article
President Donald Trump arrives at Ellington Airport in Houston, on August 27, 2026, ahead of a Republican National Committee fundraiser and a ceremony for the Artemis II astronauts at NASA Johnson Space Center. A man wears a red hat with the words ''Gulf of America'' embroidered on it. (Photo by Reginald Mathalone/NurPhoto via Getty Images)
A man wears a red “Gulf of America” hat as Donald Trump arrives in Houston’s Ellington Airport, on Aug. 27, 2026, ahead of a Republican National Committee fundraiser. Photo: Reginald Mathalone/NurPhoto via Getty Images

Seth Stern is the director of advocacy for Freedom of the Press Foundation.

Clayton Weimers is the executive director for Reporters Without Borders (RSF) North America.

Donald Trump received Chinese President Xi Jinping and the visiting Chinese press pool last week as reporters from news outlets he had banished stood outside the White House gates. At one point, Trump joked that China, the world’s leading jailer of journalists , “has the friendliest press corps.” Now, his acolytes in Alabama are taking up Trump’s cause, emulating the Chinese Communist Party’s mastery of information control at the local level by decreeing what supposedly independent public radio reporters are allowed to say.

The Alabama legislature passed a bill in late March requiring all state and local entities to refer to the Gulf of Mexico as the “Gulf of America,” in line with one of the many executive orders Trump signed on his first day in office. The legislation does not set forth penalties for noncompliance, but its sponsor, state Rep. David Standridge, called it “an opportunity to make a patriotic statement” on the House floor in January. Alabama Gov. Kay Ivey, a Trump ally, signed the bill into law.

Birmingham public radio station WBHM told its listeners in a letter from Will Dahlberg, the station’s executive director and general manager, that due to the law, which took effect October 1, the University of Alabama at Birmingham is forcing it to use the legislature’s preferred language in its reporting. The move puts it at odds with NPR, of which it is a member station; The Associated Press, which publishes the style guide used by newsrooms across the country; and other journalistic institutions. The station is licensed to the university, making its employees both state and UAB employees. Still, WBHM’s Journalism Code of Integrity emphasizes the newsroom’s independence and states that the university “should not — and cannot — expect to exert influence on WBHM’s journalists and its independent editorial process.”

Dahlberg wrote that not implementing the new language “may result in disciplinary action or termination of employees and affect WBHM’s relationship and support from UAB.” But nothing in the law requires the university to discipline or fire anyone. In other words, the university took a toothless stunt of a law, gave it teeth, and used them to take a bite out of press freedom.

Dahlberg also announced in the letter that he is stepping down from his position and leaving the station. His resignation was supposed to be effective October 16, but he posted an update that the university told him on Monday morning that he should pack his bags immediately. We cannot say for sure whether the university hastened Dahlberg’s departure due to his message to listeners and the criticism it generated, but the timing is certainly suspicious.

So why did the University of Alabama at Birmingham choose to embarrass itself rather than stand up for journalists and the First Amendment? As ridiculous as Alabama’s Gulf of America law is, its text did not tie the university’s hands. The university would not even have needed to bring a legal challenge to contest the law — it could have just let WBHM ignore it and supported its employees.

The university could have taken the position that the government choosing which words reporters must say is impracticable and burdensome.

The law only refers to “state and local entities” and makes no mention of media. It contains no penalties for noncompliance and limits its mandate to “when practicable,” giving a pass when complying would “impose an operational or financial burden.” Compelled speech is never practicable because it’s unconstitutional, but the operational and financial burdens of compliance are also significant. WBHM picks up national content from NPR and The Associated Press, for example. Are they expected to reject stories that say “Gulf of Mexico,” and fill that time with original reporting they haven’t budgeted or hired for? Or should they somehow intercept and edit those stories, likely in violation of their contractual obligations?

There is also ample Supreme Court precedent for the proposition that the government cannot put words in reporters’ mouths. In a landmark 1943 case on whether the state could force students to salute the flag or recite the Pledge of Allegiance, the court wrote that “if there is any fixed star in our constitutional constellation, it is that no official, high or petty, can prescribe what shall be orthodox in politics, nationalism, religion, or other matters of opinion or force citizens to confess by word or act their faith therein.” That means neither Standridge, Ivey, nor Trump is entitled to impose their definitions of patriotism on anyone —least of all the news media.

More recently, some conservatives celebrated 2024’s NRA v. Vullo decision, which found that public officials could not “use the power of the State to punish or suppress disfavored expression.” There is no exception for radio stations that receive public funding. In 1984, the Supreme Court struck down a ban on editorializing by public broadcasters that received federal money, holding that funding a station doesn’t entitle the government to control its journalism — a principle affirmed in March when a federal judge rejected Trump’s executive order to defund NPR and PBS.

The university could have taken the position that, for one, public radio stations are not state entities, and two, the government choosing which words reporters must say is impracticable and burdensome. Last but not least, the university should have taken the stand that the First Amendment still exists to protect us all, despite Trump’s best efforts to cast it aside.

Just as law firms , news outlets , colleges , and others that have settled with Trump to stay in his good graces chose to cave when they could have easily fought back (as others successfully have ), the University of Alabama at Birmingham made the cowardly choice to comply in advance.

If anything, its capitulation is worse, because there’s no indication that the state intended to, or could, put up a fight. Is the Alabama legislature really so committed to Standridge’s silly law that it would waste taxpayer money litigating likely unwinnable constitutional battles? Maybe, but the university could have at least waited to see how, or if, the state would respond.

It is worth noting that Ivey is an ex officio member of the University of Alabama System Board of Trustees. That same board published an annual freedom of speech report in August. It consists of a single identical paragraph, copy-pasted each year with the only change being the date and the number of “events” held on its campuses that year, concluding that “the requirements of the [free speech] Policies have been repeatedly satisfied without a single alleged violation, which further speaks to their effectiveness in appropriately achieving the Board’s objectives.”

One is left to wonder if, despite pressuring WBHM to echo Trump talking points, the university will once again copy-paste that conclusion next year.

Show HN: Made an open-source Lego AI generator

Hacker News
github.com
2026-10-02 16:00:15
Comments...
Original Article

Give an AI agent a model idea. Guide it, let it build it, and get an LDraw LEGO© model.

What you get when the building process finishes:

  • Its source code , in LDraw language .
  • Different views: 3D viewer, 3D player, VR interactive (Meta Quest 3), images...
  • Blender editable glTF file, in .glb format, metainfo as Blender's Custom Properties.
  • Chat history and agent thinking process .
  • and more... 👌

Take a look at the video :

Installation

Important

Tools used by agents in order to find suitable parts and example models take advantage of jev-rerank (I'm also the author). This is a semantic search tool with re-ranking backed by TypeSafe 's Jev System One AI model.

  • If you have a TypeSafe API key ( TYPESAFE_API_KEY ), set its value on the web app's Settings section.
  • If you don't , reranking search will not work , and agents will resort to a FTS (Full Text Search) strategy as a fallback, which could (maybe) yield worse models.

Run ldraw-nova as a web app, with Docker. You need Git and Docker .

This web app will run dockerized, and to build the Docker image, two sibling repos are required:

1. Clone both repos side by side, at the same tag , so they work together:

git clone --branch v0.6.0 https://github.com/anteloc/ldraw-nova.git
git clone --branch v0.6.0 https://github.com/anteloc/ldraw-nova-docker.git

2. Build the Docker image. The first build takes a while and needs about 5 GB of disk space:

cd ldraw-nova-docker
docker compose build

3. Start the app:

4. Open it in your browser:

  • https://localhost:8443 : needed for VR on Meta Quest 3. The certificate is self-signed, so accept the browser's warning the first time.
  • http://localhost:8765 : plain HTTP, no certificate warnings. Use it if the self-signed certificate gets in the way. VR won't work over it.

Other devices on your network can reach the app by your computer's IP instead of localhost , e.g. https://192.168.1.20:8443 from a Quest 3. The app has no login, so only run it on networks you trust.

To stop it:

Why all of this?

Well, to summarize: I did this in order to get agentic LLMs capable of designing buildable, physical things!

Finding LDraw , an assembly language (pun intended! 😜) that would be at the same time simple , low level , and executable in order to produce 3D CAD models , gave me the idea of experimenting with both ChatGPT and Claude in order to try and make them code in LDraw , same as they do with other programming languages.

To my surprise, even though this language is heavily focused on math (parts rotations, positioning...), which LLMs are usually bad at, agents did pretty well instead on initial tests, and subsequent projects also yielded good results , but never enough in order to consider generated models to be correct:

These three attempts, and quite some other experimentation, led me to the following conclusions :

💡 Conclusion 1: there is a minimum resistance path to geometry math for agents, i.e.:

  • Giving the agents tooling to generate LDraw sources would sidestep (evil!) geometry math
  • ... because they do way better at generating python code that produces math
  • ... than on producing math themselves!

💡 Conclusion 2:

  • Agents tend to do better when learning from python code that produces models
  • ... than from models themselves (LDraw's evil geometry, again...)

Then, the only thing left 🤔 was to create a python-based tooling with the required primitives, verbs, constructive vocabulary ... so agents would learn by example and do similar things on their own.

Which proved to be really hard to get right , even if vibe coding it... until GPT-6 Astra and Claude Opus 5.5 arrived... and vibe-coded it right! 🚀🚀🚀

How it works

ldraw-nova provides the tools, examples and instructions an agent needs to design models with real LDraw parts.

The process is as follows:

  1. The agent takes a prompt .
  2. Reads instructions.md and related documents to LDraw language and LEGO© models building.
  3. Plans how to build the model: required parts, submodels to be created, aesthetics...
  4. Iteratively :
    1. Renders images from the model/submodel(s)
    2. Inspects them, adjusts positioning, aesthetics... and back to rendering
  5. ... until it considers the model finished and ready to deliver!

Provided tooling helps the agent in:

  • Finding suitable parts.
  • Also, example models and submodels to start with.
  • Collision and gaps detection for placing parts correctly.
  • Headless rendering for inspecting current results.
  • and more...

The agent doesn't actually start with placing parts , except for things like e.g. prototyping and learning by altering pre-existing example models.

The way it produces models is more like:

  • Collects the required information, from experimental results, docs and planning.
  • Builds one or more plans , that fully describe the model and submodels, including its geometry, like e.g. atlas-crane.plan.json
  • And with that plan, it creates one or more generator scripts like e.g. generate.py
  • ... that when executed, produce LDraw source file(s), a very specialized 3D CAD language .
  • ... like e.g. atlas-crane.mpd

To summarize , this is like:

  • an agent creating a generator
  • ... that produces a 3D model
  • ... in an assembly language named LDraw 🤯

A compiler of sorts, so to say 🤓

flowchart TD
    agent([agent]) -- produces --> plan[plan.json]
    plan -- interpretation --> gen[generator.py]
    gen -- execution --> model[model.mpd]

Loading

Agent's informational sources

These are some of the guides and references given to the agent in order to make it a builder:

I want to… Read…
Ask an agent to generate a model Agent instructions
Improve shape, colour and detail Visual design guide
Build vehicles Vehicle workflow and examples
Build advanced spaceships Spaceship workflow and atlas
Learn a submodel and grow an atlas Build-manual workflow
Build Technic structures Structural workflow and examples
Build with mechanisms Mechanism workflow and build manuals
Find parts and reusable constructions Reference discovery and reference atlas
Organize a large model Module workflow and Copper Lane example
Understand connections and checks Geometry , snapping and validation
Look up a command or file-format rule Tool reference and LDraw rules

Development

Being this a first release , there are quite some things that still require some work:

  • VR on Meta Quest 3: model handling has many issues , performance issues.
  • Adapt for low-end agents: adapt current tooling, docs and instructions in order to improve usage by low-end models like e.g. Luna, Haiku, etc.
  • Expensive generation: currently, only expensive , high-end models, are currently capable of generating large-sized and correct models.
  • Improve efficiency: generative process is currently slow.
  • Add and improve more model families:
    • Humans and animals: minifigs
    • Technic models: machines, engines...
    • Spaceships: generated models are not very good
  • Building models from manuals: it partially works, better if page manuals are given as images.
  • Fine-grained inspection: for inspecting submodels and their step-by-step building processes.

Contributing

COMING SOON

Acknowledgements

I'd like to thank the following:

... and thanks to all of the many other LDraw creators!

NOTE: For this work, I've used many LDraw models, libraries, tools, docs... from many sources.

There is a lot amazing people that generously contributed to this, even for decades , by generously donating their finest work to the public domain and open source community.

If you think you should be included on this section, please drop me an email!

Trademarks

LEGO(R) is a trademark of the LEGO Group of companies which does not sponsor, authorize or endorse this software.

GrapheneOS has fixed the Android 17 QPR1 kernel performance regression

Hacker News
discuss.grapheneos.org
2026-10-02 15:45:06
Comments...

The Real Housewives of New York City Have 'the Worst Seder in the History of Judaism'

hellgate
hellgatenyc.com
2026-10-02 15:44:17
"You're a lying-ass snake of a bitch. You're a garden snake and I'm a fucking mongoose."...
Original Article
The Real Housewives of New York City Have 'the Worst Seder in the History of Judaism'
Sometimes we scare ourselves (Photos by Eugene Gologursky/Bravo, Canva. Collage by Hell Gate)

RHONY

"You're a lying-ass snake of a bitch. You're a garden snake and I'm a fucking mongoose."

Scott's Picks:

The disastrous dinner party is a Housewives tradition. OG RHONY cast member Dorinda Medley hosted many a demented meal at her infamous Berkshires manse in the 2010s; the seminal "Real Housewives of Beverly Hills" episode " The Dinner Party from Hell " is perhaps the single best episode, beat for beat, that Bravo has ever produced.

The dinner party at the center of the fourth episode of RHONY season 16 is much less iconic than its forebears, but at least it's something .

Give us your email to read the full story

Sign up now for our free newsletters.

Sign up

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to Hell Gate.

Your link has expired.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.

Updates to Full Disk Access in macOS

Hacker News
developer.apple.com
2026-10-02 15:37:01
Comments...
Original Article

Updates to Full Disk Access in macOS

October 2, 2026

We give developers powerful APIs to build incredible capabilities into their apps for Apple products, backed by a set of controls designed to protect users’ private data. Full Disk Access largely sidesteps these controls in order to allow backup apps to function properly on the Mac. Some developers are using Full Disk Access in ways that could put users at risk, exposing everything on their systems—including files, mail, messages, and even browsing history—without users’ full knowledge and understanding. For communication apps, this can also compromise the privacy of the people users are communicating with.

Going forward, we will introduce additional controls to ensure that users who genuinely wish to grant an app this extraordinary level of access can only do so with very explicit user action. Addressing this is critical. As AI agents become increasingly capable and autonomous, the risks associated with this level of access will grow substantially. We are committed to ensuring users clearly understand these risks before granting such access, so they can make informed decisions about their own data and privacy.

Make Tmux the OS

Hacker News
matduggan.com
2026-10-02 15:29:07
Comments...
Original Article

I recently watched the talk by Scott Jenson titled "Are we really going to use the same Desktop UX forever?" https://www.youtube.com/watch?v=V7AfAcQwLW0&t=445s . He's a great presenter, really articulate and concise. The kind of speaker that you'd gladly listen to for 3+ hours if given the chance. Which for a talk about window management is quite the compliment.

The overall point of the talk was "Apple and Microsoft aren't going to innovate anymore in the desktop OS space, so it is up to us all to decide what the desktop of the future is going to look like". I would argue that the actual situation to more nuanced than that, they're trying ideas they are just extremely conservative. As Jenson points out, a lot of what we treat as "how computers work" was a workaround for hardware we stopped having years ago. So why are we still accepting those tradeoffs?

To be clear, I'm not a UI/UX designer. I don't know what I'm doing and I lack the skill to make this. I'm making this mostly as a thought exercise and hopefully a prompt to get other people to think about this same problem. I have, however, suffered through others bad choices, which turns out to be most of the qualifications required.

TL;DR: The thing I want is "tmux is the OS", but I want tmux where a normal person could use it. I use it all the time, it's great, but is there a way to take the amazing experience of a scrollable, session-persisted, detachable, task-oriented system to normal people?

What are the high-level problems with windows management in 2026?

I think you can break the problem down to the following components:

My screen real estate is constantly changing and I have a lot more of it.

I am going from my laptop screen to my desktop monitor back to the laptop screen about 6 times a day. When I'm on my big monitor or multiple monitors, overlapping windows don't make sense because I have too much space as it is.

    1. First, one big monitor is a a different experience entirely compared to 2 smaller monitors. I don't use the two monitors as one large unbroken screen, I tend to use one of them for "less important static content" like a ToDo list or my work chat tool and my email and then the "primary" monitor for my terminal/tmux/browser which is how I do all my actual work. But windows "belong" to one or they "belong" to another. Grudin found in 2001 that people use a second monitor exactly this way, where focal work on one, peripheral glances on the other. Twenty-five years ago. The OS still doesn't know which monitor is doing one role and which the other.
    2. And when I pull the laptop out of the dock, everything collapses onto one screen and I manually drag and resize windows to claw back enough real estate to keep working. I don't know how much time this actually wastes for people, but it feels like it wastes a bunch. The OS treats a display change as a surprise, instead of as a scheduled event that happens six times a day.

Different people use windows totally differently.

We have three distinct groups of people that we are trying to design around, while most modern OS optimizes only for the first use case.

    1. Some people are window maximizers, where every window takes up most of the screen and then they use Command + Tab or the Dock or something else to switch between the bigger windows. The value of more screen real estate is mostly that this one big window can be even bigger.
    2. Some people are near maximizers, where one window takes up almost all of the screen real estate, but they'll have one or more smaller windows where they glance at things like status or chat or whatever.
    3. Finally there are people who carefully coordinate all of the windows on their screen.
    4. Almost everyone has "private windows" and "public windows". You are fine with your public windows being lined up and persisting, but you want to hide specific information in other windows from people walking by. The people I complain about are also the people I present quarterly numbers to. I never remember that when I connect my laptop to a display they're going to see my entire screen.
    5. I assumed this was just me with private vs public. But it was measured in 2004, in "Revisiting Display Space Management: Understanding Current Practice to Inform Next-generation Design." Same three types, same public/private split, twenty years ago. See how absolutely none of this is a new idea? Link .

Browsers are mini operating systems.


In 2026 everybody quietly agrees that most of your software runs in a browser tab. These web applications will cover a wide variety of use-cases that have historically been owned by local applications. My desktop OS treats a browser just like a normal application, when in fact it is closer to a virtualized OS running inside of my host OS.

    1. Tabs inside of a browser and windows of that browser contain the same level of complexity as my other applications.
    2. Tabs are associated with streams of work alongside my conventional applications. I'm writing Terraform in Vim in my terminal while referencing the Terraform docs for that provider. But the relationship between tabs and work is messy: the same docs tab pulls duty while I write the code and again while I write the ticket update. Whether a tab should be allowed to belong to two tasks at once, or whether something cheaper is going on, is the exact spot where I break with the research. I'll come back to it once you've met WindowScape.

The concept of "filesystem" is an increasingly weak concept.

My notes in Apple Notes belong to Apple Notes, not my filesystem. My texts live in iMessage's database, not my Documents folder. Teams and Slack content lives inside those applications unless I manually "bring it out," and when I do, I'm making a copy. Firefox will happily show me a PDF, but the PDF doesn't live with macOS unless I take an action to make it so.

    1. Now a lot of engineers are going to read that and think "well you cannot force all applications to use the same storage system for all of their files, are you a lunatic?!?". I'm not suggesting that, in fact the weak filesystem might be a perk. There is a CHI paper from 2004 called, and I am not making this up, "Stuff goes into the computer and doesn't come out." That was 2004 and we have the exact same problem with no real solution. Link .
    2. Here's why this is a windowing problem. Since Windows 95 and System 7 (which is basically as old as my memory goes back to), the machine has worked as a chain: something writes a file, the user opens an app on it, the app writes it back, the file gets sent to someone else, repeat. Every link in that chain assumed the file lived somewhere the OS could see. Every link is now broken in all the modern OS.

What have people already tried?

So I'm not the first person to see this problem. Some of it got solved, but in a different direction than I want. Some of it got solved, but not for normal people.

I started with the document that I kept seeing everyone else cite to. Henderson, D. Austin and Stuart K. Card. “Rooms: the use of multiple virtual workspaces to reduce space contention in a window-based graphical user interface.” ACM Transactions on Graphics (TOG) 5 (1986): 211 - 243. Interesting that even back in the 80s there was a pretty clear understanding that the current system for managing windows wasn't very good. The design they were talking about looked something like this, which is pretty advanced compared to where we are.

Henderson and Card measured window use the way operating systems people measured memory. The screen is RAM. A closed window is a page swapped out to disk. And windows, like memory pages, don't get touched at random: you sit inside a small set of them, two to ten, and that set is the task. Programs spend about 98% of their time inside one of these sets, and roughly half the cost of running happens during the 2% of time spent switching between them. Rooms' whole design was preloading the next set before you ask. They described this as reducing "knowledge faulting in the user," which is the best phrase in the literature and I intend to use it until someone stops me. Just try it out in a corporate meeting: "we need to reduce knowledge faulting in the user".

One of the more interesting side papers I found was "No Task Left Behind? Examining the Nature of Fragmented Work". Link . Mostly because it confirmed something I've long suspected, which is the single task for a long time focus on modern widowing systems isn't actually how people work. The title of the companion paper is a real participant quote: "Constant, constant, multi-tasking craziness." People were juggling around ten "working spheres" a day, minutes at a time.

The closest to what I wanted is from the paper WindowScape: A Task Oriented Window Manager. Link .

WindowScape dropped explicit grouping entirely so now every time you changed the arrangement, it took a photograph, and you went back into photographs instead of filling containers. I love their one-line diagnosis of every system before them: "requiring windows to be in a single group forces users to decide ahead of time where a new window belongs." The photograph metaphor also cracks a problem I'll get to with browser tabs: one window can appear in many photos, because, as they put it, "people understand that there can be several photos of an object with there being only one underlying object." The catch is that the photos evaporated. That's the gap that I think you'd want to solve.

For the first problem, we have more or less already solved for overlapping windows. A tiling window manager ensures that you are maximizing your screen real estate in such a way that you can switch between different layouts with no need to manually modify the windows. There are basically 2 problems with tiling window managers as they exist now.

    1. They're way too hard to use. Like an order of magnitude too hard for normal people to use. Basically if you need to start a sentence with "just open up the configuration file" shut it down the thing is over. The average tiling window manager tutorial asks you to clone a repository before it asks you to open a window.
    2. We need a scrollable tiling window manager. So basically I should be able to set up specific window configurations over here, leave it alone, then scroll to the right and do something else different, like shoving the mess into the back of a drawer. It should be infinite space to work without needing to subdivide the windows into different macOS Spaces.
    3. The scrolling also solves the private vs public window problem. I put my private stuff on the far left and then my public stuff on the right. If I want to look at the private stuff, scroll to the left as far as it will go.
    4. Thankfully this already exists with Niri: https://github.com/niri-wm/niri . It just has to be easier to use. But the tough design elements more or less already work.

Browsers as mini-operating systems: people have tried to solve this, but in the wrong direction. They made the browser more of an operating system instead of making its contents first-class citizens of the one you already have.

    1. The best two examples of this are the Arc browser and Chrome OS. But I think Arc is actually the more interesting of the two experiments. Their first innovation was to break the idea of tabs at the top and instead move it to be a sidebar model.

They also basically "took over" the concept of windows from the OS.

The diagrams above and a good write-up of Arc is available here: https://blakecrosley.com/guides/design/arc .

All of this makes sense from the perspective of "the browser is now the operating system", but I think this is a fundamentally flawed idea. If a web application is operating as an application, it should be its own window. If it is complimentary to another application, it should be a window associated with another application and be allowed to contain many tabs, to reflect the idea of the browser as the portal for all research and lookup.

I was shocked to go through the historical progress of people trying to solve this problem. I'm not the first person to try to solve this problem, I might not even be the 10,000th person.

Graveyard of Attempted Solutions

  • Windows Timeline (2017–2019): All of your activity across all of your apps sorted and organized for you.
  • Windows Sets (2018–19, canceled): apps and webpages in shared tabbed sets. Note what this actually was: the operating system grouping apps and web content into tasks, which makes it the closest thing anyone has ever shipped to what I'm about to propose.
  • macOS Stage Manager (2022): automatic task-based window grouping, widely ignored. The reason I think Stage Manager didn't scratch this itch for people is that its super designed around the iPad style flow of "there is one thing you are using at a time and you need to be able to quickly switch between them". On iPad the model is "one thing at a time, switch fast." On desktop, two or three things have to work together . It compromised toward the iPad and hit for neither.
  • KDE Activities : Everything I'm writing here is old news for KDE. They've been doing task-scoped desktop state for over a decade. Honestly this does most of what I'm writing about, it's just hyper manual and requires you to manage and set it all up. Also I love the KDE design docs, amazing stuff, worth reading. https://community.kde.org/Get_Involved/design
  • PWA install: already gives web apps their own window and no browser chrome, which is clunky but at least makes them "real" applications.

So clearly we understand there's a problem. Why haven't these taken off like wildfire? I think there's a couple of different problems here.

Timeline needed apps to opt in to reporting activity, and most never did. It also logged everything, which read as creepy (which is a problem that my idea would have too), and the UI surfaced Edge features nobody wanted, so it felt like a browser push more than a feature. Sets died somewhere between internal strategy and app-compatibility chaos where the interesting question is why nobody demanded it loudly enough to save it. Stage Manager tried to be universal and pleased nobody. KDE Activities is opt-in, and opt-in means the people who need it most never configure it.

The pattern in the graveyard: the ideas aren't wrong, they're just not defaults, or they're not done at the OS level where they can see all apps. Everything I want has to be structural.

Filesystems as a weak abstraction.

    1. Operating systems have tried to solve this, but because there's no requirement that you use the user OS filesystem to store documents, there's no consistency. MacOS has recent files, but "recent" doesn't mean anything (these are not the most recent files I have downloaded to my computer). As far as I can tell this functionality is completely broken or implemented in a way that makes zero sense to me.

I have opened, downloaded and created dozens of documents since September 23rd (I'm writing this on September 28th). I don't have a single fucking clue how MacOS populates this window. Maybe its broken. Maybe its working as intended through some criteria I don't understand.

Maybe it's personal.

But the idea is here. The ideal would be "across my entire computer and applications, what are the recent files I have interacted with" and expand that out to include emails and Teams/Slacks and everything. I should be able to see everything going on with my machine, search through it and not care if its stored inside of Slack or iCloud or whatever. Single pane of glass.

Microsoft tried it, people didn't like it, but I think the concept makes sense.

What might a solution look like?

So what has changed? Why might we be able to crack this problem now when before it was too complicated? This is where I think a local LLM might make sense, if you can figure out a way to do it where it doesn't cause more problems than it solves.

So what are we looking at here. The concept would be organizing it around the idea of tasks. When you define a task, this allows you to organize all the windows together, including specific browser tabs detached and associated with the task, not with the concept of "browser". You still get get the flexibility of defining glanceables that exist outside of the strict tiling view.

The basic window flow would look like this:

The desktop is one infinite canvas; each physical display is a viewport onto it.

You would end up with the following "transition contract" to handle my initial problem of "what about many monitors/new monitors".

  • Never Change items: horizontal order of windows, scroll position, input focus, task membership.
  • Reflows Deterministically: column widths, like a responsive layout. When the shelf narrows, the books stay in order, the rightmost ones fall off the edge of the viewport into the scroll region
  • Is state, replayed on redock : viewport aims. Two monitors means two viewports aimed at two regions with one at your focal task, one at the glanceables region. Undock: second viewport disappears; its region is one scroll gesture away.Redock: it re-aims where it was.

The OS finally learns which monitor is focal. The display receiving keystrokes ~90% of the time is focal; windows that stay visible but rarely receive focus are glanceables. That's inferable from focus telemetry and requires no eye tracking or config. And when the laptop screen is small, glanceables demote to a thin strip rather than full tiles which is, note, exactly what tmux already does.

One substrate that covers our three user types. Remember the maximizers, near-maximizers, and coordinators? On the canvas they're just column counts. A maximizer is one full-width column. A near-maximizer is one column plus the glance strip. A coordinator is N columns. Nobody gets forced into anything. This is why the design can be a default where KDE Activities was a preference. It's because the substrate doesn't impose a style, it just stops punishing whichever style you already have.

Privacy becomes a first-class flag, driven by machine state. Private is a property of a window or tab, not a region of the screen. Private content renders occluded unless the machine can vouch for safety. The rule I landed on, at the cost of my favorite bad idea (story below): derive privacy state from machine-observable facts like connected displays, or active capture sessions and never from inferred human states like attention or idleness.

Connect to a novel display and it gets a "ready to present" screen by default. You explicitly aim it: this task, this window, or extend the canvas. Known displays skip the dance. Screen sharing is the same event with no cable. You don't need the model to have discipline if the door won't open, and you don't need the user to have memory if the default can't leak.

None of this is exotic. Apple's Keynote's presenter view has been shipping for twenty years plus with slides on the projector, notes on your screen. Apple never promoted the pattern to the OS even though it makes obvious sense. I'm asking for presenter view as a desktop primitive.

Browser tabs detach into real windows and join whatever task they belong to. The "browser" stops being a place windows live.

This all makes sense until you get to "how do you organize this stuff around tasks". I think for more expert users they're going to be able to do this stuff, but part of the problem is that we don't want them to ever have to drop back down into a configuration file. A mouse and dragging stuff around is too clunky for how this would work. Ideally I should be able to ask something that has access to what information is contained inside of each one of these browser tabs and window and help me organize it in some logical flow.

That's where the local LLM comes in.

I'm calling it a "quake overlay" because I'm old and the idea of a text input that drops in from the top with a universal shortcut has always been the quake terminal dropdown to me. Younger people know the same pattern from the Discord overlay, which I have decided not to be mad about. However in the previous diagram its shown as a more friendly "Ask" box.

But the basic flow would be that you ask the LLM to assist you with organizing, it would show you what it's thinking by querying the information through a constrained MCP server with a set verb list and then gives you a preview of what the layout is going to look like. The model never touches the window manager directly. It proposes and you press y.

People rarely pick up completely novel tasks: a graphic designer spends their life in "open files from network storage, make changes, save output, paste into chat, repeat." People rarely sit down at novel display contexts: laptop-only, the desk, the conference room. People rarely attach novel displays. Novelty is rare everywhere, so the expensive one-time machinery of classify, propose, preview, confirm only runs rarely, and everything in between is deterministic replay. You organize a limited series of tasks once and reuse them forever. Check the tickets, open Vim and a browser, work the ticket, write the update, open the PR, post it for review, next ticket. You could make the config a nightmare of JSON and stop caring, because the only intended reader is a model that doesn't mind. The config file stops being an interface.

The biggest issue here would be spatial memory. People understand where things are in relationship to each other on their computer and they don't like it if they are changed. Think of it like "if I came and messed up your physical desk by moving stuff around and adjusting your chair". So I suspect you would need to enforce a strict "append, don't replace" model.

There's a name for this actually. Kirsh and Maglio call it epistemic action, where Tetris players rotate pieces more than the game requires, because rotating is how they think about the piece. Link Your window layout is thought in progress. A helper that "optimizes" it mid-thought is interrupting you.

This is a problem I encountered with just trying this idea out though. Modern LLMs are actually not very good at "touch this stuff, never touch that stuff". So unclear if this is a realistic idea or not. One boring fix: the LLM proposes, dumb deterministic code executes, and the executor refuses to touch anything pinned. You don't need the model to have discipline if the door won't open.

You'd still need a high level UI element for users and this is where I think the more inclusive concept of filesystem could come in.

Tasks become roughly comparable to Directories now. But everything flows from the initial high level concept of "Tasks" and then flows out to individual things. One application is never the task. It's "some chat window, a browser, maybe Preview." Can we try to build it?

So my hope was that here I would be able to post a Linux desktop running an example. But as it turns out (unsurprisingly) it is insanely complicated to make something like this run at all. Some of the underlying ideas work reasonably well (global terminal shortcut dropdown from the top, scroll-able tiles), but it's still pretty clunky. I'm gonna keep working on it and see if I can get something as a demo running, so if you are interested either add me to an RSS reader or just....wait I guess.

In the interest of full disclosure I think this is gonna take many weekends of work to even get a functional demo running. So I'll do my best, but be patient.

Reasons this design sucks

So in attempting to get this up and running in a Linux VM, I immediately ran into serious logical issues with my design. I think failures are often more interesting to read about than successes, so let's talk about them.

  1. The privacy-minded design runs on surveillance-shaped data. There's an irony I sat with for a while. The design promises the OS will finally know which monitor is focal which is the thing Grudin measured twenty-five years ago and the mechanism is input-focus telemetry. Watch which display gets the keystrokes. Watch which windows never receive focus. Store that over time. The privacy-first window manager wants to know more about what you're doing than your current OS does. Keep it local, keep it ephemeral but it's still a lot of behavioral data to drive this system. Just the observation that every mechanism in this post is hungry for data that has historically derailed ideas like this.
  2. My favorite privacy feature was a joke. The original plan was elegant: private stuff far left, public to the right, and when you go idle the viewport drifts back to public. Like a friend changing the subject for you. Then I ran it, and the first thing I learned is that reading is an idle state. You stop typing to read the thing, and the system starts scrolling the thing away from you. Meanwhile, a person walking by sees everything, because a "private" window on a wide monitor isn't hidden it's just further left. The feature hides content from its owner and shows it to the threat. Privacy by occlusion needs to know when a threat exists, and the threat is a pedestrian, which is undetectable.
  3. Nothing in this design ever ages out. The canvas is infinite, the rule is append-only, spatial memory is sacred so the mess grows monotonically. I have solved the electronic messy desk by buying it an infinite desk. When I first read WindowScape, I called their evaporating photographs "the gap to fix." I've changed my mind because evaporation was quietly doing the job of a garbage collector. Something needs to go dormant without moving but deciding what goes dormant is a reaping decision, and reaping is exactly what append-only forbids. It didn't take long for the infinite desktop to feel overwhelming when I tried it.

Fun experiment

I don't know if this underlying idea is a good idea, but it is liberating to stop waiting for MacOS and Windows to do something better or interesting in this space and decide to try it yourself. The graveyard up there is full of good ideas that died of distribution. We're overdue for a radical experiment that ships as a default.

The most eye-opening part of this entire experience was how much the original overlapping window design was a hack to get around fundamental hardware limitations and how soon after it was launched did people clearly see the problems. I'm surprised how often that happens in technology, where you see a problem, search for academic papers about the problem and find just a massive wealth of information from people saying "oh yeah this is 100% a problem and one we should solve soon". Maybe this post inspires someone smarter than me to solve it finally.

Muse Gadgets

Hacker News
gadgets.muse.ai
2026-10-02 15:26:54
Comments...
Original Article

Muse gadgets are open source devices you build yourself. Program an off-the-shelf ESP32 board or set up a Raspberry Pi with our SDKs, then connect Muse to your displays, buttons, sensors, actuators, and whatever else you’ve got lying on your workbench.

Muse gadgets: a Waveshare round AMOLED, an M5Stack StickS3, Muse Home Link, a Raspberry Pi and a Seeed reTerminal e-ink display

01 / ON YOUR HARDWARE

Make your own

We open sourced the code you need to make a Muse gadget. It’s built by hackers, for hackers, just for fun. Side effects of tinkering may include bricked boards, voided warranties, brownouts, or bankruptcies. Proceed at your own risk!

Get started by grabbing an SDK token and one of our featured project ideas . Then tinker and customize to your heart’s delight. Or build support for an entirely new board if that’s more your thing.

ESP32 Device SDK

Connect your ESP32 board to Muse through our open source SDK. Throw in a screen to show images, add audio in and out, or wire up other sensors.

Linux Device SDK

Turn that spare Raspberry Pi or Linux box into a Muse gadget. Hack in your own commands to let Muse handle sysadmin chores or your Home Assistant setup.

03 / BUILD TOGETHER

Meet other hackers who are building and customizing Muse gadgets in our community Discord. Get inspired, support each other, and share what you make.

38 members online now

Join Discord

04 / OPEN IT UP

Source

Explore the source code for our device SDKs and firmware. Use it as a starting point for your projects and contribute your improvements back to the community.

The device SDKs and firmware are open source under the Apache 2.0 license and provided as-is, without warranty.

Muse Gadgets on GitHub

05 / MUSE GADGET EXAMPLES

Devices shown here are made and sold by other companies, and Meta doesn’t endorse or warrant them. Names and logos belong to their owners.

"The only intuitive interface is the nipple" (2012)

Hacker News
www.greenend.org.uk
2026-10-02 15:18:12
Comments...
Original Article

The only "intuitive" interface is the nipple. After that it's all learned.

This is usually attributed to one Bruce Ediger (though at least one person thought it was originally said by Steve Jobs , a manufacturer of fine computers), and refers to the frequent (and invariably inaccurate) description of computer user interfaces as "intuitive". But in 2001, Bruce denied that it was original to him . This made me wonder what the history of the quote was.

My first stop was a dictionary of quotations , but that doesn't have it at all. Perhaps it is indeed relatively recent, then.

I can get back as far as August 1994, where Scott Francis suggests the nipple as the only intuitive interface, in response to "There really is no user interface metaphor that is truly intuitive." The idea is there, but the exact form is not. (And of course, nipples aren't metaphors, at least in this context.) To get close to the usually-quoted form of words, the earliest I can find is this one from January 1995 , by one Jay Vollmer. He said:

Actually, the only truly intuitive interface is the nipple.

Did Bruce see this post and like it, or did he see the idea somewhere else? It's not really possible to tell from Google's archive . At any rate, he started using variants of it shortly afterwards. For instance, in February 1995 :

It's an old saw, I know, but the only really, truly "intuitive" interface is the human nipple.

Slightly later the same month , we see a variant on the "it's all learned" theme from one Taylor Hutt:

I argue that no computer interface is intuitive -- none; they are all learned.

In April 1995 Bruce is back:

By this definition, the nipple is the only "intuitive" user interface.

and then later that month :

Basically, the only "intuitive" interface is the nipple. After that, it's all learned.

So perhaps Bruce does have the best claim to this form after all? Maybe. At any rate, I think I prefer his 2001 version:

There is no intuitive interface, not even the nipple. It's all learned.


RJK | Contents

Show HN: Pi pod – Run your pi coding agent in sandboxes on your own server

Hacker News
pipod.dev
2026-10-02 15:10:38
Comments...
Original Article

simple software

For me, pi is the epitome of simple, elegant software. Its maintainers embue true craftmanship into the harness. It allows for endless customization and tinkering. My pi harness truly feels like my own. It frees me from lock-in from any of the multitude of AI companies vying for my wallet and data.

At the same time, its minimal approach leaves certain aspects of a standard "agentic engineering" environment lacking. This also is its beauty, us, the enthusiastic community of pi users, can build and use our tooling to satisfy our workflow. I consider agent sandboxing, native clients, rbac session sharing, and surface automation (namely browser and native apps) to comprise the minimal agentic engineering environment that should be used for software engineers in an organization.

pi pod is an ongoing effort to consolidate all of this capability in a simple and elegant software built around and on top of pi. Using pi pod should feel no different from using the pi harness that you already have spent countless hours perfecting. pi pod aims to seamlessly run your pi in an isolated sandbox with minimal, carefully curated tooling to get out the way and give you the boring stuff you need.

The self-hosted edition is designed to be simple to run yourself. I hope that anyone with an interest in pi will take a look at the available repository and spin it up on their own. Self-hosting of course allows for true ownership and privacy, making this project yours to own and customize to your liking. Self-hosting also is a great option for businesses who want their entire agentic environment in their preferred data center.

The hosted service is coming soon. It is not available yet: there are no hosted accounts, trials, or subscriptions today. The intended v1 offer is a 7-day, card-required Standard trial with 10 active hours ($20/month after unless canceled), or paid Pro at $50/month. Intended hosted workstation pricing and trial terms are on the billing page .

I hope that you will take a look at the project, make it your own, and even endeavor to improve upon it. This has been a labor of love for pi and working with agents. Feel free to reach out with any comments or feedback on the GitHub repository or by email at evan at pipod dot dev. I hope for engaging and fruitful discussion as we work together to put form to the continually changing art of software development.

- Evan

Apple Pass Designer

Hacker News
developer.apple.com
2026-10-02 15:06:56
Comments...
Original Article

Introducing Pass Designer

Whether you’re creating passes for a local fitness center, a small music venue, a global airline, or a national coffee chain, Pass Designer makes it easy to craft and preview passes that reflect your business.

Building passes made easy

Start with templates provided by Apple or your own to capture the personality of your brand. You can use your preferred design tool to create images for passes — such as logos, backgrounds, or strip images — then bring them into Pass Designer to complete a pass design.

As you edit, Pass Designer updates the preview in real time to show how a pass will appear on iPhone and Apple Watch. The preview uses the same rendering as iOS and watchOS, so what you see in Pass Designer is exactly what customers will see on their device.

Colors and appearance

Easily adjust a pass’s background, foreground, and label colors.

Adaptable layouts

Pass Designer supports designs that take full advantage of the latest advances in Apple Wallet, as well as backward-compatible passes.

Standard fields

Each pass type includes standard fields for displaying information, and you can edit each field’s content directly in Pass Designer.

Validation

Pass Designer validates your pass as you work and alerts you to any issues such as missing required key values and unexpected definitions.

Semantic tags

For boarding passes and event tickets, semantic tags add structured data — such as event dates, venue locations, or flight details — that the system uses to enable features like Siri Suggestions, Calendar integration, and Maps directions.

Pass Designer lets you edit semantic tags directly in the UI, and you can view a semantic and non-semantic pass side by side. It can also automatically generate a backward-compatible pass structure from your semantic data that doesn’t require semantic tags, ensuring your pass works where they may not be supported.

Download the beta

Pass Designer beta requires macOS 27 or later. To download, sign in with your Apple Account. If prompted, complete the free registration and accept the Apple Developer Agreement.

View download

View documentation

Frontline Education breach exposes school district employee data

Bleeping Computer
www.bleepingcomputer.com
2026-10-02 15:01:40
Frontline Education is notifying school districts of a data breach after attackers exploited a vulnerability in third-party software to gain unauthorized access to its systems and steal employee information, including Social Security numbers. [...]...
Original Article

School hacker

Frontline Education is notifying school districts of a data breach after attackers exploited a vulnerability in third-party software to gain unauthorized access to its systems and steal employee information, including Social Security numbers.

Frontline Education is an edtech company that provides administration and workforce management software and services used by school districts.

Last night, a reader shared a data breach notification with BleepingComputer that Frontline sent to an impacted school district, stating that attackers breached its environment through a vulnerability in a third-party application.

"On August 14, 2026, our security team identified a vulnerability in a third-party software product we use that allowed unauthorized access to a portion of the environment," the notification letter reads.

"We promptly investigated the issue with the assistance of an independent cybersecurity firm, remediated the vulnerability, engaged with law enforcement, and took steps to further reinforce the security of our systems."

The company has not disclosed which third-party application was involved or when the unauthorized access first occurred.

For the notification seen by BleepingComputer, our source said all employees at the district were impacted, with exposed information including Social Security numbers, email addresses, and physical addresses.

BleepingComputer contacted Frontline Education yesterday about the breach but did not receive a reply to our email.

However, school IT administrators reported on the K12SysAdmin subreddit that district officials had begun to receive similar notifications.

One administrator initially said their superintendent and business manager received the notification on October 1 from frontline@notifications.cyberscout.com, but Frontline support had not yet confirmed whether the message was legitimate.

However, other administrators later said they had independently confirmed the breach notifications were legitimate.

"Can confirm this is legitimate. We've had verbal contact with our Frontline rep on it," one administrator reported.

One administrator also shared a copy of a Frontline notification stating that 1,210 employees associated with their district were impacted and that Social Security numbers, email addresses, and addresses were also exposed.

Frontline says it will handle notifications to affected individuals on behalf of impacted school districts unless a district opts out by October 16.

Districts that want to opt out can do so through www.frontline-transunion.com or by calling 833-516-8792. If a district opts out, Frontline says it will not provide notification services or reimburse the district for the costs of issuing its own notices.

Impacted adults are being offered two years of free credit monitoring and identity theft protection through TransUnion, while minors will be offered cyber monitoring services.

The company says it will also handle required notifications to state attorneys general and cover costs associated with individual notifications and the identity protection services.

It is unclear how many school districts or individuals were affected.

article image

Build your security blueprint for AI-powered attacks

Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.

Save your seat

Site-Blocking Will Not Defend IP, No Matter the Bill’s Name

Electronic Frontier Foundation
www.eff.org
2026-10-02 14:59:23
There has been a raft of site-blocking bills in the latest Congress, and the latest is called the “Deterring Extraterritorial Foreign Exploitation of Networks Damaging Intellectual Property” aka the “DEFEND IP Act.” The problem is that instead of “defending IP,” this bill will incentivize censorship...
Original Article

There has been a raft of site-blocking bills in the latest Congress, and the latest is called the “Deterring Extraterritorial Foreign Exploitation of Networks Damaging Intellectual Property” aka the “ DEFEND IP Act .” The problem is that instead of “defending IP,” this bill will incentivize censorship, overblocking, and bad faith attempts to block access to a website. DEFEND IP Act, and all of these site-blocking proposals, threaten the open web.

We keep seeing attempts to pass site-blocking legislation–from SOPA/PIPA in 2012 to Block BEARD , FADPA , and ACPA this year. Every one of them has at its core the rotten idea that enforcing copyrights requires building a censorship machine for websites into the architecture of the internet. This is, of course, a disaster for a free and open web. There is no way to create a mechanism for blocking access to an entire website that does not invite both deliberate abuse and lots of collateral harm to free and lawful speech.

DEFEND IP deputizes every service provider into a copyright cop, so long as a rightsholder has accused a website of copyright infringement. Let’s be clear: this isn’t about removing access to an infringing work–that already exists via the DMCA. This isn’t about getting damages from the website or the uploader. It is about making an entire website inaccessible for everyone trying to visit it.

DEFEND IP lets any rightsholder go to a court and get an order requiring service providers to block access to an entire website after alleging copyright infringement. What DEFEND IP does not have is any deterrent for someone seeking to block a website in bad faith. There are no punishments for getting a website blocked for protected speech. There are no meaningful remedies for those whose speech is vanished from the internet due to an entire website being disappeared. It creates a one-stop shop for getting an entire website–again, not an instance of infringement but an entire site hosting all sorts of user content–removed. But for those whose business, speech, or access to information is affected, there is no easy way to get the site restored.

DEFEND IP scales up the extraordinary legal structures that already exist for copyright enforcement. In doing so, it likewise scales up the problems those regimes pose to protected speech.

We see this with DMCA takedowns all the time. We see it with bad faith takedowns used to silence criticism or commentary. We see it with the voluntary use of copyright filters by sites like YouTube, where seconds of sound matching seconds of sound in another video can prevent an entire work from reaching its audience. In these existing systems, there are at least some mechanisms of challenge available to the targeted creator. DEFEND IP has none. Instead, site owners, users, or readers will have to find a lawyer and go to court and hope to challenge the order, a slow, expensive, and daunting process

Those existing systems are already frustrating for the targeted creators and users, but under DEFEND IP a whole class of people doing protected speech will find themselves deplatformed because of the actions of others

This bill is not a defense of creativity or creators. It is a way to reshape the internet by building a vast new infrastructure of censorship. Congress should put aside DEFEND IP and the failed idea of site-blocking laws, for good.

The Nintendo 64 Partner-N64 Development Kit

Hacker News
www.behindthecode.ca
2026-10-02 14:53:34
Comments...
Original Article

Why have I been blocked?

It seems that something in your recent request triggered our website’s security system. These protections are in place to keep the site safe from online threats and misuse. Various actions can cause this block, such as submitting specific words or phrases, executing a SQL command, or sending improperly formatted data.

What can I do to resolve this?

If you think you were blocked by mistake, please reach out to the site owner. Be sure to include the Request ID shown below so they can quickly look into your request and restore access if appropriate.

Microsoft doubles down on Rust

Lobsters
www.infoworld.com
2026-10-02 14:48:59
Comments...
Original Article

Microsoft has promoted the Rust programming language to Tier-1 status internally and begun integrating it with the Microsoft Visual C++ tool chain for Windows development.

After many years of watching how Microsoft develops platforms and rolls out technologies to external users, one thing is clear: anything that is important to Microsoft internally quickly becomes important to anyone building on top of Windows or Azure.

The transition from internal to external tooling invariably takes the form of new features in Visual Studio and Visual Studio Code . Often these capabilities have been in use inside Microsoft for years, and they’re ready for you to use immediately in your own software. We’ve seen it happen many times before, such as with the development of C# and TypeScript.

Microsoft’s Tier-1 programming languages

C# and TypeScript , like C++ before them, are what Microsoft internally calls its Tier-1 languages. This status rests on being backed by a complete tool chain, from editors to compilers; integration with Windows and Azure through native SDKs and optimized libraries; and compliance with Microsoft’s internal software development life-cycle requirements. When you’re shipping software for a billion users, these trappings aren’t optional extras, they’re essential.

Now there’s a new member of the Tier-1 club: the Rust language . That shouldn’t be a surprise, as Azure CTO Mark Russinovich has talked about his organization’s commitment to Rust and how Rust’s memory safety features are a key component in Microsoft’s security strategy. In addition, Microsoft was a founding member of the Rust Foundation, and has made significant investments in Rust’s Windows tooling and in Rust compilers that work as part of Microsoft’s build platform. Now Rust enjoys first-class development tooling and workflows inside Microsoft.

A recent post on the Rust Foundation website details how Microsoft is working to integrate Rust with the existing Microsoft Visual C++ (MSVC) platform. This approach will allow Microsoft to ensure what the blog post calls “seamless interoperability” between the languages, with Rust inheriting Visual C++ features as they’re delivered. Using Visual C++ with Rust will allow Rust to be used to build low-level Windows services, from drivers to the kernel itself.

We’ve already seen that in the Windows release of Coreutils , which are a Rust re-implementation of core UNIX commands. Coreutils allow you to use certain UNIX commands inside the Windows terminal, making it easy to switch between Windows and WSL (Windows Subsystem for Linux)-hosted Linux distributions. Now that Rust has crossed the Tier-1 threshold, we can expect to see more Rust-based Windows developer tooling, where we get access to low-level functionality while avoiding memory leaks.

The secret sauce: a new Rust code generator

One key to delivering production-grade Windows software with Rust is a new code generation tool for Rust’s compiler , rustc. rustc is designed to use different code-generation back ends, beyond the default LLVM . There are already variants that work with GCC and with Cranelift . If you’ve not come across Cranelift, it’s the code generator of the Bytecode Alliance , used as part of the wasmtime framework. Adding a new code generator is a matter of building a tool that connects to the Rust compiler’s APIs, takes its bytecode output, generates native code, and passes it on to your choice of build pipeline.

This is the role of Microsoft’s rustc_codegen_utc tooling. It’s designed to work with the existing MSVC stack, giving Rust access to a mature, well-tested, proven part of the Windows development platform without needing any changes to your code (compatibility is baked in). This links it directly to the back end of the Visual C++ compiler, a set of tools marshaled by Microsoft’s build system to work with complex projects that mix code and libraries.

Microsoft calls this back end the UTC, for Universal Tuple Compiler . You won’t see that name anywhere, though; the actual code is a DLL, C2.DLL. Code is generated and delivered as .OBJ files where it can then be passed to the MSVC linker and delivered as a binary executable. This allows Rust tooling to take advantage of the decades of work Microsoft has put into both its compilers and build tooling. There’s no need for Microsoft to duplicate that work, ensuring that existing security and resilience features carry forward to the Rust compiler chain. And the resulting code is compatible with C and C++ code that has been passed through the same back end.

Managing only a single compiler back end will also make it cheaper to support Rust, with no need to run two separate compiler teams that are likely to be out-of-sync with each other and with Windows SDK and kernel development. Reducing the economic impact of Rust will speed up any transition, as well as supporting what is likely to be a long-term hybrid delivery of combined Rust and C++ applications.

This approach also puts Microsoft’s Rust development on a par with platforms that use GCC and LLVM compilers, which can already use the GCC and Clang back ends with Rust, leveraging the investments into Linux and macOS application development. Code developed in Rust will get the same interoperability capabilities no matter which platform you end up targeting.

Unfortunately for the rest of us, rustc_codegen_utc is currently only available internally, unlike the other Rust compiler codegen tools, which are all open-source projects. Microsoft describes it as “production ready”, and it’s used by more than a hundred Microsoft Rust projects—an interesting pointer to the scale of Rust adoption by Redmond. It’s not the only Rust tool currently hidden away, as Microsoft’s complete Rust tool chain also includes its own builds of the core rustc compiler, standard libraries.

That doesn’t mean that rustc_codegen_utc is not going to arrive for the rest of us. Past experience suggests that it and undoubtedly other Rust tools will come with a major update to Visual Studio or Visual Studio Code. However, the timing is always hard to predict. Even when tooling that works inside Microsoft is highly suitable for external developers, there needs to be more for it to reach the rest of the world, with language servers and IntelliSense support, as well as full integration with the Visual Studio debugger platform. There’s also more to building code outside Microsoft’s offices, where we don’t have access to their custom build tools.

In discussions on the Rust compiler forums , Microsoft has indicated that it will first upstream its test and back-end infrastructure work into the wider Rust project. That implies that it may be some time before there is a public release, but also that open sourcing the project is clearly on Microsoft’s radar. Getting rustc_codegen_utc out in public is important—adoption of Rust to build third-party device drivers should help make PCs, servers, and cloud VMs more stable and reduce security risks.

Get started with Rust in VS Code

For now, Microsoft offers Rust tooling that integrates with Visual Studio Code and guidance for developing on Windows with Rust . To use Microsoft’s Rust language extension for VS Code, rust-analyzer , you’ll still need the MSVC C++ Build Tools to compile and build your code, along with a Microsoft-developed Windows crate that gives you direct access to Windows APIs. An automated process collects metadata from Windows APIs, ensuring that they are always up-to-date. Microsoft also provides documentation on how Windows API calls and structures are projected into Rust .

It’s worth getting started now with the existing Rust tools to become familiar with Windows development in a new language. That way you can start to take advantage of Rust’s memory safety features, understanding how the language works and what changes you need to make to your programming style. Then when Microsoft starts shipping its new internal tooling for the rest of us, you can start using it to build low-level code, like drivers or libraries, that need to interoperate with Microsoft’s own C++.

US arrests California man for allegedly smuggling $300m worth of computer servers to China

Guardian
www.theguardian.com
2026-10-02 14:46:29
Greg Lui allegedly used false paperwork to smuggle export-controlled gear from US to third countries and then China US authorities have arrested a man accused of smuggling more than $300m worth of computer servers to China, the Department of Justice announced. Greg Lui, 38, of California, also known...
Original Article

US authorities have arrested a man accused of smuggling more than $300m worth of computer servers to China , the Department of Justice announced.

Greg Lui, 38, of California , also known as “Yiu Kong Lui”, allegedly used false paperwork and shipments through third countries to smuggle the export-controlled computer servers to China, the justice department said in a statement on Thursday.

The servers contained “US-manufactured graphics processing units (GPUs) commonly used for Super Intelligence (SI)”, it added, using the Trump administration’s new name for artificial intelligence (AI).

“Protecting America’s national security means keeping our advanced super intelligence technology from being used to strengthen our adversaries’ military capabilities,” the US attorney for the central district of California , Bill Essayli, said.

“We will aggressively prosecute those who put our national security at risk for profit,” he added.

Lui operated a technology company called Earthmade Computer, which he used to buy and send items to China without the licenses required by the US Department of Commerce from 2023 to 2024, according to the statement.

The company shipped the items to countries including Malaysia and Singapore, where no such license is required, then illegally re-exported them to China.

If convicted, Lui could face a maximum of 50 years in prison.

skip past newsletter promotion

China, long seen as trailing the United States on cutting-edge AI, has been closing that gap fast.

The United States has long tried to slow China’s progress with export controls on the advanced chips needed for cutting-edge AI, blocking access to top-end US semiconductors such as those made by Nvidia.

Klassik Revives the KDE 3 Desktop on Modern Plasma 6

Lobsters
linuxiac.com
2026-10-02 14:34:38
Comments...
Original Article

Klassik is a new open-source project that brings the classic KDE 3 desktop experience back to life on top of the modern KDE Plasma 6 stack.

Importantly, rather than simply applying a KDE 3-inspired theme to Plasma, the project goes further by recreating several of the desktop’s original components using current technologies. Klassik is built with Qt 6 and Qt Quick and supports features such as Wayland and fractional scaling.

For some historical context, KDE 3 debuted with KDE 3.0 on April 3, 2002. The series remained active for more than six years, with its final official maintenance release, KDE 3.5.10, arriving on August 26, 2008.

Back to Klassik, the project currently includes recreations of several familiar KDE 3 panel components, including the K Menu application launcher, Quicklaunch, task manager, lock and logout buttons, and digital and analog clocks.

Klassic is a modern recreation of the classic KDE 3 desktop experience for KDE Plasma 6.
Klassic is a modern recreation of the classic KDE 3 desktop experience for KDE Plasma 6.

These applets are written in C++ and QML and use QStyle and QPainter for rendering to match the original KDE 3 appearance as closely as possible. Klassik also provides its own KDE 3-style panel containment, including support for custom panel background images.

Moreover, the default KDE 2/3 QStyle application theme has been rewritten for Qt 6, while old KWin window decorations are being ported to the modern KDecoration3 API. Currently, the KDE 2 decoration is included.

Klassik also ships with several color schemes originally available in KDE 3, along with a collection of wallpapers from the same era. Combined, these pieces allow Plasma 6 to reproduce much more of the old KDE desktop than a conventional Global Theme could provide.

Under the hood, this remains a modern Plasma environment. Building Klassik requires Qt 6, KDE Frameworks 6, and KDE Plasma 6 libraries, including PlasmaQuick, KSysGuard, and KDecoration3. In other words, the KDE 3 appearance just sits on top of today’s KDE technologies.

Installation involves cloning the project’s GitHub repository, checking out the latest tagged release, compiling it with CMake, and installing it system-wide. Once installed, Klassik can be selected from System Settings under Colors & Themes > Global Theme.

The developer specifically recommends using the latest tagged release rather than the master branch, as the latter frequently contains unfinished features that may cause problems.

For additional details, see the project’s GitHub repo .

Image credits: Klassik

Warlock ransomware breach SharePoint in water, telecom operator attacks

Bleeping Computer
www.bleepingcomputer.com
2026-10-02 14:33:01
The China-linked ransomware group Warlock targeted a water utility, a telecom provider, a regional government body, and a university by exploiting SharePoint vulnerabilities to gain initial access. [...]...
Original Article

Warlock ransomware breach SharePoint in water, telecom operator attacks

The China-linked ransomware group Warlock targeted a water utility, a telecom provider, a regional government body, and a university by exploiting SharePoint vulnerabilities to gain initial access.

​Over the past two months, the threat actor appears to have focused on countries speaking Portuguese and Spanish across Europe, Africa, and Latin America.

The gang emerged in June 2025 and gained notoriety a month later after exploiting a chain of zero-day vulnerabilities in Microsoft SharePoint known as ToolShell (CVE-2025-49704, CVE-2025-49706, CVE-2025-53770, and CVE-2025-53771).

By August, Microsoft observed state-backed hacking groups Linen Typhoon and Violet Typhoon using ToolShell exploits in attacks, along with a ransomware threat actor the company tracks as Storm-2603.

EDR killer deployed to 40 hosts

Cybersecurity company Symantec identifies the same actor as Longlegs and attributes the development of the Warlock ransomware to the group.

According to the researchers, in one intrusion that started on July 22, the threat actor deployed a tool that disabled protection software on “at least 40 hosts within about two hours.” The attacker then launched Warlock ransomware on at least 33 hosts.

After gaining access, typically by exploiting vulnerabilities in on-premises SharePoint deployments, the attacker drops a web shell designed to function across multiple SharePoint versions.

Symantec and Carbon Black researchers say that in some attacks attributed to Longlegs, an AV/EDR-killing tool was deployed via the bring your own vulnerable driver (BYOVD) technique using a signed K7RKScan driver vulnerable to CVE-2025-1055.

Analysis of the intrusion on July 22 revealed that two days after gaining initial access, the threat actor engaged in reconnaissance activity and deleted what seemed like staging artifacts.

Using VS Code's tunneling capability

The ransomware payload was staged in the domain’s SYSVOL share, a location that stores public files and is replicated across every domain controller.

This is “a known method of pushing a payload out for execution by a logon script or Group Policy object across an entire network at once, rather than one host at a time,” the researchers say .

During the attack, the main executable file for Visual Studio Code Insiders was installed as a service to enable connecting remotely to compromised machines using VS Code's built-in tunneling capability.

On one of the systems, the researchers found the open-source penetration testing framework NetExec, which helped the attacker with Active Directory enumeration, credential spraying, and remote command execution.

The final stage of the attack occurred on July 31st, after deploying the AV/EDR killer, with Warlock ransomware “appearing almost as soon as protection was disabled on each host.”

The researchers warn that ToolShell and other SharePoint vulnerabilities remain viable initial access vectors, more than a year after Warlock first emerged exploiting SharePoint flaws.

The report from the Symantec and Carbon Black threat hunters includes a set of indicators of compromise for files and infrastructure used in the attacks.

article image

Build your security blueprint for AI-powered attacks

Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.

Save your seat

AI could expose how Georgia voters cast their ballot, researchers warn

Guardian
www.theguardian.com
2026-10-02 14:32:05
A Princeton researcher found that publicly available election records could be combined with AI to link voters to their ballots When voters cast their ballots, their votes are supposed to remain secret: from their family, their neighbors, and the government. But what if artificial intelligence could...
Original Article

When voters cast their ballots, their votes are supposed to remain secret: from their family, their neighbors, and the government.

But what if artificial intelligence could make secret votes visible?

Last month, Max Springer, a postdoctoral research fellow at Princeton University’s Center for Information Technology Policy, tested the prospect of identifying voters based on their Georgia ballots with a $20 subscription to an AI large language model and data obtained from an Open Records Act request.

“Within a couple of hours, I had a pipeline to analyze and identify secret ballots across the state of Georgia and the agent told me exactly what further information it would need to identify real voters’ ballots,” Springer wrote. “At no point did the agent refuse to comply or raise concerns over implementing the exploit.”

The issue has raised alarms among Georgia’s election officials less than two weeks before early voting starts in the high-stakes midterm election that will determine control of the US Congress.

Ballot secrecy is a democratic safeguard , allowing voters to shield their selection from anyone who might try to influence or intimidate them into voting a certain way. But Springer’s test showed that he could use AI to find out who people voted for, using Georgia’s public election records.

Georgia’s state elections board held an emergency meeting on Thursday morning to discuss how artificial intelligence has made it easier to identify voters through their ballot code, leaving people more vulnerable.

Georgia’s outgoing secretary of state, Brad Raffensperger, locked down the tabulation data that someone might use to identify a voter, ordering any public release to redact the ID numbers after an election, which “substantially address[es] this flaw”, said Ben Adida, an MIT-trained cryptographer and the founder of the election technology non-profit VotingWorks, as he addressed the board on Thursday.

Election workers make a record of every voter who comes through a polling place. Voters then choose candidates on a touchscreen, which then prints out a paper ballot. The voter feeds that ballot by hand into a scanner to be tabulated.

“When a voter inserts their ballot into a tabulator, a digital record of that ballot is created. Think of it as a row in a spreadsheet. One row is one scanned ballot. It’s often called a cast vote record, or CVR for short. In order to later retrieve the CVR, the tabulator assigns a unique identifier to each CVR at the time of scanning.”

The flaw shows up in how the machine records that ballot, Adida said. The identifying is not actually random enough.

Using AI, the early-voting list and the cast-vote record file for each county, Springer, the Princeton researcher who says he’s never been to Georgia, could recover the voting order for 1.52m ballots, or 98.9% of in-person ballots in 114 of the 139 counties he examined.

“In some of Georgia’s counties, I can determine the exact ballot cast by every voter at the vote centers – ballot secrecy is entirely lost,” he wrote.

In places with few voters on any given day, like tiny Heard county, Springer said: “the agent was able to match the majority of the 650 early in-person voters to a specific ballot, and the rest to within a single swap.”

About four years ago, election security researchers had discovered a flaw in the software used by Georgia voters that could connect a voter’s identity to their secret ballot.

The fact that ballots have an identifying number assigned at all is a product of the political paranoia attending Georgia’s elections. An ID number for each ballot addresses a fear of unscrupulous election workers double-stuffing ballots through tabulator machines. But if the order of ballots can be paired with other information, like sign-in data when a voter checks into the precinct, or perhaps Flock camera imagery that shows when someone entered a building on a slow day, then that ballot can be synched up to a voter’s name.

Representatives of the secretary of state’s office say they have been shouting into the wind for help with this problem for three years. Requests for money have been bound up in the long-running political feud between Raffensperger and conservative legislators positioning themselves in the orbit of Donald Trump and 2020 election skeptics. (Raffensperger has remained a top Trump target after refusing to overturn his defeat in the state’s 2020 presidential election.)

Spokesperson Robert Sinners sent a litany of emails and phone records showing requests for an appropriation to upgrade Georgia’s voting system, to no avail. The governmental affairs committee secretary, Tim Fleming, now the Republican nominee for secretary of state, was on the receiving end of many of those requests.

“We have presented numerous plans to the legislature for this – not, not just a patch, but an entire overhaul of the software to fix up any known vulnerabilities,” Sinners said. “Guess what we were gifted with from the legislature? A big fat goose egg.”

Fleming did not respond to a request for comment.

Other states use Dominion software for their elections. Every other state has either patched the software or withholds vulnerable voting data from public inspection.

Georgia law has two competing imperatives. The Georgia constitution explicitly requires a secret ballot, and the law protecting ballot secrecy in Georgia is among the strongest in the country. However, Georgia law – with revisions driven by the election disputes of 2020 – also demands transparency for election records.

Elections board member Salleigh Grubbs raised that latter point as she circulated a draft proposal that would call for poll workers to hold ballots in a tray before voters cast them into the machine, to shuffle the printed pages out of order as a way to defeat the problem. Grubbs said that Raffensperger’s fix – redacting the ID numbers in the cast vote record – violates the law.

But county elections officials and other board members – including Janelle King, a conservative Republican – dismissed the proposal as equally problematic.

“We are asking our election officials to completely retrain everybody that has been a part of this process up until this point on how to administer the election,” King said. “I also have major concerns about dropping our ballots into a secured box and then having other individuals scan my ballot. How do I know that in the shuffling something doesn’t fall out? I mean, I just have some major concerns about whether or not we are creating more problems while we are trying to address a problem.”

With early voting around the corner, changes to processes now are “a day late and about $30m short,” Sinners added.

The Liberty Are Alive—Now Anything Is Possible

hellgate
hellgatenyc.com
2026-10-02 14:17:20
A New York team that slogged through a regular season is now suddenly on fire in the playoffs. Sound familiar?...
Original Article

Watching the New York Liberty play this season was an excruciating affair.

Impressive, signature wins would be followed up by long languid stretches of play, where no one wanted to take shots, defense was optional, and the rotations made little to no sense.

At points, and especially during the brutal July stretch which saw them lose six of seven games against both scrub teams and contenders, it looked like that was it for this particular iteration of the Liberty, the one that had won the championship only two years ago.

Give us your email to read the full story

Sign up now for our free newsletters.

Sign up

Around 2-6% of World Bank foreign aid got siphoned into crypto wallets

Hacker News
www.nber.org
2026-10-02 14:14:29
Comments...
Original Article

The 2016 Panama Papers leak tightened regulatory enforcement around money laundering and offshore banking. We investigate whether the diversion of foreign aid in developing countries led to a shift to cryptocurrency as an alternative laundering platform. We develop a disbursement-timed forensic measure of cryptocurrency activity, combining on-chain Bitcoin transactions and wallet creation, off-chain exchange records, and IP-linked web traffic, and apply it to World Bank aid disbursements covering $238 billion across the 93 recipient countries in our estimation sample during 2018-2024. Exploiting the administrative timing of aid tranche arrivals, we find sharp, short-lived surges of crypto activity at the disbursement month, driven mainly by anonymous and newly created wallets on both tax-haven and mainstream exchanges. Blockchain forensics reveal patterns consistent with the placement, layering, and integration sequence of conventional money laundering. We estimate an implied leakage of 2 to 6 cents per aid dollar, which amounts to roughly 1.7 to 4.4 billion dollars of aid diversion across the tranche arrivals we study. Capture carries no funding penalty: the four sectors where we detect it, Transport, Water and Sanitation, Social Protection, and Governance, still absorb half of subsequent World Bank funding. Cryptocurrency facilitates aid diversion, but its transparent ledgers also leave forensic traces that may help detect and recover diverted funds.

What if we stopped using GPUs? [video]

Hacker News
www.youtube.com
2026-10-02 14:09:38
Comments...

Leaderboards and speedrun.com's new terms of service

Hacker News
therun.gg
2026-10-02 14:07:05
Comments...
Original Article
Leaderboards and speedrun.com’s new terms of service

From the creator of Redis; run LLM locally with ds4

Hacker News
dwarfstar.sh
2026-10-02 14:01:16
Comments...
Original Article

DS4 · LOCAL FRONTIER INFERENCE

Run frontier open weights locally with ds4.

DwarfStar 4 is a narrow C inference engine for high-memory Mac, CUDA and ROCm machines. It supports DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next, with text and vision models, local APIs, a CLI and a native agent in one stack.

SUPPORTED: DEEPSEEK V4 / V4.1 + GLM 5.x + QWEN3.8 · MIT LICENSE · C / METAL / CUDA / ROCM · QWEN ON 64GB

PRINCIPLE · LOCAL MODEL STACK

PHASE 1 · THE GIANT

A 284-billion-parameter star

DeepSeek V4 Flash is a large mixture-of-experts model. The usual path is remote serving; ds4 starts from the opposite constraint.

PHASE 2 · THE COLLAPSE

Compressed, not lobotomized

Asymmetric quantization targets the routed experts while preserving critical paths. The model becomes practical on high-memory machines.

PHASE 3 · THE DWARF STAR

Dense, resident, yours

The local engine exposes a CLI, HTTP APIs and a native agent, all sharing the same model state and cache.

How the collapse works →

SCROLL ▾

DESIGN CHOICES

Local frontier inference, narrow on purpose

Not a generic GGUF runner. ds4 follows a small, opportunistic set of model families and validates each supported layout end to end.

CORE 01

Asymmetric 2-bit quantization

Compress the routed experts, keep critical shared paths precise. That is how the supported routed-MoE builds fit their target machines.

CORE 02

KV cache as a disk citizen

Save long prefixes to SSD and resume by prompt hash. Restarts do not have to mean full re-prefill.

CORE 03

One engine, three interfaces

Use ./ds4 for chat, ./ds4-server for local APIs and ./ds4-agent for persistent coding sessions.

  • SSD STREAMING
  • TENSOR PARALLELISM
  • SESSION BATCHING
  • DSPARK + MTP
  • VISION INPUT

ARCHITECTURE

How the ds4 stack fits together

Project GGUFs, a self-contained engine and agent-facing interfaces, checked against official model outputs.

MODEL DeepSeek V4 / V4.1 GLM 5.x · Qwen3.8 supported GGUF layouts only asymmetric 2-bit + imatrix load ENGINE ds4 engine written in C metal · cuda · rocm parallel · batch · speculate KV cache RAM ⇄ SSD · survives restarts serve ./ds4 interactive CLI ./ds4-server OpenAI + Anthropic API ./ds4-agent native coding agent personal → distributed memory classes

RUNTIME MAP · SIMPLIFIED. SEE ARCHITECTURE NOTES FOR THE FULL DRAWING.

RUN IT

Run ds4 in three steps

Download the project GGUF, build for your backend, then start the CLI or server. Generic GGUF files are not the target.

STEP 1 · FETCH THE WEIGHTS

ds4 · zsh

$ git clone https://github.com/antirez/ds4
$ cd ds4 && ./download_model.sh ds4f-q2

STEP 2 · BUILD FOR YOUR BACKEND

ds4 · zsh

$ make
$ make cuda-spark

STEP 3 · TALK TO IT

ds4 · zsh

$ ./ds4
$ ./ds4-server --ctx 100000

FIT CHECK

ds4 hardware fit: local, streamed and distributed

Pick your platform and memory: get a conservative starting path and understand which execution modes apply.

✓ Runs well

V4 Flash Q2 is the baseline. At 128 GB, GLM 5.3 Q2 and Qwen Q4 also fit; V4.1 Q2 streams from SSD.

./download_model.sh ds4f-q2 && make

REF · M5 MAX 128GB · 32K CTX: 34.4 T/S GEN · 557 T/S PREFILL

Estimates from the ds4 benchmark table . Full guide in Hardware and Installation .

BENCHMARKS

ds4 benchmarks: prefill and generation

Reference rows from upstream. Read prefill and generation separately, especially for long-context agent workloads.

Machine Context Prefill t/s Generation t/s
M5 Max, 128 GB q2 · 2,048 tok 790.2 39.4
M5 Max, 128 GB q2 · 65,536 tok 398.5 27.6
DGX Spark, 128 GB q2 · 2,048 tok 825.8 18.1
DGX Spark, 128 GB q2 · 65,536 tok 823.0 13.8

All benchmarks →

API & AGENTS

Use ds4 from Codex, Claude Code and OpenCode

ds4-server speaks OpenAI and Anthropic-style APIs, so local coding agents can connect to your own machine with a base URL.

Own your local AI inference.

Start with the quickstart, check the hardware matrix, then connect your editor, agent or API client to the local server.

Driving the GDEH0154D67 e-paper display with Rust

Lobsters
sgt.hootr.club
2026-10-02 13:55:35
Comments...
Original Article

A while ago I bought a SQFMI Watchy . It's an open source e-paper watch driven by an ESP32 that I specifically bought to hack on: the firmware is open source and you can very easily create your own watchface by cloning the firmware and directly modifying the code! And I did create my watchface, which I now recognize looks kinda bad.

My Watchy in 2023 with its default case and band, running a modified version of the stock firmware

I had just bought a coffee table book on arcade videogame typefaces and I ended up picking the typeface for Passing Shot , an arcade top-down tennis game from 1988, for the watch face. It looks a lot better with color, but I had spent a lot of time encoding those glyphs into 1s and 0s and damn if I wasn't gonna use it in my crappy watch face.

I wanted to do a lot of things with this firmware: rewrite the menu, have it receive notifications from my phone via BLE, set configurable alarms... There was only one problem: the firmware is written in C++ and uses Arduino libraries. I found the documentation and quality of the code for those libraries lacking. I didn't understand how the firmware interacted with the display and the sensors. I couldn't use tagged unions. The solution? Say it with me: Rewrite👏It👏In👏Rust!

Rust on the ESP32

At the time of receiving my Watchy, Rust support for ESP32 microcontrollers was in an experimental stage and developing rapidly. At first I gave up the experiment because the HAL didn't even allow me to put the microcontroller in deep sleep, which is very important for a smartwatch!

Once that was added, I struggled with getting Espressif's forks of LLVM and rustc running on my NixOS desktop. At some point I had this very janky setup where I aliased gcc and cargo and rustc to tiny shell scripts that brought up Espressif's Docker container, running that one command and stopping right after. Eventually I managed to get this working on Nix by fetching the precompiled binaries; I still haven't figured out how to make an overlay with custom rustc and LLVM using nixpkgs.

Throughout all this the esp-rs crates kept changing the API on each release (they hadn't reached a 1.x release yet) so every time I upgraded I had to take an hour or two to fix all the compilation failures and figure out where the new APIs were and ensure that all the dependencies were on the correct version. Riveting work!

On top of all that, recently esp-hal dropped support for ESP32 chips < v3.0 , and my Watchy runs on a revision 1.0 chip. I managed to get around it by setting an environment variable in my .cargo/config.toml but it doesn't exactly inspire confidence.

Despite all this I still wanted to get my Rust firmware running on the Watchy. The first step was to implement drivers for the external devices: the accelerometer, RTC, and e-paper display.

The display driver

Implementing drivers was a much smoother experience: I wanted the firmware to be async through Embassy , so all I had to do in the driver crates was to pull embedded_hal_async as a dependency and write my code against its interfaces. There are crates published on crates.io for all of the devices, but they were either incomplete or not async, so I had to write my own versions. I'll focus on the display driver here.

My revision of the Watchy (2.0) uses the GDEH0154D67 e-paper display , driven by the SSD1681 display driver . I pored over the datasheets and GxEPD2 's source code and eventually ended up with a working driver for the panel.

I'd never touched a driver or a microcontroller before this so it took me many bursts of work spread over several years (with many moons passing between each burst) to figure everything out. Turns out embedded development is not so daunting as it looks! If you've worked with sockets and binary protocols before you'll find that communicating with this controller over SPI is basically the same: you write a byte corresponding to a command number, and then some data serialized according to what the datasheet tells you. The driver code itself is boring: it simply maps each SPI command described in the datasheet to a method and models data as type-safely as I could manage.

When I started testing the driver I got some very weird results! I more or less followed Watchy's code, but I wanted to experiment on the code that updates the display, because I didn't understand it and I wanted to see if I could get it to run any faster with some tweaks. A full display update would get me an empty display, and partial updates looked corrupted, save for the first one after a full update. What's going on!?

My Watchy today, with its much thinner case and an orange cloth band, running my Rust firmware and displaying very bad ghosting after a partial update

Corrupted updates

When you update the e-paper display you can choose between display mode 1 and 2:

  • Mode 1 is a full display update: the whole display flashes black and white and you get a clean and artifact-free image, though it takes a few seconds to update. E-paper devices generally do a full update when you boot them up or every once in a while to clean up the image.
  • Mode 2 is a "partial" update: the driver tries to only move the pixels that have changed since the last update, which is much faster than a full update. This works best when the frame hasn't changed much since the last update and usually produces artifacts that kind of look like the ink of the previous page bleeding into the next on real paper.

The SSD1681 display driver supports both monochrome black/white and 3-color black/white/red e-paper panels, like the ones you see on labels in some supermarkets. To support 3-color displays it exposes two 1-bit framebuffers: one for the black and white pixels and one for red, which I'll refer to here as b/w RAM and red RAM .

The GDEH0154D67 panel is monochrome, so it repurposes red RAM as a "previous frame" buffer for partial updates, while the b/w RAM contains the current frame. As I understand it, a partial update essentially diffs the current and previous frame buffers to produce the appropriate waveforms to transition the ink from one color to the next, so when you're driving the display you have to make sure that the previous frame buffer actually reflects the pixels that are currently on the display.

Understanding that should be enough to drive the display. My Display::draw function roughly looked like this:

  1. Initialize the display
  2. Write a frame to b/w RAM
  3. Update the display
    • For a full update, use mode 1
    • For a partial update, use mode 2
  4. Write the same frame to red RAM for the next partial update

This... didn't have the effect I thought it would. A full update showed a stale frame from before I had flashed the new firmare , the first partial update looked clean, but subsequent updates would corrupt the image like in the photo above. Clearly I had something wrong.

Ping-pong

While writing the driver I came across an option that the datasheet calls "ping-pong": basically when the option is on, a partial update will also swap the b/w and red RAM so that b/w ram points to the contents of red RAM and vice-versa. I thought this looked useful and wondered whether I could use it on this firmware; the datasheet mentioned that this option was disabled by default.

It took me a while to realize that this option is turned on in the OTP memory of the Watchy. It's kind of obvious in hindsight, but I had to draw a diagram and write some pseudocode to understand it correctly. Let me show you how it works in pseudo-C.

Since b/w and red RAM are not stable anymore, let's call the two RAMs RAM0 and RAM1. After a software reset (which is part of the startup sequence of the display, according to the datasheet), b/w RAM points to RAM0 and red RAM points RAM1.

typedef uint8_t ram[5000];

static ram *bw_ram;
static ram *red_ram;

void software_reset() {
    bw_ram = &RAM0;
    red_ram = &RAM1;
}

On a full update I wrote to b/w RAM, but in hindsight this didn't make much sense: partial updates are driven by the previous frame being in red RAM, so when you're doing a full update you'd have to write to both RAMs to leave the red one in place for a partial update. It makes much more sense to drive full updates exclusively from red RAM, and indeed, this turned out to fix the issue with the stale frame I was seeing earlier.

void full_update() {
    draw_on_panel(*red_ram);
}

On a partial update, as I mentioned before, b/w RAM is diffed with red RAM, which is expected to contain the previous frame. Like in subtraction, the order of the operands is important here. If ping-pong is enabled, the RAMs are also swapped.

void partial_update() {
    draw_diff_on_panel(*bw_ram, *red_ram);
    swap(&bw_ram, &red_ram);
}

Hence, the partial update sequence kind of looks like this:

uint8_t *F0 = prev_frame;
uint8_t *F1 = current_frame;

// We'll assume red RAM already contains the previous frame.
assert(memcmp(*red_ram, F0, sizeof(ram)) == 0);

write(bw_ram, F1);
partial_update();

// The pointers are now swapped, so bw_ram now points
// to the previous frame and red_ram to the current.
assert(memcmp(*bw_ram, F0, sizeof(ram)) == 0);
assert(memcmp(*red_ram, F1, sizeof(ram)) == 0);

If you're doing several partial updates in a row, ping-pong is useful because it saves you from writing to red RAM after you've written to b/w.

On a software reset though, the pointers are reset to their original values, meaning that a partial update will diff against a stale frame (unless we partially updated the display an even number of times before the reset).

software_reset();

assert(memcmp(*bw_ram, F1, sizeof(ram)) == 0);
assert(memcmp(*red_ram, F0, sizeof(ram)) == 0);

This explains the partial update corruption: after the first partial update, red RAM pointed to RAM0, and after a software reset it points back to RAM1, so I basically kept overwriting RAM0 while RAM1 was only written to during a full update.

write(bw_ram, F2);
partial_update(); // draw_diff_on_panel(F2, F0);

The fix turned out to be simple. I split the Display::draw function into two, with Display::draw_full only writing to red RAM:

  1. Initialize the display
  2. Write frame to red RAM
  3. Update the display using mode 1

...and Display::draw_partial writing to b/w RAM twice:

  1. Initialize the display
  2. Write the frame to b/w RAM
  3. Update the display using mode 2
  4. Write the frame to b/w RAM (which used to be red RAM, and will be red RAM after a reset)

The second write ensures that both RAMs now contain the current frame, so we always get a valid transition both with and without a software reset. This be optimized away through some internal bookkeeping, or maybe I could avoid the software reset when I know the Watchy isn't coming from an "unknown" state. Perhaps I'll explore those options in the coming days.

Future plans

Now that I finally have everything working I think I'll just implement the basics and enjoy wearing my epaper watch around, after it's been confined to a drawer for so much time. I expect you'll have some questions for me.

May I see the watchface?

Here you go. It uses the Upheaval font, whose bytes I got from this repo , and it's running 100% pure Rust! (Except for the bootloader, I think, but I decided that it doesn't count.)

The watch as in the previous picture, with 100% less partial update corruption and a battery voltage reading on the display

May I see the code?

It's yours, my friend , as long as you respect the terms of the license it comes with. It's still kind of messy so don't expect much outside of the drivers!

Is this AI slop?

Nope. I used it a little bit to understand some things about EPDs but I wrote all the code, and all of this post, with my own hands and using my own brain.

I'll take further questions through email or comment sections, if you find this post while browsing a site that has one of those. Thank you for reading!

There's a New Sea Spider in Town

Hacker News
nautil.us
2026-10-02 13:48:54
Comments...
Original Article

What has hairy feet, legs dotted with eggs, and a mouth loaded with tentacles like a Sarlacc from Star Wars ?

That’s how researchers from the University of British Columbia describe seven such creatures from the Salish Sea in a new paper published in the journal Organisms Diversity & Evolution . They include five known species plus a couple of new ones, marking the first discovery in a century of new sea spiders in the region.

As ancient arachnids closely related to scorpions, spiders, and horseshoe crabs, sea spiders share some traits such as exoskeletons, jointed legs, and chelicerae-type mouthparts adapted for piercing or grabbing prey. But sea spiders also rank as weird for the way they sport up to 12 legs, breathe through their skin, and spawn eggs from their legs, traits they’ve honed during their 500 million years of ocean life.

Read more: “The Mystery of 111,000 Spiders Living in a Giant Subterranean Web”

“A lot of the shallow water species are quite small, and so they’ll live as parasites on larger organisms, and they use this weird apparatus called a proboscis to suck the juices out of their host,” explained lead author Cormac Toler-Scott.

The researchers collected specimens from sites that represented the range of sea spider habitats in the Salish Sea, from depths of 0 to about 60 feet, which would be considered shallow waters for sea spiders. Using scanning electron microscopy and DNA analysis, they classified the sea spiders by species, detecting the two previously undescribed ones in the process.

Callipallene pilosuspedes (“hairy feet”) joins 36 other species of Callipallene worldwide, which tend to be slender and without claws on the male’s “ovigers,” or the special appendages used by males to pluck eggs off females and stick them to their own body for rearing. It has a triangular mouth bounded by three lips and covered with sensory hairs that give it the Sarlacc appearance, albeit in miniature—after all, C. pilosuspedes is only about a centimeter long.

The three lips of the C. pilosuspedes . Credit: Cormac Toler-Scott

The other new species, Tanystylum kiixin , gets its name from an indigenous village near the discovery site. Its legs are too short and stiff for good hygiene, so it gets covered with debris. Amongst the body hairs and debris, the researchers found even tinier single-celled parasites, such that the sea spiders maintain ecosystems on their bodies.

Which is yet another reason why Toler-Scott and his fellow researchers consider these water-based arachnids “horrifyingly fascinating.”

Enjoying Nautilus ? Subscribe to our free newsletter .

Lead Image: Cormac Toler-Scott

Congress Has Another Site-Blocking Bill, And This One Targets VPNs

Electronic Frontier Foundation
www.eff.org
2026-10-02 13:41:08
Congress is taking another run at site-blocking, a deeply flawed concept that would undermine basic internet infrastructure. Rep. Darrell Issa (R-CA) has introduced the American Copyright Protection Act (ACPA), H.R. 10364, a bill that would give copyright owners a new legal tool to block Americans’ ...
Original Article

Congress is taking another run at site-blocking, a deeply flawed concept that would undermine basic internet infrastructure. Rep. Darrell Issa (R-CA) has introduced the American Copyright Protection Act (ACPA), H.R. 10364 , a bill that would give copyright owners a new legal tool to block Americans’ access to foreign websites accused of copyright infringement.

The basic idea is all too familiar , and it’s still dangerous. A copyright owner first asks a court to label a foreign website a “foreign piracy site.” Once that happens, the copyright owner could seek orders requiring internet service providers, DNS providers, and—new and explicit in this bill—VPN providers to take “commercially reasonable steps” to stop their users in the United States from accessing those sites. The decision to label a website as a “foreign piracy site” can happen without the accused site even showing up in court to defend itself.

ACPA Goes Further Than Other Site-Blocking Proposals

In some ways, the ACPA is even worse than a site-blocking legislation introduced last year , the Foreign Anti-Digital Piracy Act (FADPA) , which EFF also opposed. That bill at least excluded companies that provide only VPN services, as well as providers that offer DNS resolution exclusively through encrypted DNS protocols. The ACPA drops those protections. In fact, the bill explicitly includes VPNs among the service providers that can be ordered to block access to a website.

The bill also broadens the definition of a “piracy site.” Last year’s site blocking bill covered sites with “no commercially significant purpose or use” other than infringement. ACPA changes that to sites with “only limited commercially significant purpose or use” beyond infringement. In other words, under ACPA, even a website with legitimate commerce going on could still be labeled a “foreign piracy site” and ultimately blocked for all Americans.

Better Process Still Doesn’t Fix The Problem

The ACPA includes some procedural protections, such as requiring service providers that could be subject to a blocking order to receive legal notice and an opportunity to respond. The bill also requires courts to consider the potential harm to other websites and internet users before ordering intermediaries to block websites. It further requires the copyright owner to post a bond, in an amount determined by the court, sufficient to cover the costs and damages incurred by any service provider found to have been wrongfully enjoined. The bill also provides a mechanism for operators or users of third-party online services affected by erroneous blocking to seek compensation after the fact in certain circumstances. Finally, a site operator can ask a court to rescind its designation as a “foreign piracy site.”

These safeguards are significant and positive changes, but they don’t solve the basic, and severe, due process problem. The initial decision to label a website a “foreign piracy site” can still be made without the site operator appearing to defend itself. The court can appoint a “special master,” which is an independent expert who helps the judge evaluate evidence, to review the copyright owner’s case—but that step is not required. In any case, a special master  is not a lawyer who actually represents the accused website, nor the users whose access to information and speech may be affected.

We know what site-blocking looks like when it’s put into practice. Supporters of site-blocking like to point to its use in other countries. But what we’re seeing in other countries is serious collateral damage to lawful websites. In Italy, 510 benign, non-streaming websites , including a Catholic convent and a telehealth platform, were blocked by the country’s “Piracy Shield” program. In Spain, a site-blocking system blocked more than 550,000 domains during soccer broadcasts, including sites belonging to Greenpeace and Harvard University.

Congress Should Reject Site-Blocking Proposals

More than a decade ago, Congress abandoned SOPA and PIPA after internet users pushed back against site-blocking and other threats to the open internet. We shouldn't start building that infrastructure now.

ACPA adds some safeguards, but those don’t fundamentally change what Congress is being asked to create: a system for blocking Americans’ access to entire websites at the request of copyright owners. By explicitly bringing VPNs into that system, the bill also reaches into basic tools that people use to access the internet safely and privately. Adding somewhat better procedures to a bad idea doesn’t turn it into a good idea.

Podcast: The FBI Was Hacked. We’ve Seen the Data

403 Media
www.404media.co
2026-10-02 13:37:35
The massive FBI hack; a company wants to add facial recognition onto Flock; and Muse in some cases is actually just people....
Original Article

The massive FBI hack; a company wants to add facial recognition onto Flock; and Muse in some cases is actually just people.

Podcast: The FBI Was Hacked. We’ve Seen the Data
Image: 404 Media.

We start this week with Joseph’s stories on the massive FBI hack. 404 Media broke this news and continued to cover the breach of data on all FBI employees and their spouses. After the break, Jason tells us how a surveillance company is telling cops it wants to add facial recognition tech to Flock camera data. In the subscribers-only section, Jason and Joseph chat all about Muse, Meta’s new agentic AI, and how humans are actually behind some of the tasks. Because of course.

Listen to the weekly podcast on Apple Podcasts , Spotify , or YouTube . Become a paid subscriber for access to this episode's bonus content and to power our journalism. If you become a paid subscriber, check your inbox for an email from our podcast host Transistor for a link to the subscribers-only version! You can also add that subscribers feed to your podcast app of choice and never miss an episode that way. The email should also contain the subscribers-only unlisted YouTube link for the extended video version too. It will also be in the show notes in your podcast player.

About the author

Joseph is an award-winning investigative journalist focused on generating impact. His work has triggered hundreds of millions of dollars worth of fines, shut down tech companies, and much more.

Joseph Cox

‘Just like my big brothers’: boy with cerebral palsy passes GCSEs with eye-tracking technology

Guardian
www.theguardian.com
2026-10-02 13:18:06
Patrick McCabe’s herculean feat shows limits and possibilities of special educational needs and disabilities system Patrick McCabe is instantly recognisable in the early autumn sunshine, his wheelchair surrounded by fellow pupils at Fallibroome Academy in Macclesfield, Cheshire as they move from les...
Original Article

P atrick McCabe is instantly recognisable in the early autumn sunshine, his wheelchair surrounded by fellow pupils at Fallibroome Academy in Macclesfield, Cheshire as they move from lesson to lesson, chatting, laughing and joking around.

His extraordinary GCSE success this summer has become the stuff of legend. The 16-year-old has cerebral palsy and it took four months for him to complete his exams using eye-tracking technology. Now Patrick has started his A-levels – French, business and media studies – and fresh challenges lie ahead.

In person, Patrick is impish, thoughtful, articulate, witty, intelligent and honest. Using his augmentative and alternative communication (AAC) device, he tells the Guardian school has been a good experience overall. Born in Switzerland, his family moved back to England when he was six. Since then, he has gone to local mainstream schools, where he wanted to learn like everyone else.

“The best thing has been that, just like my big brothers, I’ve been able to learn and develop skills and gain knowledge of the world around me,” he said. “Even if some days I didn’t always want to go [to school], I think the opportunity to learn has helped me to grow and develop an independent mind.”

It has not been easy. “The most difficult things have been building meaningful relationships, and communication barriers do exist,” said Patrick. “People are not always familiar with Eyegaze technology and are not always willing or able to give it a chance in terms of friendships.”

Patrick posing with Jo, who has her arm around him
Patrick with his mother Jo. Photograph: Christopher Thomond/The Guardian

In many ways, Patrick epitomises the Labour government’s ambitions for students with special educational needs and disabilities (Send). Forthcoming reforms are built on an inclusive approach , keeping children and young people in local – rather than special – schools with all the necessary expertise and support on hand.

But Patrick’s story also illustrates the scale of change required if education and assessment in England are to be truly inclusive. His experience shows that exams are far from accessible to all, his family and school say vital expertise is not always available and – crucially – funding is never enough.

“All we wanted was for Patrick to have the opportunity to demonstrate his knowledge and be given something resembling a level playing field compared to other students,” said the headteacher, Ross Martland. The school and Patrick’s mother, Jo, battled with the exam boards to try to ensure he was given a fair chance.

In the end, it doesn’t sound very fair at all. While his peers spent the first part of the year revising, Patrick’s exams got under way on 2 February and went on every morning until 5 June. It required a herculean effort from him, which left him exhausted, frustrated and sometimes ill.

Patrick at school
‘I felt I had to work harder to express what I know,’ said Patrick about using the eye-tracking technology. Photograph: Christopher Thomond/The Guardian

A 30-mark question that might form just one element of an exam could take up to six hours. Patrick was supported throughout by CandLE, an organisation that helps students using AAC technology to access education across the UK. That he passed all seven GCSEs and achieved the grades he needed to continue into sixth form is testament to his determination.

“Revising for my GCSEs was very exhausting because it took extra time and effort to revise in the first place,” said Patrick. “I felt I had to work harder to express what I know, even when I understand the subject. This was tiring, both physically and mentally. I needed extra breaks to rest, but I still wanted to do my best to show what I had learned.”

skip past newsletter promotion

Martland is fully supportive of the government’s Send ambitions. “I think the plans are right. Morally and values-wise, it’s absolutely right.” The issue is funding, he said, adding: “I’m concerned it will add significant pressures to schools if it’s not resourced properly.”

Throughout his time at Fallibroome, Patrick has worked with a devoted team of teaching assistants who support him not only with his learning in class but also his personal care. “We are incredibly proud of Patrick and all of our Send students,” said Rachel Crew, one of the teaching assistants working with him.

Patrick says the transition to sixth form is going well. He’s shed the bottle green uniform in favour of a smart blue suit and he says the school is doing everything it can to support him.

He says his mum has been “amazing” and he thinks students with physical disabilities who use eye-tracking technology should be assessed in a more appropriate way, so they don’t face additional barriers. But it is not only exams that worry him. “My concerns remain around friendships, if I’m honest,” he says.

So what’s Patrick planning next? After A-levels, he wants to study media at university. “I find it interesting and I’d like to go into the industry.” His mum tentatively suggests nearby Salford. Patrick’s looking at Bournemouth, 200 miles away. Is he daunted? No chance. “I feel excited about the experience.”

Mystery Function

Hacker News
codeset.ai
2026-10-02 13:03:08
Comments...
Original Article

© 2026 codeset · Codeset, Lda

Early access. 2,000 free credits, no card required.

ICC judge on what U.S. sanctions mean for her and global courts

Hacker News
www.npr.org
2026-10-02 13:02:31
Comments...
Original Article
Exterior of the International Criminal Court building in The Hague.

The International Criminal Court in The Hague on Sept. 23. Remko de Waal/ANP/AFP via Getty Images hide caption

Remko de Waal/ANP/AFP via Getty Images

If you want to see how far U.S. sanctions can reach, ask Kimberly Prost's Alexa.

After the United States sanctioned her, the Canadian judge at the International Criminal Court began losing access to services from companies in different countries. Her credit cards were also canceled.

"You start to lose services from American companies, Amazon, Google, potentially Apple, whatever it may be," she told Morning Edition . "Completely randomly, though, you don't know when something's going to happen."

One of those moments came during an ordinary moment at home.

"I went home and Alexa wouldn't talk to me," she told NPR's Steve Inskeep.

Her credit cards, issued in the Netherlands and Canada, were canceled. She has managed to keep some bank accounts, but says her banking services are extremely limited and even small transfers can be blocked.

The sanctions themselves are American. Still, businesses abroad may comply because they have ties to the U.S.

"We have this massive overcompliance by companies," she said.

One of the most serious examples is her health insurance. Prost receives coverage through AXA, the French company that provides insurance for the court. She says the company has refused to pay her claims even though it is not legally required to follow the U.S. sanctions.

"It's a business decision," she said.

Her phone still works — "so far" — but some email, cloud and other technology services are limited.

She has lived with the restrictions for more than a year.

Why was she sanctioned?

Prost said her sanction stems from Trump's executive order targeting ICC officials involved in cases concerning Israel and Gaza or Afghanistan.

Her own case goes back six years.

At the time, Prost was temporarily serving on the ICC's Appeals Chamber when prosecutors sought authorization to open an investigation into alleged crimes connected to the war in Afghanistan. Prost and four other judges unanimously approved the request.

The investigation covered alleged crimes involving the Taliban and ISIS and included what Prost described as "a small component related to the United States."

That U.S. component was never pursued and was later deprioritized by prosecutors, she said.

"And it is for that decision that I'm sanctioned," Prost said.

Her role was judicial, not investigative. She said the appeals judges were required to decide whether prosecutors had the proper legal basis and jurisdiction to proceed — similar, she said, to a judge authorizing a search warrant.

"There is no investigation related to the American component in Afghanistan, and I have nothing to do with the Israel-Gaza situation," Prost said.

Her designation came amid a much wider confrontation between Washington and the ICC that intensified after the Oct. 7, 2023, Hamas-led attack on Israel and the ensuing war in Gaza.

For Prost, the broader concern is what sanctions like these mean for judges on international courts making decisions from the bench.

"We are judges, and we are being told that rather than deciding our cases on the basis of the facts and the law before us, as we are duty bound to do, we're being asked to decide the cases or not decide them in accordance with the wishes of the U.S. government," she said.

By August 2026, nine of the court's 18 judges were under U.S. sanctions, along with both deputy prosecutors, its former prosecutor and one staff member.

The Trump administration has threatened to go further, announcing what it called a "whole-of-government campaign" to "systematically dismantle" the ICC over its investigations involving Americans and Israelis.

NPR reached out to the State Department for more information about whether and when the administration plans to impose additional sanctions on the court. The department did not respond.

Prost is challenging the sanction in U.S. court. In June, she and fellow ICC judges Reine Alapini-Gansou and Solomy Balungi Bossa sued President Trump and other administration officials in federal court in New York, arguing that the measures were unlawful.

The work goes on in The Hague

For now, the sanctions have changed how she lives outside the courtroom, but not the work she continues to do inside it.

Her current case involves alleged crimes tied to the jihadist takeover of Timbuktu in Mali, including persecution, torture and inhumane treatment.

"I will never quit my job as a judge," Prost said. "These methods that they're using, these coercive measures, are completely futile."

She said she and other sanctioned judges and prosecutors continue to do their jobs "objectively, independently and in accordance with the oath that we've taken."

The radio interview was produced by Ava Pukatch, and the web version was edited by Treye Green.

STS-51-F Abort-to-Orbit (1985)

Hacker News
en.wikipedia.org
2026-10-02 12:56:47
Comments...
Original Article

From Wikipedia, the free encyclopedia

STS-51-F

Experiments in Challenger ' s payload bay

Names Space Transportation System -19
Spacelab 2
Mission type Astronomical observations
Operator NASA
COSPAR ID 1985-063A Edit this at Wikidata
SATCAT no. 15925 Edit this on Wikidata
Mission duration 7 days, 22 hours, 45 minutes, 26 seconds
Distance travelled 5,284,350 km (3,283,540 mi)
Orbits completed 127
Spacecraft properties
Spacecraft Space Shuttle Challenger
Launch mass 114,693 kg (252,855 lb)
Landing mass 98,309 kg (216,734 lb)
Payload mass 16,309 kg (35,955 lb)
Crew
Crew size 7
Members
Start of mission
Launch date 21:00, July 29, 1985 (UTC) (5:00 pm EDT )
Launch site Kennedy , LC-39A
Contractor Rockwell International
End of mission
Landing date August 6, 1985, 19:45:26 UTC (12:45:26 pm PDT )
Landing site Edwards , Runway 23
Orbital parameters
Reference system Geocentric orbit
Regime Low Earth orbit
Perigee altitude 312 km (194 mi)
Apogee altitude 320 km (200 mi)
Inclination 49.49°
Period 90.90 minutes
Instruments
  • Carbonated Beverage Dispenser Evaluation
  • Infrared telescope (IRT)
  • Instrument Pointing System (IPS)
  • Plasma Diagnostics Package (PDP)
  • Shuttle Amateur Radio Experiment

STS-51-F mission patch

Front row (seated): C. Gordon Fullerton , Roy D. Bridges Jr.
Back row (standing): Anthony W. England , Karl G. Henize , F. Story Musgrave , Loren W. Acton , John-David F. Bartoe

← STS-51-G (18)

STS-51-I (20) →

STS-51-F (also known as Spacelab 2 ) was the 19th flight of NASA 's Space Shuttle program and the eighth flight of Space Shuttle Challenger . It launched from Kennedy Space Center , Florida , on July 29, 1985, and landed eight days later on August 6, 1985.

While STS-51-F's primary payload was the Spacelab 2 laboratory module, the payload that received the most publicity was the Carbonated Beverage Dispenser Evaluation , which was an experiment in which both Coca-Cola and Pepsi tried to make their carbonated drinks available to astronauts. [ 1 ] A helium-cooled infrared telescope (IRT) was also flown on this mission, and while it did have some problems, it observed 60% of the galactic plane in infrared light. [ 2 ] [ 3 ]

During launch, Challenger experienced multiple sensor failures in its Engine 1 Center SSME , which led to it shutting down. As a result, the shuttle had to perform an " Abort to Orbit " (ATO) emergency procedure. It is the only Shuttle mission to have carried out an abort after launching. As a result of the ATO, the mission was carried out at a slightly lower orbital altitude.

Crew seat assignments

[ edit ]

Seat [ 5 ] Launch Landing
Seats 1–4 are on the flight deck.
Seats 5–7 are on the mid-deck.
1 Fullerton
2 Bridges
3 Henize
4 Musgrave
5 England
6 Acton
7 Bartoe
Aborted launch attempt at T−3 seconds on July 12, 1985
The control panel of the Shuttle on the STS-51-F mission, showing the selection of the Abort-to-Orbit (ATO) option

STS-51-F's first launch attempt on July 12, 1985, was halted with the countdown at T−3 seconds after main engine ignition, when a malfunction of the number two RS-25 coolant valve caused an automatic launch abort. Challenger launched successfully on its second attempt on July 29, 1985, at 17:00 p.m. EDT , after a delay of 1 hour 37 minutes due to a problem with the table maintenance block update uplink.

At 3 minutes 31 seconds into the ascent, one of the center engine's two high-pressure fuel turbopump turbine discharge temperature sensors failed. Two minutes and twelve seconds later, the second sensor failed, causing the shutdown of the center engine. This was the only in-flight RS-25 failure of the Space Shuttle program . Approximately 8 minutes into the flight, one of the same temperature sensors in the right engine failed, and the remaining right-engine temperature sensor displayed readings near the redline for engine shutdown. Booster Systems Engineer Jenny M. Howard acted quickly to recommend that the crew inhibit any further automatic RS-25 shutdowns based on readings from the remaining sensors, [ 6 ] preventing the potential shutdown of a second engine and a possible abort mode that may have resulted in the loss of crew and vehicle (LOCV). [ 7 ] [ failed verification ]

The failed RS-25 resulted in an Abort to Orbit (ATO) trajectory, whereby the shuttle achieved a lower-than-planned orbital altitude. The plan had been for a 385 km (239 mi) by 382 km (237 mi) orbit, [ 4 ] but the mission was carried out at 265 km (165 mi) by 262 km (163 mi) . [ 8 ]

Attempt Planned Result Turnaround Reason Decision point Weather go (%) Notes
1 12 Jul 1985, 3:30:00 pm Scrubbed — Technical 12 Jul 1985, 3:29 pm ​ (T−00:00:03) Pad abort: malfunction in SSME #2 coolant valve shutdown of all three main engines. [ 9 ] [ 10 ]
2 29 Jul 1985, 5:00:00 pm Success 17 days 1 hour 30 minutes Launched after 1 hour 37 minute delay to resolve issue with table maintenance block update uplink. At T+343 seconds, SSME #1 shut down leading to ATO (Abort to Orbit). [ 8 ]
The Plasma Diagnostics Package (PDP) grappled by the Canadarm
Space art for the Spacelab 2 mission, showing some of the various experiments in the payload bay
Tony England drinks soda in space
A view of the Sierra Nevada mountains and surroundings from Earth orbit, taken on the STS-51-F mission

STS-51-F's primary payload was the laboratory module Spacelab 2. A special part of the modular Spacelab system, the " igloo ", which was located at the head of a three-pallet train, provided on-site support to instruments mounted on pallets. The main mission objective was to verify performance of Spacelab systems, determine the interface capability of the orbiter, and measure the environment created by the spacecraft. Experiments covered life sciences , plasma physics , astronomy , high-energy astrophysics , solar physics , atmospheric physics and technology research . Despite mission replanning necessitated by Challenger ' s abort to orbit trajectory, the Spacelab mission was declared a success.

The flight marked the first time the European Space Agency (ESA) Instrument Pointing System (IPS) was tested in orbit. This unique pointing instrument was designed with an accuracy of one arcsecond . Initially, some problems were experienced when it was commanded to track the Sun , but a series of software fixes were made and the problem was corrected. In addition, Anthony W. England became the second amateur radio operator to transmit from space during the mission.

Spacelab Infrared Telescope

[ edit ]

The Spacelab Infrared Telescope (IRT) was also flown on the mission. [ 3 ] The IRT was a 15.2 cm (6.0 in) aperture helium-cooled infrared telescope, observing light between wavelengths of 1.7 to 118 μm . [ 3 ] It was thought heat emissions from the Shuttle would corrupt long-wavelength data, however it still returned useful astronomical data. [ 3 ] Another problem was that a piece of mylar insulation broke loose and floated in the line-of-sight of the telescope. [ 3 ] IRT collected infrared data on 60% of the galactic plane. [ 2 ] (see also List of largest infrared telescopes ) A later space mission that experienced a stray light problem from debris was Gaia astrometry spacecraft launch in 2013 by the ESA—the source of the stray light was later identified as the fibers of the sunshield, protruding beyond the edges of the shield. [ 11 ]

Carbonated Beverage Dispenser Evaluation

[ edit ]

In a heavily publicized marketing experiment, astronauts aboard STS-51-F drank carbonated beverages from specially designed cans from Cola Wars competitors Coca-Cola and Pepsi . [ 12 ] According to Acton, after Coke developed its experimental dispenser for an earlier shuttle flight, Pepsi insisted to American president Ronald Reagan that Coke should not be the first cola in space. The experiment was delayed until Pepsi could develop its own system, and the two companies' products were assigned to STS-51-F. [ 13 ]

Blue Team tested Coke, and Red Team tested Pepsi. As part of the experiment, each team was photographed with the cola logo. Acton said that while the sophisticated Coke system "dispensed soda kind of like what we're used to drinking on Earth", the Pepsi can was a shaving cream can with the Pepsi logo on a paper wrapper, which "dispensed soda filled with bubbles" that was "not very drinkable". [ 13 ] Acton said that when he gives speeches in schools, audiences are much more interested in hearing about the cola experiment than in solar physics . [ 13 ] Post-flight, the astronauts revealed that they preferred Tang , in part because it could be mixed on-orbit with existing chilled-water supplies, whereas there was no dedicated refrigeration equipment on board to chill the cans, which also fizzed excessively in microgravity .

The Plasma Diagnostics Package (PDP), which had been previously flown on STS-3 , made its return on the mission, and was part of a set of plasma physics experiments designed to study the Earth's ionosphere . During the third day of the mission, it was grappled out of the payload bay by the Remote Manipulator System ( Canadarm ) and released for six hours. [ 14 ] During this time, Challenger maneuvered around the PDP as part of a targeted proximity operations exercise. The PDP was successfully grappled by the Canadarm and returned to the payload bay at the beginning of the fourth day of the mission. [ 14 ]

In an experiment during the mission, thruster rockets were fired at a point over Tasmania and also above Boston to create two "holes" – plasma depletion regions – in the ionosphere. A worldwide group of geophysicists collaborated with the observations made from Spacelab 2. [ 15 ]

An eggshell and the bone of a baby Maiasaura (a hadrosaurid dinosaur from Cretaceous North America ), were brought along on the mission by Acton. They became the first dinosaur fossils to have ever been brought into space. [ 16 ]

Challenger landed at Edwards Air Force Base , California , on August 6, 1985, at 12:45:26 p.m. PDT . Its rollout distance was 2,612 m (8,570 ft) . The mission had been extended by 17 orbits for additional payload activities due to the Abort to Orbit. The orbiter arrived back at Kennedy Space Center on August 11, 1985.

The mission insignia was designed by Houston, Texas , artist Skip Bradley. Space Shuttle Challenger is depicted ascending toward the heavens in search of new knowledge in the field of solar and stellar astronomy, with its Spacelab 2 payload. The constellations Leo and Orion are shown in the positions they were in relative to the Sun during the flight. The nineteen stars indicate that the mission is the 19th shuttle flight.

One of the purposes of the mission was to test how suitable the Shuttle was for conducting infrared observations, and the IRT was operated on this mission. [ 17 ] However, the orbiter was found to have some drawbacks for infrared astronomy, and this led to later infrared telescopes being free-flying from the Shuttle orbiter. [ 17 ]

  1. ↑ "9 Weird Things That Flew on NASA's Space Shuttles - Final Shuttle Missions and NASA's Space Shuttle Souvenirs - NASA Shuttle Program" . Space.com. July 7, 2011 . Retrieved February 5, 2022 .
  2. 1 2 "Archived copy of Infrared Astronomy From Earth Orbit" . Archived from the original on December 21, 2016 . Retrieved December 10, 2016 .
  3. 1 2 3 4 5 Kent, S. M.; Dame, T. M.; Fazio, G. (September 1, 1991). "Galactic Structure from the Spacelab Infrared Telescope. II. Luminosity Models of the Milky Way" . The Astrophysical Journal . 378 : 131. Bibcode : 1992ApJS...78..403K . doi : 10.1086/170413 . ISSN 0004-637X .
  4. 1 2 3 Public Domain This article incorporates text from this source, which is in the public domain : "Space Shuttle Mission STS-51F Press Kit" (PDF) . NASA. 1985. Archived (PDF) from the original on July 18, 2024 . Retrieved March 1, 2014 .
  5. ↑ "STS-51F" . Spacefacts . Retrieved February 26, 2014 .
  6. ↑ Travis, Matthew (October 16, 2012). NASA STS-51F space shuttle launch, SSME shutdown and mid-ascent Abort To Orbit (ATO) - July 29, 1985 . Retrieved October 2, 2022 – via YouTube.
  7. ↑ Public Domain This article incorporates text from this source, which is in the public domain : Welch, Brian (August 9, 1985). "Limits to inhibit" (PDF) . Space News Roundup . Houston, Texas: NASA Lyndon B. Johnson Space Center. pp. 1, 3. Archived from the original (PDF) on March 22, 2009 . Retrieved January 10, 2010 .
  8. 1 2 Public Domain This article incorporates text from this source, which is in the public domain : Legler, Robert D.; Bennett, Floyd V (September 1, 2011). "Space Shuttle Missions Summary" (PDF) . NASA Scientific and Technical Information Program Office. Archived (PDF) from the original on October 21, 2020.
  9. ↑ Public Domain This article incorporates text from this source, which is in the public domain : "STS-51F Launch attempt #1" . NASA. Archived from the original on April 27, 2020.
  10. ↑ "Radio Coverage of STS-51F launch attempt 1" . AP. Archived from the original on November 23, 2021.
  11. ↑ "STATUS OF THE GAIA STRAYLIGHT ANALYSIS AND MITIGATION ACTIONS" . ESA. December 17, 2014 . Retrieved February 5, 2022 .
  12. ↑ Pearlman, Robert (May 31, 2001). "A Brief History of Space Marketing" . Space.com. Archived from the original on February 14, 2009 . Retrieved March 24, 2014 .
  13. 1 2 3 "Loren Acton: The Coke and Pepsi Flight" . Air & Space/Smithsonian . Smithsonian Institution . November 18, 2010. Archived from the original on April 12, 2022.
  14. 1 2 Public Domain This article incorporates text from this source, which is in the public domain : "STS-51F National Space Transportation System Mission Report" . NASA Lyndon B. Johnson Space Center. September 1985. p. 2 . Retrieved March 1, 2014 .
  15. ↑ "Elizabeth A. Essex-Cohen Ionospheric Physics Papers" . 2007 . Retrieved February 5, 2022 .
  16. ↑ Chure, D. (2009). "dino bones in space – was it a PR thing" . Cleveland Museum of Natural History. Archived from the original on November 8, 2011 . Retrieved November 12, 2011 .
  17. 1 2 "The Space Review: From Skylab to Shuttle to the Smithsonian" . thespacereview.com . October 16, 2017 . Retrieved February 5, 2022 .

Wikimedia Commons logo

Wikimedia Commons has media related to STS-51-F

.

Dutch Computer Museums

Hacker News
aresluna.org
2026-10-02 12:52:39
Comments...
Original Article

Marcin Wichary

16 June 2022 / 49 tweets / 100 photos

Dutch computer museums

This is an archive of a Twitter thread from 2022. I have since deleted my Twitter account.


I visited three different Dutch computer museums last week, and they were so great I wanted to tell you about them a bit more.

1.7K / 354 / 16 Jun 2022 / 21:50

1. The first was HomeComputerMuseum in Helmond, which was sort of a “living room museum” – over 500 computers, mostly from the 1980s onwards, many of them running and available to play or interact with.

252 / 15 / 16 Jun 2022 / 21:51

(I talked to one of the employees and they said they could power up even more of them! But they found mixed success having computers running that people had no relationships or prior experience with.)

150 / 4 / 16 Jun 2022 / 21:51

2. The second one was Bonami SpelComputer Museum in Zwolle. This was more of a “warehouse museum,” but what a warehouse!

147 / 2 / 16 Jun 2022 / 21:52

This time around, *thousands* of computers surrounded me. I might have never seen as many computers under one roof before. The collection also included big iron mainframes, typewriters, calculators, and tons more videogames.

118 / 1 / 16 Jun 2022 / 21:54

The place was a bit messy, and less interactive, although there were some machines ready to play with…

106 / 1 / 16 Jun 2022 / 21:55

…and a separate arcade with dozens of arcade games, including a solid group of late 1980s Atari games, very close to my heart.

122 / 3 / 16 Jun 2022 / 21:55

3. The last place was a literal barn in the middle of nowhere. (Yes, even a small country like the Netherlands has a middle of nowhere.)

In that barn, a certain DEC enthusiast gathered a lot of Digital machines from the 1970s and 1980s.

206 / 21 / 16 Jun 2022 / 21:56

(Many of them aren’t working yet – it’s a project for a few years from now, after retirement – but at least we got to play some Zork.)

119 / 16 Jun 2022 / 21:56

I enjoyed those museums immensely, each one in isolation, but even more so as an unexpected troika. (I visited the second one because it was close to the first one, and I learned about the third one overhearing the conversation during an earlier visit.)

96 / 16 Jun 2022 / 21:57

First of all, it’s really fun that museums can feel so different – one an extension of your apartment, one that felt like that ending scene in Indiana Jones, one a DEC datacenter amidst corn fields.

(Or not corn! I have no idea.)

117 / 4 / 16 Jun 2022 / 21:58

Secondly, I really love seeing computing museums outside of America, because the regional artifacts are sometimes the most interesting and surprising.

88 / 1 / 16 Jun 2022 / 21:58

It’s fun to see keyboards with non-English legends, or creative (but necessary!) attempts to localize some keys.

116 / 3 / 16 Jun 2022 / 21:59

I saw the amazing Aesthedes – a unique, specialized graphic editor from the early 1980s, with one of the biggest keyboards ever.

266 / 50 / 16 Jun 2022 / 22:00

(At some point I counted 514 keys, and that was before realizing in addition to the membrane keys, there was also a regular keyboard hiding under the console.)

100 / 1 / 16 Jun 2022 / 22:00

Why so many keys? Many of the functions you recognize from today’s graphic programs were right here, as separate keys instead of onscreen buttons:

128 / 9 / 16 Jun 2022 / 22:01

What’s more, the museum in Helmond is restoring this computer, and I got to play with it briefly! I managed to make it hang in no time simply by switching to the text tool, but the prospect of using it properly one day and exploring its strange UI was very exciting.

88 / 16 Jun 2022 / 22:02

I also saw the E.T.-looking Dutch Holborn computers for the first time, at two separate museums!

182 / 17 / 16 Jun 2022 / 22:03

And various displays related to the ill-fated Philips CD-i:

88 / 1 / 16 Jun 2022 / 22:04

And here are some more interesting things I’ve seen in the three museums.

43 / 16 Jun 2022 / 22:04

This very 1980s chording keyboard had a pretty weird mnemonic:

124 / 11 / 16 Jun 2022 / 22:05

We actually got to connect it to an old PC and I typed some things on it! (And even saw the unfinished left-handed prototype.)

79 / 4 / 16 Jun 2022 / 22:05

Speaking of chording, I actually have this one-handed data entry device (with three Shifts under three longest fingers) myself, but I have never seen it in such a pristine condition!

90 / 4 / 16 Jun 2022 / 22:06

And speaking of data entry, this keyboard looks so delicious – but also now you know why German ergonomic laws forbid shiny keys owing to reflections from overhead office lights.

97 / 5 / 16 Jun 2022 / 22:06

*Severance theme starts playing*

135 / 5 / 16 Jun 2022 / 22:07

APL remains one of the most interesting-looking programming languages ever made. And this might be my favourite keycap shape.

80 / 2 / 16 Jun 2022 / 22:07

I have heard of many Selectric-based computers (like the last photo), but I have never seen one based on what looks like the keyboard from a Smith Corona typewriter! I *think* that’s what it is?

63 / 16 Jun 2022 / 22:08

Decades before the colourful iMacs (here in all the liveries!), there was a very green and gorgeous Lorenz teletype:

78 / 16 Jun 2022 / 22:08

This is a Korean NES. I don’t know what is the story behind it.

71 / 5 / 16 Jun 2022 / 22:09

This is a clone of Apple II with a very curious label, and Google seems silent on the subject.

66 / 1 / 16 Jun 2022 / 22:09

“It’s more powerful than Apple IIe,” a.k.a. the classic Chinese Education Computer.

63 / 16 Jun 2022 / 22:10

A lot of time in the world of DEC computers was apparently spent explaining the numpad keys.

53 / 16 Jun 2022 / 22:10

Elsewhere, this might be the best rendition of arrows keys I’ve ever seen.

79 / 2 / 16 Jun 2022 / 22:11

I know I’m supposed to hate this keyboard, but I actually love it.

65 / 16 Jun 2022 / 22:11

I know I was supposed to be grossed out by these keyboards, but I actually loved them, too.

50 / 16 Jun 2022 / 22:11

(So did @PixelAmbacht, who actually arranged for the visit in the DEC barn. Thank you!)

55 / 16 Jun 2022 / 22:12

Speaking of bad keyboards, here are some ZX Spectrum clones…

52 / 2 / 16 Jun 2022 / 22:12

But I particularly loved this ZX Spectrum case with all sorts of documentation taped all over it.

50 / 1 / 16 Jun 2022 / 22:13

On the other hand, the glue on this glue-on ZX81 keyboard has seen better decades.

44 / 16 Jun 2022 / 22:13

I am not sure why this battered 800XL was on display, but at least I learned that the special keys on the right used the same switches as regular keys.

41 / 1 / 16 Jun 2022 / 22:13

At least this Commodore 64 was cut through intentionally.

59 / 4 / 16 Jun 2022 / 22:14

I got to experience some classic 1980s creator nostalgia by designing in Print Shop (!), and then Deluxe Paint (!!) on an Amiga (!!!).

64 / 16 Jun 2022 / 22:14

And then I moved straight to the 1990s.

51 / 16 Jun 2022 / 22:15

I loved finding all sorts of random ephemera I barely understood.

54 / 16 Jun 2022 / 22:15

And I have *finally* found a key that straight up admits that (BREAKING) Enter and Return are one and the same!

88 / 4 / 16 Jun 2022 / 22:15

Not to mention an exquisite version of the same key on a really cool Japanese teletype.

53 / 1 / 16 Jun 2022 / 22:16

But no key was more glorious than this one – one I’ve never seen before!

87 / 6 / 16 Jun 2022 / 22:16

I figured out what it does later that day, but to hear that story, you have to subscribe to my newsletter:
Shift Happens newsletter

49 / 1 / 16 Jun 2022 / 22:17

And that’s it! If you interested in more photos, they are here: Flickr: Marcin Wichary

Thanks for reading!

85 / 1 / 16 Jun 2022 / 22:17

GitLab warns of critical RCE vulnerability in AI Gateway service

Bleeping Computer
www.bleepingcomputer.com
2026-10-02 12:20:05
GitLab warned customers today to immediately patch a critical AI Gateway vulnerability that could let attackers run arbitrary commands on vulnerable instances. [...]...
Original Article

GitLab

GitLab warned customers today to immediately patch a critical AI Gateway vulnerability that could let attackers run arbitrary commands on vulnerable instances.

AI Gateway is a service that gives access to AI-native GitLab Duo features. While GitLab operates its own cloud-based AI Gateway instance used by GitLab.com, GitLab Self-Managed, and GitLab Dedicated, users can also deploy their own self-hosted instances on GitLab Self-Managed through GitLab Duo Self-Hosted.

Tracked as CVE-2026-90970 , this security flaw stems from an improper neutralization weakness and can let attackers with basic privileges and Duo Agent Platform access execute arbitrary commands on unpatched instances.

"GitLab has remediated an issue in the GitLab AI Gateway that, under certain conditions, could have allowed an authenticated user with Duo Agent Platform access to escape the prompt template sandbox via a specially crafted flow configuration, leading to arbitrary command execution on the AI Gateway," the company explained in a Friday advisory.

GitLab released versions 19.2.4, 19.3.2, and 19.4.1 to address this vulnerability for Self-Hosted AI Gateway users and said that customers using a GitLab-hosted AI Gateway are already protected and do not need to take action.

"These versions contain a critical security fix for GitLab Self-Hosted AI Gateway, and we strongly recommend that all GitLab Self-Managed customers with GitLab Self-Hosted AI Gateway installations update to one of these versions immediately," it said. "We have conducted targeted outreach to Self-Hosted AI Gateway customers prior to this release post with this guidance."

GitLab added that it reached out to those who host their own AI Gateway before disclosure and urged users to upgrade vulnerable instances as soon as possible.

Last month, GitLab also patched a maximum severity path traversal vulnerability (CVE-2026-85706) in GitLab Community Edition (CE) and Enterprise Edition (EE) that allows unauthenticated attackers to read sensitive data such as credentials and other secrets from vulnerable servers.

One day later, the U.S. Cybersecurity and Infrastructure Security Agency (CISA) added CVE-2026-85706 to its list of actively exploited flaws and gave federal agencies three days to secure their systems as mandated by Binding Operational Directive (BOD) 26-04.

Since November 2021, CISA has tagged five GitLab vulnerabilities abused in the wild, including one exploited by ransomware gangs.

GitLab's DevSecOps platform has over 30 million registered users and is used by over 50% of Fortune 100 companies, including Nvidia, Lockheed Martin, T-Mobile, Goldman Sachs, Airbus, and UBS.

article image

Build your security blueprint for AI-powered attacks

Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.

Save your seat

Behind the Blog: Freaking Out

403 Media
www.404media.co
2026-10-02 12:02:30
This week, we discuss inbox slop, Claude cults, and more....
Original Article

This is Behind the Blog, where we share our behind-the-scenes thoughts about how a few of our top stories of the week came together. This week, we discuss inbox slop, Claude cults, and more.

JOSEPH: This is something I’ve been thinking about for a while. Jason and Emanuel covered iLands, the AI agents that email people offering to do work. My thing isn’t exactly the same, but similar.

For months, I’ve been noticing that many people email me ‘tips’ (I put in quotation marks because they’re sometimes not tips, but more, my Instagram account was banned, please help me) that are CLEARLY written by AI. I can tell this because the AI responses seem to keep using the same or similar words that it believes would apply to journalists, or are how journalists speak, but as a real, human journalist for more than ten years, I know that journalists don’t actually speak like that.

This post is for paid members only

Become a paid member for unlimited ad-free access to articles, bonus podcast content, and more.

Subscribe

Sign up for free access to this post

Free members get access to posts like this one along with an email round-up of our week's stories.

Subscribe

Already have an account? Sign in

Pizza and Pho? That's That Morris Park Magic

hellgate
hellgatenyc.com
2026-10-02 11:57:31
VPho and Pizzeria in the Bronx features an unlikely but delightful pairing....
Original Article

In Morris Park in the Bronx, on an otherwise unremarkable stretch of Williamsbridge Road a few blocks from the 5 train, there's an unusual (maybe even magical?) place that serves extremely satisfying renditions of two very different, but equally appealing dishes: Vietnamese pho and classic New York pizza.

It's called VPho & Pizzeria, and Elizabeth Lee, the cofounder and acting manager, explained how it came to life.

"We took over a pizza parlor about three or four years ago," she told Hell Gate. "It's our first restaurant so we just stuck to pizza in the beginning, but then we decided to spice up the menu by slowly adding in Vietnamese food. We were a little afraid at first because there are not a lot of Vietnamese people in this area, but we thought, let's take a risk! "

Give us your email to read the full story

Sign up now for our free newsletters.

Sign up

Victory! Court Rejects Government Effort to Dismiss Social Media Surveillance Lawsuit

Electronic Frontier Foundation
www.eff.org
2026-10-02 11:54:42
Judge Allows Social Media Surveillance Lawsuit Against Trump Administration to Move ForwardNEW YORK — A lawsuit filed by three labor unions against the Departments of State and Homeland Security for their viewpoint-based surveillance and suppression of protected expression online can move forward, a...
Original Article

NEW YORK — A lawsuit filed by three labor unions against the Departments of State and Homeland Security for their viewpoint-based surveillance and suppression of protected expression online can move forward, a federal judge ruled yesterday.

On October 1, 2026, Judge Alvin K. Hellerstein of the U.S. District Court for the Southern District of New York rejected the government’s motion to dismiss the lawsuit. The case was filed in October 2025 on behalf of the United Automobile Workers (UAW), Communications Workers of America (CWA), and American Federation of Teachers (AFT). The Electronic Frontier Foundation (EFF), Muslim Advocates (MA), and the Media Freedom & Information Access Clinic (MFIA) represent the labor unions.

This decision is a victory: The Court held that claims that the government’s social media surveillance program is harming the unions’ members, as well as hampering the ability of the unions to associate with their members and potential members, can move forward.

The Court ruled that: "This threat of adverse immigration consequences, under a government whose harsh immigration crackdowns has been heavily publicized and reported on, is certainly enough to 'deter a person of ordinary firmness from the exercise of First Amendment rights.' It is objectively reasonable that noncitizens would limit their expression of disfavored viewpoints under the [Challenged Surveillance Program] given the credible threat of adverse immigration action from the Government."

"The freedom of Plaintiffs' members to speak, associate, and appear publicly is not incidental to union work, but rather is the mechanism through which unions recruit, organize, communicate, and bargain," the Court further explained. "A program alleged to silence members and drive them from the unions' rolls therefore strikes at the unions' representational function itself, which is the 'grounds that bring [their] membership together.'"

Since taking power, the Trump administration has created a mass surveillance program to monitor constitutionally protected speech by noncitizens lawfully present in the U.S. Using AI and other automated technologies, the program surveils the social media accounts of visa and green card holders with the goal of identifying and punishing those who express viewpoints the government disfavors. The surveillance program has been paired with a public intimidation campaign—silencing not just noncitizens with immigration status, but also the families, coworkers, and friends with whom their lives are integrated.

In October 2025, UAW, CWA, and AFT sued the Departments of State and Homeland Security, alleging that this viewpoint-based surveillance program violates the First Amendment and the Administrative Procedure Act.

"No one should have to fear government surveillance or retaliation against their immigration status for expressing their views or participating in their union. We're pleased the Court has allowed this challenge to move forward and will continue fighting to protect the rights of everyone to speak, organize, and advocate without fear," said UAW President Shawn Fain .

"This is a victory for working people, for the labor movement, and for our democracy," said CWA President Claude Cummings Jr . "Our very freedom is under attack by the Trump administration's online surveillance program, and today's decision is a critical first step toward affirming our freedom to speak, to protest, to organize without fear of government retaliation. These essential freedoms underpin our union rights to join together and fight to improve our working conditions. CWA is a fighting union, and our members remain ready to stand together to protect our rights and our freedoms."

"Today’s decision is a critical step toward vindicating our Constitutional right to freedom of speech and rejecting the Trump Administration’s cynical attempts to criminalize and punish those who disagree with them," said AFT President Randi Weingarten . "Government surveillance to monitor the 'opposition' is a tool of dictators that erodes the democratic principles this country was founded on. We will continue to remain vigilant in defending our 250-year-old rights—not just for our members, but for all Americans."

"Our plaintiff-unions have members that have wholly changed the way they interact with social media—including limiting their engagement with union content—because of the government's social media surveillance program," said EFF Senior Staff Attorney Lisa Femia . "Many have stopped posting online together, and have even stopped engaging in offline activities, for fear of being scrutinized or targeted related to immigration benefits. We are pleased that the Court has agreed to let the case proceed, and allow unions and their members to seek justice for infringement of their rights."

"Today’s ruling is an important step forward in holding the government accountable for its ever-expansive online surveillance program that silenced non-citizens, stoking fear that exercise of their protected First Amendment rights could result in unfavorable treatment on their immigration applications or worse." said Sadaf Hasan, Staff Attorney at Muslim Advocates . "We will keep fighting until all non-citizens are able to freely associate, organize, and speak out without the looming threat of visa revocation and immigration enforcement simply because the government dislikes their views."

"Defendants' attempt to evade accountability on specious jurisdictional grounds was rightly rejected by the Court," said Nick Jones , a student in the Media Freedom & Information Access Clinic . "We are excited to see the case now proceed to the merits, where we expect to prevail as well.”

For the ruling: https://www.eff.org/document/uaw-v-dos-opinion-order-denying-motion-dismiss

For more about the litigation: https://eff.org/cases/united-auto-workers-v-us-department-state

Contacts:
Electronic Frontier Foundation: press@eff.org
Muslim Advocates: melissa@muslimadvocates.org

Supabase Is Acquiring Turso

Hacker News
supabase.com
2026-10-02 11:43:03
Comments...
Original Article

AI is enabling builders to create an immense amount of software. Today, agents are spinning up millions of databases to power the prototypes, explorations, dashboards, and apps they're building. This new pattern, at this massive scale, requires an evolution in database infrastructure. Turso is joining Supabase to build this evolution.

For existing users, nothing changes. Supabase will continue building around Postgres, while Turso will continue its work on SQLite. And together, we’re creating something new that will bring the Supabase experience to every agent.

Supabase is already launching over one million databases per week. As agents build more software, we believe database demand will outpace the world’s current capacity to support them. So we need infrastructure that can scale to meet this growth and is suited to how we build with agents.

Agents should be able to create a database as easily as creating a file, with just as little concern about cost. For smaller workloads, that shouldn’t require provisioning a dedicated machine every time. Databases should be cheap to create, available on demand, and have a clear path to production when needed.

SQLite is well suited to these small, on-demand workloads. Postgres is what you want as your application scales. We want builders to have the same developer experience from prototyping to production.

Turso has built an architecture for exactly this pattern. They rebuilt SQLite in Rust and created a cloud platform where a single server can manage millions of databases, loading them when needed and suspending them when they’re not.

This architecture makes it possible to provision a database for each agent on demand, whether on Turso Cloud or in customers’ own clouds, as Superhuman, Sauna.ai, CTO.new, and Mastra do.

We share a vision for how infrastructure needs to evolve to support agentic AI. Turso will continue operating, with a clear path into the broader Supabase ecosystem as workloads grow.

Turso and Supabase have a lot in common: a commitment to open source, a focus on developer experience, and an appetite to tackle hard infrastructure problems.

We’re excited to welcome Glauber Costa and Pekka Enberg to Supabase with the rest of the Turso team. Glauber will lead this agentic infrastructure effort.

Our mission is to store the world’s data, and agents are going to create a lot more of it. This gets us closer to that goal.

A 20-year-long permanent cookie: America.gov and tracking

Hacker News
www.biometricupdate.com
2026-10-02 11:30:24
Comments...
Original Article
Timed out getting readerview for https://www.biometricupdate.com/202610/america-gov-launches-with-privacy-pledge-as-login-gov-code-raises-tracking-questions

One month coding with GLM 5.3 Flash

Hacker News
wagtail.org
2026-10-02 11:29:15
Comments...
Original Article

Screenshot of AgentsView tokens usage for September 2026 over 2B tokens

Zooming in on the models split specifically:

Tree map of agents token usage over September 2026, with half going to GLM 5.3 Flash

The goal was to spend the whole month on GLM 5.3 Flash pictured in teal. Here’s what went well:

  • Successfully spent the first half of the month on just that model.
  • That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).

The second half of the month didn’t go so well, with 1B tokens going to other models.

Unexpected hurdles

The cost of vibe coding

We’re pretty transparent that our experimental Wagtail MCP server is a vibe-coded prototype . Vibe coding isn’t quite what we normally aspire to, but for a prototype it’s spot on. Unfortunately there are still consequences to it. I chose the 'wrong' model for the prototype, and we spent 450M tokens / $150 / 5kWh of energy use almost overnight. The MCP server itself works well and we now have a great demo of the capabilities, so it’s not for nothing:

Nonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.

Infrastructure woes

Another unexpected hurdle was infrastructure availability issues. We’ve written extensively about comparing inference providers . Our choices work really most of the times, but it turns out they’re very popular, and do not have the same capacity as the big labs who hoard all the GPUs. We noted degradation with the performance of GLM 5.3 Flash in particular, most likely because of it being so high up the Pareto frontier of relevant models for our work.

Scatter plot of AI models, with the drawn pareto frontier, GLM 5.3 Flash in the top left

This meant having to switch to other similar models (DeepSeek V4.1 Flash, Qwen 3.8 Flash). Which is very simple to do, but nonetheless unexpected!

The cost of experimentation and R&D

Last but not least, beyond using one model for day-to-day engineering, it felt essential to keep experimenting with a wide range of models, keeping up with what providers are releasing. This is particularly essential as we start to benchmark models’ performance on Wagtail tasks, where we need data across a wide range of models. Sneak peek of our benchmark:

Data table of 14 AI models reporting their accuracy as a percentage, energy use in Wh, Cost in $. Top of the table is DeepSeek V4.1 Flash with 95% accuracy and 14.9Wh energy use, $0.09 per task

It’s much easier to guide people towards leaner options with this kind of concrete data. And for us to make those options even more viable with agent skills, or our new CLI prototype , which is intended to work well with agents.

Takeways and what to do next

So technically this challenge was a failure. Only 50% usage on the target model, 1B out of 2B tokens. About 35 kWh of energy use instead of 10. But we did learn a lot, which is crucial for the current moment. Reflecting on this for October, here’s what will make it work:

  1. Constant, local usage measurement and reporting . Looking not just at tokens but also energy use and spend, and ideally how well this all leads to concrete positive outcomes.
  2. Budgeting for experimentation, not just day-to-day tasks . Making more concerted decisions about which prototypes are worth building, and how.
  3. Better prompt selection and multi-agent techniques . Orchestrator vs. scout vs. implementer vs. reviewer agents. Bounded goals. Not rocket science but certainly one more thing to learn.
  4. Keep pushing for more efficient techniques and models . The Jev-style decision diffusion models look very promising if they can run so efficiently. Latest flagship models also look like a step in the right direction on that front.

For day-to-day developer work, it’s totally viable to focus on one or two flash-tier cheap models. A viable target is probably that the majority of AI inference work should be done with such efficient models, measured in cost or energy use rather than meaningless tokens. That’s the goal for October! You should try it too, you’ll learn a lot in the process.


And come say hi at Wagtail Space 2026 in November to hear how that all pans out!

You might also like

Power approval set to delay Oracle's Wisconsin AI datacenter

Hacker News
www.theregister.com
2026-10-02 11:24:54
Comments...
Original Article

on-prem

Grid connection awaits regulatory approval, putting planned 2027 customer delivery at risk

Oracle's Wisconsin AI datacenter could miss its planned 2027 start date because the transmission infrastructure needed to power it is awaiting regulatory approval.

Research company Aterio warns that Project Lighthouse, being developed by Vantage in Port Washington with Oracle as tenant, faces a real risk of slipping beyond Big Red's target of delivering capacity to customers in the second half of next year.

Separately, Oracle reportedly sent a force majeure notice to a Blue Owl Capital subsidiary developing Project Jupiter in New Mexico. Bloomberg reported that Oracle was seeking to defer payments if the campus failed to come online as planned, rather than withdraw from the project. Oracle says the development remains on schedule.

Aterio said of this and the Wisconsin facility: "The regulatory record points to the potential for dates later than company guidance, and Oracle's stretched finances, together with its force majeure notice on Project Jupiter in New Mexico this week, raise the potential for slippage on its largest campuses."

Oracle has been offered the opportunity to comment.

The Register has already reported that Oracle could face more than $100 million a year in financing costs to guarantee the power commitments behind the Wisconsin datacenter campus, which is rated at 902 MW of computing load and 1.3 GW of total electricity demand.

The immediate obstacle is American Transmission Company's (ATC) application to the Public Service Commission of Wisconsin (PSC) for approval to build the campus's grid connection.

ATC is planning more than $2 billion in transmission infrastructure projects tied to datacenter developments across Wisconsin, ranging from high-voltage power lines and substations to upgrades to existing facilities.

Construction of the transmission infrastructure requires PSC approval.

"The problem is power. The campus cannot run until ATC connects it to the grid, and ATC cannot start building until the PSC approves its application," Aterio said.

The PSC closed the original case after withdrawing its finding that ATC's application was complete. ATC refiled in September, restarting the approval process.

Aterio's earliest scenario puts partial power in October 2027 and full supply in August 2028, although it considers that timetable highly unlikely.

Its base case puts partial power in December 2027 and full supply in October 2028. An extended regulatory review could push those dates to June 2028 and April 2029.

Even December 2027 power delivery would leave little room for testing and commissioning before Oracle's deadline.

"In practice, a datacenter does not enter service as soon as first grid power arrives. Each building's electrical and cooling systems have to be energized, tested and commissioned before it can carry customer load. That will take time. We therefore expect the campus's load to ramp up mainly from the first quarter of 2028," Aterio said.

Meanwhile, the Financial Times reported that Tencent had signed a $7 billion deal to lease around 100,000 AI processors from Oracle.

For context, Oracle reported $55.7 billion in total capital expenditures for its fiscal 2026, which closed at the end of May. Its total revenue was $67.4 billion. Capex guidance for the 2027 financial year is up to $95 billion. ®

America's Largest Producer of Coal Is Operating Under an Expired Permit

Hacker News
insideclimatenews.org
2026-10-02 11:21:42
Comments...

US sanctions Tren de Aragua gang members in ATM hacks crackdown

Bleeping Computer
www.bleepingcomputer.com
2026-10-02 11:20:53
The U.S. Treasury Department has sanctioned eight members of the Venezuelan gang Tren de Aragua (TdA) for their role in the theft of millions of dollars in ATM jackpotting attacks across the United States. [...]...
Original Article

Treasury Department

The U.S. Treasury Department has sanctioned eight members of the Venezuelan gang Tren de Aragua (TdA) for their role in the theft of millions of dollars in ATM jackpotting attacks across the United States.

In jackpotting attacks, criminals deploy malware (e.g., ATMii , ATMitch , GreenDispenser , Alice , RIPPER , Skimer , SUCEFUL , and Ploutus ) on bank and credit union automated teller machines (ATMs), which allows them to empty the ATMs of cash and delete evidence via an attached USB keyboard or the built-in PIN pad.

The list of designated individuals includes Anibal Alexander Canelon Aguirre (known as "Prometheus") and six of his associates: Eric Gabriel Cardenas Arzola, Jose Dario Galeano Bazurto, Anthony Wuiliam Hernandez Guerrero, Carlos Javier Martinez Armenta, Oscar Leonardo Martinez Pirona, and Alejandro Mejia Castillo.

Aguirre has allegedly developed the Ploutus malware used in ATM hacks and has been on the FBI's list of Ten Most Wanted Fugitives since March.

"While Prometheus's network is based in Mexico and Venezuela, the scheme targets U.S.-based automated teller machines (ATMs), where criminals deploy malicious software (malware) to force ATMs to dispense cash. The stolen funds are then laundered and transferred to TdA members in various countries," the Office of Foreign Assets Control (OFAC) said .

"These actions form part of a sustained, whole-of-government campaign that has resulted in over 30 actions against more than 300 individuals and entities tied to transnational criminal organizations since 2025."

The Treasury designated TdA as a Transnational Criminal Organization in July 2024, and the Department of State designated it as a Foreign Terrorist Organization in February 2025.

Suspects installing malware in ATM jackpotting attacks
Suspects installing malware in ATM jackpotting attacks (Treasury Department)

​According to TRM Labs, the Treasury also added seven TRON addresses to its Specially Designated Nationals and Blocked Persons List (SDN List), which "have received approximately USD 6.1 million in total inflows since March 2022 [..] and sent funds to other TdA-associated addresses."

As of August 2025, TdA gang members have stolen $40.73 million from U.S. financial institutions across over 1,500 alleged ATM jackpotting attacks, according to OFAC's estimates.

Since October 2025, the U.S. Justice Department has charged 98 suspects linked to the TdA gang and involved in ATM jackpotting schemes, who are now facing maximum prison terms ranging from 20 to 335 years each.

After a wave of arrests targeting members of the Tren de Aragua Venezuelan criminal organization, the FBI warned in February that criminals had stolen over $20 million in 2025 in a massive surge of ATM hacking incidents.

Most recently, in early September, five Venezuelan nationals pleaded guilty after failing to install malware in ATM jackpotting attempts in Wamego and Manhattan, Kansas.

article image

Build your security blueprint for AI-powered attacks

Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.

Save your seat

The Four Horsemen of Agentic Coding

Hacker News
distantprovince.substack.com
2026-10-02 11:19:56
Comments...
Original Article

Agentic coding is undeniably very useful. It’s also very bad. Or rather, it’s having some horrible effects on us, our craft, and our relationships with each other. I can feel it in my bones, and a lot of others can too. But whenever I try to explain what exactly is wrong, I find myself waving my hands wildly and jumping from point to point. What do you mean, what’s wrong? So many things are wrong!

Fine, here’s a list. 4 problems with no solution in sight–or, as I call them, the four horsemen of agentic coding.

LLM-generated code has a strong smell that makes codebases repulsive to humans.

LLMs aren’t humans, and the way they write code is... different. Not strictly worse, but different for sure. We quickly came up with a name for it—slop—and it undoubtedly carries a negative connotation.

For plain prose, humanity seems to be converging on the opinion that AI prose is bland at best and an insult at worst. For code, due to its instrumental nature, the debate is far from over. A lot of people latch on to the idea that very soon we won’t have to read the code at all. “Remember, it’s the worst the models will ever be,” they say.

And yet, the latest generation of models is surprisingly underwhelming. They’re certainly smart, but they seem to be moving towards “unemployable smart” rather than “inspiring smart.” Claude now famously communicates entirely through word salad, and Astra writes in a bizarre competitive code-golfy style , incomprehensible to normal humans. Sloppiness turned out to be a surprisingly persistent property of LLM output and at this point looks like a signature move rather than a growing pain.

Which is very bad news for people who still want to poke around in the codebase and maintain some familiarity with it, especially in a team setting. Once agents are allowed, they very quickly take over. What used to be a shared space for humans where each team member contributed became an AI wasteland where people don’t want to spend time.

But what if we don’t have to spend time in these slop caves? Why spend your time there when you can be a design mastermind and reign over a swarm of loyal agents from the comfort of the chat interface?

Engineers become alienated from code and, as a result, care less.

Software engineering used to be a fairly... kinesthetic field of work. Maybe not to the same degree as woodworking, but we loved our tools too, both real and virtual. Text editors had a cult-like following. Popular color themes and programming fonts had their own fanbases. Split keyboards were a hot thing once, and there was no shortage of keyboard sound testing videos.

We produced code directly, by hand. It’s only natural that this resulted in a strong sense of ownership of the fruits of our labor. Sure, this feeling varied depending on the context, but even if you just assembled something from a set of building blocks, you would feel something . IKEA certainly knows that !

Agentic coding has dramatically increased the distance between engineers and the product of their work. This gap is so large that none of our kinesthetic tools can cross it. We don’t experience code directly anymore: we send an agent with a vague set of instructions and then skim through its report. Hell, even typing is optional—you can just yap at your phone while driving.

So naturally, we become more detached. We care less. We lose touch with our craft. And it shows. It shows in poorly thought-out features that receive no pushback. It shows in band-aid bugfixes that don’t address the root cause. It shows in how unfulfilled we feel at the end of the day. For software to spark joy, it has to be imbued with its creator’s love.

Elsie Driggs, “Pittsburgh”

AI erodes existing skills, impedes learning, and doesn’t offer any meaningful skill progression.

I think at this point it shouldn’t be controversial to say that deskilling is real. There is no shortage of people reporting their wits slipping away after prolonged use of AI. Which makes sense: a lot of things in life are use-it-or-lose-it, and professional skills are no exception. Yes, you’ll always know how to ride a bike, and I myself will die before I forget how to exit Vim, but anything beyond that is fair game.

I’ve made the point before that we’re currently experiencing a historical anomaly. We have all those seasoned engineers who spent years doing things the hard way and can now hold LLMs exactly the right way to be insanely productive. But unless we figure out how learning works in the age of AI, this supply won’t last long.

One common piece of advice is that if you’re at the beginning of your career, you simply shouldn’t use AI. And I completely agree, but it’s important to understand that this isn’t a viable strategy for humanity as a whole. Everyone will not just .

We can’t be blind to the fact that the introduction of an “easy button” has led to a tectonic shift in incentives to learn. And if I know one thing about humans, it’s that they follow incentives.

Another common response to the deskilling argument is that LLM usage comes with its own learning curve. You just need to upskill and learn the intricacies of downloading markdown files! I’m not buying it. The learning curve is basically non-existent.

A printing press or a chainsaw at least requires some skill to operate, but the fancy new magic box? You don’t have to know how to read or write. A caveman could build a B2B SaaS and be on time for his weekly podcast appearance. And if you think your custom agent orchestration workflow is an important feat of engineering, just you wait.

Agents weaken the social fabric of engineering teams by reducing interpersonal communication, increasing self-reliance, and hindering skill expression.

Slack channels used to be crowded with people. Anything that your rubber duck couldn’t answer would go into a team chat. Now the chat is a pale shadow of its former self. Everyone is too busy talking to their agents. Each of us has received a magic familiar that follows us around like we’re some kind of Disney protagonists. It’s very smart, available 24/7, and will never report you to HR. The rubber duck can finally talk back!

In a certain sense, being more autonomous is virtuous. Why distract colleagues with your silly nonsense? The problem is, helping other people with silly nonsense is our thing . It’s how we build and maintain relationships. You need to walk the path; otherwise, it will be overtaken by grass.

I’m convinced that AI has dramatically reduced interpersonal communication. The average team chat right now consists of agent threads and an occasional LLM-generated manuscript followed by a humble “what do you guys think?”

Agentic coding also changes the way we see other people. Teams used to be sort of like parties of adventurers: a Git wizard, a Rust witch, a bard who is really into mechanical keyboards. There was plenty of room for skill expression, and the diversity of expertise was celebrated.

Now, an average team consists of a Claude operator, a Codex whisperer, and a guy with AI psychosis who insists on committing .md files. “This is John—you’ll want to see how he prompts ,” said nobody ever.

And that’s true: nobody wants to see how John prompts. There is something repulsive and fundamentally unsexy about talking to an agent. It feels fine when you are doing it, but others... ugh, get a chat room. Their prompts are cringe; their agents are weird. Like looking at someone else’s TikTok feed.

So here we are, still peers in the org chart, but growing more and more distant over time. I used to praise your Vim skills and Advent of Code solutions. These days you prompt Claude all day, and I do too. Are we even a team anymore?

So that’s my list. And I’ll be honest with you, I have no idea what to do with any of this. I can only hope that eventually we’ll move forward in this cursed discussion. However beneficial AI is, we pay dearly for it with our curiosity, craftsmanship, and social connection.

Discussion about this post

Ready for more?

Mozilla shutting down Solo AI website creator

Hacker News
support.soloist.ai
2026-10-02 11:19:22
Comments...
Original Article

We made the difficult decision to shut down Solo. Rest assured, we are here to help and to make the closure, and the migration of your existing site to another hosting provider, as seamless as possible.

How long do I have?

All Solo websites and data will be deleted on November 30, 2026 . There will be no way to recover your site, account, or data after that date.

How do I download my website data?

  1. Log in to Solo .
  2. Click the menu at the top right, and select Account Settings .
  3. In the Account Settings popup, click Download Website Data .

You'll receive a ZIP file containing an HTML export of your website and a CSV file with links to your image assets, as well as your contact form and newsletter submissions.

Image source files are not included in the ZIP. Download each image individually from the links in the CSV file before the shutdown date on November 30, 2026 .

What are some alternative hosting providers?

Wix , Squarespace , and WordPress are popular website hosts. You can also consider AI website builders such as Bolt or Lovable .

What if I have a Solo Pro subscription?

If you have an active Solo Pro or Grow subscription, we will issue a prorated refund for the unused portion of your current billing period, calculated as of October 1, 2026 . The refund goes back to your original payment method, and you don't need to do anything to receive it. Your subscription will not renew.

What if I purchased a Solo domain?

If you purchased a domain through Solo Domains (not IONOS), you will need to migrate it to Name.com , our domain partner. To do so:

  1. Create a Name.com account . If you already have one, log in.
  2. Verify the email address on your Name.com account. Name.com requires a verified account before it can accept the transfer. See Name.com's guide to account and domain verification .
  3. Locate your unique Account Code , displayed in the top-right corner of the Name.com dashboard.
  4. Email the Account Code to support@soloist.ai . We will transfer your Solo Domains account to Name.com and email you when it's complete.
  5. Add your billing information to Name.com. If you don't, Name.com will remind you as your domain renewal date approaches.

Where can I get more assistance?

If you have any further questions, please don't hesitate to contact us at support@soloist.ai .

AI Makes Me Sad

Hacker News
mondobe.com
2026-10-02 11:18:00
Comments...
Original Article

AI Makes Me Sad

October 2nd, 2026

Am I allowed to say that? Because it does.

Maybe it's because I'm writing this at 2:00 AM, which is about the only time I let myself experience sadness for a prolonged period of time, but every time I see a post on Twitter or LinkedIn about a new development in AI, I feel a sense of hopelessness.

And I'm saying this as someone who is absolutely in a position to benefit from AI. If I added a few vibecoded projects to my resume, the gaping maw of every startup, enterprise company, and FAANG behemoth would open to me. Just the other day, I sat near a table of my fellow TAs as they discussed their plans to work at OpenAI after graduation, eyeing $500k salaries.

But I don't want to do that. I don't want to prompt Codex for a living, or Claude for that matter. Maybe it's because I don't even really know what that looks like. As sure as I know anything, I know what it sounds like. It sounds like a loss of craft, a loss of control, a loss of creativity. It sounds like Charlie Chaplin in Modern Times , the scene where he's so enthusiastic about screwing hex nuts in a factory that he screws the buttons on people's shirts, and then he goes and gets himself caught in the conveyor belt, snaking through the gears and screwing in the hex nuts inside the machine. It sounds like people around me are throwing themselves at the machine, volunteering to make themselves obsolete for $500k a year. If you genuinely, genuinely , believed artificial superintelligence was coming, why would you want to work at the company that would be in the best position to replace its employees?

It makes me sad that I missed the boat on startups. I think I could have been a pretty formidable startup founder had I been born five or ten years earlier. I'm good at web stuff, I'm good at Rust, I can even do OS stuff. I watched all the YC Stanford lectures this summer while translating tests from one language to another as an intern. Now there's no such thing as a startup, there are only "AI startups", where you build a company that somebody else can write a prompt for in 30 seconds. And if they can't write a prompt for it now, they'll be able to next year with GPT-7, or whenever AI hardware fabrication gets better.

It makes me sad that AI evangelists are so quick to adopt the language of online incels, and that AI-generated art and music have become the defining aesthetic of fascism in the 2020s while AI-generated writing went from risible to a mark of stupidity to something reluctantly accepted since it sits at the fringes of everything useful on the Internet like a kind of digital mold.

The overwhelming weight of it all feels so stifling. It feels like there's no way to stop it. I can't imagine a Twitter (sorry, an X) where every post isn't filled with verified bots spewing LLM nonsense in the replies. I can't imagine an article about some random topic on some random webpage, designed to maximize SEO for some banner ad I won't see because I block ads, being written by a human. We would have to rebuild all our institutions.

The part where I'm not (that) sad anymore

Honestly, though, it doesn't sound that hard to rebuild all our institutions. Twitter is only 20 years old, LinkedIn is only 23. Most of these were in need of a refresh anyway. I don't know much about the Dark Ages, but I know that Muslims and Irish monks kept the candle of knowledge burning while everyone else ran around crusading. In a similar way, I know that there are a lot of people who feel a lot of the same sadness that I do, even if they don't express it as sadness. I like to think even a lot of AI evangelists feel at least a little sad, like Paul Graham, because he's a Lisp man from way back, and programming language fanatics usually care about humans writing programs.

Another point of hope: the technology industry moves at unfathomable speeds, so we'll start to see a way out of our "dark age" within only a few years as our institutions start to mature and improve. In fact, I suspect that AI will have a fixed role in our society before the end of my natural life. Whether that role is a small one or a major one, I'm not sure. But I know I will feel at least a little happier to have some finality about the whole thing.

I think this will require some painful shifts in our industry, but I also think that this will be a satisfying kind of pain, at least for people like me, like the burn you feel when working out. It feels extremely strange to be looking forward to a bubble bursting. It makes me feel like the AI evangelists who look forward to artists filling out McDonald's applications. But I do. It makes me very hopeful. When I imagine that scenario, I can conjure no emotion besides glee. I desperately want to work in an industry that makes sense, even if a lot of people right now would rather work in an industry that sells nothing to nobody for untold sums of money.

I'm sorry. I want to go back to the Twitter thing because I use Twitter a lot. And, seriously, there are so many verified accounts that just post AI-generated replies to popular posts, and we all just ignore them. It's the most insane, Kafkaesque thing. Absolutely nobody likes it. It's an eyesore. And we're supposed to be living in the most powerful age for AI. You're telling me that nobody at SpaceX/Xai/Grok/SuperGrok/whatever can write a classifier to ban these accounts? I can press a button on any post and be given a full explanation by Grok, a button which I've pressed 2 dozen times on accident and never on purpose. You can't run a decision model for $0.0000001 cents whenever someone hits "Post" and realize that a reply makes no sense?

Recently, I've seen a lot of people point out the fact that billions (trillions?) in AI spend has translated to absolutely no benefit for most consumers of everyday apps. Obviously, people who use ChatGPT get the benefit of having their question answered. But in the world where development is getting faster and cheaper, I can't think of a single piece of software I use that has gotten better. Everything is exactly the same, just with a button to ask a chatbot. Somebody has to notice this eventually.

I want an industry that makes sense. Right now, we have (most of) my undergrad classmates' dream industry. They drink constantly [1] , skip lectures, scrape by on their homework with AI, complain about classes, grind Leetcode to get internships, and eventually get their diploma. It must feel like walking around in a literal dream, where your footsteps have no weight, like Wile E. Coyote near a ledge. They commiserate with their friends, "That class is so hard. I just asked Chat." I know because I hear them two seats away from me. Now they get to work a job where they commiserate to an coding agent, "this task is so hard. Do this for me." How come they get to look forward to relief while I feel dread? And that just makes me feel selfish. They're only human too. They is how they've been told to navigate life. For some reason, somebody forgot to tell me. I want an industry that makes sense.

I guess I've been talking a lot about sense. It's 4 AM now, so I guess I'm not making a lot of sense myself. If I wake up tomorrow and the president has signed an injunction on all AI inference, including local models, would I be sadder than I was going to sleep? I don't know. Just ignore me. I'm the homeless guy. I have no home in this industry. Codex, make me a home by the time I wake up.


[1] Seriously, if you ever thought Gen Z doesn't drink, look no further than Georgia Tech. I'm in multiple group chats and Discord servers where the subject discussed most often is the date, location, and intensity of the next planned drinking session.

GitHub repository landing pages now show an accessibility tab, if provided

Lobsters
ericwbailey.website
2026-10-02 11:15:11
Comments...
Original Article

My last official contribution to GitHub was something I’ve wanted for a long time: Writing the code to enable displaying an ACCESSIBILITY.md file’s contents on the repository landing page . This content lives in the same tab component that the README, License, Code of Conduct, Security, and other important information is surfaced.

For example, I’m using an ACESSIBILITY.md file located in ./github to communicate the accessibility statement on my a11y-webring.club repository . Here’s an image of it in action:

A GitHub repository landing page, scrolled down to the tab list of special repository files. The tab labeled 'Accessibility' is active, displaying an accessibility statement. Other tabs are Readme, Code of Conduct, Contributing, MIT License, and Security. Above the tab strip are some repository files. The visible content of the accessibility statement reads, 'a11y-webring.club strives to be AA WCAG 2.2 compliant, and is committed to creating and maintaining an accessible, inclusive environment. It is intended to be able to be used by everyone. What we are doing: The following initiatives are how this webring attempts to be accessible. We are: Guided by a Code of Conduct that outlines expected behaviors. Using time-tested, stable and interoperable technology based on open standards to help ensure our content can be accessed by the widest range of devices as possible. Running automated and manual checks to test for accessibility issues. Hosting our code on a public repository, allowing anyone with the interest and capability to inspect and modify it. Striving to keep our interactions and user interface unambiguous and easy to understand. Striving to keep our download size small and memory footprint light. Supporting magnified and zoomed displays, as well as custom typefaces and themes potentially set by someone in their browser.' Cropped screenshot.

The file only appears if it is supplied, so it is not a required part of creating or maintaining a repository. However, it is my hope that the act of providing one becomes more commonplace the same way providing those other special kinds of files are.

Another broader hope I have is that this promotion of content helps to normalize accessibility as a consideration and practice in some small way—something that helps to send a signal of a mature software project.

I’m excited to see how people use this new addition to the platform. I would also like to extend a huge thank you to Jan Maarten and Maria Lamardo for their help getting this effort across the finish line.

How accurately calibrated is Jev?

Hacker News
maximumeffort.substack.com
2026-10-02 11:12:41
Comments...
Original Article

Jev is TypeSafe’s new “System One” classifier model . The name is inspired by Daniel Kahneman’s book Thinking, Fast and Slow , in which he distinguishes between fast, instinctive, System One thinking and slower, conscious, System Two thinking. 1

In this analogy, Large Language Models (LLMs) like Claude or GPT are System Two models and “Decision Models” like Jev and its predecessors ( e.g. Laya ) are System One.

The transformer architecture is the backbone of modern LLMs, and it is an extremely flexible, general paradigm for learning most tasks. However, modern LLMs are irreducibly stochastic, autoregressive, and their output style (freeform text) is not well suited to classification tasks. For example, let’s say I made a call to an LLM, something like claude(“2 + 2 = ?”) ? We expect 4 , but as a string, an integer, a float…? With Jev, you would instead call something like jev("2+2=?", "3 : int, 4 : int, 5 : int") , and you’d receive 4 , correctly typed as an integer.

Jev, in my understanding, takes a pretrained transformer in all its generality and bolts a classifier on the end of it. This way, the classifier doesn’t have to be trained on any specific task and can use new context immediately , but it still acts as a classifier. You give it context and a multiple-choice question, and it gives you a probability distribution over the multiple choices. 2 It’s also extremely cheap, the entire below series of experiments cost less than $4.00.

However, Jev doesn’t just pick an answer, it gives you a probability distribution over the possible answers. How accurate is that probability distribution?

I decided to check this on questions where the answer is well understood. For example:

A classical particle of mass m is embedded in a system at thermodynamic equilibrium with temperature T . What is its velocity v?

The answer is a probability distribution over v, and specifically, the Maxwell-Boltzmann distribution :

Wikipedia: “For a system containing a large number of identical non-interacting, non-relativistic classical particles in thermodynamic equilibrium, the fraction of the particles within an infinitesimal element of the three-dimensional velocity space d 3 v , centered on a velocity vector v with magnitude v , is given by [the above distribution].” m is mass
Also from Wikipedia: “The speed probability density functions of the speeds of a few noble gases at a temperature of 298.15 K (25 °C). The y -axis is in s/m so that the area under any section of the curve (which represents the probability of the speed being in that range) is dimensionless.”

A fun demo here: did you know you can physically generate a Maxwell-Boltzmann distribution with a motor and some balls? Video Here 3

GPT-6 Astra and Claude Opus 5.5 were used for implementing these experiments, writing the templated prompts, API calls, etc. I’ve also been experimenting with Opus 5.5’s ability to make plots, and am very impressed so far.

So, I picked a list of physically relevant distributions, and had GPT-6 Astra and Claude Opus 5.5 write a series of prompt templates, to which the answers should produce probability distributions.

I am not asking Jev for a probability distribution per se . I am asking it for a “choice” over a finite set (binned ranges of a continuous parameter, usually). Jev returns a typed decision with its internal probability for each bin. If Jev is well-calibrated, its output probabilities should match the physically correct probability distribution function.

In total, I chose 10 candidate distributions, 5 prompt templates per distribution, and 20 variations of each prompt (changing, for example, the ambient temperature for each call), which gives 1,000 settings. The answer bins are fixed for each template and do not change between draws.

Here’s an example prompt, with state giving the context, instructions the task, and criteria a set of bins of the continuous parameter over which Jev returns a probability distribution

Some example prompts templates. Prompts + Variations written by Claude Opus 5.5.

Distributions: Gaussian, Lorentzian, Maxwell, Gamma, Exponential, Rayleigh, Uniform, Poisson, Binomial, Boltzmann.

Bins: For each continuous template we set one physical range, wide enough for the widest law among its 20 draws (except for the Lorentzian that has long tails), and divided it into equal-width bins (except for the Lorentzian, where the last bin was open-ended).

  • Signed answers (Gaussian, Lorentzian): 48 bins of width R/24 on [−R, R], plus “below −R” and “R or more”.

  • Nonnegative answers (Maxwell, Gamma, Exponential, Rayleigh): 49 bins of width w starting at 0, plus “49 w or more”.

  • Uniform positions: 50 bins on a range that contains all draws

  • Poisson: the counts 0–48 and “49 or more”

  • Binomial: 0–49, with N = 49 trials

  • Boltzmann: the listed energy levels

Here’s a really lovely figure that Opus 5.5 made showing the method visually.

Methods. A) The basic structure of the prompts to Jev (example). B) Jev’s reported probabilities (example) vs. correct distribution. C) An example of how I’ve chosen to display the difference—a ratio of Jev predictions to theory, by percentile. This rolls quite a bit of Jev’s errors into the boundary bins, hence the very common “spiked edges” structure

Jev’s answers were scored by Total Variation across all possible choices, per draw.

\( \mathrm{TV}(q, p) = \frac{1}{2}\sum_{i=1}^{K} \lvert q_i - p_i \rvert \)

For K bins, q is Jev’s response probability and p the integral of the correct pdf in that bin. TV = 0 is perfect agreement, TV = 1 indicates totally disjoint probability mass.

So how well calibrated is Jev? Not well. If Jev were to completely punt on the answer, and spread the probability evenly over all possible choices, it would score a mean TV of 0.546. But Jev scores 0.518. For Uniform it is 0.77 vs. 0.39 for a flat guess, and for Poisson 0.65 vs. 0.64. This makes me deeply suspicious of methods like JevEval as automated judges of LLM answers (not to pick on this, it’s a good idea, but the distributions are not well calibrated for very well known problems).

It has a very noticeable failure mode.

Straightforward test of calibration. A prompt with a physics question, to which the answer is a draw from a probability distribution (Jev was not given the distribution for this test), was given to Jev, as well as a set of binned multiple choice answers. Answers were scored by Total Variation A) gives an example of a median TV Jev response for each distribution family, B) plots the “relative overestimate” (Jev/theory) as a ratio, and C) the total variation (error, effectively), for each family (lower is better).

Jev has a strong tendency towards distributions that are too peaky, with additional difficulty in smoothly vanishing tails. Jev tends to assign nonzero weight to the tail bins (which have near-zero probability mass in the correct distribution). The Lorentzian is actually pretty good, which is, I suspect, a result of it being peaky and fat-tailed to begin with.

Now, the reason it is good at identifying the peak of the distribution is very likely that the peak appears in the prompt. E.g. for gaussian distributions:

An isolated emission line has center -14.033 MHz above a reference laser and half width at half maximum 6.8364 MHz. Its broadening comes solely from an exponentially decaying excited state.

There isn’t really any other way to specify the problem without giving the mean, or some other characteristic statistic, but Jev then, naturally, just chooses ‘from -16 to -12 MHz’ (p = 0.74) as its output choice. In other words, a lot of these tests can be confused with copying tests, and Jev does. For Maxwell, Rayleigh, and Gamma distributions, where the peak is not one of the parameters that defines the distribution but is instead derived, it finds the peak in ~20% of settings, and its probability distribution is very flat.

Also bizarre is for the Rayleigh and Gamma distributions, it does produce quite credible uniform distributions (I suppose just indicating its uncertainty), but on the uniform distribution, it produces a near delta function around the 50th percentile!

I collected some more plots of Best, Median, and Worst plots as measured by Total Variation to take a look at in the following plots:

The first three columns show representative plots of the best, median, and worst of the trials, per distribution family, measured by TV. The Percentile Average (fourth column) gives a log-scale representation of the ratio between Jev and theory, and the P-P plots Jev’s cumulative probability vs. the theoretical cumulative probability distribution, averaged over all 100 settings for that family. The shading on the Percentile Average and the P-P plot is the middle 80%, and perfect agreement is the dashed black diagonal.

The calibration of the uniform distribution in particular was so surprising(ly bad), with a TV of nearly 0.8 (!), that I decided to look for any prior literature on this.

Kanta Hayashi and Yu Xi Chau both look into something similar and find similarly ‘peaky’ or overconfident choices over nominally uniform distributions:

I asked Jev , TypeSafe AI’s new decision model, to call a fair die roll it could not see. Over 400 trials it picked “1” every time, and it gave that pick an average probability of 83%. It was right 19% of the time, which is chance. [Hayashi, Jev Does Not Play Dice ]

I started with tests that should not require much interpretation. A fair six-sided die gives each face a probability of 16.67 percent. A fair coin gives heads and tails 50 percent each. I asked Jev for its probabilities repeatedly, rather than asking software to sample the die or coin. It assigned a mean probability of 90.01 percent to face 1 and 93.23 percent to heads . [Chau, Jev is fast. It still cannot flip a fair coin. ]

Gu et al. find that LLMs generally (and Jev has the transformer front end…) are also poor at sampling probability distributions, and generally introduce bias.

However, Baldelli et al. find “that probabilistic calibration can be improved through fine-tuning” but that “the gains sometimes reduce downstream capability, especially arithmetic reasoning, with costs varying by model.”

Interesting! A friend of mine pointed out that “mode-seeking” or “mean-seeking” behavior is actually a commonly studied property of machine learning algorithms, particularly in models trained on KL divergence-type losses . Perhaps there’s something there?

This also reminds me of a figure from the GPT-4 technical report:

LLM confidence in the correctness of their answers is very well calibrated for pre-trained models, but the post-training ( i.e. fine-tuning or RLHF) appears to ruin this calibration.

Next, I test whether Jev can accurately pick the distribution appropriate to the same set of problems.

The state is the scenario followed by its question; the instructions ask “Which distribution is appropriate for the answer to this question?”. Gamma with k = 1 is exactly exponential, so both options are correct in that row.

Yes, with extremely high accuracy. The sole exception is problems that require a Gamma distribution, where Jev picked “exponential” in 40/100 settings. However, in 20 of those cases, the Gamma distribution actually reduces to the exponential, so those are correctly assigned.

So Jev does actually know which distributions are correct, it just fails to produce them, and instead prefers concentrating its probability mass on a single bin. This could be downstream of an inability to do math—for example, if you’re given that Maxwell-Boltzmann distribution I mentioned in the first section, and asked for the mean, it is calculable from the distribution, the mass, and the temperature, but the math is multi-step and not trivial.

So, can Jev do math?

Opus 5.5 proposed this tiered ladder of mathematical capability tests, starting with “can Jev copy” (yes) and moving through addition, multiplication, reciprocals, all the way up to the type of calculations (level 8) required to actually answer the Maxwell-Boltzmann questions we asked earlier.

Opus-proposed experiment to check mathematical capavilities
Some example prompt-response pairs for math ability checks.
Summary statistics of prompt-response pairs for math ability checks.

Surprisingly, Jev is pretty decent at math… until it’s not. It goes from being really quite close on even relatively hard math (square roots) to unsure when asked to combine more steps.

But, what exactly is it unsure about? Two hypotheses come to mind—perhaps it’s bad at unit conversions? Perhaps it’s bad at multi-step problems where #steps > 2? Opus suggests that it also might be bad at exponent tracking, which I doubt, but worth checking! We also have to be careful that we’re not biasing the model in the way that we bin its multiple choice answers.

Summary statistics of prompt-response pairs.

Well, there we go. It doesn’t appear to be especially worse at unit conversions, instead it appears Jev just breaks down at multi-step arithmetic.

I can sort of understand why this might be—a transformer is a unidirectional, multi-layer object that has to evolve non-recurrent mechanisms for doing computation. Since Jev has no chain-of-thought scaffolding to ‘save’ intermediate results, it must perform multi-step arithmetic internally, and the model may simply not be deep enough for that to work. Consider this my hand-wavey guess at an explanation.

Some caveats: Jev appears to be REALLY bad at tracking powers of ten (Opus was right!) and you can maybe sort of argue that it prefers answers closer to the correct answer in most cases (the distribution is middle-heavy).

Note: I further checked that I’m not biasing the model too much with answer distributions by shifting the bins up by half a decade—this moved the center of Jev’s distribution by < 0.08 of a decade.

I thought I’d put an addendum here, as this was my first foray into allowing frontier models (Opus 5.5 and Astra 6) to assist with design of experiments.

They are very good at experiment design in the abstract and absolutely awful at catching fatal errors in implementation. For a huge fraction of the original experiments they did, they had put the answer IN the prompt and then reported the data as if the calibration of the model over the uniform distribution had improved! They didn’t do anything wrong , the experiment was faithful to the naive stated intent, and the calibration did, in fact, improve. But the models completely failed to recognize that we had accidentally moved from testing Jev’s calibration to testing Jev’s copying ability.

I think this is a general failure mode of frontier models. I very rarely see the necessary spontaneous metacognition to reread an experiment and think “Hmm, is this testing what I think i’m testing?”

Still an enjoyable experience, but I’m glad that my training in experimental science is still worth something, for now. Also, has anyone else noticed that Claude became British when 5.0 came out? It says “centred” instead of “centered,” and “colour” instead of “color” now. Weird.

Some funny screenshots of my Claude Code session:

Discussion about this post

Ready for more?

[$] Beyond the &

Linux Weekly News
lwn.net
2026-10-02 11:11:11
Rust has a number of kinds of smart pointers, both in the standard library and defined by users. Still, some operations that are possible with built-in references are not possible to perform with user-defined smart pointers. Tyler Mandry, lead of the Rust project's language team, spoke at RustCon...
Original Article
The page you have tried to view ( Beyond the & ) is currently available to LWN subscribers only. Reader subscriptions are a necessary way to fund the continued existence of LWN and the quality of its content.

If you are already an LWN.net subscriber, please log in with the form below to read this content.

Please consider subscribing to LWN . An LWN subscription provides numerous benefits, including access to restricted content and the warm feeling of knowing that you are helping to keep LWN alive.

(Alternatively, this item will become freely available on October 15, 2026)

Brooklyn Will Soon Have One of the Nation’s Only Middle Schools for LGBTQ Kids

hellgate
hellgatenyc.com
2026-10-02 11:10:18
Miss Major Middle, opening in Fort Greene next year, is "an opportunity to show something different."...
Original Article

In late August, on a sunny upper floor of an artsy office building in Fort Greene, a three-person crew laid the foundation for a schoolhouse. With tool belts strapped to their waists, they hammered donated planks into the floor of what will soon become Miss Major Middle , a tuition-free middle school geared towards queer, trans, and ally students, opening next year.

The school's kind-eyed cofounder and co-principal, Joji Florence, excitedly showed Hell Gate around the 1,100-square-foot space, pointing out newly built bins for students' portfolios and the enclosed outdoor space where the first class of sixth graders can have lunch, as well as a studio where they can take classes in dance, movement, and theater. Florence said students will also learn in nearby parks and the Brooklyn Navy Yard, where the school has studio space and site-specific units planned.

"The way we're thinking about our kiddos and the school is that we have this space, and the extension of our space is the community. It sounds pretty basic, but when you're in schools for a long time, it is...not basic," Florence explained. "Schools get rooted. They get insular out of safety issues, out of a lot of things. But we have an opportunity to not do that."

Florence, a former math teacher, began dreaming up Miss Major Middle a decade ago, enrolling in graduate school, shadowing other principals, and meeting with parents for brainstorming sessions to figure out how to do it. Named for the pioneering Black trans author and activist Miss Major Griffin-Gracy , who gave the project her blessing and was involved until her death in 2025 , the school will accept its first class of sixth graders in 2027, with seventh and eighth grades added in 2028 and 2029. Florence wanted to apply their experience working in transformative schools—a model that lets the way students learn guide the curriculum—to create a haven that would support marginalized students, who can be ignored or bullied in more traditional school settings.

Miss Major Middle was always meant to be tuition-free. In 2024, it attempted —and failed—to obtain one of the state's "zombie" charters. Instead, it will function as a "public micro middle school" in partnership with Community School District 13 under New York City Public Schools' homeschooling designation. This means that parents register with the City under its homeschooling rules and, with Miss Major Middle's assistance, must submit quarterly reports to the City's Department of Education showing they are keeping up with its requirements. (The DOE did not respond to Hell Gate's questions by publication time, including about how many schools currently operate under this system.)

The school room at Miss Major Middle (Courtesy Joji Florence)

Give us your email to read the full story

Sign up now for our free newsletters.

Sign up

JetBrains Air: Building a System of Products for Agentic Software Development

Lobsters
blog.jetbrains.com
2026-10-02 11:03:41
Comments...
Original Article

Agentic AI AI JetBrains New Products News Releases

Read this post in other languages:

AI can produce code. Organizations still have to produce software. Agentic development is changing how software gets made, but it hasn’t changed what it costs to be wrong.

Six months ago, we began publicly experimenting with agentic development environments. Around the same time, we introduced JetBrains Central as an open control and execution system for agent-driven development. We subsequently began rolling out JetBrains Central CLI , shared context, cloud agents, automations, governance, and AI cost controls for teams and organizations.

Today, we are bringing this work together as JetBrains Air: an open, coherent system of products for developers, teams, and organizations, inside and beyond JetBrains IDEs. It is multi-surface and multi-service. Each product solves a distinct problem, but the products work better together.

JetBrains Air marks a significant expansion in what JetBrains is building for. For 26 years, we have focused primarily on the individual developer workbench. Now, we are building for the wider system through which agentic work is initiated, executed, coordinated, reviewed, and governed.

The IDE remains important to JetBrains’ future. The era in which the whole software development system can be contained in one window is ending. As part of our continued investment, we are now bringing the foundational agentic experience into JetBrains IDEs, giving professional developers an environment where they can work effectively with agents while understanding, changing, and verifying the resulting code. JetBrains Air connects the wider system developing around it.

That system is based on a core belief that the future of agentic development will be multi-vendor. No single model, agent, or service will be right for every developer, team, or task.

From one product to an open system of products

This strategic shift has a practical consequence: JetBrains Air cannot be just another agent or development environment. It must connect products for individual work, team coordination, organizational control, context, and process automation – and remain open to the tools and agents developers choose, including those JetBrains does not build.

JetBrains Air includes products that are available today alongside others that will be introduced as the system develops:

  • Air in JetBrains IDEs – a complete agentic development experience for directing and orchestrating agents and verifying their work inside JetBrains IDEs.
  • Air Teams – a new way to coordinate and automate software-delivery workflows involving developers and autonomous agents.
  • Air Governance (formerly JetBrains Central) – organizational policy, visibility, auditability, cost management, and accountability for AI-assisted and agent-driven development.

Junie is JetBrains’ coding agent for professional software development. It will be supported across all Air surfaces.

Air in JetBrains IDEs gives developers the environment to direct agents and verify their output using JetBrains’ code intelligence. Air Teams turns individual agent activity into coordinated team workflows. Air Governance makes that activity visible, governable, and accountable across the organization.

But an open system cannot stop at JetBrains’ own products. The Agent Client Protocol ( ACP ) standardizes the connection between an IDE and an agent’s full harness, including its planning, logic, tools, model routing, and observability. Through the ACP Registry, developers can discover and run a growing range of compatible agents while continuing to work inside JetBrains IDEs.

Air Governance is designed to extend visibility and cost governance across providers and the different tools through which agentic work takes place. This means developers can choose the agent, model, or service suited to the task without forcing the organization to give up context, visibility, or control.

Together, the Air products allow work to move between developers, agents, tools, and environments without losing the context and controls surrounding it.

Individual adoption has moved faster than organizational infrastructure

Since March, our products have progressed significantly, but so has our understanding of what agentic development requires.

Developers have been adopting agents faster than organizations can build the infrastructure around them. Agent capabilities have advanced, and different models and agents have proven useful for different tasks. However, the context, coordination, governance, and cost management surrounding them have not kept pace.

For many developers, agents are already delivering practical value. At the organizational level, the economics are much harder to prove. The costs surface elsewhere – in review, rework, security, infrastructure, and spend.

Which agents can access company code? Where can data go? Which output requires human review? What happened while an agent was working remotely? Who approved the resulting change, and how was it verified?

Fragmentation at this level isn’t just irritating. It makes software development harder to understand, measure, and govern at exactly the point when more of the work is being delegated.

The bottleneck is shifting with the work

Code that’s obviously wrong gets caught quickly. That part of the system still works. The harder problem is code that’s almost right: plausible, capable of passing a superficial check, but quietly carrying a bad assumption or architectural inconsistency that won’t surface until it’s expensive.

As agents take on more of the execution, the bottleneck shifts from producing change to understanding, verifying, and owning it. Code becomes cheaper to generate but more expensive to verify. Agent activity becomes easier to start but harder to coordinate, audit, and explain.

And while the work can be delegated, accountability cannot. An agent will not get the call at 3:00 am when something breaks. The responsibility for what ships still belongs to the people and organizations that ship it.

This is why control becomes harder, not easier, as AI improves. A more capable model may produce better output. It does not establish organizational policy, preserve provenance, provide cost visibility, or decide who accepts responsibility for the resulting change.

The future is multi-vendor

Multi-vendor support is a foundational design principle of JetBrains Air, shaping how the system is being built from the outset.

We don’t believe this market will consolidate any time soon. Models vary in what they’re good at, and rankings change every few months. Teams inside the same company already make different choices, and they are often right to do so. Standardizing on one AI vendor today means making a multi-year commitment in a market that won’t look the same next quarter.

Keeping the options open is the reasonable thing to do. The problem is what openness currently costs. Every new model, agent, or service an organization adds takes away a little more visibility into its own development work. Context doesn’t carry over between tools. Spend can’t be attributed. Policies have to be rebuilt for each service.

Organizations should not have to choose between using the best available tools and understanding what is happening inside their own engineering. That trade-off exists because nothing in the current stack was built to sit above several vendors at once.

This is the work JetBrains has taken on. We build our own agent, and we intend to make it excellent. But JetBrains Air does not require customers to use ours, and our strategy does not depend on which model provider leads the rankings this quarter. We have no reason to make the ecosystem smaller than it is.

What we can offer instead is one place to run, see, govern, and account for agentic development across every model, agent, and service – for the developer, the team, and the organization.

Supporting multiple models and agents is the floor, not the ceiling. The part that matters is what sits above them: shared context, one set of policies, a single cost view, and a record of what happened, regardless of which vendor produced the change.

Why JetBrains?

Multi-vendor choice solves only part of the problem. Agents also need reliable software intelligence.

JetBrains brings 26 years of engineering intelligence to the problem, helping developers understand the structure and behavior of complex software, not simply generate more of it. That deterministic code intelligence provides a foundation for making agentic work more reliable, efficient, and understandable across different models and agents. We are seeing promising results from giving AI agents access to deterministic code intelligence.

This is an economic advantage as well as a technical one. Agents spend time and money rediscovering information the codebase already contains. An agent that can retrieve that knowledge is cheaper and more accurate than one that has to reconstruct it. Because intelligence does not belong to one model, the benefit can extend across supported agents and services.

We are also going through the same transition as the organizations we build for, adopting agents internally, redesigning workflows, and learning where individual productivity gains translate into better software delivery and where they simply move work elsewhere.

What comes next

JetBrains Air will develop through a rolling series of releases. We will be explicit about what customers can use now, what is entering preview, and what remains part of our longer-term direction.

Over time, JetBrains Air will extend further into mobile and remote experiences, allowing people to initiate, monitor, review, and continue agentic work as it moves between environments. The goal is not to reproduce the IDE on every surface. We are making the right context and controls available wherever decisions need to be made.

We will also bring JetBrains’ intelligence into more agentic workflows. This includes richer context drawn from code, architecture, repositories, runtime behavior, and organizational knowledge, as well as better ways to route work between developers, models, agents, and services.

More work will be triggered by repository events, schedules, and delivery processes rather than by a developer opening an editor and issuing a prompt. JetBrains Air will provide the intelligence, oversight, and human control these workflows require across surfaces and services.

We will not name future products before their scope and availability are ready to be confirmed. With each release, we will explain what works, how it connects, and what’s still in progress.

Where JetBrains Air is going

The companies that succeed in adopting AI will not necessarily be those that generate the most code or deploy the most agents. They will be those that can expand experimentation without losing quality, context, cost discipline, or human understanding.

JetBrains Air is our commitment to building for that reality. It expands JetBrains from the developer workbench into a system of products connecting developers, agents, teams, and organizations.

The goal is not more code. It is software that developers, teams, and organizations can understand, verify, and stand behind.

Subscribe to JetBrains Blog updates

Discover more

OpenAI’s Medicare attack has exposed Australia’s ‘tech debt’. Fixing it could bring a big bill for taxpayers

Guardian
www.theguardian.com
2026-10-02 11:00:13
Home affairs department orders all federal government agencies to conduct review of ‘legacy technology’ amid fallout from AI agent hacks Get our breaking news email, free app or daily news podcastThe Australian government faces significant “tech debt” that could bring a big bill for taxpayers after ...
Original Article

T he Australian government faces significant “tech debt” that could bring a big bill for taxpayers after the OpenAI Medicare breach , as government agencies will need to fortify their defences against future attacks by AI agents.

This week, the home affairs department ordered all federal government agencies to conduct a “legacy technology stocktake” that requires a plan for each agency to “reduce legacy technology systems” to a level within the agency’s risk tolerance and appetite, the direction stated .

OpenAI this week revealed an internal agent had gained non-public access to the Services Australia Medicare statistics portal during a training task seeking information on government spending on skin conditions in Victoria. The agent was able to run commands, retrieve internal files, credentials, and write files.

Sign up for the Breaking News Australia email

While OpenAI has apologised to Australia for the incident, it has served as a wake up call for the federal government, with the government-wide review now under way.

The finance minister, Katy Gallagher, asked her department whether some of the A$160m funding allocated to the agency in the last budget for cyber upgrades can be accelerated.

The statistics portal, Gallagher told reporters last month, is a “legacy system.”

“It dates back decades.”

Services Australia will be far from alone in managing legacy systems. They can – but not all do – present a security risk for businesses and government as they age and vendors cease providing new security updates.

Prof Salil Kanhere, a University of New South Wales cybersecurity and AI expert said the age of the system alone does not tell an agency whether it needs replacing.

“A 15-year-old system that is properly supported, patched and properly isolated would perhaps present less risk than even a newer system that might not be properly maintained,” he said.

Those older systems with vulnerabilities are often known to human attackers, but AI agents persistent in looking for holes in a system may be able to discover them quicker.

Rogue AI hacks government system for first time - The Latest

Gartner, a technology analysis firm stated in a note to clients released after the Medicare hack that “technical debt, not a rogue AI agent attack” represented the greatest threat to legacy systems.

“Agentic AI’s interactions with [government] resources will greatly increase,” the firm stated. “Underinvestment is no longer sustainable and agencies should urgently prioritise funding in light of AI-driven risks.”

In the Australian government’s commonwealth cybersecurity posture in 2025 report released in February this year , 59% of federal agencies and departments reported their ability to implement the “essential eight” measures to reduce cyber risk was being affected by use of legacy technologies.

The essential eight includes requirements to patch applications and operating systems, using multi-factor authentication, and other security measures. Of those agencies being hindered by legacy tech, 34% blamed insufficient dedicated funding, while 18% said it was due to a lack of a viable replacement.

Prof Yang Xiang, from Monash University’s department of software systems and cybersecurity, said the government stocktake was “very necessary” and there was urgency to needing to audit all government systems.

“The agents bring significant changes in terms of the speed of getting into – hacking into – the system,” he said. “The cost to launch an attack is much reduced to launch an attack and with the help of agents, it is fairly easy for the hacker to launch a very large scale atack against any systems.”

skip past newsletter promotion

Clearing the tech debt could prove costly for the federal government, but Xiang said agencies should identify priority systems for replacement. Kanhere said high-risk systems should take priority.

“You do the high risk stuff first, I think it is absolutely needed, and then put the perimeter around [other systems],” he said. “It is possible to do it quite systematically once they have a good understanding of what needs to be done.”

Some states have audited their legacy technology and have invested hundreds of millions to rectify the issues.

A Victoria government cybersecurity audit of its IT servers found 25% of the operating systems used by servers were no longer supported by the vendor, with 48% in extended support.

In a South Australian audit report of legacy ICT systems published in June, of the ten agencies reviewed, nearly half of the 11,602 hardware devices or appliances were determined to be legacy devices. Almost one quarter of the operating systems and applications were also determined to be legacy.

In one example, the Department for Child Protection’s case management system is now over 15 years old, and has limited vendor support.

The SA government has allocated $325.6m over the past three budgets in part to address legacy technology.

“Frontline workers spend significant time managing system limitations and maintaining records, reducing the time available to support vulnerable children and families.”

A 2025 Queensland government audit of IT systems found more than half of the 57 systems audited are at the end of life. Many of the systems identified by the government in 2012 as needing to be replaced were still in operation in 2025, including a patient administration system at Queensland Health, a forensic register at the Queensland Police Service, and a trust accounts system for young people in detention.

The Queensland government in the 2025 budget allocated $1bn over four years for IT investment, including replacing or updating legacy systems.

The Australian Cyber Security Centre said in recent guidance that the most effective way to mitigate risks associated with legacy IT is to replace it, and where it cannot be replaced, then legacy technology should potentially be segregated or isolated from the broader department network to restrict access to the rest of the department.

‘These guys are just coming from nothing’: questions over multibillion-dollar Firmus float amid datacentre backlash

Guardian
www.theguardian.com
2026-10-02 11:00:12
‘AI factories’ developer on track to launch the second-biggest IPO in Australian history, but investors warn its forecasts are ‘a little bit of a fairytale’Get our breaking news email, free app or daily news podcastFrom her balcony in Launceston, Kayla Thompson can see the buzz of construction at wh...
Original Article

F rom her balcony in Launceston, Kayla Thompson can see the buzz of construction at what will soon be one of Australia’s first AI factories. Her teenage stepchildren get an even clearer view from their classroom window.

Like many of her neighbours, Thompson didn’t realise Firmus Technologies had an ambitious plan to build datacentres in Tasmania until after construction began. The Australian company was on a mission, and preparing to list on the ASX later this month after an anticipated initial public offering designed to raise $7bn from investors.

It is on track to be the second-largest IPO on record behind Telstra’s $14bn share sale in 1997, and one of the top five in the world this year.

Its massive 104-megawatt datacentre at St Leonards, worth an estimated $2.1bn, was given fast-tracked approval by the city of Launceston in September last year without a public hearing.

Thompson was furious. She joined a protest group that believes Tasmanians were kept in the dark by the company.

“Their consultation was nonexistent until they had it shoved in their face that they made a mistake,” says Thompson. “By then, it was too late.

“The more we researched this company, the less appropriate it seemed that size of facility should be voted on by a local council with limited community awareness.”

In June, Firmus realised it had a problem. Community backlash was threatening two other proposed datacentres in the state’s north-east . With an IPO months away, it organised dozens of public consultation sessions with a pre-prepared message.

View of the building of the Firmus AI datacentre in St Leonards, Launceston, Tasmania
The $2.1bn Firmus AI datacentre being built at St Leonards in Launceston. Photograph: Sarah Rhodes/The Guardian

“They did a bit of a mea culpa and said we handled that badly and we should have done something more,” says Joe Zadravec, who lives close to another proposed Firmus datacentre.

“But they only did that after they were caught.”

A Firmus spokesperson confirms the company overhauled its community engagement program after receiving strong feedback from Launceston locals.

“The feedback we receive will continue to help shape our approach,” they say.

Soaring valuation

As Firmus prepares to ask prospective shareholders for billions of dollars in the heat of a global AI boom, its Tasmanian missteps offer an unfiltered look at the operational risks behind the pitch.

A year ago, as local councillors approved the five-hectare AI factory in Launceston, Firmus was worth just under $2bn in a private funding round.

As plans emerged to build liquid-cooled datacentres across the Asia-Pacific alongside AI chip maker Nvidia, which acts as both an investor and hardware supplier, that valuation surged.

In less than 12 months, the implied valuation leaped from $2bn toward a targeted $40bn-plus ASX debut. Now, investment analysts are dangling figures as high as $100bn , a price tag that would make Firmus twice the size of Telstra.

Sign up for the Breaking News Australia email

One investment manager who viewed the company’s draft prospectus tells Guardian Australia the valuation “keeps randomly compounding when nothing has really changed”.

“I’m not saying it’s not worth anything, it’s just not worth the valuation they’re trying to get it away at,” the investment manager says.

The majority of its ambitious pipeline is unbuilt, with the company now operating just two facilities, one in Melbourne and the other in Singapore.

The draft prospectus shows Firmus is forecasting $5bn in annual earnings once its development pipeline progresses, which would put it in the top tier of ASX earners.

“This is a capital-hungry business that needs to keep raising debt and/or equity to fund the losses, and meanwhile we’re being told this will have $5bn of earnings within a few years – it’s a little bit of a fairytale,” the investment manager says.

A Firmus spokesperson declined to comment when asked series of questions by Guardian Australia regarding the feasibility of its earnings forecast and valuation.

The Firmus IPO is a pitch to investors that they can own a stake in the future, with echoes of the speculative allure of Elon Musk’s SpaceX.

A man walking a dog passes a sign that says ‘No Goodman Data Centre’ outside a home in Lane Cove, Sydney, Australia
Goodman Group has withdrawn plans for a datacentre near a school in Lane Cove, Sydney, amid local opposition. Photograph: Hollie Adams/Reuters

But while Firmus relies on the prompt construction of datacentres to justify its valuation, the US, which is further along in the AI datacentre boom, offers a sobering look at the obstacles ahead.

A comparable datacentre offering was recently delayed in the US after bankers struggled to find buyers willing to support a sought-after valuation of US$50bn or more, the New York Times reports .

Part of the problem is that datacentres are hitting community resistance, with US cities Tulsa, New Orleans and Birmingham among those implementing temporary bans on permits and construction.

skip past newsletter promotion

In Sydney, infrastructure company Goodman Group has just withdrawn its plans for a datacentre near a public school amid fierce local opposition, in a clear warning to the industry they must win over local communities.

Community anger typically stems from the scale of the datacentres, their strain on power grids and water supplies, heavy use of diesel backup generators, and the speed of approvals that have left residents feeling shut out.

Just hype?

Firmus is racing to get its multibillion-dollar float away amid opposition to fast-track datacentre approvals, leaving new investors at risk of holding the bag if the excitement proves to be mostly hype.

Rob Talevski, the chief executive of Webull Securities Australia, says there’s a danger retail investors become the “exit strategy” for Firmus’s institutional investors.

New York trading houses Blackstone and Jane Street are among the early backers who may look to the public float to offload their stakes at high prices.

Firmus is offering what is known as a small free float, whereby the company’s founders and early investors initially retain most of the company’s shares.

This can create a temporary scarcity of stock that initially drives the share price higher, giving those backers the ability to offload their holdings at inflated prices before their stock floods the market.

An entrance to the Firmus datacentre in St Leonards, Launceston, Tasmania
An entrance to the St Leonards datacentre. Photograph: Sarah Rhodes/The Guardian

“The question retail might have to ask is, are they basically the liquidity exit strategy for some of the early investors?” says Talevski.

“That’s the bit that probably is not spoken about enough, and that’s where the question marks are for a retail investor.”

There are no escrow arrangements in place for major backers like Blackstone, leaving institutional investors free to sell down their stakes immediately after it lists.

The founders – Oliver Curtis, who has spent time in prison for insider trading , his cousin Tim Rosenfield, and Curtis’s former brother-in-law Jonathan Levee – can begin selling some of their holdings as early as six months after listing if the stock price rallies.

Richard Hemming, the editor at Under The Radar Report, says there are so many interested parties pushing for a successful IPO, including brokers, investment banks and early institutional investors, that the risks to retail investors have been understated.

“Telstra is one of the biggest companies in Australia and these guys are just coming from nothing and saying that they’re going to be bigger than Telstra – it’s gobsmacking,” says Hemming.

Critics of the Firmus datacentre in St Leonards gather to observe construction work
Critics of the St Leonards datacentre gather to observe building work. Photograph: Sarah Rhodes/The Guardian

Back in Tasmania, Firmus has turned to a tested method of countering bad press: sponsoring major sporting teams.

The company’s name will now be carried by the Tasmanian Devils AFL team, Cricket Tasmania, and the JackJumpers and Jewels basketball teams, in what Emma Sherry, the dean of RMIT’s school of management, says is a classic case of “sportswashing”.

“This works until it doesn’t. The goodwill relies on a bet that the community loves the sport enough to forgive the potentially problematic relationship,” she says.

Firmus believes that’s unfair. They say the company was founded in Tasmania, which is home to its staff who are passionate about sport.

“Our sporting partnerships are not a substitute for genuine and thoughtful community engagement,” the spokesperson says.

So far, the deals have done little to ease concerns of vocal critics.

“We are already distrusting of AI in a general sense,” Thompson says. “Why would we trust this company?”

Holes (1996-2025)

Hacker News
plato.stanford.edu
2026-10-02 11:00:03
Comments...
Original Article

Holes are an interesting case study for ontologists and epistemologists. On the one hand, naive descriptions of the world treat holes as objects of reference and quantification, on a par with ordinary material objects (‘This hole has no lid’; ‘That cheese has seven holes’). Moreover, we often appeal to holes to account for causal interactions (‘The hole in the roof caused a leak’), or to explain the occurrence of certain events (‘The water spilled because the bucket has a hole’). All this provides initial evidence for treating holes as real entities. On the other hand, holes appear to be much more elusive than ordinary objects, ontologically as well as epistemologically; after all, they are a paradigm example of nothingness. So one may wonder whether such evidence is to be taken seriously. Should we instead think that all talk about holes is a misleading façon de parler ? Are holes real things, or are they mere entia representationis , as-if entities, fictions?

1. Problems

‘A hole?’ the rock chewer grunted. ‘No, not a hole’, said the will-o’-the-wisp despairingly. ‘A hole, after all, is something. This is nothing at all’. —Ende (1979/1983: 19)

Hole representations—no matter whether veridical—appear to be commonplace in human cognition. Not only do people have the impression of seeing holes; they also form a corresponding concept, which is normally lexicalised as a noun in ordinary languages. (Some languages even discriminate different types of hole, distinguishing e.g. between superficial hollows , inner cavities , and see-through perforations .) Moreover, data from developmental psychology and the psychology of perception confirm that infants and adults are able to perceive, count, and track holes just as easily as they perceive, count, and track paradigm material objects such as cookies and tins (Giralt & Bloom 2000; Horowitz & Kuzmova 2011; Porter et al. 2016). These facts do not prove that holes and material objects are on equal psychological footing, let alone on equal metaphysical footing. But they indicate that the concept of a hole is of significant salience in the common-sense picture of the world, specifically of the spatiotemporal world.

If holes are entities of a kind, then, they appear to be spatiotemporal particulars, like cookies and tins and unlike numbers or moral values. They appear to have a determinate shape, a size, and a location. (‘These things have birthplaces and histories. They can change, and things can happen to them’, Hofstadter & Dennett 1981: 6–7.) On the other hand, if holes are particulars, then they are not particulars of the familiar sort. For holes appear to be immaterial : every hole has a material “host” (the stuff around it, such as the edible part of a donut) and it may have a material “guest” (such as the liquid filling a cavity), but the hole itself does not seem to be made of matter. Indeed, holes seem to be made of nothing , if anything is. And this gives rise to a number of conundrums. For example:

  1. It is difficult to explain how holes can in fact be perceived. If perception is grounded on causation, as Locke urged ( Essay , II-viii-6), and if causality has to do with materiality, then immaterial bodies cannot be the source of any causal flow. So a causal theory of perception would not apply to holes. Our impression of perceiving holes would then be a sort of systematic illusion, on pain of rejecting causal accounts of perception. (On the other hand, if one accepts that absences can be causally efficacious, as urged by Lewis 2004, then a causal account could maintain that we truly perceive holes; see Sorensen 2008 and 2015 along with Siegel 2009, Farennikova 2013, and Calabi 2019.)
  2. It is difficult to specify identity criteria for holes—more difficult than for ordinary material objects. If holes are immaterial, we cannot account for the identity of a hole via the identity of any constituting stuff. But neither can we rely on the identity conditions of the hole’s material host, for we can imagine changing the host—partly or wholly, gradually or abruptly—without affecting the hole. And we cannot rely on the identity conditions of its guest, for it would seem that we can empty a hole of whatever might partially or fully occupy it and leave the hole intact. (Indeed, both “host” and “guest” are relative notions. Doesn’t their proper application presuppose that we already know how to identify holes in the first place? See Meadows 2015.)
  3. It is equally difficult to account for the mereology of holes. Take a card and punch a hole in it. You have made one hole. Now punch again next to it. Have you made another hole? In a way, yes: now the card is doubly perforated. But what prevents us from saying that we still have one hole, though a hole that comes in two disconnected parts? After all, material objects can be disconnected: a bikini, your copy of the Recherche , a token of the lowercase letter ‘i’. Perhaps holes may be disconnected, too? If so, perhaps we have just punched a single, disconnected hole? (Casati & Varzi 2004, 2021)
  4. It is also difficult to assess the explanatory relevance of holes. Arguably, whenever a physical interaction can be explained by appealing to the concept of a hole, a matching explanation could be offered that invokes only material objects and their properties. (That the water ran out of the bucket can be explained by appealing to various facts about fluid dynamics, together with an accurate account of the physical and geometric conditions of the bucket.) Aren’t these latter explanations enough?

Further problems arise from the ambiguous status of holes in figure-ground displays. On the one hand, we have robust evidence that their perception is influenced by such factors as convexity (Arnheim 1948), contour relationship (Bozzi 1975), grouping (Nelson & Palmer 2001), or pattern entropy (Gillam & Grove 2011), and we understand how these factors relate to others, such as texture and shading (Bertamini & Helmy 2012). On the other hand, while it appears that the shapes of holes can be recognized by humans as accurately as the shapes of ordinary objects, the area seen through a hole typically belongs to the background of its host, and there is evidence to the effect that background regions are represented as shapeless (Rubin 1915). So what would the shape of a hole be, if any? (Bertamini & Croucher 2003; Palmer et al. 2008; Nelson et al. 2014; Bertamini & Casati 2015.)

2. Theories

These difficulties—along with some form of horror vacui —may lead a philosopher to favor ontological parsimony or revisionism over naive realism about holes. A number of options are available:

  1. One may hold that holes do not exist at all, arguing that all truths putatively about holes boil down to truths about holed objects (Jackson 1977: 132) or, more generally, that all sentences that seem to imply the existence of holes can be paraphrased by sentences that lack the implication but could in principle be used for all the same purposes as the original (van Inwagen 2014). This view calls for a systematic way of producing the relevant paraphrases. For instance, the sentence ‘There is a hole in the carpet’ can be treated as a mere grammatical variant of ‘The carpet is perforated’; the sentence ‘There are three round holes in that piece of cheese’ as a variant of ‘That piece of cheese is triply-roundly-holed’, etc. (Challenges: Can every quantification over holes be analyzed this way? Consider: ‘There are as many holes in my cheese as crackers on my plate’; Lewis & Lewis 1970: 207. Can a language be envisaged that contains all the necessary shape predicates? Consider: ‘There are two interlocking toroidal cavities in that piece of cheese’; Casati & Varzi 1994: 180. Can every hole-referring noun-phrase be de-nominalized? Consider: ‘The hole in the tooth was smaller than the dentist’s finest probe’; Geach 1968: 12.)
  2. One may hold that holes do exist, but they are nothing over and above the regions of spacetime at which they are found (Wake et al. 2007). Surely they are not just regions of space, for holes can move—as happens any time you move a donut—whereas regions of space cannot. However, as regions of spacetime, holes can be said to move in virtue of having different temporal parts follow one another in different places. (Challenges: Is this view compatible with a relationist metaphysics of spacetime? See Hestevold 2020, ch. 6. Is it compatible with the possibility that holes undergo rotational motion? Take the donut and spin it clockwise. Now take a wedding ring, put it inside the hole in the donut, and spin it counterclockwise. It would seem that the two holes are spinning in opposite directions. Since the relevant temporal part of the little hole is a spatiotemporal part of the bigger one, would it be spinning in both directions? See Lewis & Lewis 1970: 208.)
  3. One may hold instead that holes are qualified portions of spacetime (Miller 2007). There would be nothing peculiar about such portions as opposed to any others that we would not normally think of as being occupied by ordinary material objects, just as there would be nothing more problematic, in principle, in determining under what conditions a certain portion counts as a hole than there is in determining under what conditions it counts as a dog, a statue, or whatnot. (Challenge: What if there were truly unqualified portions of spacetime, in this or some other possible world? Would there be truly immaterial entities inhabiting such portions, and would holes be among them?)
  4. One might also hold that holes are ordinary material beings: they are neither more nor less than superficial parts of what, on the naive view, are their material hosts (Lewis & Lewis 1970; Mollica 2022). For every hole there is a hole-lining and for every hole-lining there is a hole; on this conception, the hole is the hole-lining. (Challenge: This calls for an account of the altered meaning of certain predicates or prepositions. Would holes surround themselves? Would filling a hole amount to filling space outside it? Would expanding the hole-lining amount to enlarging the hole?)
  5. Alternatively, one might hold that holes are parts of their material hosts, viz. “negative” parts (Hoffman & Richards 1985; Spröte & Fleming 2013). On this account, a donut would be a sort of hybrid mereological aggregate—the mereological sum of a positive pie and the negative bit in the middle. (Again, this calls for an account of the altered meaning of certain modes of speech. For instance, making a hole would amount to adding a part , and splitting a donut open would amount to removing a part , contrary to ordinary usage.)
  6. Yet another possibility is to treat holes as “disturbances” of some sort (Karmo 1977). On this view, a hole is always found in some object (its “medium”) in the same sense in which a knot may be found in a rope or a wrinkle in a carpet, and in each case the relevant in -relation would entail ontological dependence. (Linguistically, this feature of disturbances is naturally reflected in the corresponding have -constructions: ‘The cheese has a hole’, ‘The rope has a knot’, ‘The carpet has a wrinkle’; see Moltmann 2019. Metaphysically, however, the status of disturbances calls for refinements. Simons 1987: 308 suggests construing them as Husserlian moments that continuously change their fundaments, but this seems to suit knots and wrinkles better than holes.)
  7. Finally, it may be held that holes are not the particulars they seem to be. Perhaps they are properties, viz. spatial qualities of their owners (O’Shaughnessy 2000: 333), or ways things are (Meadows 2013), or relational entities whose fundamental mode of being is being-in (McDaniel 2010). Perhaps holes are ways the world is, viz. non-abstract, localized absences that can serve as truth-makers for negative existentials or false-makers for positive existentials (Martin 1996). Or perhaps holes are genuine negative facts, viz. anti-instantiations of material properties by material things: an absence of F-stuff constitutes a hole insofar as it is surrounded by the presence of F-stuff (Barker & Jago 2012).

On the other hand, the possibility remains of taking holes at face value, which is to say as bona fide immaterial particulars. Any such undertaking would have to account, not only for the general features mentioned in section 1, but also for a number of additional peculiarities (Casati & Varzi 1994). Among others:

  1. Holes are ontologically parasitic: they are always in something and cannot exist in isolation. (‘There is no such thing as a hole by itself’, Tucholsky 1931: 100; ‘Take away the thing, and the hole goes too’, Heath 1967: 524). This is why there cannot exist a world consisting only of holes (Coggins 2010: 71).
  2. Holes are fillable. You don’t necessarily destroy a hole by filling it up (except perhaps when the filling is homogeneous with the hole-lining); you don’t create a new hole by removing the filling.
  3. Holes are locationally gracious. When you fill or put something in a hole, the hole does not squeeze to the side; it shares its location with its guest (and if the guest has a hole, you get two co-located entities of the same kind).
  4. Holes are mereologically structured. They have parts and can bear part-whole relations to one another (though not to their hosts and not to their guests).
  5. Holes are topologically assorted. Superficial hollows are distinguished from internal cavities; straight perforations are distinguished from knotted tunnels.

As is often the case, the choice between all these alternatives—whether holes are to be subjected to Ockham’s razor, reduced to other entities, or taken at face value—will depend on one’s general metaphysical inclinations (Lewis & Lewis 1996). It may also depend on controversial details concerning the nature of space and spacetime, e.g. whether they are relational or substantival (Braddon-Mitchell & Miller 2015). It is, more generally, an instance of the sort of decision philosophers have to make when they scrutinize the ontology inherent in the common-sense picture of the world and the concepts, words, and purposes through which it is described and apprehended.